Linker for double-strand sequence determination
An adapter with a covalently closed end and linker improves sequencing accuracy by enabling accurate sequencing of both strands on amplification-free platforms, addressing the challenges of sequencing artifacts and low read coverage in nanopore sequencing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KEYGENE NV
- Filing Date
- 2024-04-04
- Publication Date
- 2026-05-01
Smart Images

Figure 2026513875000050 
Figure 2026513875000051 
Figure 2026513875000052
Abstract
Description
[Technical Field]
[0001] This invention relates to the field of molecular biology, more specifically to the fields of genomics and gene research. In particular, this invention relates to the field of (deep) sequencing. Novel methods and means for improving the accuracy of deep sequencing, preparing libraries, and reducing the complexity of nucleic acid samples are disclosed. [Background technology]
[0002] A crucial element of genetic research is the sequencing analysis of known and unknown nucleotide sequences. However, current sequencing methods have limitations in their ability to accurately determine the sequences of nucleic acid molecules. For example, in genotyping and SNP calling, it is essential to be able to distinguish between genuine nucleotides (variants) and sequencing artifacts, especially when read coverage is low.
[0003] Long-read and amplification-free sequencing platforms, such as nanopore sequencing, are emerging. Longer reads are particularly preferred in whole-genome sequencing and mapping. It has been shown that long sequence reads significantly improve the quality of de novo genome assembly, especially for species with large and / or complex genomes. Furthermore, long-read sequencing technologies have the potential to resolve long repeats, polyploidy, and haplotypes, and facilitate the identification of genetic elements associated with complex traits through the detection of copy number variations (CNVs).
[0004] The nanopore sequencing technology platform can generate readouts on the size of hundreds of kilobases, and potentially exceed megabase-sized readout lengths. This platform relies on single-stranded nucleic acid molecules passing through tiny protein channels (nanopores) embedded in an electroresistive membrane. The single molecule entering the nanopore causes a characteristic disturbance in the electric current. By measuring this disturbance, DNA or RNA molecules can be characterized.
[0005] PacBio's single-molecule sequencing technology can generate highly accurate reads up to 25kb in length by reading both strands of a single (circular) dsDNA molecule multiple times. The platform relies on the detection or incorporation of fluorescently labeled nucleotides, with the incorporation duration recorded and measured. The incorporation duration and fluorescence signal are used to determine which nucleotides were incorporated and, if present, what modifications were present in those nucleotides. Because the original DNA strand is measured, amplification errors do not affect the resulting sequencing accuracy.
[0006] Element Biosciences' Avidity and MGI are examples of short-read sequencing platforms. Element Biosciences' Avidity-based chemiactivation involves attaching cyclic library molecules to a low-binding surface, where a rolling circle reaction is used to amplify the molecules. Sequencing is based on synthetic sequencing, where incorporated nucleotides are fluorescently labeled after the DNA strand has been extended.
[0007] The MGI DNBSEQ platform utilizes DNBSEQ technology, which circulates sequencing library molecules by ligating both ends of a DNA strand using a connecting oligo, and then amplifies the circulated molecules by rolling circle amplification (RCA). The resulting product is called DNA nanoballs (DNBs) and is bound to a patterned array. The patterned array of DNBs is used as input for sequencing using combinatorial probe anchor synthesis (cPAS) technology, which includes primer hybridization to the adapter region of the DNBs, incorporation of fluorescently labeled dNTP probes, washing, and imaging.
[0008] In this technical field, there is still a need to improve the read accuracy of sequencing methods, particularly the accuracy of long-read and / or amplified free sequencing methods. In this field, there is a particular need to improve the read sequencing accuracy of nanopore sequencing techniques. [Overview of the project]
[0009] The present invention can be summarized by the following embodiments.
[0010] Embodiment 1. An adapter that is at least partially double-stranded, comprising an open end and a closed end, the closed end being covalently closed by a linker, the linker preferably not composed of a deoxyribonucleotide.
[0011] Embodiment 2. The adapter according to Embodiment 1, wherein the linker includes a portion of general formula (L1), and 5' and 3' are used to represent binding sites to nucleotides that form the ends of the adapter: [ka] During the ceremony, m is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10; n is either 0 or 1; r 1 and r 2 These are either H in each case, or together form a bridge portion -(O) 0~1 -(CH2) 1~6 -forms, Here, preferably, if n is 1, then m is 0, or if m is not 0, then n is 0. Preferably the linker is a C3 spacer (m is 0, n is 1, r 1 H, r 2 (where is H), spacer 9 (m is 3, n is 0), spacer 18 (m is 6, n is 0), or one or more 1',2'-dideoxyribose (m is 0, n is 1, r 1 and r 2 The group consisting of (which together form a bridging portion -O-(CH2)2-) is selected.
[0012] Embodiment 3. The adapter according to Embodiment 1 or 2, wherein the adapter includes staggered open ends.
[0013] Embodiment 4. The adapter according to any one of the preceding embodiments, wherein the adapter is formed of a single-stranded nucleic acid molecule containing a linker.
[0014] Embodiment 5. The adapter according to any one of the preceding embodiments, wherein the adapter contains an identifier sequence.
[0015] Embodiment 6. a) providing a sample containing a double-stranded nucleic acid molecule and the adapter according to any one of Embodiments 1 to 5; b) ligating the adapter to the ends of the double-stranded nucleic acid molecule, thereby generating a double-stranded nucleic acid molecule having one open end and one closed end; c) sequencing both strands of at least a part of the double-stranded nucleic acid molecule, preferably by a single sequencing reaction performed on an amplification-free sequencing platform, to generate a double-stranded read; d) generating a consensus sequence from the double-stranded read and determining the sequence within the double-stranded nucleic acid molecule A method for determining a target sequence in a double-stranded nucleic acid molecule, comprising:
[0016] Embodiment 7. Step b) is b1) ligating the adapter to both ends of the double-stranded nucleic acid molecule, thereby closing both ends of the nucleic acid molecule; b2) optionally exposing the sample to an exonuclease; b3) cleaving the closed double-stranded nucleic acid molecule, thereby generating a double-stranded nucleic acid molecule having one open end and one closed end The method according to Embodiment 6, comprising the (sub)steps of:
[0017] Embodiment 8. The method according to Embodiment 6 or 7, wherein the nucleic acid molecule is provided by fragmentation of a longer nucleic acid molecule, and the longer nucleic acid molecule is preferably a genomic nucleic acid molecule.
[0018] Embodiment 9. The method according to Embodiment 8, wherein fragmentation is carried out by restriction endonuclease and / or site-specific endonuclease digestion.
[0019] Embodiment 10. The method according to Embodiment 9, wherein a restriction enzyme and / or site-directed endonuclease generates a single-stranded overhang in a double-stranded nucleic acid molecule.
[0020] Embodiment 11. The method according to any one of Embodiments 6 to 10, wherein the nucleic acid molecule provided in step a) includes a single-stranded overhang, and the adapter includes an overhang that can be ligated to the overhang of the nucleic acid molecule.
[0021] Embodiment 12. The method according to any one of Embodiments 7 to 11, wherein the closed double-stranded nucleic acid molecule of step b3) is cleaved by a site-specific endonuclease or restriction endonuclease, preferably the site-specific endonuclease is an RNA-induced CRISPR nuclease or a TALEN.
[0022] Embodiment 13. The method according to any one of Embodiments 6 to 12, wherein the method is performed on multiple samples, preferably with multiple samples pooled before step c).
[0023] Embodiment 14. The method according to any one of Embodiments 6 to 13, wherein the sequencing adapter is ligated to the open end of the double-stranded nucleic acid after step b) and before step c).
[0024] Embodiment 15. The method according to any one of Embodiments 6 to 14, wherein the amplification-free sequencing platform in step c) is nanopore sequencing, preferably nanopore selective sequencing. [Brief explanation of the drawing]
[0025] [Figure 1]This is a schematic diagram of library preparation of λDNA HindIII digests. In step a, an index adapter (shaded area) is ligated to a fragment of λDNA HindIII digest; the diagram shows a fragment containing a PvuI restriction site, with each strand indicated by a black bar and the restriction site by a white bar. After DNA repair and A addition, the indexed fragment is closed in step b by ligation with a linking adapter (spotted area, with linker). In step c, the unclosed fragment is removed by exonuclease treatment. Subsequently, the closed fragment is opened by PvuI restriction (step d), an adapter for nanopore sequencing is added (white bar at the open end of the fragment, with helicase motor protein enclosed in a white circle) (step e), and the fragment is sequenced (step f). [Figure 2] This is a schematic diagram of λDNA restriction by sgRNA-induced (red arrow) CRISPR-Cas9. The three DNA fragments generated by CRISPR-Cas9 (17.5Kb, 15Kb, and 12.5Kb) each contain one recognition site for restriction enzyme BSaI or XbaI (blue arrows). The lengths of the fragments after BSaI or XbaI digestion are shown by dashed lines, and the Cos-terminal fragment is shown by a dotted line. [Figure 3-1] Figure 3: Tapestation profile showing the normal generation of DNA fragments of expected length after CRISPR-Cas9 digestion. (A) Gel-based visualization and (B) Electrophoresis report are shown. [Figure 3-2] (As stated above.) [Modes for carrying out the invention]
[0026] definition Various terms relating to the methods, compositions, uses, and other aspects of the present invention are used throughout this specification and the claims. Unless otherwise indicated, such terms have their common meanings in the art to which the present invention pertains. Other specifically defined terms are to be interpreted in a manner consistent with the definitions set forth herein. Similar or equivalent methods and materials described herein may be used in carrying out the tests of the present invention, but preferred materials and methods are described herein.
[0027] The methods for carrying out the prior art used in the present invention will be apparent to those skilled in the art. The practices of the prior art in molecular biology, biochemistry, computational chemistry, cell culture, recombinant DNA, bioinformatics, genomics, sequencing, and related fields are well known to those skilled in the art and are described, for example, in the following references: Sambrook et al. Molecular Cloning. A Laboratory Manual, 2nd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989; Ausubel et al. Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1987 and periodic updates; and the series Methods in Enzymology, Academic Press, San Diego.
[0028] "A," "an," and "the": These singular terms refer to multiple objects unless otherwise specified in the context. Therefore, for example, a reference to "cell" can include a combination of two or more cells.
[0029] In specification usage, the term "about" is used to describe and account for small variations. For example, the term may refer to ±(+ or -) 10% or less, e.g., ±5% or less, ±4% or less, ±3% or less, ±2% or less, ±1% or less, ±0.5% or less, ±0.1% or less, or ±0.05% or less. Furthermore, quantities, ratios, and other numerical values may also be presented in range form. Such range forms are used for convenience and conciseness and should be understood flexibly to include numerical values explicitly specified as range limits, but should also be understood to include all individual numerical values or subranges that fall within that range, as if each numerical value and subrange were explicitly specified. For example, a ratio in the range of about 1 to about 200 should be understood to include not only the explicitly stated limits of about 1 and about 200, but also individual ratios such as about 2, about 3, and about 4, and subranges such as about 10 to about 50, about 20 to about 100, etc.
[0030] As used herein, the term “adapter” means a single-stranded, double-stranded, partially double-stranded, Y-shaped, or hairpin-shaped nucleic acid molecule that is preferably chemically synthesized, has a limited length, such as about 10 to about 200, or about 10 to about 100 base pairs, or about 10 to about 80, or about 10 to about 50, or about 10 to about 30 base pairs, and may be attached to, preferably ligated to, the ends of other nucleic acids, such as one or both strands of a double-stranded DNA molecule. The double-stranded structure of the adapter may be formed by two different oligonucleotide molecules that form base pairs with each other at least partially, or by a single, optionally modified oligonucleotide chain. Optionally, the single oligonucleotide chain is modified to include a linker. Preferably, at least one end of the adapter is designed to fit with the end of a nucleic acid. The fitting end may be a blunt end or a single-stranded overhang. The binding ends of the adapter may be designed to fit with overhangs resulting from cleavage by restriction enzymes and / or site-specific nucleases, and to be optionally ligated, or they may be designed to fit with overhangs generated after addition in non-template extension reactions (e.g., 3'-A addition).
[0031] "and / or": The term "and / or" refers to a situation in which one or more of the described cases may occur individually or in combination with at least one of the described cases, up to all of the described cases.
[0032] As used in relation to nucleic acids or nucleic acid reactions, "amplification" refers to an in vitro method for creating copies of a specific nucleic acid, such as a target nucleic acid or tagged nucleic acid. Numerous methods for amplifying nucleic acids are known in the art, and amplification reactions include transcription-mediated amplification methods such as polymerase chain reaction, ligase chain reaction, strand displacement amplification, rolling circle amplification, NASBA (e.g., U.S. Patent No. 5,409,818), loop-mediated amplification methods (e.g., "LAMP" amplification using loop-forming sequences as described in U.S. Patent No. 6,410,278), and isothermal amplification. The nucleic acid to be amplified may include DNA or RNA, be composed of or derived from DNA or RNA, or be a mixture of DNA and RNA including modified DNA and / or RNA. The products resulting from the amplification of nucleic acid molecules or groups of molecules (i.e., "amplification products") may be either DNA or RNA, a mixture of both DNA and RNA nucleosides or nucleotides, or modified DNA or RNA nucleosides or nucleotides, regardless of whether the starting nucleic acid is DNA, RNA or both.
[0033] A “copy” may, but is not limited to, a sequence that has complete sequence complementarity or complete sequence identity with respect to a particular sequence. Alternatively, a copy may not necessarily have complete sequence complementarity or identity with this particular sequence, and some degree of sequence diversity may be permitted. For example, a copy may include nucleotide analogs such as deoxyinosine or deoxyuridine, intentional sequence alterations (such as sequence changes introduced via primers containing sequences that are hybridizable but not complementary to a particular sequence), and / or sequence errors that occur during amplification.
[0034] In this specification, the term “complementarity” is defined as the sequence identity of a sequence to a fully complementary chain (e.g., a second or reverse chain). For example, a 100% complementary (or fully complementary) sequence is understood herein to have 100% sequence identity with a complementary chain, and a sequence that is, for example, 80% complementary is understood herein to have 80% sequence identity with a (fully) complementary chain.
[0035] "Includes": This term is to be interpreted as comprehensive and unrestricted, and not exclusive. Specifically, this term and its variations mean that the specified features, steps, or components are included. These terms should not be interpreted as excluding the existence of other features, steps, or components.
[0036] "Construct," "nucleic acid construct," or "vector": This refers to an artificial nucleic acid molecule obtained from the use of recombinant DNA technology, which can be used to deliver exogenous DNA into host cells, often with the aim of expressing the DNA region contained in the construct within the host cell. The vector backbone of the construct may be, for example, a plasmid into which a (chimeric) gene has been incorporated, or, if a suitable transcriptional regulatory sequence (e.g., an (inducible) promoter) already exists, only the desired nucleotide sequence (e.g., a coding sequence) is incorporated downstream of the transcriptional regulatory sequence. The vector may contain further genetic elements such as selectable markers and multiple cloning sites to facilitate their use in molecular cloning.
[0037] In the use of this specification, the terms “double-stranded” and “double helix” refer to two complementary sequences that form a base pair, i.e., a hybrid together. A “double-stranded” sequence may be the result of base pairing of two complementary polynucleotides, or it may be the result of base pairing of two complementary sequences within a single (modified) polynucleotide chain. Complementary nucleotide chains are also known in the art as reverse complementary chains. In this specification, two complementary chains may also be described as forward and reverse chains, or first and second chains.
[0038] In the usage herein, the term “effective amount” refers to the amount of a biological activator sufficient to produce the desired biological effect. All enzymes used in the methods disclosed herein are used in an effective amount. For example, in some embodiments, the effective amount of exonuclease refers to the amount of exonuclease sufficient to induce, for example, the cleavage of unprotected or “open” nucleic acids. As will be understood by those skilled in the art, the effective amount of a drug may vary depending on various factors such as the drug used, the conditions under which the drug is used, and the desired biological effect, for example, the degree of nuclease cleavage to be detected.
[0039] “Exemplary”: This term means “serving as an example, case, or illustration,” and should not be construed as excluding other configurations disclosed herein.
[0040] "Expression": This refers to the process by which appropriate regulatory regions, particularly DNA regions operably linked to promoters, are transcribed into RNA, which is then translated into proteins or peptides.
[0041] In this specification, "guide sequence" is understood as a sequence that guides an RNA or DNA guide endonuclease to a specific site on an RNA or DNA molecule. In the context of gRNA-CAS complexes, "guide sequence" is further understood in this specification as a portion of sgRNA or crRNA necessary to target the gRNA-CAS complex to a specific site on double-stranded DNA.
[0042] The gRNA-CAS complex is a CAS protein, also referred to herein as CRISPR endonuclease or CRISPR nuclease, which is complexed or hybridized with a guide RNA, the guide RNA may be crRNA, a combination of crRNA and tracrRNA, or sgRNA.
[0043] "Identity" and "similarity" can be readily calculated using known methods. "Sequence identity" and "sequence similarity" can be determined by aligning two peptide sequences or two nucleotide sequences using a global or local alignment algorithm, depending on the lengths of the two sequences. Sequences of similar lengths are preferably aligned using a global alignment algorithm (e.g., Needleman-Wunsch) that optimally aligns the sequences over their entire length, while sequences of significantly different lengths are preferably aligned using a local alignment algorithm (e.g., Smith-Waterman). Sequences may then be referred to as "substantially identical" or "essentially similar" if they share at least a certain minimum percentage of sequence identity (as defined below) (for example, when optimally aligned by the GAP or BESTFIT program using default parameters). GAP uses the Needleman and Wunsch global alignment algorithm to align two sequences over their entire length (total length), maximizing the number of matches and minimizing the number of gaps. When two sequences have similar lengths, global alignment is preferably used to determine sequence identity. Generally, default GAP parameters of gap formation penalty = 50 (nucleotides) / 8 (proteins) and gap elongation penalty = 3 (nucleotides) / 2 (proteins) are used. For nucleotides, the default weight matrix is nwsgapdna, and for proteins, the default weight matrix is Blosum62 (Henikoff & Henikoff, 1992, PNAS 89, 915-919).Sequence alignment and sequence identity percentage scores may be determined using computer programs such as the GCG Wisconsin Package, Version 10.3, available from Accelrys Inc., 9685 Scranton Road, San Diego, CA 92121-3752 USA, or using open-source software such as the EmbossWIN version 2.10.0 programs "needle" (using the global Needleman-Wunsch algorithm) or "water" (using the local Smith-Waterman algorithm), using the same parameters as GAP described above, or using default settings (for both "needle" and "water," and for both protein and DNA alignment, the default gap start penalty is 10.0, the default gap stretch penalty is 0.5; the default score matrix is Blosum62 for protein and DNAFull for DNA). If the overall lengths of the sequences differ significantly, local alignment using the Smith-Waterman algorithm or similar is preferred.
[0044] Alternatively, the percentage of similarity or identity may be determined by searching public databases using algorithms such as FASTA or BLAST. Therefore, the nucleic acid and protein sequences of the present invention can be further used as "query sequences" to perform searches against public databases, for example, to identify other family members or related sequences. Such searches can be performed using the BLASTn and BLASTx programs (version 2.0) from Altschul, et al. (1990) J.Mol.Biol.215:403-10. A BLAST nucleotide search can be performed using the NBLAST program, with a score of 100 and a word length of 12, to obtain nucleotide sequences homologous to the nucleic acid molecule of the present invention. A BLAST protein search can be performed using the BLASTx program, with a score of 50 and a word length of 3, to obtain amino acid sequences homologous to the protein molecule of the present invention. For comparative purposes, Gapped BLAST may be used to obtain gap alignment, as described in Altschul et al., (1997) Nucleic Acids Res. 25(17):3389-3402. When using the BLAST and Gapped BLAST programs, the default parameters of each program (e.g., BLASTx and BLASTn) may be used. Please refer to the National Center for Biotechnology Information website (http: / / www.ncbi.nlm.nih.gov / ).
[0045] The term "nucleotide" includes, but is not limited to, naturally occurring nucleotides, including deoxyribonucleotides and ribonucleotides, which are the basic components of DNA and RNA. The basic structure of a nucleotide is a nitrogen (purine or pyrimidine) base, a pentose sugar, and a phosphate group. The nitrogen base (or nucleic acid base) may be guanine, cytosine, adenine, thymine, and uracil (G, C, A, T, and U, respectively). The term "nucleotide" is further intended to include not only known purine and pyrimidine bases, but also moieties containing other modified heterocyclic bases. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated ribose, or other heterocyclic structures. Furthermore, the term "nucleotide" may include moieties containing haptens or fluorescent labels, and may also include other sugars, not just conventional ribose and deoxyribose sugars. Modified nucleosides or nucleotides include modifications to the sugar moiety, such as one or more hydroxyl groups being substituted with halogen atoms or aliphatic groups, or being functionalized with ethers, amines, etc.
[0046] The terms “nucleic acid,” “polynucleotide,” and “nucleic acid molecule” are used herein synonymously to describe polymers of any length, such as more than about 2 nucleotides, more than about 10 nucleotides, more than about 100 nucleotides, more than about 500 nucleotides, more than 1000 nucleotides, up to about 10,000 nucleotides, or more, which may be produced enzymatically or synthetically (e.g., deoxyribonucleotides or ribonucleotides) (e.g., PNA as described in U.S. Patent No. 5,948,902 and the references cited herein). Nucleic acids may hybridize with naturally occurring nucleic acids in sequence-specific manner similar to two naturally occurring nucleic acids, for example, to participate in Watson-Crick base pair interactions. Furthermore, nucleic acids and polynucleotides may be isolated (and optionally subsequently fragmented) from cells, tissues, and / or body fluids. Nucleic acids may be, for example, genomic DNA (gDNA), mitochondrial DNA, cell-free DNA (cfDNA), DNA from a library, and / or RNA from a library.
[0047] In the use of this specification, the terms “nucleic acid sample” or “sample containing nucleic acid” refer to any sample containing nucleic acid, wherein the sample is a substance or mixture of substances containing one or more double-stranded nucleic acid molecules, typically in liquid form but not necessarily so. At least one nucleic acid molecule may contain the sequence of interest. The nucleic acid sample used as a starting material in the method of the present invention may be obtained from any source, such as a whole genome, a collection of chromosomes, a single chromosome, one or more regions from one or more chromosomes, a transcriptome, or a selection of transcribed genes, and may be purified directly from a biological source or purified from a laboratory source, such as processed or chemically modified nucleic acid. The nucleic acid sample may be obtained from the same individual, which may be human or another species (e.g., plant, bacterium, fungus, algae, archaea, etc.), or from different individuals of the same species, or from different individuals of different species. For example, the nucleic acid sample may be from cells, tissues, biopsies, bodily fluids, genomic DNA libraries, cDNA libraries, and / or RNA libraries.
[0048] In this specification, the term “target sequence” is understood to mean a nucleotide sequence that is subject to analysis or sequencing, i.e., a nucleotide sequence whose nucleotide order is evaluated by a sequencing method. Known sequencing methods that can be used in the methods of the present invention include, for example, nanopore sequencing, PacBio single-molecule sequencing, MGIDNBSEQ, and Element Biosciences Avidity sequencing. A preferred sequencing method is amplification-free sequencing, which is preferably at least one of nanopore sequencing and PacBio single-molecule sequencing. A preferred sequencing method is nanopore sequencing. Alternatively, or in addition thereto, the sequencing method may provide epigenetic information relating to the target sequence. The target sequence includes, but is not limited to, any gene sequence preferably located in a cell, such as a chromosome (or part thereof), a gene (or part thereof), or a non-coding sequence within or adjacent to a gene. The sequence of interest may be any sequence in a nucleic acid sample, such as a gene, gene complex, locus, pseudogene, regulatory region, highly repeating region, polymorphic region, or a portion thereof. The sequence of interest may be a region containing genetic or epigenetic diversity that suggests a phenotype or disease. The sequence of interest may be subject to further analysis or measures, preferably, but not limited to, copying, amplification, sequencing, and / or other procedures for nucleic acid investigation. Preferably, the sequence of interest is a short or longer sequence of single-stranded nucleotides (i.e., polynucleotides) of double-stranded DNA, where the double-stranded DNA further includes a complementary strand containing a sequence complementary to the sequence of interest. Preferably, the double-stranded DNA is genomic DNA (gDNA) and / or cell-free DNA (cfDNA). The target sequence may be located in organelle genomes such as chromosomes, episomes, transcripts, mitochondrial or chloroplast genomes, or in genetic material that can exist independently of the genetic material itself, such as infectious virus genomes, plasmids, episomes, or transposons. The target sequence may be located within the coding sequence of a gene, or within a transcribed non-coding sequence, such as a 5' untranslated region, a 3' untranslated region, or an intron.The target nucleic acid sequence may be present in a double-stranded nucleic acid or a single-stranded nucleic acid. The target sequence may be, but is not limited to, a sequence that exhibits polymorphism, such as SNPs, or a sequence suspected of exhibiting polymorphism. The target sequence may be a sequence unknown in the art, for example, through de novo sequencing.
[0049] In the use of this specification, the term “oligonucleotide” refers to a single-stranded polymer of nucleotides, preferably about 2 to 200 nucleotides in length, or up to 500 nucleotides. Oligonucleotides may be synthesized or enzymatically produced, and in some embodiments, they are about 10 to 50 nucleotides in length. Oligonucleotides may contain ribonucleotide monomers (i.e., oligoribonucleotides) or deoxyribonucleotide monomers. Oligonucleotides may be, for example, about 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 100, 100 to 150, 150 to 200, or about 200 to 250 nucleotides in length.
[0050] "Plant" refers to the entire plant, or parts of a plant such as tissues or organs obtained from a plant (e.g., pollen, seeds, gametes, roots, leaves, flowers, flower buds, anthers, fruits, etc.), as well as any derivatives thereof, and offspring obtained from such plants by self-pollination or cross-pollination. Non-limiting examples of plants include African eggplant, Allium, Korean thistle, asparagus, barley, beet, bell pepper, bitter melon, ground cherry, gourd, cabbage, canola, carrot, cassava, cauliflower, celery, chicory, green bean, wild lettuce, cotton, cucumber, eggplant, endive, fennel, gherkin, grape, chili pepper, lettuce, corn, melon, oilseed seed Examples of crops and cultivated plants include rapeseed, okra, parsley, parsnips, pepino, pepper, potatoes, pumpkins, radishes, rice, loofah, arugula, rye, snake gourd, sorghum, spinach, loofah, pumpkin, sugar beet, sugarcane, sunflower, tomatillo, tomatoes, tomato rootstock, cruciferous vegetables, watermelon, winter melon, wheat, and zucchini.
[0051] "Plant cells" include protoplasts, gametes, suspension cultures, microspores, and pollen grains, either individually or within tissues, organs, or other organisms. Plant cells may also be part of multicellular structures such as callus, meristematic tissue, plant organs, or explants.
[0052] A "primer" refers to a single-stranded synthetic nucleotide molecule that can prime DNA synthesis. DNA polymerase cannot synthesize DNA without a primer: it can only extend an existing DNA strand in a reaction that directs the order in which nucleotides are assembled using the complementary strand as a template. Nucleotides are incorporated into the complementary DNA strand from the 3' end of the primer, which has hybridized to the complementary DNA strand, using the complementary strand as a template. Primers can be amplification primers used in amplification reactions, including but not limited to polymerase chain reaction (PCR), or sequencing primers used for DNA sequencing.
[0053] A "protospacer sequence" is a sequence within the guide RNA that can be recognized or hybridized with the guide sequence, more specifically with the crRNA, or, in the case of sgRNA, with the crRNA portion of the sgRNA.
[0054] An "endonuclease" is an enzyme that binds to a target or recognition site and hydrolyzes at least one strand of a double-stranded DNA or RNA molecule. Endonucleases are understood herein as site-specific endonucleases, and the terms "endonuclease" and "nuclease" are used synonymously herein. Restriction endonucleases are understood herein as endonucleases that hydrolyze both strands of a double helix simultaneously, introducing a double-strand break into the DNA. "Nicking" endonucleases are endonucleases that hydrolyze only one strand of a double helix, resulting in the DNA molecule being "nicked" rather than cleaved.
[0055] In this specification, an "exonuclease" is defined as an enzyme that cleaves one or more nucleotides from the (open) end (exo) of a polynucleotide.
[0056] "Reducing complexity" or "reducing complexity" is understood herein as reducing the complexity of a nucleic acid sample, such as a sample derived from genomic DNA, cfDNA derived from a liquid biopsy, or an isolated RNA sample. Reducing complexity results in the enrichment of one or more specific nucleic acids contained in the composite starting material, preferably containing the sequence of interest, and / or the generation of a subset of the sample, where the subset contains or consists of one or more specific nucleic acids, preferably containing the sequence of interest contained in the composite starting material, while preferably not containing the sequence of interest, and the amount of nonspecific nucleic acids in the starting material, i.e., before reducing complexity, is reduced by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% compared to before reducing complexity.
[0057] Complexity reduction is generally performed before further analytical or methodological steps, such as amplification, barcoding, sequencing, and epigenetic diversity determination. Preferably, the complexity reduction is reproducible, meaning that, as opposed to random complexity reduction, the same, or at least comparable, subset is obtained when the same sample is subjected to complexity reduction using the same method.
[0058] Examples of methods for reducing complexity include, for example, AFLP® (Keygene NV, the Netherlands; see, for example, European Patent No. 0534858), optional priming PCR amplification, capture probe hybridization, methods described by Dong (see, for example, International Publication No. 03 / 012118 and International Publication No. 00 / 24939), and indexed linking (Unrau P. and Deugau KV (1994) Gene). Methods described in 145:163-169), International Publication No. 2006 / 137733; International Publication No. 2007 / 037678; International Publication No. 2007 / 073165; International Publication No. 2007 / 073171; U.S. Patent No. 2005 / 260628; International Publication No. 03 / 010328; U.S. Patent No. 2004 / 10153; genome splitting (see, for example, International Publication No. 2004 / 022758); sequential gene expression analysis (SAGE; see, for example, Velcurescu et al., 1995, see above; and Matsumura et al., 1999, The Plant Journal, vol.20(6):719-726); and modification of SAGE (see, for example, Powell, 1998, Nucleic Acids See Research, vol.26(14):3445-3446; and Kenzelmann and Muehlemann, 1999, Nucleic Acids Research, vol.27(3):917-918), MicroSAGE (see, for example, Datson et al., 1999, Nucleic Acids Research, vol.27(5):1300-1307), Massively Parallel Signature Sequence Determination (MPSS; see, for example, Brenner et al., 2000, Nature Biotechnology, vol.18:630-634; and Brenner et al., 2000, PNAS, vol.97(4):1665-1670), Self-Subtraction cDNA Library (Laveder et al.,2002, Nucleic Acids Research, vol.30(9):e38), real-time multiplex ligation-dependent probe amplification (RT-MLPA; see, e.g., Eldering et al., 2003, vol.31(23):el53), high coverage expression profiling (HiCEP; see, e.g., Fukumura et al., 2003, Nucleic Acids Research, vol.31(16):e94), general-purpose microarray systems disclosed by Roth et al. (Roth et al., 2004, Nature Biotechnology, vol.22(4):418-426), transcriptome subtraction methods (see, e.g., Li et al., Nucleic Acids Research, vol.33(16):el36), and fragment presentation (see, e.g., Metsis et al., 2004, Nucleic Acids See Research, vol.32(16):el27 for further details.
[0059] "Sequence" or "nucleotide sequence": This refers to the sequence of nucleotides in a nucleic acid, or the sequence of nucleotides within a nucleic acid. In other words, any sequence of nucleotides in a nucleic acid may be called a sequence or nucleic acid sequence. For example, a target sequence is the sequence of nucleotides contained in one strand of a DNA double helix.
[0060] As used herein, the term “sequencing” refers to a method for obtaining the identity of at least 10 consecutive nucleotides of a polynucleotide (e.g., identity of at least 20, at least 50, at least 100, or at least 200 or more consecutive nucleotides). Sequencing is preferably performed by at least one of amplification-free sequencing, preferably nanopore sequencing, and PacBio single-molecule sequencing. Sequencing is preferably performed by nanopore sequencing, such as those commercialized by Oxford Nanopore Technologies (ONT), or by electron detection-based methods, such as Ion Torrent technology commercialized by Life Technologies. Preferably, the next-generation sequencing method is nanopore sequencing, preferably nanopore-selective sequencing.
[0061] "Nanopore selective sequencing" is understood herein to mean the selective sequencing of a single molecule in real time using nanopore sequencing techniques such as Oxford nanopore or Ontera, mapping a streaming nanopore current signal or base call to a reference sequence to eliminate non-target sequences. Depending on the generated data, the sequencing instrument is guided to either continue sequencing the nucleic acid or to stop, thereby removing the nucleic acid from the sequencing pore by reversing the polarity of the voltage across a particular pore for a short enough time to eliminate the non-target molecule and make the pore available for new sequence readings. Examples of nanopore-selective sequencing methods are described herein by reference in Payne et al., 2020 (Nanopore adaptive sequencing for mixed samples, whole exome capture and targeted panels, February 3, 2020; DOI:10.1101 / 2020.02.03.926956) and Kovaka et al., 2020 (Targeted nanopore sequencing by real-time mapping of raw electrical signal with UNCALLED, February 3, 2020; doi:10.1101 / 2020.02.03.931923).
[0062] In the context of the present invention, "double-stranded nucleic acid molecule" may refer to a short range, a longer range, or a selected portion of a nucleic acid. Before performing the method of the present invention, the double-stranded nucleic acid molecule may be contained within a larger nucleic acid molecule, for example, within a larger nucleic acid molecule present in the sample being analyzed. Preferably, the double-stranded nucleic acid molecule contains the sequence of interest.
[0063] "Target sequence" is defined herein as a sequence present in a double-chain acid molecule as defined herein, which is recognized by at least one of the nucleases and nicasses as defined herein.
[0064] In some embodiments, the “multiple” or “set” of nucleic acid molecules used in the methods of the present invention include one or more sequences of interest selected for determination. Optionally, such sets consist of structurally or functionally related nucleic acid molecules. Nucleic acid molecules in the context of the present invention may include, but are not limited to, DNA, RNA, BNA (crosslinked nucleic acid), LNA (locked nucleic acid), PNA (peptide nucleic acid), methylated DNA such as morpholino nucleic acid, glycol nucleic acid, threose nucleic acid, epigenetically modified nucleotides, and mimics and combinations thereof, including both natural and non-natural, artificial, or non-standard nucleotides.
[0065] Detailed explanation The inventors have discovered a method for improving sequencing read accuracy, particularly a method for improving sequencing read accuracy in long-read sequencing using an adapter comprising one open end and one closed end. In a first embodiment, an adapter is provided which is at least partially double-stranded and has one end of the adapter covalently closed by a linker. The adapter of the present invention is also referred to herein as “adapter comprising a linker,” “linking adapter,” “linker adapter,” or “adapter comprising a closed end.” Preferably, the adapter is used in a method for determining a desired sequence within a double-stranded nucleic acid molecule. Preferably, the adapter can be bound to a nucleic acid molecule containing the desired sequence.
[0066] The double-stranded structure of the adapter is preferably formed by a single-stranded oligonucleotide modified to include an (internal) linker, preferably a linker as defined herein. Thus, the modified single-stranded oligonucleotide is also provided. At least a portion of the upstream (5') and downstream (3') sequences of the linker are complementary sequences, resulting in the oligonucleotide folding into a double-stranded (base-paired) structure, thereby positioning the linker at one end of the double-stranded structure, as shown herein as the "closed end" of the adapter, as opposed to the opposite end of the adapter, which is shown herein as the "open end" of the adapter. The open end of the adapter is preferably suitable for ligation to a nucleic acid molecule containing the sequence of interest, and is preferably 5'-phosphorylated. The complementary sequences flanking the linker can flank the linker directly (or "immediately"). Alternatively, one or more nucleotides may be located between the complementary sequences and the linker.
[0067] The modified single-stranded oligonucleotide forming the adapter may include elements A, B, C, D, and E in order from the 5' end to the 3' end, where - Element C is a linker, preferably a linker as defined herein; - Elements B and D are complementary nucleotide sequences of equal length (also referred to in the art as "reverse complementary"), i.e., these complementary sequences together form a double-stranded structure of the adapter. Elements B and D may each contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 or more nucleotides. Elements B and D may each contain about 500, 300, 200, 150, 100, 80, 70, 60, or less than about 50 nucleotides. Preferably, elements B and D each contain about 10-200, 10-100, 10-80, 10-50, or about 10-30 nucleotides. Preferably, elements B and D directly sandwich linker element C; -Elements A and E are optional elements, preferably containing or consisting of a sequence of one or more nucleotides. Optionally, a blunt-ended adapter can be formed by the absence of elements A and E. Optionally, if either element A or E is present and the other is absent, the adapter will have an overhanging 5' or 3' end, respectively, and will also be shown as an alternating-end adapter or a sticky-end adapter. The overhang may consist of 1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 13, 14, or 15 nucleotides. Preferably, element A is absent and element E consists of a thymidine nucleotide, forming an adapter with a 3'-T overhang.
[0068] Preferably, the modified single-stranded oligonucleotide forming the adapter of the present invention includes or consists of the following elements in order from the 5' end to the 3' end: -B, C, and D; -A, B, C, and D; or -B, C, D, and E.
[0069] Preferably, these elements are covalently linked. Those skilled in the art will readily understand that each of elements A, B, D, and E may include an identifier sequence (e.g., a sample identifier, molecular identifier, or UMI and / or signature sequence as defined herein), a primer binding sequence, and / or a restriction element recognition site. Preferably, these functional domains are contained within the double-stranded structure formed by elements B and D. Preferably, such restriction enzyme recognition sites are recognized by enzymes capable of restricting double-stranded DNA.
[0070] Preferably, the single-stranded oligonucleotide comprises the bases adenine (A), guanine (G), cytosine (C), thymine (T), and / or uracil (U), except for element C. Optionally, the single-stranded oligonucleotide comprises the bases adenine (A), guanine (G), cytosine (C), and / or thymine (T). The single-stranded oligonucleotide may comprise DNA, RNA, or a combination of RNA and DNA. Preferably, except for element C, the single-stranded oligonucleotide consists of DNA, RNA, or a combination of RNA and DNA. The oligonucleotide preferably does not contain at least one of modified DNA, modified RNA, peptide nucleic acid (PNA), locked nucleic acid (LNA), glycerol nucleic acid (GNA), cross-linked nucleic acid (BNA), and / or threose nucleic acid (TNA), except for element C. The oligonucleotide preferably consists of DNA, except for element C, i.e., deoxyribonucleotides A, G, C, and / or T. Optionally, one or more deoxyribonucleotides of the single-stranded oligonucleotide of element A, B, D, or E, preferably element B and / or D, contain one or more natural modifications such as methylation groups.
[0071] Preferably, the sequences directly (or adjacently) flanking the linker are complementary sequences. In other words, preferably, elements B and D directly flank linker element C as defined herein. Preferably, about 10–200, 10–100, 10–80, 10–50, or about 10–30 nucleotides immediately upstream and immediately downstream of the linker are complementary sequences. Preferably, the adapter does not form a hairpin structure and does not contain a hairpin structure. A hairpin structure is understood in the art as a single-stranded (unpaired) nucleotide chain between two complementary nucleotide chains. Preferably, at least about 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or about 100% of the nucleotides in the adapter are double-stranded. Preferably, the adapter is a complete double-stranded (base-paired) structure, with the optional exclusion of the adapter's open end. Optionally, the adapter may include staggered (or "sticky") open ends (in other words, the adapter may include either element A or E). Alternatively, the adapter may include a blunt open end (in other words, the adapter may not include either element A or E). Optionally, the adapter may include a 3'-T overhang (in other words, the adapter may include or consist of elements B, C, D, and E, where element E is a single nucleotide that is thymidine).
[0072] Preferably, the open end of the adapter fits to the end of a nucleic acid molecule containing the desired sequence. The fitting end may be a blunt end or a single-stranded overhang. For example, if the nucleic acid molecule contains a 3'-A overhang, the open end of the adapter preferably contains a 3'-T overhang. Similarly, if the nucleic acid molecule is obtained by enzymatic digestion and leaves overhangs of 1, 2, 3, 4, or 5 or more nucleotides, the adapter preferably contains overhangs of 1, 2, 3, 4, or 5 or more nucleotides, complementary to the overhangs of the nucleic acid molecule. In other words, preferably, the adapter contains overhangs that can fit to and ligate to the end of a restriction fragment prepared using restriction enzymes. The restriction fragment is a nucleic acid molecule produced by digestion with a restriction endonuclease. The other (closed) end of the adapter cannot be ligated to a nucleic acid molecule or to the adapter.
[0073] The adapter is preferably limited in length. The length of the adapter is preferably at least sufficient to form a double-stranded structure including a linker at one end of the adapter. Preferably, the adapter forms a double-stranded structure under conditions, preferably optimal conditions, for annealing and / or ligating the adapter to a nucleic acid molecule containing the sequence of interest. The length of the double-stranded adapter is preferably at least about 5, 10, 15, 20, or at least about 25 base pairs. Preferably, the length of the double-stranded adapter is about 10-200, 10-100, 10-80, 10-50, or about 10-30 base pairs. Therefore, preferably, the modified single-stranded oligonucleotide forming the double-stranded adapter is about 20-400, 20-160, 20-100, or about 20-60 nucleotides. The linker is preferably located in or near the center of the modified single-stranded oligonucleotide. Therefore, preferably, if the single-stranded oligonucleotide contains n nucleotides, the linker is located at the 1 / 2n position. Selectively, the linker is located at position 1 / 2n plus x nucleotides, where x is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. Selectively, the linker is located at position 1 / 2n minus x nucleotides, where x is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0074] Element C can be any linker known in the art. Various types of linkers are commercially available. Preferably, the linker is not composed of deoxyribonucleotides. Preferably, the linker is not composed of nucleotides. The linker preferably does not contain nucleobases, where "nucleobase" is also referred to as "nitrogenous base" in the art. Preferably, the linker of the adapter of the present invention is a linker that does not contain nucleobases. In other words, the linker of the present invention does not contain any of adenine, cytosine, guanine, thymine, and uracil. In the context of this description, the linker connects the 3'-OH of one strand of the double-stranded adapter to the 5'-O-P(O)2-OH of the other strand of the double-stranded adapter. In other words, the linker is disposed within the sugar-phosphate backbone of a single-stranded oligonucleotide that can fold back to form the adapter, i.e., between the sugar-phosphate backbones of elements B and D. Those skilled in the art know how to select an appropriate linker.
[0075] Preferred linkers contain a moiety of general formula (L1), where 5' and 3' are used to represent the binding sites to the nucleotides that form the ends of the adapter:
Chemical formula
[0076] Preferably, when n is 1, m is 0. Preferably, when m is not 0, n is 0. Preferably r 1 and r 2 are each H. r 1 and r 2 together form a crosslinking moiety -(O) 0~1 -(CH2) 1~6-If it forms, it is preferably -(CH2) 2~4 -or -O-(CH2) 1~3 And more preferably, (CH2) 2~3 -or -O-(CH2) 1~2 The most preferred form is -(CH2)3- or -O-(CH2)2-, where the oxygen-containing bridged portion is always preferred over analogues of the same length. Preferably, when oxygen is present, r 1 It will be linked to the position marked with an asterisk.
[0077] A more preferred linker includes a portion of general formula (L2), (L3), (L4), or (L5) having the characteristics described for (L1): [ka]
[0078] In a preferred embodiment, the linker comprises or consists of one or more parts having a general formula (L1) selected from the following (reference names in parentheses): n is 0 as shown in general formula (L2), where preferably m is 3 as shown in formula (L7) (spacer 9, also referred to herein as / iSpC9 / or triethylene glycol spacer) or preferably m is 6 as shown in formula (L8) (spacer 18, also referred to herein as / iSpC18 / or hexaethylene glycol spacer); m is 0; n is 1 as shown in general formula (L3), where preferably r 1 and r 2 These are H, as shown in formula (L6) (C3 spacer, also referred to herein as / iSpC3 / or phosphoramidite spacer); n is at least 1; r 1 and r 2These together form a crosslinked portion -O-(CH2)2- as shown in formula (L4), preferably formula (L5), and more preferably as shown in formula (L9), where m is 0 and n is 1 (dSpacer, also referred to herein as / idSp / or 1',2'-dideoxyribose spacer); Particularly preferred linkers include or consist of a portion of formula (L6), (L7), or (L8) or (L9): [ka]
[0079] In some embodiments, the linker is a spacer 9 (having structural formula (L7)) or a spacer 18 (having structural formula (L8)). In some embodiments, the linker is a C3 spacer (having structural formula L6) or a dSpacer (having structural formula (L9)). In some embodiments, the linker is a spacer 9, a spacer 18, or a C3 spacer.
[0080] Preferably, the linker of the adapter of the present invention comprises or consists of up to 6, 5, 4, 3, 2, or 1 portion as defined herein. Preferably, the linker comprises or consists of up to 1 portion of formula (L2), where m is up to 6, 5, 4, 3, 2, or 1. Alternatively, or in addition thereto, the linker may comprise or consist of up to 6, 5, 4, 3, 2, or 1 portion of formula (L9).
[0081] Linkers can be routinely used to incorporate single-stranded oligonucleotides to form double-stranded adapters, and those skilled in the art will recognize which chemical reactions are suitable for this purpose. For example, a phosphoramidite reaction using a 4,4'-dimethoxytrityl (DMT) protecting group can be used where appropriate. Further guidance is available in manuals from commercial linker suppliers such as "integrated DNA technologies, Inc." (Coralville, Iowa, USA). Multiple parts can be used together to form a linker. Preferably, a linker comprises multiple consecutive identical or different parts of any one general formula (L1), where each 5' end (as shown) of a linker is bonded to the 3' end (as shown) of an adjacent linker, for example, one, two, three, four, five, six, seven, eight, nine, ten or more consecutive identical or different parts of general formula (L1), preferably one or two consecutive parts of general formula (L1), forming a linker.
[0082] For example, two or more parts of (L1) can be used consecutively, in which case each 5' (as shown) of the linker is connected to the 3' (as shown) of the adjacent part. In some embodiments, a single part of general formula (L1) is used. In other embodiments, two such parts are used consecutively. In this embodiment, the linker may have general formula (L12). [ka]
[0083] In other embodiments, three such sections are used in succession. In this embodiment, the linker may have a general formula (L13). [ka]
[0084] In other embodiments, more than two consecutive portions are used, in which case the 5' end of each subsequent portion is connected to the 3' end of the preceding portion, as shown above, to connect portions 1 and 2 of (L13). Optionally, four, five, or six consecutive such portions are used. Preferably, one, two, or three consecutive such portions are used. Optionally, a single such portion is used as a linker.
[0085] Preferably, the linker comprises multiple consecutive portions of general formula (L5), where each 5' end (as shown) of the linker is coupled to the 3' end (as shown) of the adjacent linker, for example, one, two, three, four, five, six, seven, eight, nine, or ten or more consecutive portions of general formula (L5), preferably one or two consecutive portions of general formula (L5), forming the linker.
[0086] The linker may contain or consist of one or more consecutive dSpacer portions, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more dSpacer portions. Optionally, the linker may contain or consist of 1, 3, or 5 dSpacer portions, i.e., the linker is dSpacer, dSpacer-dSpacer-dSpacer, or dSpacer-dSpacer-dSpacer-dSpacer-dSpacer. Preferably, the linker contains a maximum of 6, 5, 4, 3, 2, or 1 dSpacer portion. The linker may have structural formula (L10) or (L11), preferably structural formula (L10). [ka]
[0087] A preferred linker has a structural formula selected from the group consisting of (L6), (L7), (L8), (L9), (L10), and (L11), and is preferably selected from the group consisting of (L6), (L8), (L9), and (L10). The linker preferably has structural formula (L6) or (L10).
[0088] Preferably, the linker comprises multiple consecutive portions of general formula (L9), where each 5' end (as shown) of the linker is coupled to the 3' end (as shown) of the adjacent linker, for example, 1, 2, 3, 4, 9, 6, 7, 8, 9, 10 or more consecutive portions of general formula (L9), preferably 1 or 2 consecutive portions of general formula (L9) forming the linker. Preferably, the linker has general formula (L9) or (L10).
[0089] The adapter of the present invention may include an identifier sequence or “barcode” or “tag” for tagging a double-stranded nucleic acid molecule. Therefore, tagging of the double-stranded nucleic acid molecule is preferably achieved by (adapter)hybrid formation and / or (adapter)ligation. The identifier is preferably at least one of a sample identifier, a unique molecular identifier (UMI), and / or a signature sequence as defined herein. Additionally or alternatively, further adapters may be bound or ligated to a nucleic acid molecule to be sequenced. Such further adapters may also include an identifier sequence. The identifier sequence preferably remains covalently bound to the nucleic acid until (including) the sequencing step. Optionally, the nucleic acid to be sequenced includes a combination of identifier sequences. For example, at least one identifier may be introduced by ligation of the adapter of the present invention, and at least one identifier may be added to the open end of the nucleic acid, for example, together with the introduction of a sequencing lymer binding site and / or by binding a sequencing adapter. When a Y-shaped adapter is ligated to the open end of a nucleic acid (containing the adapter) to be sequenced, there may be two identifiers, each located on a different single-stranded arm of the Y-shaped adapter. The nucleic acid constituting the desired sequence may be labeled with three different identifiers in combination with identifiers that may be introduced by ligation of the adapter of the present invention. The functions of the identifiers may differ. For example, the identifier introduced by ligation of the adapter of the present invention may be a sample identifier, and the identifier(s) introduced to the open end may be molecular identifiers, and vice versa. Such identifier(s) may also be useful for tracing back to the template molecule after the amplification process of the labeled nucleic acid molecule. The UMI may be a separate sequence within the adapter of the present invention. The adapter of the present invention may include a sample identifier and a UMI. Additionally or alternatively, the UMI and / or sample identifier may be appended to the nucleic acid molecule containing the desired sequence before ligation of the adapter of the present invention.The UMI and / or sample identifier may be a portion of the primer to be annealed (e.g., a portion of the 5'-terminal overhang of the primer), and / or a further adapter (e.g., an "indexing adapter") that is ligated to a nucleic acid molecule containing the sequence of interest before ligation of the adapter of the present invention. The adapter of the present invention may subsequently be ligated to a further adapter containing the identification sequence, or to a sequence introduced by the primer, and / or the adapter of the present invention may be ligated to the other end of the nucleic acid molecule containing the sequence of interest. Optionally, two indexing adapters may be ligated to a nucleic acid molecule containing the sequence of interest before ligation of the adapter of the present invention and before ligating to a potential sequencing adapter, i.e., one indexing adapter may be at either end of the nucleic acid molecule. Preferably, the indexing adapter has open ends at both ends of the adapter.
[0090] The sample identifier may be a sequence of nucleic acid molecules linked to a specific sample. For example, the adapter used in the method herein may include an identifier sequence specific to a particular sample, and each additional sample may be processed using the adapter having the identifier sequence specific to the additional sample. The processed samples are subsequently pooled, and the resulting sequences may be assigned to a specific sample using the sample identifier sequence. In a preferred embodiment, multiple samples are processed in parallel by ligating the adapter of the present invention in parallel to nucleic acid molecules present in multiple samples and performing the step of pooling the samples before sequencing.
[0091] A UMI is a substantially unique sequence, tag, or barcode specific to a nucleic acid molecule, preferably completely unique, i.e., (substantially or completely) unique for each nucleic acid molecule present in a particular sample. A UMI may have random, pseudo-random, partially random, or non-random nucleotide sequences. A UMI can be used to uniquely identify the original molecule from which a sequencing read is derived. For example, reads from amplified nucleic acid molecules can be combined into a single consensus sequence from each original nucleic acid molecule. As stated above, a UMI may be completely or substantially unique. In this specification, completely unique means that every (adapter-linked) nucleic acid molecule in a sample provided by the method of the present invention contains a unique tag that is different from all other tags contained in further (adapter-linked) nucleic acid molecules used in the method of the present invention. In this specification, substantially unique means that each adapter-linked nucleic acid molecule provided by the method of the present invention contains a random UMI, but a small proportion of these adapter-linked nucleic acid molecules may contain the same UMI. Preferably, a substantially unique molecular identifier is used when there is very little probability that identical molecules containing the same sequence would be tagged with the same UMI. Preferably, the UMI is completely unique with respect to a particular sequence of a nucleic acid molecule. The UMI is preferably long enough to ensure this uniqueness. In some embodiments, less unique molecular identifiers (i.e., substantially unique identifiers as shown above) may be used in combination with other identification techniques to ensure that each nucleic acid molecule is uniquely identified during the sequencing process. The UMI may be used to confirm that a sequence reading originates from a particular nucleic acid molecule, and the UMI may be used to construct a consensus sequence for a single original nucleic acid molecule. Additionally or alternatively, the UMI may be used to confirm, for example, that a particular amplicon originates from a particular original nucleic acid molecule or fragment. The UMI may therefore be ligated or annealed before amplification, where the amplification step may be performed before ligating the adapter of the present invention to the resulting nucleic acid amplicon molecule or fragment.
[0092] The length of the identifier sequence may range from approximately 2 to 100 nucleotides or more, preferably between approximately 4 and 16 nucleotides. The identifier sequence may be a continuous sequence or may be divided into multiple subunits. These subunits may be present in a single adapter and / or primer, or in separate adapters and / or primers. For example, if the nucleic acid molecule to be sequenced is sandwiched between two adapters, each of these two adapters may contain a subunit of the identifier sequence. To obtain a consensus sequence, the sequence readings obtained by the method of the present invention may be grouped based on combining the information of each of the two identifier subunits.
[0093] Preferably, the identifier sequence does not contain two or more consecutive identical bases. Furthermore, preferably, there is a difference of at least two bases, and more preferably at least three bases, between identifier sequences.
[0094] In a further embodiment, provided herein are methods for determining a sequence in a double-stranded nucleic acid molecule, preferably a method for determining a desired sequence in a double-stranded nucleic acid molecule. Preferably, the method is: a) A step of providing a sample comprising a double-stranded nucleic acid molecule and an adapter comprising a linker as defined herein; b) Ligate an adapter to the end of a double-stranded nucleic acid molecule, thereby generating a double-stranded nucleic acid molecule having one open end and one closed end; c) A step of sequencing at least a portion of both strands of a double-stranded nucleic acid molecule in a single sequencing reaction to generate a double-stranded read; d) A step of generating a consensus sequence from the double-stranded reading and determining the sequence within the double-stranded nucleic acid molecule. Includes.
[0095] The nucleic acids to be sequenced (sequences) are, optionally, genomic sequences and / or genomes from an individual such as a human, plant, animal, insect, or microorganism (e.g., virus or bacterium). Therefore, the methods defined herein may be considered, for example, at least one of the following: - Methods for genotyping; -Methods for haplotyping; and - A method for determining the (genome) sequence.
[0096] Step a) of the method of the present invention provides a sample containing a double-stranded nucleic acid molecule and an adapter as described herein, i.e., an adapter with one end closed by a linker. The provided double-stranded nucleic acid molecule preferably contains the sequence of the choice. The sample containing the double-stranded nucleic acid molecule may be of any origin, such as an individual such as a human, animal, plant, or microorganism, and the double-stranded nucleic acid may be any kind, such as cellular, cell-free, or synthetic double-stranded nucleic acid, e.g., genomic DNA, chromosomal DNA, artificial chromosome, plasmid DNA, or episomal DNA, cDNA, RNA, mitochondria, or an artificial library such as BAC or YAC. Preferably, the nucleic acid molecule present in the sample used as the starting material of the method of the present invention is any one of DNA such as genomic DNA, chromosomal DNA, organelle DNA, nuclear DNA, mitochondrial DNA, artificial chromosome, plasmid DNA, episomal DNA, cDNA, and RNA. Preferably, the DNA is chromosomal DNA and preferably endogenous to the cell. The DNA may also be cDNA reverse-transcribed from RNA, where the RNA may originate from any source, such as a human, animal, plant, or microorganism. Preferably, the sample is a plant sample.
[0097] The nucleic acid molecule is preferably a long nucleic acid molecule and is provided, for example, by cell lysis and optionally by organelle lysis. The nucleic acid molecule used in the method of the present invention may have a size of at least about 50 kb, 100 kb, 150 kb, 200 kb, 300 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, or at least about 1000 kb (1 Mb). The nucleic acid used in the present invention may be a high molecular weight (HMW) nucleic acid or a very high molecular weight (uHMW) nucleic acid. HMW nucleic acids may have a length of at least 10 kb. uHMW nucleic acids may have a length of at least 1 Mb. The nucleic acid molecules used in the method of the present invention may have a size of at least 1.1 Mb, 1.3 Mb, 1.5 Mb, 1.7 Mb, 2 Mb, 2.5 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, or at least about 10 Mb.
[0098] Alternatively, a nucleic acid molecule, preferably a long nucleic acid molecule, may be fragmented first to obtain the double-stranded nucleic acid provided in step a). Therefore, in one embodiment, a step of fragmenting a longer nucleic acid molecule may precede step a) of the method provided herein. Fragmentation is preferably fragmentation of a genomic nucleic acid molecule.
[0099] Those skilled in the art are familiar with means of fragmenting nucleic acid molecules, and the present invention is not limited to any particular means of fragmenting nucleic acid molecules. The fragmented nucleic acid is preferably fragmented genomic DNA. DNA, in particular genomic DNA, can be fragmented using any suitable method known in the art. Methods of DNA fragmentation include, but are not limited to, enzymatic digestion and mechanical force.
[0100] Non-exclusive examples of using mechanical force to fragment nucleic acid molecules include acoustic shearing, atomization, sonication, spot shearing, needle shearing, and the use of French pressure cells.
[0101] Preferably, fragmentation is carried out by restriction enzyme and / or site-specific endonuclease digestion. Enzymatic digestion for fragmenting nucleic acid molecules includes, but is not limited to, endonuclease restriction. For example, enzymatic digestion used in AFLP® and / or sequence-based genotyping techniques may further reduce the complexity of the nucleic acid sample. Those skilled in the art know which enzymes should be selected for DNA fragmentation. As a non-limiting example, at least one high-frequency cutter and at least one low-frequency cutter may be used for fragmenting a nucleic acid sample. The high-frequency cutter is preferably, but is not limited to, MseI and has a recognition site of about 3-5 bp. The low-frequency cutter is preferably, but is not limited to, EcoRI and has a recognition site of more than 5 bp.
[0102] In certain embodiments, particularly when the sample contains or originates from a relatively large genome, it may be preferable to use a third enzyme, which is a low-frequency or high-frequency cutter, to obtain a larger set of shorter restriction fragments.
[0103] The enzymatic digestion steps are not limited to specific restriction endonucleases. The endonuclease may be a type II endonuclease such as EcoRI, Msel, or Pstl. In certain embodiments, type IIS or type III endonucleases, i.e., endonucleases whose recognition sequence is located away from the restriction site, may be used, including but not limited to Acelll, AIwI, AIwXI, Alw26I, Bbvl, BbvII, Bbsl, Bed, Bce83I, Bcefl, Bcgl, Binl, Bsal, Bsgl, BsmAI, BsmFl, BspMI, EarI,EciI, Eco3ll, Eco57I, Esp3I, Faul, Fokl, Gsul, Hgal, HinGUII, Hphl, Ksp632I, MboII, Mmel, MnII, NgoVIII, PIeI, RIeAI, Sapl, SfaNI, TaqJI, and ZthllIII. The restriction fragment may be a blunt end or a protruding end, depending on the endonuclease used.
[0104] In a preferred embodiment, at least one of the recognition sites of the high-frequency cutter and the low-frequency cutter is located within or near the target sequence, for example, the recognition site of the high-frequency cutter or the low-frequency cutter is located at a position of approximately 0 to 10000, 10 to 5000, 50 to 1000, or approximately 100 to 500 bases from the target sequence.
[0105] Additionally or alternatively, nucleic acid molecules may be digested using a site-specific nuclease, preferably at least one of CRISPR nuclease, zinc finger nuclease, TALEN, and meganuclease.
[0106] Preferably, restriction enzymes and / or site-specific endonucleases generate single-stranded overhangs in the double-stranded nucleic acid molecule. The generated single-stranded overhangs may provide adapter ligation or become part of the primer binding site. Optionally, the ends of the nucleic acid molecule provided in step a) may be repaired, preferably by polishing the ends of the provided nucleic acid molecule. Thus, in the case of adapter ligation to (optionally fragmented) nucleic acids, the method of the present invention may include a step of polishing the (optionally fragmented) nucleic acid beforehand. Polishing reactions are well known in the art, and those skilled in the art will readily understand how to perform polishing reactions using, for example, Klenow fragments of DNA polymerase, T4 DNA polymerase, and / or mung bean nuclease.
[0107] Additionally or alternatively, nucleic acid molecules containing the desired sequence may be obtained from, or as a result of, amplification of nucleic acid molecules or fragments.
[0108] Additionally or alternatively, the nucleic acid molecule provided in step a) may be polished to create a blunt end, followed by the addition of a 3'-A alternating overhang, which is then preferably partially or completely ligated into a double-stranded adapter in step b), where one end of the adapter is covalently closed by a linker. Similarly, the addition of the 3'-A overhang, i.e., the addition of a deoxyadenosine nucleotide to the 3' end of the (polished) nucleic acid provided in a), may be achieved using any conventional method known to those skilled in the art. The nucleic acid molecule containing the 3'-A overhang may subsequently be ligated into a suitable adapter containing a 3'-T overhang (preferably the adapter of the present invention), which has a deoxythymidine nucleotide overhang at its 3' end. Thus, the method of the present invention may optionally include the step of adding an A tail to an optionally fragmented and / or optionally polished nucleic acid molecule before annealing the adapter to the (optionally fragmented) nucleic acid. A-tail addition reactions are well known in the art, and those skilled in the art can easily understand how to carry out an A-tail addition reaction, for example, by using a Krenow fragment (exo-).
[0109] A nucleic acid sample may contain multiple double-stranded nucleic acid molecules. These multiple nucleic acid molecules may originate from at least one of the same organism, the same tissue, the same cell, the same organelle, and / or the same (fragmented) molecule. Optionally, the sequences of all double-stranded nucleic acid molecules are determined by the method provided herein. Alternatively, the sequences of some of the double-stranded nucleic acid molecules are determined. Preferably, the sequences (of some) of a sample with reduced complexity of double-stranded nucleic acid molecules are determined. In one embodiment, the method may include a step of concentrating the provided sample for the double-stranded nucleic acid molecules of interest, preferably those containing the sequence of interest. Therefore, the step of reducing the complexity of the sample may precede step a) of the method of the present invention.
[0110] Additionally or alternatively, the complexity of the sample may be reduced between steps b) and c) of the method of the present invention. This may be achieved by specifically attaching adapters, including linkers as defined herein, to both ends of a double-stranded nucleic acid containing the sequence of interest, thereby closing both ends of the nucleic acid. Optionally, both ends of the nucleic acid are closed at one end by a first adapter consisting of a linker as defined herein and the other end by a second adapter consisting of a hairpin structure, preferably the second adapter does not include the linker. Since the remaining nucleic acid molecule or fragment has at least one open end, the closed fragment can subsequently be enriched by degrading the remaining nucleic acid molecule or fragment using an exonuclease. The closed fragment can subsequently be opened, for example, using a programmable endonuclease, preferably a restriction enzyme that recognizes a sequence present only in one of the linking adapters forming the closed end, and / or by mechanical shear stress induced, for example, by Megaruptor®, so that the two nucleic acids each have one end closed by a covalent bond and one open end. In assay designs where different adapters are ligated to either end of a fragment (for example, one end of the fragment is closed by a ligating adapter containing a restriction enzyme recognition site, and the other end of the fragment is closed by a ligating adapter that does not contain the restriction enzyme recognition site, or the first adapter contains a linker-closed end, and the second adapter does not contain a linker-closed end), it should be noted that those skilled in the art will know how to do so. For example, by using different restriction enzymes and / or by adding different nucleotides to a particular end, fragments with different ends may be created, thereby creating a fragment with one end that fits only one of two different adapters, and the other end that optionally fits only the other of the two different adapters. Alternatively, or in addition thereto, if two different adapters fit (and can ligate) the same end of a fragment, the adapters may be “mixed” before ligating to the bound adapter.As further shown herein, prior to ligation, when different adapters are mixed in a constant ratio with a constant amount of fragment, fragments containing both adapters ligated at both ends are produced at a constant percentage. Subsequently, sites for amplification and / or sequencing may be added to the open end. For example, a sequencing adapter may be attached to the open end. The adapter may introduce a sequence that enables sequencing of the fragment on an amplification-free sequencing platform, where preferably the sequencing platform is at least one of nanopore sequencing and PacBio single-molecule sequencing. Preferably, the adapter introduces a leader sequence at the 5' end of the nucleic acid molecule of the present invention, and the leader sequence includes a motor protein for nanopore sequencing.
[0111] The specific attachment of an adapter including a closed end as defined herein to both ends of a double-stranded nucleic acid containing the sequence of interest can be achieved by excising the sequence of interest using one or more (programmed) endonucleases capable of generating an overhang that fits with the ligation side of the linker-containing adapter (i.e., the adapter of the present invention).
[0112] Another method for selectively sequencing a subset of nucleic acid molecules and / or fragments in a sample is, in step b), to attach an adapter as defined herein to one end of a double-stranded nucleic acid containing the sequence of interest, and simultaneously attach a sequencing primer binding site to the other end of the nucleic acid molecule. Such primer binding sites can be specifically introduced, for example, by prime editing and / or specific adapter ligation to a suitable overhang.
[0113] The current methods disclosed herein can also be used, for example, in sequence-based genotyping (SBG) techniques for polyploid cells. SBG techniques are described in more detail, for example, in International Publication No. 2007 / 114693, International Publication No. 2006 / 137733, and International Publication No. 2007 / 073165, which are incorporated herein by reference. The SBG techniques described in the art may be modified by attaching an adapter containing a linker described herein to a fragmented nucleic acid sample.
[0114] In step b) of the method of the present invention, an adapter containing a linker is added to a nucleic acid containing the sequence of interest. The nucleic acid may be (partially) double-stranded nucleic acid or single-stranded nucleic acid. Preferred single-stranded nucleic acid molecules are RNA molecules or single-stranded DNA molecules, preferably RNA or cDNA molecules, and preferably cDNA molecules. The adapter containing the linker preferably has alternating ends, preferably a 3' overhang, so that the ends of the adapter can be annealed to nucleotides located at the ends, preferably the 3' end, of the single-stranded nucleic acid containing the sequence of interest. Therefore, preferably, the adapter containing the linker is designed to be annealed to the 3' end of the single-stranded nucleic acid molecule. Optionally, about 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleotides in the 3' overhang can be annealed to the 3' end of the single-stranded nucleic acid molecule. Preferably, at least the last 1, 2, 3, 4, 5, 6, 7, 8, 9 or last 10 nucleotides of the 3' overhang can be annealed to the 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides, respectively, of the 3' end of the single-stranded nucleic acid. After annealing the adapter to the single-stranded nucleic acid, the adapter can be ligated to the single-stranded nucleic acid molecule. Additionally or alternatively, the single-stranded nucleic acid can be made double-stranded ("packed" or "extended"), for example, using a suitable DNA polymerase or reverse transcriptase. A preferred DNA polymerase is a high-fidelity DNA polymerase. Subsequently, the resulting double-stranded nucleic acid molecule may be sequenced as described herein after optionally ligating a second (or further) adapter, as described herein.
[0115] Preferably, the linker-containing adapter is attached to a (partially) double-stranded nucleic acid molecule. The linker-containing adapter may be attached to at least one end of the provided double-stranded nucleic acid molecule using any conventional means known to those skilled in the art, such as adapter ligation, hybridization, annealing, and / or tagmentation.
[0116] When both ends of a nucleic acid molecule are closed by an adapter containing a linker, the molecule lacks free "terminal" nucleotides and is therefore protected from degradation by exonucleases. In the method detailed herein, optionally, all nucleic acid molecules present in a particular nucleic acid sample are closed by a linker at at least one end, becoming double-stranded nucleic acid molecules having at least one closed end.
[0117] Linkers are added by hybrid formation and / or ligation of adapters. In particular, adapters containing a linker at one end of a double-stranded structure can be ligated to a double-stranded molecule. Therefore, preferably, step b) includes the step of attaching an adapter to a provided nucleic acid molecule, the adapter being covalently closed at one end by a linker.
[0118] The linker adapter is at least partially double-stranded, but may include a single-stranded portion. Preferably, the linker is at one end of the adapter, and the other end of the adapter may include a single-stranded portion. The optional single-stranded portion of the adapter preferably includes at its 3' end a section that can hybridize to the nucleic acid molecule provided in step a) of the method described herein. The single-stranded portion of the adapter may hybridize to a single-stranded overhang of the nucleic acid molecule, preferably to a 3' overhang of the nucleic acid molecule. Optionally, the remaining single-stranded portion of the annealed adapter may be subsequently filled and fully double-stranded using a polymerase such as, but not limited to, Klenow polymerase (known to those skilled in the art as having 5'→3' polymerase activity and 3'→5' exonuclease activity but lacking 5'→3' exonuclease activity), or Bst polymerase (known to those skilled in the art as a DNA polymerase derived from Bacillus stearothermophilus, having 5'→3' polymerase activity and strand displacement activity but lacking 3'→5' exonuclease activity). The hybridized adapter may be ligated to the 5' end of the opposite strand nucleic acid molecule with which the adapter hybridized, either before or after double-stranding. Alternatively, or in addition to the above, the extended 3' end of the nucleic acid molecule may be ligated to the 5' end of the polynucleotide forming the adapter.
[0119] A linker adapter that is at least partially double-stranded may be ligated to a nucleic acid molecule provided in step a) as defined herein. Preferably, at least 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the nucleotides in the adapter are double-stranded. The adapter may be 100% or "completely" double-stranded. The adapter may be made completely double-stranded after being ligated to a nucleic acid molecule, for example by filing the single-stranded portion of the adapter using DNA polymerase.
[0120] Preferably, at least one end of a partially double-stranded linker adapter can be ligated to the nucleic acid molecule provided in step a). Therefore, preferably, at least one end of the double-stranded ends of the adapter may be a blunt end, an alternating end, or a “sticky” end. Preferably, the adapter includes an alternating end. Preferably, the end of the adapter ligated to the nucleic acid molecule has an open end that fits with the end of the nucleic acid molecule.
[0121] The other end of the linker adapter is closed by the linker and cannot be ligated to a nucleic acid molecule or further adapter. Means for designing and constructing adapters used in the methods provided herein are well known to those skilled in the art, and the methods provided herein are not limited to the design and / or construction of a particular adapter. The adapter is at least partially double-stranded, preferably including a closed double-stranded end, which is closed by a linker. The adapter may have further double-stranded structures on or near the side to be ligated to the nucleic acid molecule in step (a) of the present invention. Preferably, the adapter is a substantially completely double-stranded nucleic acid structure except for the linker and nucleotide overhangs on the ligable side of the adapter. The ligable ends of the adapter may be blunt-ended or alternating-ended so that they can be ligated to compatible blunt-ended or alternating-ended nucleic acid molecules, respectively. In this specification, compatible means that the overhangs of the adapter and the nucleic acid molecule can hybridize with each other. The nucleic acid fragments produced by restriction enzyme digestion may each contain a 3' or 5' overhang that can hybridize with a complementary (compatible) 3' or 5' adapter overhang.
[0122] Adapters that are at least partially double-stranded can be ligated at their open ends to the phosphorylated 5' end of a nucleic acid molecule. For example, DNA fragments produced by restriction enzyme fragmentation generally contain a phosphorylated 5' end and are therefore readily ligated. Additionally or alternatively, adapters that are at least partially double-stranded can be ligated at their open ends to the 3' end of a nucleic acid molecule's strand, provided their 5' end is phosphorylated. Preferably, the adapter is formed by a single-stranded oligonucleotide containing a linker, which folds into a double-stranded structure containing a linker at its (closed) end. The 5' and 3' ends of the oligonucleotide form the open ends (i.e., the adapter) of the double-stranded structure. The 5' end of the oligonucleotide is preferably phosphorylated.
[0123] The linker may be located precisely in the center of the modified oligonucleotide forming the adapter of the present invention, and the resulting double-stranded structure may include a blunt open end. Alternatively, the linker may not be located precisely in the center, and the resulting double-stranded structure may include an open end with a 3' or 5' overhang. Preferably, in the case of an overhang adapter, after ligating to a matching nucleic acid to be sequenced, the linker may be located precisely in the center of the double-stranded structure of the ligated adapter-ligated nucleic acid. If at least one strand of a partially double-stranded adapter is ligated to the nucleic acid molecule of step (a) of the present invention, preferably at its 3' end, the opposite strand of the adapter is filled, thus producing a fully double-stranded adapter, preferably a fully double-stranded ligated adapter-ligated nucleic acid. Here again, preferably, the linker may be located precisely in the center of such a double-stranded ligated adapter-ligated nucleic acid. Filling of the single-stranded sequence, i.e., the production of the double-stranded sequence, can be carried out using a conventional polymerase, such as, but not limited to, Klenow polymerase or BST polymerase. The preferred polymerase is BST polymerase.
[0124] Optionally, the adapters are ligated to both sides of the nucleic acid molecule, and at least one of the adapters includes a linker-closed end, in other words, one of the adapters is the linking adapter of the present invention. Optionally, the other of the two adapters does not include a linker-closed end. The non-linker adapter may have a single-stranded overhang at its non-ligatable end for, for example, amplification primer binding, sequencing primer binding, and / or identification. Optionally, the non-linker adapter may introduce a sequence that enables sequencing of the fragment on an amplification sequencing platform, which is preferably performed on at least one of a nanopore sequencing and a PacBio single-molecule sequencing platform. Optionally, both strands at the non-ligatable end of the non-linker adapter are single-stranded or non-complementary, thereby forming a Y-shaped adapter, where one or both strands at the non-complementary end of the adapter may include a sequence for amplification primer binding, sequencing primer binding, nanopore sequencing, and / or identification. Preferably, the 5' end overhang on the non-ligate side of the non-linker adapter contains a leader sequence comprising a motor protein for nanopore sequencing.
[0125] Adapter ligation may be carried out using conventional methods known to those skilled in the art, and the present invention is not limited to a specific ligation method or ligation enzyme (effective amount) (of ligase). Preferably, to facilitate ligation, the adapter includes an (open) end that fits with at least one end of the nucleic acid molecule. Preferably, the nucleic acid molecule provided in step a) includes a single-stranded overhang, and the adapter provided in step b) includes an overhang that can be ligated to the overhang of the nucleic acid molecule.
[0126] Optionally, the fragmentation step of step a) to provide nucleic acids and the subsequent adapter ligation step b) may be combined into a single step, for example, by tagmentation using a transposase enzyme. When a repair step is performed after the transposase reaction, most nucleic acid molecules will contain adapters at their ends. In this embodiment, the adapters in step b) are ligated by tagmentation, preferably using a Tn5 transposase. The transposase randomly cleaves long nucleic acid molecules into shorter ones, and adapters can be ligated on both sides of the cleavage point. Tagmentation, or "transposase-mediated fragmentation and tagging," is a process well known to those skilled in the art, as illustrated, for example, in the Nextera® workflow. The adapters may contain sequences adapted for use in a tagmentation reaction. Preferably, the adapters used in a tagmentation reaction further include a linker at one end. After the tagmentation reaction, a repair step may be performed, preferably on both sides, so that all, or substantially all, of the resulting nucleic acid molecules contain adapters. Therefore, nucleic acid molecules containing ligation adapters, obtained by optional tagmentation, may be repaired to remove single-strand breaks. Such a repair step can be carried out using any conventional means known in the art.
[0127] Adapters such as those provided herein, i.e., adapters including a linker-closed end, may be attached to only one end of the nucleic acid molecule or to both ends. In a non-limiting example, a linker may be added to one end of the nucleic acid molecule using a first adapter including a linker-closed end and an optional second adapter for each adapter pair that does not include a linker-closed end, preferably the second adapter including two open ends.
[0128] Alternatively, the second adapter may be, for example, an adapter comprising a linker-closed end as defined herein, or a hairpin structure, in which case the second adapter comprises an enzyme recognition site and / or a target sequence for a site-specific nuclease. As further detailed herein, after ligating the first and second adapters to each end of the provided nucleic acid molecule, the closed fragments can be treated with the respective enzymes to cleave the second adapter and obtain a fragment having one open end and one closed end.
[0129] The first and second adapters can be combined or “mixed” before the combined adapters are linked to the provided nucleic acid molecule, the first adapter including a linker-closed end and the second adapter not including a linker-closed end. The second adapter preferably includes two open ends. Optionally, the first and second adapters may differ only in the presence or absence of a linker. Alternatively, the second adapter may be a sequencing adapter, preferably for amplification-free sequencing. The second adapter may be an adapter for nanopore sequencing or PacBio single-molecule sequencing. The second adapter may preferably be a sequencing adapter for nanopore sequencing. As a non-limiting example, using a combination of adapters in which about 50% are first adapters as defined herein and about 50% are second adapters as defined herein, about 50% of the nucleic acid molecules will have a first adapter linked to one end of the same nucleic acid molecule and a second adapter linked to the other end. Therefore, step b) may include step bi) combining ("mixing") the first and second adapters as defined herein, and step bii) binding the combined adapters to a provided nucleic acid molecule, preferably the first adapter being bound to one end of the nucleic acid molecule and the second adapter being bound to the other end of the same nucleic acid molecule. Preferably, the adapter combination has an excess of the adapters of the present invention. The adapter combination may consist of a first to second adapter (w / w) ratio of about 1:1, or about 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or about 100:1. Alternatively, the adapter combinations may be configured with a first-to-second adapter (w / w) ratio of approximately 1:2; 1:3; 1:4; 1:5; 1:6; 1:7; 1:8; 1:9; 1:10; 1:11 or approximately 1:12.
[0130] Alternatively, or in addition to the above, an adapter containing a linker may be ligated to only one end of the provided nucleic acid molecule. As a non-limiting example, the adapter may include an end, preferably a 3' end, that can hybridize to only one end of the provided nucleic acid molecule. For example, if the nucleic acid molecule contains different ends such as a blunt end and a sticky end, the ligable ends of the adapter may each include a blunt end or a (fitting) sticky end, and as a result, it may be ligated to only one end of the nucleic acid molecule. Similarly, the nucleic acid molecule may contain two different overhangs, and the adapter may fit to only one of these overhangs. As a result, the adapter may be ligated to only one end of the provided nucleic acid molecule. Optionally, an adapter containing a linker at one end may be attached to both ends of the provided nucleic acid molecule.
[0131] The adapter used in the methods defined herein preferably does not contain a recognition site for a restriction endonuclease or site-specific endonuclease that may be used in step b3) of the methods provided herein.
[0132] Step b) of the method of the present invention involves generating a double-stranded nucleic acid molecule having one open end and one closed end, and adding an adapter to the end of the double-stranded nucleic acid molecule provided in step a), wherein one end of the adapter is closed by a linker. In other words, the nucleic acid molecule provided in step a) is modified to generate a double-stranded nucleic acid molecule having one open end and one closed end. The addition of the linker closes at least one end of the double-stranded nucleic acid molecule, providing a double-stranded nucleic acid molecule having one open end and one closed end. The ends of a double-stranded nucleic acid in which the 3' terminal nucleotide of each upper strand is covalently bonded to the 5' terminal nucleotide of each lower strand are annotated herein as “closed ends”. Similarly, the ends of a double-stranded nucleic acid in which the 5' terminal nucleotide of each upper strand is covalently bonded to the 3' terminal nucleotide of each lower strand are also annotated herein as “closed ends”. Therefore, “closed end” is understood herein as the end of a double-stranded nucleic acid where the terminal nucleic acids from the opposite strand are covalently bonded to each other, and “open end” is understood herein as the end of a double-stranded nucleic acid where the terminal nucleic acids from the opposite strand are not covalently bonded to each other.
[0133] The ends of a double-stranded nucleic acid are closed by a linker, preferably a linker as described herein. Therefore, preferably, the double-stranded nucleic acid is closed using a non-natural chemical moiety; that is, preferably, the double-stranded nucleic acid is not closed by a hairpin consisting of (natural) nucleotides. Preferably, at least one closed end of the double-stranded nucleic acid is obtained by ligating an adapter as described herein to the nucleic acid molecule.
[0134] If the adapter is ligated to only one end of the nucleic acid molecule provided in step a), the resulting double-stranded nucleic acid molecule will have one open end and one closed end.
[0135] Optionally, the adapters described herein may be ligated to both ends of the provided nucleic acid molecule. In this embodiment, the resulting double-stranded nucleic acid molecule will have two closed ends. Therefore, step b) may include the following (sub)steps: -Step b1) The step of linking adapters described herein, i.e., adapters closed at one end by a linker, to both ends of a double-stranded nucleic acid molecule. The resulting double-stranded nucleic acid molecule is closed at both ends by linkers. -Optional step b2) Contacting the nucleic acid sample with an exonuclease. The exonuclease may digest the remaining nucleic acid molecules containing at least one open end, thereby enriching the nucleic acid molecules in the sample that contain closed ends on both sides. - Step b3) A step of cleaving a double-stranded nucleic acid molecule having two closed ends, thereby generating a double-stranded nucleic acid molecule having one open end and one closed end.
[0136] This method may include an optional step b2) of contacting a nucleic acid molecule with an effective amount of exonuclease. The exonuclease may digest any nucleic acid molecule that does not have two closed ends, i.e., has one or two open ends. Such nucleic acid molecules include, but are not limited to, nucleic acid molecules without adapters, nucleic acid molecules with one or two adapters having open ends, and / or cleaved nucleic acid molecules having one open end and one closed end.
[0137] Therefore, the sample provided in step a) may contain multiple nucleic acid molecules, and in step b1) a portion of the nucleic acid molecules is covalently closed by adapter ligation, i.e., a portion of the nucleic acid molecules contains two closed ends. Therefore, the optional step b2) is a step of removing the open-end double-stranded nucleic acid molecules by exposing the sample to an exonuclease. Nucleic acid molecules with two closed ends are protected from degradation, while unprotected fragments are degraded, resulting in the enrichment of nucleic acid molecules containing the desired sequence or a reduction in complexity. Therefore, in one embodiment, the method of the present invention employs an approach to remove undesirable (non-targeted) portions of the nucleic acid sample. As a non-limiting example, the adapter in step b) may be ligated to a nucleic acid molecule having selective alternating overhangs produced, for example, by enzymatic digestion. The molecule containing the adapter thus has both ends closed, and the exonuclease treatment in step b2) may digest nucleic acid molecules that do not have both ends closed. Therefore, the exonuclease treatment in step b2) may enrich nucleic acid molecules containing closed ends.
[0138] The exonuclease may be exonuclease I, III, V, VII, VIII, or related enzymes, or any combination thereof. Exonuclease III recognizes a nick and extends the nick to the gap until a fragment of ssDNA is formed. Exonuclease VII can degrade this ssDNA. Exonuclease I also degrades ssDNA. ExoIII and ExoVII are a preferred combination of exonucleases for use in step b2) of the method described herein.
[0139] Exonuclease V can degrade ssDNA and dsDNA in both the 3'-to-5' direction and the 5'-to-3' direction. Therefore, in a preferred embodiment, the exonuclease in step b2) of the method described herein is an exonuclease that can degrade ssDNA and dsDNA in both the 3'-to-5' direction and the 5'-to-3' direction, and is preferably exonuclease V.
[0140] Further information regarding methods for degrading non-target sequences is provided in U.S. Patent Publication No. 2014 / 0134610, which is incorporated herein by reference in its entirety for all purposes.
[0141] Step b2) is preferably carried out under conditions (time, temperature, reaction buffer, enzyme concentration, etc.) sufficient for the exonuclease to degrade substantially all unprotected nucleic acid molecules. Preferably, step b2) is carried out under conditions and for a time sufficient for the exonuclease to degrade all unprotected nucleic acid molecules. Step b2) is preferably carried out at a temperature of about 20 to 80°C, preferably about 37°C, for about 1 minute to about 12 hours, preferably 30 minutes.
[0142] After exonuclease treatment in step b2), the exonuclease may be inactivated by at least one of proteinase treatment or thermal inactivation, such as, but not limited to, proteinase K. Such techniques are standard in the art, and methods for inactivating exonucleases are readily understood by those skilled in the art. A preferred inactivation step is to heat the sample at a temperature of about 50 to 90°C, preferably about 75°C, for about 1 to 120 minutes, preferably about 10 minutes.
[0143] Optionally, both ends of double-stranded nucleic acid molecules are closed. Molecules with closed ends are insensitive to 5' or 3' modifying enzymes. As shown above, an optional step of exonuclease treatment can be added to remove nucleic acid molecules that may not be covalently closed at both ends. Subsequently, (covalently closed) nucleic acid molecules can be selectively opened, for example, using a (effective amount) of limiting endonuclease or site-specific endonuclease. All nucleic acid molecules are still present in the reaction mixture, but only those cleaved in the final opening reaction can be used in subsequent (sequencing) processes by selectively making these open fragments sequenceable, for example, by ligating a sequencing adapter to the open ends. Alternatively, the opened fragments may be degraded by exonuclease treatment, and the unopened nucleic acid molecules may be enriched and further processed. For example, these unopened molecules may be opened in a second round of selective opening using a site-specific endonuclease that targets these unopened molecules.
[0144] The methods described herein may further include step b3) cleaving a closed double-stranded nucleic acid molecule to generate a double-stranded nucleic acid molecule having one open end and one closed end. Herein, “cleavage” means the generation of a double-strand break. The double-strand break may be created by the use of an (endo)nuclease or by the use of two nickases cleaving opposite strands. The double-strand break may result in a blunt open end of the nucleic acid molecule. After cleavage, the cleaved nucleic acid molecule may have one open blunt end and one closed end. Alternatively, the double-strand break may generate alternating open ends of the cleaved nucleic acid molecule. After cleavage, the cleaved nucleic acid molecule may have one open alternating end and one closed end. A double-stranded nucleic acid molecule having two closed ends may be cleaved at a target sequence located on one of the bound adapters and / or at a target sequence located on the double-stranded nucleic acid molecule.
[0145] In one embodiment, an adapter for use in the method provided herein, optionally further adapters further include a restriction enzyme recognition site or target sequence for a site-specific nuclease between the linker (or hairpin) and the portion of the adapter for ligation to the nucleic acid molecule. Preferably, such an adapter is located only on one side of the nucleic acid molecule of the present invention, while an adapter located on the opposite side of the molecule includes a linker but lacks the restriction enzyme recognition site or target site. After step b1) closing both ends of the nucleic acid molecule, the nucleic acid molecule is brought into contact with a restriction enzyme or site-specific nuclease to obtain a double-stranded nucleic acid molecule having one closed end and one open end.
[0146] One end of a nucleic acid molecule containing two closed ends may be opened by fragmentation and / or tagmentation. Preferably, the nucleic acid molecule in step b3) is cleaved by a site-specific endonuclease or restriction endonuclease. Optionally, all nucleic acid molecules containing two closed ends are cleaved by an endonuclease in step b3) and subsequently sequenced in step c) of the method provided herein. Optionally, only a portion or "subset" of the nucleic acid molecule contains the target sequence recognized by the endonuclease. Thus, optionally, only a portion of the nucleic acid molecule containing two closed ends is cleaved by an endonuclease and subsequently sequenced in step c) of the method provided herein. In this way, selective opening of the nucleic acid molecule enriches the nucleic acid molecule containing the desired sequence. Preferably, the closed nucleic acid molecule contains a single sequence targeted by the endonuclease. Optionally, the nucleic acid molecule may contain the target sequence multiple times; for example, the nucleic acid molecule may contain the target sequence one, two, three, four, five, six or more times.
[0147] Nucleic acid molecules containing closed ends may be cleaved by a restriction endonuclease. Any sequence-specific endonuclease may be suitable for use in the method described herein. The endonuclease may be a so-called “restriction endonuclease” or “restriction enzyme,” for example, a type I, type II, type III, type IV, or type V restriction endonuclease. A preferred restriction endonuclease is a type II restriction endonuclease, preferably type IIP or type IIS. If the fragmentation in step a) is performed by cleaving the DNA with a restriction enzyme, the enzyme used in step b3) is preferably a different restriction endonuclease.
[0148] Nucleic acid molecules may be cleaved by site-specific nucleases. Site-specific nucleases may be selected from the group consisting of TALENs, RNA-induced CRISPR nucleases, zinc finger nucleases, and meganucleases. Preferably, the site-specific nuclease is a TALEN or an RNA-guided CRISPR (short repeat palindromic sequence clustered at regular intervals) nuclease.
[0149] Preferably, the CRISPR nuclease is a type V CRISPR nuclease such as Cas9 (e.g., the protein of SEQ ID NO: 1 encoded by SEQ ID NO: 2, or the protein of SEQ ID NO: 3), or Cpf1 (e.g., the protein of SEQ ID NO: 4 encoded by SEQ ID NO: 5), or Mad7 (e.g., the protein of SEQ ID NO: 6 or 7), or an derived protein having at least about 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the aforementioned protein over its entire length. Preferably, the site-specific nuclease is a type II CRISPR nuclease, preferably a Cas9 nuclease. Those skilled in the art know how to obtain site-specific nucleases (e.g., TALEN or CRISPR nucleases) used in the methods of the present invention.
[0150] Prior art has included numerous reports on the design and use of CRISPR nucleases. For example, see the review by Haeussler et al. on the combined use of guide RNA design and CAS protein (originally derived from S. pyogenes) (J Genet Genomics. (2016) 43(5):239-50. Doi:10.1016 / j.jgg.2016.04.008.), or the review by Lee et al. (Plant Biotechnology Journal (2016) 14(2) 448-462). Optionally, site-specific nucleases are CRISPR nucleases, which are either nickases or (endo)nucleases.
[0151] The site-specific nuclease used in the method of the present invention may include or consist of a whole type II or type V CRISPR nuclease, a variant thereof, or a functional fragment thereof. Preferably, the site-specific nuclease is a Cas9 protein.The Cas9 protein is found in the bacteria Streptococcus pyogenes (SpCas9; NCBI reference sequence NC_017053.1; UniProtKB-Q99ZW2), Geobacillus thermodenitrificans (UniProtKB-A0A178TEJ9), Corynebacterium ulcerous (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); and Spiroplasma syrphidicola (NCBI Ref:NC_021284.1); Prevotella intermedia (NCBI Ref:NC_017861.1); Spiroplasma taiwanense (NCBI Ref:NC_021846.1); Streptococcus iniae (NCBI Ref:NC_021314.1); Belliera baltica (NCBI Ref:NC_018010.1); Psychroflexus torquis (Psychroflexus torquisl) (NCBI Ref:NC_018721.1); Streptococcus thermophilus (NCBI It may also be derived from Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: YP_002344900.1), or Neisseria meningitidis (NCBI Ref: YP_002342100.1).These include Cas9 mutants that become nickases by having an inactivated HNH or RuvC domain homolog to SpCas9 (e.g., SpCas9_D10A or SpCas9_H840A), or Cas9s that become nickases by having equivalent substitutions at positions corresponding to D10 or H840 in the SpCas9 protein.
[0152] The site-specific nuclease may be Cpf1, for example, Cpf1 from the genus Acidaminococcus;UniProtKB-U2UMQ6, or derived therefrom. The mutant may be a Cpf1 nickase having an inactivated RuvC domain or NUC domain, where the RuvC domain or NUC domain no longer possesses nuclease activity. Those skilled in the art are well aware of the techniques available in the art, including site-directed mutagenesis, PCR-mediated mutagenesis, and whole-gene synthesis enabling inactivated nucleases such as inactivated RuvC or NUC domains. An example of a Cpf1 nickase having an inactivated NUC domain is Cpf1 R1226A (see Gao et al. Cell Research (2016) 26:901-913, Yamano et al. Cell (2016) 165(4):949-962). In this variant, there is a conversion from arginine to alanine (R1226A) in the NUC domain, and the NUC domain is inactivated.
[0153] The site-specific nuclease may be a nuclease about half the size of CRISPR-CasΦ or Cas9, or may be derived from them. CRISPR-CasΦ uses a single crRNA for nucleic acid targeting and cleavage, as described, for example, in Pausch et al (CRISPR-CasΦ from huge phages is a hypercompact genome editor, Science (2020); 369(6501): 333-337).
[0154] An active, partially inactive, or inactive site-specific nuclease, preferably an active, partially inactive, or inactive CRISPR nuclease complex, may play a role in inducing a fusion functional domain at a specific site within a nucleic acid molecule determined by a guide RNA. Thus, the site-specific nuclease may be fused to a functional domain. Preferably, such a functional domain is an endonuclease domain. In one embodiment, the inactive CAS protein used in the method defined herein (e.g., dCas9, dCpf1) is preferably fused to a restriction enzyme such as Fok1 or Clo51, as described in International Publication No. 2014 / 144288, International Publication No. 2016 / 205554, Tsai et al. Nat Biotechnol. 2014 Jun;32(6):569-576, or Cheng et al. Biotechnol J. 2022 Jul;17(7):e2100571, all of which are incorporated herein by reference. Preferably, the fusion protein has a sequence that is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identical to any one of sequence numbers 8 to 10.
[0155] A nucleic acid molecule containing two closed ends may be cleaved by a guide RNA-CAS complex. The guide RNA guides the complex to a defined target site within the double-stranded nucleic acid molecule, also referred to as a protospacer sequence. The guide RNA contains a sequence for targeting the site-specific nuclease complex to a protospacer sequence, preferably near, on, or within the target sequence within the double-stranded nucleic acid molecule.
[0156] The guide RNA may be a single guide (sg)RNA molecule, a combination of crRNA and tracrRNA as separate molecules (e.g., in the case of Cas9), or a crRNA molecule alone (e.g., in the case of Cpf1 and CasΦ). Optionally, the guide RNA may be a single guide (sg)RNA (e.g., in the case of Cas9) or a crRNA alone (e.g., in the case of Cpf1 and CasΦ).
[0157] The guide RNA used in the methods provided herein may include a sequence within a double-stranded nucleic acid molecule, preferably the sequence of interest or its vicinity, preferably a sequence that can hybridize to the sequence of interest as defined herein. The guide RNA preferably includes a nucleotide sequence that is fully complementary to the sequence located in the double-stranded nucleic acid molecule, and optionally fully complementary to the sequence located in interest, i.e., the sequence of interest may include a protospacer sequence.
[0158] Alternatively, or in addition to the above, nucleic acid molecules containing two closed ends may be cleaved with a site-specific endonuclease, where the site-specific endonuclease is a TALEN. TALENs are well known to those skilled in the art and are constructed by fusing a TAL effector DNA-binding domain (TALE) to an effector domain, preferably a (non-specific) DNA cleavage domain, such as a FokI cleavage domain. As used herein, the terms “transcriptional activator-like effector,” “TALE,” or “TAL effector DNA-binding domain” refer to a protein containing a DNA-binding domain that comprises a highly conserved 33-34 amino acid sequence and a highly variable two-amino acid motif (repeat variable two residues, RVD). RVD motifs are known to determine binding specificity to nucleic acid sequences and can be modified to specifically bind to desired DNA sequences according to methods well known to those skilled in the art (see, for example, International Publication No. 2010 / 079430, International Publication No. 2011 / 072246, and International Publication No. 2015027134, the entire contents of which are incorporated herein by reference). Because of the simple relationship between amino acid sequences and DNA recognition, it has become possible to engineer specific DNA-binding domains by selecting combinations of repeat segments containing appropriate RVDs.
[0159] In the use of this specification, the terms “transcription activator-like element nuclease” or “TALEN” refer to a modified nuclease consisting of a transcription activator-like effector DNA-binding domain (TALE) linked to a DNA cleavage domain, such as a Fokl domain. Several modular assembly schemes for constructing recombinant TALE constructs have been reported and are known to those skilled in the art.
[0160] A transcription activator-like effector nuclease (TALEN) is a fusion of a restriction enzyme cleavage domain, preferably a FokI domain, and a DNA-binding transcription activator-like effector (TALE) repeat array. Other useful endonuclease domains may include, for example, HhaI, HindIII, NotI, BbvCI, EcoRI, BglII, and AlwI. Preferably, the cleavage domain is the FokI domain.
[0161] The method detailed herein is provided, further comprising step c) sequencing both strands of at least a portion of a double-stranded nucleic acid molecule in a single sequencing reaction to produce a double-stranded read. The double-stranded nucleic acid obtained after step b) of the method comprises one open end and one closed end. Using a single sequencing reaction, both strands are sequentially and directly sequenced to produce a double-stranded read.
[0162] Prior to sequencing, an additional adapter may be attached to the nucleic acid molecule. The additional adapter is preferably attached to the open end of the double-stranded nucleic acid molecule. The additional adapter may be an adapter suitable for amplification and / or sequencing, and preferably the additional adapter is a sequencing adapter. Therefore, preferably, after step b) and before step c), the sequencing adapter is ligated to the open end of the double-stranded nucleic acid. The additional adapter may be a sequencing adapter containing a functional domain that enables deep sequencing. Preferably, the additional adapter may introduce a sequence that enables sequencing on an amplification-free sequencing platform, where preferably the sequencing platform is at least one of nanopore sequencing and PacBio single-molecule sequencing. The additional adapter may be a sequencing adapter containing a functional domain that enables nanopore sequencing such as Oxford Nanopore Technologies (ONT), Ontera sequencing, or Complete Genomics sequencing. Preferably, the additional adapter enables Oxford Nanopore Technologies (ONT) sequencing, preferably nanopore-selective sequencing.
[0163] Therefore, preferably, the further adapter includes at least one sequencing primer binding site, and / or the further adapter includes at least one amplification primer binding site. The further adapter may include at least two sequencing primer binding sites, and / or the further adapter may include at least two amplification primer binding sites. The further adapter may be a single-stranded, double-stranded, partially double-stranded, Y-shaped, or hairpin-shaped nucleic acid molecule. Preferably, the further adapter is a (partially) double-stranded adapter including two open ends. Optionally, the further adapter includes an identifier sequence, preferably an identifier sequence as defined herein.
[0164] After step b), and before sequencing at least a portion of both strands of the nucleic acid molecule, and preferably before ligating any further adapters of any choice, the double-stranded nucleic acid molecule, including one open end and one closed end, may be repaired, preferably by polishing the ends of the nucleic acid molecule. Thus, before sequencing and before ligating any further adapters of any choice, the method provided herein may optionally include the step of polishing the open end of the nucleic acid molecule.
[0165] Additionally or alternatively, a nucleic acid molecule having one open end and one closed end may be modified to include an A overhang, preferably to facilitate ligation to a further adapter, in which case the further adapter preferably includes a T overhang. Therefore, the method of the present invention may optionally include the step of adding an A tail to the polished nucleic acid molecule before annealing the further adapter to the open end of the nucleic acid molecule.
[0166] Step c) of the method of the present invention involves sequencing a double-stranded nucleic acid molecule having one open end and one closed end to generate a double-stranded readout. A portion of the nucleic acid molecule may be sequenced, and the sequenced portion includes a portion of the forward strand and a portion of the complementary reverse strand.
[0167] Double-strand reading is a sequencing method that involves reading the sequence of the forward strand of a double-stranded nucleic acid molecule followed by the sequence of the reverse strand of the same double-stranded nucleic acid molecule. A double-stranded nucleic acid molecule containing one open end and one closed end is denatured, and as a result the forward and reverse strands form a single-strand template, which is linked by the closed end of the nucleic acid molecule (see Figure 1). Therefore, the linker that forms the closed end is located between the sequence of the forward strand and the sequence of the reverse strand. Preferably, the linker that forms the closed end is in the center of the single-strand template.
[0168] Linkers preferably provide a signal that may be emitted during sequencing. Detection of such signals has been described in the Art, for example, with respect to amplification-free sequencing platforms. For example, Tse et al. (Proc Natl Acad Sci USA. 2021; 118(5): 1-11) disclose the detection of modified cytosines in a DNA strand using a PacBio platform. Furthermore, nanopore sequencing generates a raw output of current intensity (picoamperes) over time (milliseconds). This output signal is often referred to as the “squiggle.” Linkers may also provide a modified signal that is generated during real-time sequencing or when there is no signal for a predetermined period (see Figure 1, which is illustrated by a straight line of the squiggle signal), and is referred to herein as the “signature sequence lag.” Preferably, the signature sequence lag is minimal or absent. Preferably, the location of the linker is known. Additionally or alternatively, the location of the linker may be marked using a signature sequence present within adapter elements flanking the linker element. Therefore, it is possible to accurately determine at what point in the readout the forward strand ends and the readout of the reverse strand begins. In other words, the linker position can be used to quickly identify the target sequence and its complementary sequence within the nucleic acid molecule being sequenced. Thus, the use of a linker may significantly improve the accuracy of double-strand readout calling.
[0169] A single-stranded nucleic acid molecule or "open double-stranded nucleic acid molecule" is sequenced in a single sequencing reaction. The resulting double-stranded reads preferably contain the following consecutive sequences: a first sequence read containing (part of) the forward strand and an optional signature sequence at the end of the first sequence read; a signature sequence lag (very small or missing); and a second sequence read containing (part of) the sequence of the reverse strand read (optionally including the reverse of the signature sequence at the beginning of the second sequence read).
[0170] Preferably, at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or at least 95% of the nucleic acid molecule is sequenced. Preferably, the entire nucleic acid molecule may be sequenced. Therefore, preferably, the entire forward strand of the nucleic acid molecule (providing the first sequence block of the sequence reading) and the entire reverse strand of the nucleic acid molecule (providing the second sequence block of the sequence reading) are sequenced in a single sequencing reaction.
[0171] Preferably, the prepared nucleic acid molecules are deep-sequenced. Preferably, the sequencing in step c) of the method provided herein is an amplification-free sequencing method. The sequencing is preferably at least one of nanopore sequencing and PacBio single-molecule sequencing. Preferably, the sequencing in step c) of the method provided herein is nanopore sequencing. Examples of nanopore sequencing technologies include, but are not limited to, Oxford Nanopore sequencing technologies (e.g., GridION, MinION). (See, for example, Logsdon et al., Nat. Rev. Genet. 2020; 21(10): 597-614).
[0172] The prepared nucleic acid molecules can be sequenced by nanopore-selective ("Read Until") sequencing. In nanopore-selective sequencing, during real-time sequencing, the generated data (either a DC signal or base calls converted from these current signals) is compared to one or more reference sequences. If the set number of nucleotides in the target sequence or the amount of signal matches the reference sequence, sequencing continues; otherwise, the current is reversed, thereby removing nucleic acid from the pore and making the pore available for sequencing a new nucleic acid. The set number of nucleotides may be at least the first 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of the nucleic acid read. The one or more reference sequences may be a number of different sequences. Preferably, each of these reference sequences is at least 50, 60, 70, 80, 90, 92, 93, 94, 95, 96, 97, 98, 99, or 100% identical to the sequence of the nucleic acid molecule obtained by the method of the present invention. In one embodiment, each reference sequence is at least 50, 60, 70, 80, 90, 92, 93, 94, 95, 96, 97, 98, 99, or 100% identical to the sequence of a particular subset of nucleic acid molecules obtained by the method of the present invention. One advantage of selectively sequencing a particular subset by nanopore selective sequencing is that different subsets may be sequenced in different sequencing runs using a library of prepared nucleic acid molecules having one open end and one closed end.
[0173] The sequence read lengths obtained by long-read sequencing in the method of the present invention may be at least about 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 200 kb, 300 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, or 10 Mb.
[0174] The method provided herein further includes step d) generating a consensus sequence from a double-stranded read. A double-stranded read includes the sequence of the forward strand, an optional sequence lag, and the sequence of the reverse strand of the same double-stranded nucleic acid molecule. The resulting forward and reverse strand sequences can be combined, or "assembled," into a consensus sequence. Optionally, the adapter provided herein includes a signature sequence (or barcoded nucleotide sequence) that can be used to filter against true double-stranded reads and / or to mark the end of a forward read and the beginning of a reverse read before assembling the reads into a consensus sequence.
[0175] A bioinformatics analysis protocol is also provided, which utilizes prior knowledge of the location of the linker and / or signature sequence to analyze the double-stranded readout and thereby accurately determine the sequence of interest, and a computer medium instructed to perform the bioinformatics analysis protocol is also provided. Optionally, a standard double-stranded analysis or split double-stranded analysis protocol may be used to analyze the sequence of interest, but preferably modified to detect the signature sequence within a double-stranded nucleic acid molecule having one open end and one closed end, obtained by step b) of the method of the present invention.
[0176] Conventional double-strand sequencing protocols detect two separate sequences (i.e., forward and reverse strands) and subsequently join these two sequences into a consensus sequence. Split sequencing protocols detect a single sequence containing both the forward and reverse strands. The forward and reverse strands are then split and combined with the consensus sequence. Optionally, bioinformatics sequencing protocols, preferably split sequencing protocols, can be adapted to the type of linker used to close the double-stranded molecule. Optionally, bioinformatics sequencing protocols, preferably split sequencing protocols, can be adapted to the type of signature sequence used to mark the location of the linker within the nucleic acid being sequenced. In other words, the protocol can be trained on where to split the sequence into forward and reverse strands to generate the consensus sequence.
[0177] Therefore, the method of the present invention improves the sequencing accuracy to a modal accuracy of preferably about Q30. An accuracy of Q30 corresponds to the probability of an incorrect base call occurring once in 1000 times, i.e., a base call accuracy of 99.9%. The consensus sequence may be, but is not limited to, a known genome sequence, and may be arbitrarily aligned or compared with a known nucleotide sequence. As a non-limiting example, if the same nucleotide mutation is detected in both the forward and reverse strands, the likelihood that the mutation is a true mutation and not, for example, a sequencing artifact increases significantly. Therefore, by generating a consensus sequence from a double-stranded readout, the sequence within a double-stranded nucleic acid molecule is determined preferably with high accuracy, preferably with a modal accuracy of about Q30.
[0178] Preferably, the method is performed on multiple samples. Optionally, the method of the present invention is multiplexed, i.e., applied simultaneously to multiple nucleic acid samples, such as at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 500, 1000 or more nucleic acid samples. Thus, the method may be performed on multiple samples in parallel, where “in parallel” is understood to mean substantially simultaneously, but with each sample processed in a separate reaction tube or container.
[0179] Additionally or alternatively, one or more steps of the methods provided herein may be performed on multiple pooled samples. The multiple samples are preferably pooled before step c). Optionally, the multiple samples are preferably pooled before step b). To trace double-stranded nucleic acid molecules back to the original sample, the nucleic acid molecules may be tagged with identifiers before pooling the samples. Such identifiers may be any detectable entity, such as, but not limited to, radioactive or fluorescent labels, but preferably a specific nucleotide sequence or combination of nucleotide sequences, and preferably a defined length. Additionally or alternatively, the samples may be pooled using sophisticated pooling strategies, such as, but not limited to, 2D and 3D pooling strategies, so that after pooling, each sample is contained in at least two or three pools. By using the coordinates of each pool containing the double-stranded nucleic acid molecules, a specific nucleic acid molecule can be traced back to the original sample. The multiple samples may be pooled before step c) and / or before step b).
[0180] During one or more steps of the methods provided herein, the nucleic acid sample may be purified and / or the reactive enzyme may be inactivated. For example, by including a purification step such as an AMPure bead-based purification process, complexes, enzymes, free nucleotides, possible free adapters, and possible small, unrelated nucleic acid molecules may be removed. The nucleic acid molecules may be recovered after purification and subjected to further processing and / or analysis, such as single-molecule sequencing.
[0181] The optional purification step is protease K treatment. Alternatively, or in addition to this, the purification may include the following steps: I. The step of exposing a nucleic acid sample to one or more solid supports that specifically and effectively bind nucleic acid molecules; optionally, II. Washing one or more solid supports and eluting nucleic acid molecules from one or more solid supports.
[0182] Examples of solid supports include, but are not limited to, AMPure beads. After purification, at least one purified nucleic acid molecule is obtained.
[0183] The method of the present invention may further include, for example, a size selection step for removing non-ligated adapters or adapter-adapter pairs. Optionally, the size selection step is performed before step c) of the method of the present invention. Alternatively, further purification, inactivation, and / or size selection steps are omitted.
[0184] In one embodiment, the method provided herein is a sequencing method that is free from amplification and / or cloning steps. Reducing the amplification step is beneficial because epigenetic information (e.g., 5-mC, 6-mA, etc.) is lost in the amplicon. Further amplification introduces variations into the amplicon (e.g., due to errors during amplification), resulting in a nucleotide sequence that no longer reflects the original sample.
[0185] In one embodiment, provided herein is a parts kit for carrying out the method described herein. Preferably, the parts kit is used in the manner defined herein. Preferably, the parts kit includes at least one or more adapters as defined herein, which have an open end and a closed end, the ends of which are closed by a linker. The multiple adapters may be combined in one vial or may be present in separate vials, for example, the adapters in one vial may include the same identifier sequence, preferably the same sample identifier sequence. Preferably, the adapters present in separate vials may include different identifier sequences, preferably different sample identifier sequences.
[0186] Alternatively, or in addition to the above, the component kit may include at least one or more sequencing adapters. Multiple sequencing adapters may be combined in one vial or present in separate vials; for example, the sequencing adapters in one vial may contain the same identifier sequence, preferably the same sample identifier sequence. Preferably, the sequencing adapters present in separate vials may contain different identifier sequences, preferably different sample identifier sequences. The component kit may also include one or more reagents for performing the ONT sequencing reaction. Therefore, the component kit may include at least one of the following: -One or more vials including an adapter with an open end and a closed end, the closed end being closed by a linker; and -One or more vials including further adapters as defined herein for sequencing, preferably for ONT sequencing reactions.
[0187] Preferably, the parts kit may further include at least one of the following: - One or more vials containing restricted endonuclease; - Vials containing a gRNA-CAS complex as defined herein, or vials containing one or more constructs encoding it; -One or more vials containing gRNA for complexing with a CRISPR-CAS protein to form a gRNA-CAS complex, and a further vial containing the CRISPR-CAS protein or a construct encoding it; and - A further vial containing one or more exonucleases.
[0188] Preferably, the kit comprises at least 2, 4, 10, 20, 30, or 50 vials, each containing one or more gRNAs as defined herein. Preferably, the volume of any vial in the kit does not exceed 100 mL, 50 mL, 20 mL, 10 mL, 5 mL, 4 mL, 3 mL, 2 mL, or 1 mL.
[0189] The reagents may be in lyophilized form or present in a suitable buffer solution. The kit may also include other components necessary for carrying out the invention, such as buffer solutions, pipettes, microtiter plates, and instructions. Such other components of the kit of the invention are known to those skilled in the art.
[0190] Although the present invention has been described in general terms, it can be more easily understood by referring to the following examples, which are provided for illustrative purposes and are not intended to limit the present invention. [Examples]
[0191] Example 1 Using a sequencing library employing the adapter of the present invention, we demonstrated that the method detailed herein actually improves the accuracy of nanopore deep sequencing. λDNA-HindIII digest (New England Biolabs; N3012; 500 ug / ml) was prepared as the template DNA.
[0192] In short, adapters with covalently closed ends and open ends were added to both ends of a λDNA-HindIII digested nucleic acid fragment for ligation, obtaining a covalently closed nucleic acid fragment. ExoV treatment removed any unclosed nucleic acids. Subsequent PvuI digestion opened one end of the covalently closed λDNA-HindIII digested fragment containing the PvuI restriction enzyme recognition site, obtaining a nucleic acid molecule with one open end and one closed end. After adding a sequencing adapter to the open end, the nucleic acid molecule was sequenced on a nanopore deep sequencing platform. A detailed summary of the experiment is as follows.
[0193] λDNA-HindIII digestion and indexed adapter ligation In this example, six different linkers were tested for the construction of ligated (or "closed") adapters (see "Ligated Adapter Ligation" below). To multiplex the sequence libraries created by each of these six ligated adapters, an indexed adapter was used to affix a sample barcode to the λDNA-HindIII digest fragment of each sample, and the index was used as the sample barcode. To prepare six different indexed adapters (Integrated DNA Technologies, Inc.), two single-stranded oligonucleotides were mixed at equimolar concentrations to form the upper and lower strands of the following indexed double-stranded adapters, respectively.
[0194] Indexed adapter 1: Upper chain: AAGGTTAACACAAAGACACCGACAACTTTCTTCAGCACCTC (Sequence ID 11) Lower chain: AGCTGAGGTGCTGAAGAAAGTTGTCGGTGTCTTTGTGTTAACCTT (Sequence ID 12) Indexed adapter 2: Upper chain: AAGGTTAAACAGACGACTACAAACGGAATCGACAGCACCTC (Sequence ID 13) Lower chain: AGCTGAGGTGCTGTCGATTCCGTTTGTAGTCGTCTGTTTAACCTT (Sequence ID 14) Indexed adapter 3: Upper chain: AAGGTTAACCTGGTAACTGGGACACAAGACTCCAGCACCTC (Sequence ID 15) Lower chain: AGCTGAGGTGCTGGAGTCTTGTGTCCCAGTTACCAGGTTAACCTT (Sequence ID 16) Indexed adapter 4: Upper chain: AAGGTTAATAGGGAAACACGATAGAATCCGAACAGCACCTC (Sequence ID 17) Lower chain: AGCTGAGGTGCTGTTCGGATTCTATCGTGTTTCCCTATTAACCTT (Sequence ID 18) Indexed adapter 5: Upper chain: AAGGTTAAAAGGTTACACAAACCCTGGACAAGCAGCACCTC (Sequence ID 19) Lower chain: AGCTGAGGTGCTGCTTGTCCAGGGTTTGTGTAACCTTTTAACCTT (Sequence ID 20) Indexed adapter 6: Upper chain: AAGGTTAAGACTACTTTCTGCCTTTGCGAGAACAGCACCTC (Sequence ID 21) Lower chain: AGCTGAGGTGCTGTTCTCGCAAAGGCAGAAAGTAGTCTTAACCTT (Sequence ID 22).
[0195] All sequences in this specification are shown from the 5' end to the 3' end.
[0196] After mixing at equimolar concentrations, each indexed adapter was diluted to a concentration of 4 pmol / μL with nuclease-free water. For indexed adapter ligation, 1 μg of λDNA-HindIII digest was used as the starting material. Prior to indexed adapter ligation, 38 μL of 10 mM Tris / HCl pH 7.5 was added to 2 μL of λDNA-HindIII digest, incubated at 85°C for 10 minutes, and then placed on ice. Table 1 shows the reaction mixtures for each indexed adapter; the reaction was performed in 5 series for each type of indexed adapter, producing a total of 30 reaction mixtures.
[0197] [Table 1]
[0198] The mixtures were incubated at 37°C for 3 hours. All five reaction mixtures were pooled for each type of indexed adapter, purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science; 100-265-900), and eluted with 30 μL of nuclease-free water. DNA concentrations were measured using Qubit BR (Invitrogen, Qubit® dsDNA BR assay kit, Q32853): Sample indexed with barcode 1: 19.5 ng / μL (585 ng in 30 μL), Sample indexed with barcode 2: 14.9 ng / μL (447 ng in 30 μL), Sample indexed with barcode 3: 21.0 ng / μL (630 ng in 30 μL), Sample indexed with barcode 4: 19.0 ng / μL (570 ng in 30 μL), Sample indexed with barcode 5: 19.7 ng / μL (591 ng in 30 μL), Sample indexed with barcode 6: 20.8 ng / μL (624 ng in 30 μL).
[0199] Polishing / Adding A Indexed λDNA HindIII digested fragments were polished, and reagent mixtures were prepared for each indexed sample in Table 2, and A addition was performed.
[0200] [Table 2]
[0201] The reaction mixture was incubated at 20°C for 30 minutes using a thermal cycler, followed by incubation at 65°C for 30 minutes. Subsequently, the product was purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science, 100-265-900) and eluted with 10 μL of nuclease-free water.
[0202] Linking adapter ligation Six differently modified 5'-terminal phosphorylated single-strand oligonucleotides were ordered by IDT, and (3'-T overhang) ligation adapters numbered 1-6 have the following structure (5' to 3' direction): Connecting adapter 1: GCAATACGTAACTGAACGAAGTACATT / iSpC3 / AATGTACTTCGTTCAGTTACGTATTGCT, Here, GCAATAGTAACTGAACGAAGTACATT is shown herein as Sequence ID No. 23, Here, AATGTACTTCGTTCAGTTACGTATTGCT is shown herein as Sequence ID No. 24, The 3' terminal T residue of sequence number 23 is linked to the 5' terminal A residue of sequence number 24 by a C3 spacer ( / iSpC3 / ).
[0203] Connecting adapter 2: GCAATACGTAACTGAACGAAGTACATT / iSpC9 / AATGTACTTCGTTCAGTTACGTATTGCT, Here, GCAATAGTAACTGAACGAAGTACATT is shown herein as Sequence ID No. 25, Here, AATGTACTTCGTTCAGTTACGTATTGCT is shown herein as Sequence ID No. 26, The 3' terminal T residue of sequence number 25 is linked to the 5' terminal A residue of sequence number 26 by a C9 spacer ( / iSpC9 / ).
[0204] Connecting adapter 3: GCAATACGTAACTGAACGAAGTACATT / iSpC18 / AATGTACTTCGTTCAGTTACGTATTGCT, Here, GCAATAGTAACTGAACGAAGTACATT is shown herein as Sequence ID No. 27, Here, AATGTACTTCGTTCAGTTACGTATTGCT is shown herein as Sequence ID No. 28, The 3' terminal T residue of sequence number 27 is ligated to the 5' terminal A residue of sequence number 28 by a C18 spacer ( / iSpC18 / ).
[0205] Connecting adapter 4: GCAATACGTAACTGAACGAAGTACATT / idSp / AATGTACTTCGTTCAGTTACGTATTGCT. Here, GCAATAGTAACTGAACGAAGTACATT is shown herein as Sequence ID No. 29, Here, AATGTACTTCGTTCAGTTACGTATTGCT is shown herein as Sequence ID No. 30, The 3' terminal T residue of SEQ ID NO: 29 is linked to the 5' terminal A residue of SEQ ID NO: 30 by a 1',2'-dideoxyribose spacer ( / idSp / ).
[0206] Connecting adapter 5: GCAATACGTAACTGAACGAAGTACATT / idSp / idSp / idSp / AATGTACTTCGTTCAGTTACGTATTGCT, Here, GCAATAGTAACTGAACGAAGTACATT is shown herein as Sequence ID No. 31, Here, AATGTACTTCGTTCAGTTACGTATTGCT is shown herein as Sequence ID No. 32, The 3' terminal T residue of SEQ ID NO: 31 is linked to the 5' terminal A residue of SEQ ID NO: 32 by three 1',2'-dideoxyribose spacer moieties ( / idSp / idSp / idSp / ).
[0207] Connecting adapter 6: GCAATACGTAACTGAACGAAGTACATT( / idSp / idSp / idSp / idSp / idSp / )AATGTACTTCGTTCAGTTACGTATTGCT, Here, GCAATAGTAACTGAACGAAGTACATT is shown herein as Sequence ID No. 33, Here, AATGTACTTCGTTCAGTTACGTATTGCT is shown herein as Sequence ID No. 34, The 3' terminal T residue of SEQ ID NO: 33 is linked to the 5' terminal A residue of SEQ ID NO: 34 by five 1',2'-dideoxyribose spacer segments ( / idSp / idSp / idSp / idSp / idSp / ).
[0208] Each of these six modified oligonucleotides was diluted to a concentration of 50 μM, as shown in Table 3.
[0209] [Table 3]
[0210] The mixture was heated at 90°C for 3 minutes, then slowly cooled to 4°C at 0.01°C per second in a thermal cycler, controlling annealing to fold the modified oligonucleotide back into a double-stranded structure containing a linker at one end, i.e., a ligation adapter was formed. The ligation adapter was diluted to 5 μM and ligated to an indexed fragment with repaired ends and A added. As shown in Table 4, the reaction mixture was prepared for each indexed sample to ligate ligation adapter 1 to a fragment containing indexed barcode 1, ligation adapter 2 to a fragment containing indexed barcode 2, and so on.
[0211] [Table 4]
[0212] The reaction mixture was incubated at 21°C for 1 hour. The sample was purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science, 100-265-900), and eluted with 10 μL of nuclease-free water. DNA concentrations were measured using Qubit HS: Sample indexed with barcode 1: 27.6 ng / μL (276 ng per 10 μL), Sample indexed with barcode 2: 23.0 ng / μL (230 ng per 10 μL), Sample indexed with barcode 3: 32.0 ng / μL (320 ng per 10 μL), Sample indexed with barcode 4: 28.2 ng / μL (282 ng per 10 μL), Sample indexed with barcode 5: 35.2 ng / μL (352 ng per 10 μL), Sample indexed with barcode 6: 30.0 ng / μL (300 ng per 10 μL).
[0213] ExoV processing Each product of the ligation adapter was treated with exonuclease V to remove unclosed polynucleotides. For each indexed sample, the reaction mixture was prepared as shown in Table 5.
[0214] [Table 5]
[0215] The reaction mixture was incubated at 37°C for 2 hours, followed by incubation at 65°C for 10 minutes. The sample was purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science, 100-265-900), and eluted with 20 μL of nuclease-free water. Sample DNA concentrations were measured using Qubit HS: Sample index barcode 1: 3.66 ng / μL (73.2 ng in 20 μL, 26.5% of the input volume), Sample index barcode 2: 5.24 ng / μL (104.8 ng in 20 μL, 45.6% of the input volume), Sample index barcode 3: 5.04 ng / μL (100.8 ng in 20 μL, 31.5% of the input volume), Sample index barcode 4: 5.58 ng / μL (111.6 ng in 20 μL, 39.6% of the input volume), Sample index barcode 5: 5.58 ng / μL (111.6 ng in 20 μL, 31.7% of the input volume), Sample index barcode 6: 5.18 ng / μL (103.6 ng in 20 μL, 34.5% of the input volume).
[0216] PvuI restriction ExoV-treated samples were treated with PvuI to open covalently closed HindIII fragments containing PvuI restriction sites, forming polynucleotide fragments with open and closed ends. For each indexed sample, the reaction mixture was prepared as shown in Table 6.
[0217] [Table 6]
[0218] The samples were incubated at 37°C for 1 hour, purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science; 100-265-900), and eluted with 15 μL of nuclease-free water.
[0219] Polishing / Adding A PvuI restriction fragments were polished, and A addition was performed for each indexed sample using the reagent mixtures shown in Table 7.
[0220] [Table 7]
[0221] Using a thermal cycler, the reaction mixture was incubated at 20°C for 30 minutes, followed by incubation at 65°C for 30 minutes. Subsequently, the product was purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science, 100-265-900), and eluted with 7.5 μL of nuclease-free water.
[0222] Native barcode ligation After end repair and A addition, the native barcode adapter was ligated to the PvuI restriction fragment. For each sample, the native adapter contained the same barcode as the indexed adapter, so that the fragment had the same index at both the closed and open ends. For this reason, a reagent mixture was prepared for each indexed sample as shown in Table 8.
[0223] [Table 8]
[0224] The sample was incubated at 21°C for 20 minutes. Subsequently, 2 μL of EDTA (Oxford Nanopore Technologies; SQK-NBD112.24) was added and the sample was mixed. Subsequently, the samples were pooled and 2 volumes of AMPure PB beads were added to purify the product according to the manufacturer's instructions (Pacific Bio Science; 100-265-900), and eluted with 12.5 μL of nuclease-free water.
[0225] ONT library preparation The pooled samples were prepared for ONT sequencing using the ligation sequencing kit SQK-LSK112.24 (Oxford Nanopore Technologies). The reaction mixture was prepared in a 1.5 ml Eppendorf DNA LoBind tube as shown in Table 9.
[0226]
Table 9
[0227] The sample was incubated at 20°C for 10 minutes. Subsequently, 0.4 volumes of AMPure PB beads were added according to the manufacturer's ONT instructions (Oxford Nanopore Technologies; SQK-NBD112.24), and the product was purified using 250 μL of Small Fragment Buffer (SFB) (Oxford Nanopore Technologies; SQK-NBD112.24) for bead washing, and eluted with 15 μL of EB (Oxford Nanopore Technologies; SQK-NBD112.24).
[0228] Oxford Nanopore Technologies sequencing and data processing The processed samples were sequenced for 72 hours on a MinION Mk1B portable sequencing device (Oxford Nanopore Technologies) using a Flow Cell R10.4 equipped with real-time base calling (MinKNOW v22.05.5, (Oxford Nanopore Technologies)). Re-base calling of the reads after sequencing was performed using the SUP base calling option of MinKNOW v22.05.5 (Oxford Nanopore Technologies) to create the most accurate sequence data.
[0229] Generation of duplex sequences was performed using the command-line version of Duplex Tools v0.2.20 (Oxford Nanopore Technologies), in which duplex reads were identified and prepared for recovery of simplex base calls from base calls and linked reads. Next, duplex reads were created using the command-line version of the standard ONT base calling software.
[0230] Results and Conclusions As a result of sequencing the pooled barcoded samples, a total of 976,000 reads corresponding to 0.6 gigabases of sequence were obtained. A total of 424,727 sequence data reads (0.484 gigabases) were assigned to one of the six barcodes used, i.e., de-duplicated. For each of the six barcoded datasets obtained, the amount of duplex sequences identified was shown as a percentage of the total nucleotide amount using the command-line version of the standard ONT base calling software (normal duplex analysis) and the command-line version of Duplex Tools v0.2.20 split-duplex sequencing (split-duplex analysis), respectively. The results are shown in Table 10 below.
[0231]
Table 10
[0232] A clear effect related to the modifications present in the linking adapter is observed. The highest percentage of double-stranded sequence nucleotides is observed when one or three consecutive "idSp" linkers are used. A single idSp linker produced a total of 34.21% of double-stranded sequence nucleotides, while a 3x idSp (idSp-idSp-idSp) linker produced 35.5% of double-stranded sequence nucleotides.
[0233] As a result, the linked adapter approach significantly increased the number of dual-chain reads by up to 35%, and dramatically improved accuracy.
[0234] Example 2 In this example, sgRNA-induced digestion with IDT Alt-R® HiFi Cas9 Nuclease V3 (Integrated DNA Technologies, 1081058) was applied instead of restriction enzymes in the initial fragmentation step to generate more uniformly dispersed DNA fragments from λDNA. This improved the comparison of double-strand sequence generation between covalently closed and open linear DNA fragments of the same size. Lambda DNA was cleaved with Cas9RNP (Cas9-guideRNA riboprotein complex) using the four sgRNAs shown in Table 11.
[0235] [Table 11]
[0236] In sgRNA / Cas9 digestion, a gRNA-Cas9 ribonucleoprotein (RNP) complex was formed first, followed by contact with λDNA. gRNA-Cas9 RNPs were prepared using 14-component reaction groups shown in Table 12. The mixture was incubated at room temperature for 10 minutes to allow for complex formation.
[0237] [Table 12]
[0238] To digest λDNA with gRNA-Cas9 RNP, 14 reaction mixtures described in Table 13 were prepared. The mixtures were incubated at 37 °C for 1 hour for digestion, followed by inactivation of Cas9 at 80 °C for 10 minutes.
[0239]
Table 13
[0240] After incubation, the 14 reaction mixtures were pooled, purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science; 100-265-900), and eluted with 94 μL of nuclease-free water. As shown in Figures 2 and 3, the digestion yielded six λDNA-fragments (three major fragments and three CoS-terminal fragments, three of the major fragments each constituted one restriction enzyme recognition site for BsaI or XbaI (New England Biolabs; R053 and R0145)). The purified Cas9RNP-cleaved λDNA was used as starting material for preparing a "first digestion sample" starting from digestion with BsaI and XbaI as shown below (see Table 14), and also as starting material for preparing a "second digestion sample" starting from end polishing and A addition (see Table 17).
[0241] First digestion sample For digestion of Cas9RNP-cleaved λDNA with BsaI and XbaI, 5 reaction mixtures shown in Table 14 were prepared. The mixtures were incubated at 37 °C for 1 hour and then incubated at 65 °C for 10 minutes.
[0242]
Table 14
[0243] After digestion, the 5-component reaction mixture was pooled and purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science; 100-265-900), and then eluted with 48 μL of nuclease-free water.
[0244] For end-polishing of the digested fragments and A addition, the reaction mixture shown in Table 15 was prepared in five batches.
[0245] [Table 15]
[0246] The reaction mixture was incubated at 20°C for 30 minutes, followed by incubation at 65°C for another 30 minutes. After end polishing and addition of A, the five reaction mixtures were pooled, purified using 1 volume of AMpure cleanup, and eluted with 61 μL of nuclease-free water.
[0247] To prepare samples for ONT sequencing using the ligation sequencing kit SQK-LSK114 (Oxford Nanopore Technologies), the sequencing adapter was ligated to end-polished and A-added BsaI / XbaI-cleaved λDNA fragments. For this purpose, the reaction mixture was prepared in 1.5 ml Eppendorf DNALoBind tubes, as shown in Table 16.
[0248] [Table 16]
[0249] The mixture was incubated at 20°C for 20 minutes. Subsequently, 0.4 liters of Ampure PB beads were added according to the manufacturer's ONT instructions (Oxford Nanopore Technologies; SQK-LSK114). The product was purified by washing the beads with 250 μL of Small Fragment Buffer (SFB) (Oxford Nanopore Technologies; SQK-LSK114) and eluted with 15 μL of EB (Oxford Nanopore Technologies; SQK-LSK114).
[0250] Second digestion sample For end polishing and A addition of gRNA-Cas9 RNP-cleaved λDNA, the reaction mixture shown in Table 17 was prepared in 5-column sets.
[0251] [Table 17]
[0252] The reaction mixture was incubated at 20°C for 30 minutes, followed by incubation at 65°C for 30 minutes. Subsequently, the product was pooled and purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science, 100-265-900), and eluted with 40 μL of nuclease-free water.
[0253] These fragments were then covalently bound using idSp ligation adapters. Prior to ligation, the ligation adapters were diluted in the mixture as shown in Table 18 and denatured at 90°C for 3 minutes, followed immediately by intramolecular regeneration of the palindromic sequence on ice (0°C).
[0254] [Table 18]
[0255] Next, ligation was performed using the Blunt / TA ligase master mix (New England Biolabs, M0367) to ligate the connecting adapter 4 from Example 1, covalently closing the fragment at the unprotected end. For this purpose, five reaction groups shown in Table 19 were prepared.
[0256] [Table 19]
[0257] The mixture was incubated at 21°C for 1 hour, then pooled and purified using 2 volumes of AMpure cleanup, followed by elution with 30 μL of nuclease-free water. Subsequently, reaction mixtures shown in Table 20 were prepared to remove open-end DNA fragments by exonuclease digestion.
[0258] [Table 20]
[0259] The mixture was incubated at 37°C for 90 minutes, followed by 65°C for 10 minutes. The product was purified with AMpure (2 volumes) and eluted with 10 μL of nuclease-free water. The three covalently closed fragments were then digested with a mixture of BsaI and XbaI to open them again (Table 21).
[0260] [Table 21]
[0261] The mixture was incubated at 21°C for 1 hour; then inactivated at 65°C for 10 minutes. The sample was purified with AMpure (2 volumes) and eluted with 10 μL of nuclease-free water. Polishing and A addition of the open fragments were performed according to Table 17. A 1-volume AMpure cleanup was performed again, and after elution with 61 μL of nuclease-free water, the fragments were prepared for ONT sequencing using the ligation sequencing kit SQK-LSK114 (Oxford Nanopore Technologies). The reaction mixture for sequencing adapter ligation was prepared in 1.5 ml of Eppendorf DNA LoBind tubes HI, as shown in Table 16. The sample was incubated at 20°C for 20 minutes. Subsequently, following the manufacturer's ONT instructions (Oxford Nanopore Technologies; SQK-LSK114), 0.4 volumes of AMpure PB beads were added, and the product was purified by washing the beads with 250 μL of Small Fragment Buffer (SFB) (Oxford Nanopore Technologies; SQK-LSK114), followed by elution with 15 μL of EB (Oxford Nanopore Technologies; SQK-LSK114).
[0262] Libraries prepared from two λDNA samples (the first and second digestion samples) both contain fragments of similar size suitable for sequencing. Here, the first sample is expected to generate mainly simplex sequence reads (unlinked fragments), while the second sample is expected to have an increased amount of double-strand reads due to the covalent closure of fragments from one end (linked fragments).
[0263] Oxford Nanopore Technologies: Sequencing and Data Processing The sequencing and data processing performed by Oxford Nanopore Technologies were carried out as shown in Example 1, that is, both standard double-strand analysis and split double-strand analysis were performed on both the unbound and bound fragment sequence libraries.
[0264] [Table 22]
[0265] Results and conclusions The results are summarized in Table 22. Approximately 38 Gb of sequence nucleotides were obtained from the unbound fragment library, while 18 and 34 sequence nucleotides were obtained from the bound fragment library. The amount of normal double-strand reads decreased from 11.23% in the unbound library to 7.98% in the bound fragment library. However, the inventors observed that split double-strand reads increased from 1.23% in the unbound fragment library to 37.72% in the bound fragment library. Combined, the inventors observed 47% double-strand sequence reads in the bound fragment library, which approaches the theoretically achievable absolute maximum of double-strand sequence reads (50%). The fact that most of these double-strand reads were identified by split double-strand analysis tools provides compelling evidence that these double-strand reads were generated as a result of physically bound fragments. This demonstrates that covalently binding DNA fragments for double-strand sequencing using the bundling adapter described in the present invention is highly efficient.
[0266] Example 3 Using a sequencing library with the adapter of the present invention, we demonstrated that the method detailed herein actually improves sequencing accuracy on the MGI DNBSEQ-G400 technology platform. A PCR marker fragment (New England Biolabs; N3234; 300 μg / ml) was prepared as template DNA.
[0267] In short, a ligation adapter with a covalently closed end and another ligation adapter with a restriction enzyme site (NotI:GCGGCCGC) with a covalently closed end were attached to both ends of a PCR marker fragment in a 5:1 ratio for ligation, yielding a covalently closed nucleic acid fragment. ExoV treatment removed any unclosed nucleic acids. Subsequent NotI-HF® digestion opened one end of the covalently closed PCR marker fragment, yielding a nucleic acid molecule with one open end and one closed end. After attaching a phosphorylated sequencing adder to the open end, the nucleic acid molecule was prepared for sequencing on the MGIDNBSEQ-G400 technology platform. A detailed summary of the experiment is as follows: Prior to polishing / A addition, 220 μL of PCR marker fragments (50 bp, 150 bp, 300 bp, 500 bp, and 766 bp fragments) were purified by adding 1.5 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science; 100-265-900), and eluted with 111 μL of nuclease-free water. DNA concentration was measured using Qubit BR (Invitrogen; Qubit® 1xdsDNABR assay kit, catalog no. Q33266): PCR marker fragments: 470 ng / μL (51.7 μg in 110 μL).
[0268] Polishing / Adding A PCR marker fragments were polished, and reagent mixtures from Table 23 were prepared for each sample (30 samples), and reagent A was added.
[0269] [Table 23]
[0270] Using a thermal cycler, the reaction mixture was incubated at 20°C for 30 minutes, followed by incubation at 65°C for 30 minutes. Subsequently, the sample was pooled and precipitated with 2 volumes of 100% ethanol (JT Baker, catalog number: 8025.1000) and 1M NaAc (Sigma Aldrich, catalog number: 25022-2.5KG-R) to a final concentration of 0.1M NaAc. The product was washed with 1 volume of 70% ethanol and eluted with 46 μL of nuclease-free water. The concentration of the PCR marker fragment was measured using the Qubit® 1xdsDNA BR assay kit (Invitrogen; Q33266): End-polished and A-added PCR marker fragment: 567 ng / μL (25.5 μg in 45 μL).
[0271] adapter ●The first connecting adapter used in this experiment is the "connecting adapter 4" shown in Example 1 above (where the spacer is a 1',2'-dideoxyribose spacer ( / idSp / )). ●The second connecting adapter used in this experiment includes a NotI restriction area (underlined section): [ka] Here, GCAATAGTAACTGAACGAAGTACATT is shown herein as Sequence ID No. 39, Here, AATGTACTTCGTTCAGTTACGTATTGCGCGGCCGCT is shown herein as Sequence ID No. 40, The 3' terminal T residue of SEQ ID NO: 39 is linked to the 5' terminal A residue of SEQ ID NO: 40 by a 1',2'-dideoxyribose spacer ( / idSp / ).
[0272] Each of these modified single-stranded oligonucleotides was diluted to a concentration of 50 μM, as shown in Table 24.
[0273] Adjustment adapter
[0274] [Table 24]
[0275] After heating the mixture at 90°C for 3 minutes, it was slowly cooled to 4°C at 0.01°C per second in a thermal cycler, and by controlling the annealing process, the modified oligonucleotide was folded back into a double-stranded structure containing a linker at one end, i.e., a linking adapter was formed.
[0276] Linking adapter A ligation adapter (50 μM) was used to ligate end-repaired and A-added fragmented PCR marker fragments in a 10-fold excess. The reaction mixture was prepared as shown in Table 25.
[0277] [Table 25]
[0278] The reaction mixture was incubated at 21°C for 1 hour. The samples were pooled and purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science, 100-265-900), and eluted with 31 μL of nuclease-free water. The concentration of the ligation adapter product was measured using the Qubit 1xdsDNA BR assay kit (Invitrogen, Q33266): ligation adapter product: 206 ng / μL (6.2 μg in 30 μL).
[0279] ExoV processing To remove unclosed polynucleotides, the ligation product was treated 8-fold with exonuclease V. The reaction mixture was prepared as shown in Table 26.
[0280] [Table 26]
[0281] The reaction mixture was incubated at 37°C for 1.5 hours, followed by incubation at 65°C for 10 minutes. The samples were pooled and purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science, 100-265-900), and eluted with 26 μL of nuclease-free water. The concentration of the ExoV-treated PCR marker product was measured using the Qubit® 1xdsDNA HS assay kit (Invitrogen: catalog number Q33231): ExoV-treated PCR marker product: 145 ng / μL (3.63 μg in 25 μL).
[0282] NotI-HF (registered trademark) restrictions ExoV-treated PCR marker products were treated with 5x NotI-HF® to open the covalently bound PCR marker fragment containing the NotI-HF® restriction site, yielding a polynucleotide fragment with both open and closed ends. The reaction mixture was prepared as shown in Table 27.
[0283] [Table 27]
[0284] After incubating the samples at 37°C for 1 hour, they were incubated at 65°C for 20 minutes, pooled, purified by adding 2 volumes of AMPure PB beads according to the manufacturer's instructions (Pacific Bio Science; 100-265-900), and eluted with 21 μL of nuclease-free water. The concentration of NotI-HF® restriction PCR marker product was measured using the Qubit® 1xdsDNA HS assay kit (Invitrogen; catalog number: Q33231): NotI-HF® restriction PCR marker product: 123 ng / μL (2.46 μg in 20 μL).
[0285] Preparation of NotI-HF® restriction PCR marker products for generating DNA nanoballs (DNBs) Using the MGIEasy FS DNA Library Preparation Kit, version 2.1 (MGI; catalog number 1000006987) kit module, NotI-HF® restriction PCR marker products for MGI DNB preparation were prepared according to the manufacturer's instructions (MGIEasy FS DNA Library Preparation Kit User Manual version B4).
[0286] Custom-made (Y-shaped) MGI adapter 5'-Phosphated Index Custom Oligonucleotides (Underlined is the index) (Top) 5' to 3' direction: [ka] Custom (5'-phosphorylated) oligonucleotides in the 5' to 3' direction (bottom): GAACGACATGGCTACGATCCGACTT (Sequence No. 42)
[0287] Custom 5'-phosphorylated single-stranded MGI oligonucleotides were annealed as shown in Table 28.
[0288] [Table 28]
[0289] These mixtures were heated at 90°C for 3 minutes and then rapidly cooled on ice.
[0290] End repair and A-tail addition The NotI-HF® restriction PCR marker fragments were subjected to end repair and A-tailing by preparing the reagent mixtures shown in Table 29 according to the manufacturer's instructions (MGIEasy FS DNA Library Preparation Kit User Manual, Version B4, Chapter 3 Library Construction Protocol; Section 3.3 End Repair and A-tailing).
[0291] [Table 29]
[0292] The samples were incubated at 37°C for 30 minutes using a thermal cycler, followed by incubation at 65°C for 15 minutes.
[0293] Custom MGI Adapter Ligation Custom adapters were ligated to end-repaired and A-tail-added PCR marker fragments by preparing the reagent mixtures shown in Table 30 according to the manufacturer's instructions (MGIEasy FS DNA Library Preparation Kit User Manual, Version B4, Chapter 3 Library Construction Protocol; Section 3.4 Adapter Ligation).
[0294] [Table 30]
[0295] The sample was incubated at 37°C for 30 minutes. Following the manufacturer's instructions (MGIEasy FS DNA Library Preparation Kit User Manual, Version B4, Chapter 3 Library Construction Protocol; Section 3.5 Adapter Ligation), the product was purified by adding 20 μL of TE buffer and 200 μL of AMPure XP beads (AMPure XP bead-based reagent, Beckman Coulter; catalog number A63881), and eluted with 24 μL of TE buffer (Ambion, catalog number AM9858). The concentration of the product was measured using the Qubit® 1x dsDNA HS assay kit (Invitrogen; catalog number Q33231): custom MGI adapter ligation product 14.3 ng / μL (329 ng in 23 μL).
[0296] Degeneration and single-strand cyclic formation 25 μL of TE buffer (Ambion, catalog number AM9858) was added to 23 μL of custom MGI adapter ligation product. The product was heated in a thermocycler at 95°C for 3 minutes. After the reaction was complete, the product was placed on ice for 2 minutes and then briefly centrifuged. Single-strand circulation was performed on the denatured PCR marker ligation product according to the manufacturer's instructions (MGIEasy FS DNA Library Preparation Set User Manual, version B4, Chapter 3 Library Construction Protocol; Section 3.10 Single Strand Circularization).
[0297] [Table 31]
[0298] The samples were incubated in a thermocycler at 37°C for 30 minutes and then placed on ice.
[0299] enzymatic digestion The cyclization product was enzymatically digested according to the manufacturer's instructions (MGIEasy FS DNA Library Preparation Kit User Manual, Version B4, Chapter 3 Library Construction Protocol; Section 3.11 Enzymatic Digestion).
[0300] [Table 32]
[0301] After incubating the sample in a thermocycler at 37°C for 30 minutes, 7.5 μL of digestion stop buffer (MGIEasy Circularization Module; MGI; catalog number 1000005260) was added. The solution was transferred to a new 1.5 mL centrifuge tube, and 170 μL of AMPure XP beads (AMPure XP bead-based reagent, Beckman Coulter, catalog number A63881) were added to the enzyme digestion product. Purification was performed according to the manufacturer's instructions (MGIEasy FS DNA Library Preparation Kit User Manual, version B4, Chapter 3 Library Construction Protocol; Section 3.12 Enzymatic Digestion Product Cleanup). The concentration of the purified enzyme digestion product was analyzed using the Qubit® ssDNA assay kit (Invitrogen, catalog number Q10212). Enzymatic digestion product (ssDNA library): 13.6 ng / μL (408 ng in 30 μL).
[0302] Creation of DNA nanoballs (DNBs) The enzyme digestion products were prepared for sequencing on the MGI DNBSEQ-G400 technology platform according to the manufacturer's instructions (High-throughput (rapid) sequencing set - DNBSEQ-G400RS user manual, version 8.0) (FCS PE150, MGI; catalog number 1000016982). Reaction mixture (1) was prepared as shown in Table 33.
[0303] [Table 33]
[0304] The mixture was placed in a thermal cycler, and the hybridization reaction was carried out using the thermal cycler settings shown in Table 34.
[0305] [Table 34]
[0306] Subsequently, DNB mixture 2 was prepared as shown in Table 35.
[0307] [Table 35]
[0308] The mixture was placed in a thermal cycler, and the hybridization reaction was carried out using the thermal cycler settings shown in Table 36.
[0309] [Table 36]
[0310] The reaction was stopped by adding 20 μL of stop DNB reaction buffer. The product concentration was measured using the Qubit™ ssDNA assay kit (Invitrogen, catalog number Q10212): DNB library: 3.33 ng / μL
[0311] DNB Loading DNB loading was performed according to the manufacturer's instructions (High-throughput (rapid) sequencing set - DNBSEQ-G400RS User Manual, version 8.0, Chapter 6.2.2 MGIDLR-200RS DNB loading). To prepare the flow cell for DNB loading, the flow cell was equilibrated at room temperature for 60 minutes to 24 hours. The DNB loading mix was prepared as shown in Table 37.
[0312] [Table 37]
[0313] The sealing gasket and equilibration flow cell were placed in the MGIDL-200H portable DNB loader high, and 30 μL of DNB was added to the fluid inlet. The flow cell was left on the bench for 30 minutes before use.
[0314] Preparation of sequencing reagent cartridges, sequencing, and data processing. The processed samples were sequenced using a DNBSEQ-G400RS (catalog number 900-0000493-00) according to the manufacturer's instructions (High-throughput (rapid) sequencing set - DNBSEQ-G400RS user manual, version 8.0). The sequencing reagent cartridge was prepared according to the manufacturer's instructions (Chapter 7, Preparing the sequencing reagent cartridge, High-throughput (rapid) sequencing set - DNBSEQ-G400RS user manual, version 8.0). The sequencing parameters on the MGI-G400 were set according to the manufacturer's instructions (Chapter 8, Preparing the sequencing reagent cartridge, High-throughput (rapid) sequencing set - DNBSEQ-G400RS user manual, version 8.0). Demultiplexing was performed offline according to the manufacturer's instructions (SplitBarcode manual, version 2.0).
[0315] Results and conclusions The sequenced reads obtained by paired-end sequencing, i.e., read 1 and read 2, were paired: read 1 (obtained by sequencing the DNB and containing the following elements in the 5' to 3' direction: PCR marker fragment, first half of the ligation adapter sequence, linker, second half of the reverse ligation adapter sequence, and reverse PCR marker fragment) and read 2 (obtained by sequencing a strand copied from the DNB by a strand substitution polymerase and containing the same sequence as read 1, compared, and a consensus read was constructed by selecting the base called with the highest quality (highest Q value or Phred quality score)). It was observed that at the modification sites used in the ligation adapter, "A" or "T" nucleotides were incorporated by the strand substitution polymerase used to generate the DNA ball.
[0316] When each read and the consensus read were validated against known adapter sequences (see Table 38), the resulting consensus sequence was found to be of higher quality than two separate single reads. In practice, as shown in column "C" of Table 38, all miscalls can be overcome by selecting the base call with the highest Q value or Phred quality score from each read.
[0317] As shown above, each of the individual reads (reads 1 and 2) also consists of a double helix of PCR fragment sequence and linker adapter sequence, but with reverse complementary orientation after the linker. Therefore, the quality of these single reads can be improved using the double base calls in each of these individual reads. In fact, all miscalls within each read can be overcome by selecting the called base with the highest Q value or Phred quality score (Table 39) from these double-helix sequences within the read, and assuming the read length is sufficient to sequence the complete double-helix sequence, the invention has demonstrated that it enables high-quality single reads and renders paired-end sequencing obsolete.
[0318] Therefore, in conclusion, the linker can be used with the MGI sequencing platform to generate high-quality double reads (within the read in the case of single-ended sequences, and between reads in the case of paired-ended sequences). Thus, the method described herein can improve the accuracy of sequencing (or "base calling") regardless of the platform used to obtain the sequencing reads. Furthermore, other sequencing platforms, such as the PacBio platform, also use the same or similar (strand-substitution) DNA polymerase. This example demonstrates that the DNA polymerase used for DNA ball generation in MGI sequencing can pass through the linker portion, and therefore, it can be reasonably expected that this method will also work with other sequencing platforms such as PacBIO and Element Biosciences sequencing platforms.
[0319] Table 38
[0320] Table 39
Claims
1. An adapter that is at least partially double-stranded, wherein the adapter includes an open end and a closed end, the closed end being covalently closed by a linker, and the linker is not composed of a deoxyribonucleotide.
2. The aforementioned linker is of the general type (L1) 【Chemistry 1】 The adapter according to claim 1, comprising a portion wherein 5' and 3' are used to represent binding sites to nucleotides that form the ends of the adapter, During the ceremony, m is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10; n is either 0 or 1; r 1 and r 2 Each of these is either H, or together they form a bridge portion - (O) 0~1 - (CH 2 ) 1~6 - forms, Preferably, when n is 1, m is 0, or when m is not 0, n is 0. Preferably, the linker is selected from the group consisting of a C3 spacer, a spacer 9, a spacer 18, or one or more 1',2'-dideoxyribose, and the C3 spacer is a portion represented by the general formula (L1) (where m is 0, n is 1, r 1 H, r 2 (is H), The spacer 9 is the part of the general formula (L1) (wherein m is 3 and n is 0), The spacer 18 is the part of the general formula (L1) (wherein m is 6 and n is 0), The 1',2'-dideoxyribose is a part of the general formula (L1) (where m is 0, n is 1, and r 1 and r 2 together form a cross-linking moiety -O-(CH 2 )) 2 - to form), an adapter.
3. The adapter according to claim 1 or 2, wherein the adapter includes staggered open ends.
4. The adapter according to any one of claims 1 to 3, wherein the adapter is formed by a single-stranded nucleic acid molecule including the linker.
5. The adapter according to any one of claims 1 to 4, wherein the adapter includes an identifier array.
6. a) Providing a sample comprising a double-stranded nucleic acid molecule and an adapter according to any one of claims 1 to 5; b) Ligate the adapter to the end of the double-stranded nucleic acid molecule, thereby generating a double-stranded nucleic acid molecule having one open end and one closed end; c) The step of sequencing at least a portion of both strands of the double-stranded nucleic acid molecule in a single sequencing reaction preferably performed on an amplified free sequencing platform to generate a double-stranded read; d) A step of generating a consensus sequence from the double-strand reading and determining the sequence within the double-stranded nucleic acid molecule. A method for determining a target sequence in a double-stranded nucleic acid molecule, including [specific example].
7. Step b) is, b1) Ligate the adapter to both ends of the double-stranded nucleic acid molecule, thereby closing both ends of the nucleic acid molecule; b2) The step of optionally exposing the sample to an exonuclease; b3) The step of cleaving the closed double-stranded nucleic acid molecule to generate a double-stranded nucleic acid molecule having one open end and one closed end. The method according to claim 6, comprising the (sub)step.
8. The method according to claim 6 or 7, wherein the nucleic acid molecule is provided by fragmentation of a longer nucleic acid molecule, and the longer nucleic acid molecule is preferably a genomic nucleic acid molecule.
9. The method according to claim 8, wherein the fragmentation is carried out by restricted endonuclease and / or site-specific endonuclease digestion.
10. The method according to claim 9, wherein the restriction enzyme and / or site-specific endonuclease generates a single-stranded overhang in the double-stranded nucleic acid molecule.
11. The method according to any one of claims 6 to 10, wherein the nucleic acid molecule provided in step a) includes a single-stranded overhang, and the adapter includes an overhang that can be ligated to the overhang of the nucleic acid molecule.
12. The method according to any one of claims 7 to 11, wherein the closed double-stranded nucleic acid molecule of step b3) is cleaved by a site-specific endonuclease or restriction endonuclease, preferably the site-specific endonuclease is RNA-induced CRISPR nuclease or TALEN.
13. The method according to any one of claims 6 to 12, wherein the method is performed on a plurality of samples, preferably the plurality of samples are pooled before step c).
14. The method according to any one of claims 6 to 13, wherein the sequencing adapter is ligated to the open end of the double-stranded nucleic acid after step b) and before step c).
15. The method according to any one of claims 6 to 14, wherein the amplification-free sequencing platform in step c) is nanopore sequencing, preferably nanopore selective sequencing.