NGS library preparation using covalently closed nucleic acid molecule ends

Adapters with protelomerase recognition sites are used to covalently close nucleic acid ends, addressing inefficiencies in library preparation by enhancing enrichment and reducing degradation, thus facilitating accurate sequencing.

JP7766029B2Active Publication Date: 2025-11-07KEYGENE NV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022535868
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-20
Filing Date
2020-12-17
Publication Date
2025-11-07
Estimated Expiration
2040-12-17

AI Technical Summary

Technical Problem

Existing methods for nucleic acid library preparation are inefficient in enriching for fragments of interest and often require separate steps for fragmentation and adapter addition, leading to random nucleic acid fragments and potential degradation.

Method used

The use of adapters containing protelomerase recognition sites, which are ligated to nucleic acid molecules and then cleaved by protelomerase to covalently close the ends, protecting them from exonuclease degradation and allowing for targeted enrichment and sequencing.

Benefits of technology

This method enables efficient enrichment of nucleic acid fragments of interest while minimizing degradation, facilitating accurate sequencing and reducing complexity in nucleic acid samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007766029000002
    Figure 0007766029000002
  • Figure 0007766029000001
    Figure 0007766029000001
Patent Text Reader

Abstract

The present invention relates to an adapter comprising a protelomerase recognition sequence, preferably a TeIN protelomerase recognition sequence. The adapter of the present invention can be used to prepare a nucleic acid molecule library. The present invention also relates to a method for producing a nucleic acid molecule library using one or more adapters comprising a protelomerase recognition sequence. The adapter can be contacted with a protelomerase to cleave and close the ends of the adapter. The closed-end adapter is protected, for example, from exonuclease treatment. The method of the present invention further relates to amplification and sequencing methods using adapters having protelomerase recognition sequences.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is in the field of genetic research, and more particularly in the field of targeted nucleic acid isolation, e.g., for sequence analysis and processing of nucleic acid samples. New methods and means for library preparation and for reducing the complexity of nucleic acid samples are disclosed. [Background technology]

[0002] A key element of genetic research is the sequence analysis of defined DNA loci, for example, to genotype known variants or to identify sequence changes or variants. Such analyses often need to be performed in a multiplexed manner, e.g., a particular set of loci needs to be analyzed in a large number of samples.

[0003] An ideal assay would be flexible in terms of the number of samples and loci required to be screened, highly accurate, and amenable to various sequencing platforms. When analyzing a subset of nucleic acids from a collection of fragments, there is often a need to enrich for fragments of interest. Enrichment can be achieved by selecting (e.g., purifying or amplifying) the target nucleic acid or by removing unwanted nucleic acids. Ideally, the enrichment step does not involve amplification. For example, U.S. Patent Application Publication No. 2014 / 0134610 describes a complexity reduction method that uses a type II restriction enzyme to fragment the nucleic acids in a sample, followed by ligation of protective adapters and subsequent decomposition of all uncaptured nucleic acids using an exonuclease. International Publication No. WO 2016 / 028887 improves on this method by using a programmable endonuclease, i.e., a CRISPR-endonuclease, to fragment the nucleic acids of the sample.

[0004] For most applications, the first step in next-generation sequencing (NGS) is library preparation. Library preparation for NGS can be performed using a variety of protocols. For long-read sequencing libraries using the PacBio platform, hairpin adapters are ligated to the ends of nucleic acid molecules. These hairpin adapters can be added and exonuclease treatment can be used to remove all unadapted molecules and generate sequencing reads that span multiple passes of the input nucleic acid molecule. The latter allows for the creation of highly accurate consensus sequences of the sequenced nucleic acid molecules. Addition of hairpin adapters involves multiple steps, starting with optional fragmentation of the input nucleic acid molecule, followed by end-treating the fragment ends and adding 3' A cohesive (or "sticky") ends. Optionally, a repair step can be performed during this end-treating step to remove damaged positions (e.g., nicks) within the nucleic acid molecule.

[0005] Instead of these separate steps, the fragmentation step and the adapter addition step can be combined into one step (tagmentation) using a transposase enzyme. Tagmentation is widely used, for example, in Illumina Nextera and Oxford Nanopore Technologies (ONT) rapid library preparation protocols. When a repair step is performed after the transposase reaction, most nucleic acid fragments will contain adapters at their ends. However, fragmentation or tagmentation generates fairly random nucleic acid fragments. Summary of the Invention

[0006] It is an object of the present invention to provide a novel method for preparing a library of nucleic acid molecules, e.g., for subsequent sequencing and / or cloning, which preferably includes a step of enriching the library for nucleic acids of interest. [Brief explanation of the drawings]

[0007] [Figure 1] Figure 1 shows Agilent 2100 Bioanalyzer DNA analysis using the DNA12000 kit. The left side of the box shows a DNA size marker including fragment lengths. The left box shows the amplicon (approximately 1050 bp) used as input in the experiment, without (left) and with (right) exonuclease (ExoV) treatment. The middle box shows the amplicon to which a TeIN adapter was ligated, without (left) and with (right) ExoV treatment. The right box shows the amplicon in which the ligated TeIN adapter was treated with TeIN protelomerase, with (left) and without (right) ExoV treatment. The results demonstrate that TeIN treatment of adapter-ligated amplicons provides protection from ExoV degradation. DETAILED DESCRIPTION OF THE INVENTION

[0008] The present invention can be summarized in the following numbered embodiments:

[0009] Embodiment 1. An adapter that is at least partially double-stranded and includes a protelomerase recognition sequence, preferably a TeIN protelomerase recognition sequence.

[0010] Embodiment 2. The adapter of embodiment 1, wherein the adapter further comprises an identifier sequence.

[0011] Embodiment 3. The adapter of embodiment 1 or 2, wherein the adapter comprises at least one sticky end.

[0012] Embodiment 4. A method for preparing a library of nucleic acid molecules, comprising: a) providing a sample comprising at least a first and a second nucleic acid molecule, wherein the first nucleic acid molecule comprises a first target sequence that is not present in the second nucleic acid molecule, and optionally the second nucleic acid molecule comprises a second target sequence; b) ligating the adaptor defined in any one of embodiments 1 to 3 to the ends of the first and second nucleic acid molecules to provide adaptor-linked nucleic acid molecules; c) contacting the adaptor-ligated nucleic acid molecules with a protelomerase, preferably a TeIN protelomerase, to cleave the adaptor-ligated nucleic acid molecules and covalently close the cleaved ends, resulting in first and second nucleic acid molecules comprising closed-ended ends; d) cleaving a first nucleic acid molecule comprising a closed-ended end at a first target sequence to provide a first nucleic acid comprising one open-ended end and one closed-ended end; A method comprising:

[0013] Embodiment 5. The method of embodiment 4, wherein the sample of step a) comprises a first and a second nucleic acid molecule and a plurality of further nucleic acid molecules.

[0014] Embodiment 6. The method of embodiment 4 or 5, wherein the first nucleic acid molecule of step d) is cleaved by a programmable nuclease or a restriction endonuclease.

[0015] Embodiment 7. The method of embodiment 6, wherein the programmable nuclease is an RNA-guided CRISPR nuclease.

[0016] Embodiment 8. The method according to any one of embodiments 4 to 7, wherein the first and second nucleic acid molecules of step a) are prepared by fragmentation, preferably by fragmentation of genomic nucleic acid molecules.

[0017] Embodiment 9. The method of embodiment 8, wherein the adapters in step b) are ligated by tagmentation.

[0018] Embodiment 10. The method of any one of embodiments 4 to 9, comprising a step c1) of exposing the sample to an exonuclease after obtaining a nucleic acid molecule comprising a closed-end end in step c) but before cleaving the first nucleic acid molecule comprising a closed-end end in step d).

[0019] Embodiment 11. The method of any one of embodiments 4 to 9, comprising, after obtaining a first nucleic acid molecule comprising one open end and one closed end in step d), step e) of exposing the sample to an exonuclease.

[0020] Embodiment 12. The method of embodiment 11, comprising step f) cleaving a second nucleic acid molecule comprising a closed-ended end at a second target sequence to result in a second nucleic acid comprising one open-ended end and one closed-ended end.

[0021] Embodiment 13. The method of any one of embodiments 4 to 12, comprising step g) ligating a further adapter to the open end of the first nucleic acid molecule or optionally the second nucleic acid molecule, the open end of the first nucleic acid molecule comprising one open end and one closed end, wherein the further adapter comprises at least one of an amplification primer binding site and a sequence primer binding site, and optionally an identifier sequence.

[0022] Embodiment 14. The method of any one of embodiments 4 to 13, wherein the nucleic acid molecule library is prepared from a plurality of samples, preferably the plurality of samples are pooled, preferably before step c), step d), step e), step f), or before step g).

[0023] Embodiment 15. The method of embodiment 13, wherein the samples are pooled after step g).

[0024] Embodiment 16. The method of any one of embodiments 4 to 13, wherein the adaptor-ligated nucleic acid molecule is repaired in step b) to remove the single-strand break before contacting the molecule with TeIN protelomerase in step c).

[0025] Embodiment 17. A method for amplifying a library of nucleic acid molecules, comprising: Preparing a library of nucleic acid molecules as defined in any one of embodiments 13 to 16; i) a first primer and optionally a second primer that anneal to the first nucleic acid molecule comprising one open end and one closed end obtained in step d); ii) a first primer and optionally a second primer that anneal to a second nucleic acid molecule comprising one open end and one closed end obtained in step f); iii) a first primer and optionally a second primer that anneals to the further adapter defined in step g); and iv) a combination of a first primer defined in i) or ii) and a second primer defined in iii); amplifying the library of nucleic acid molecules using at least one of A method comprising:

[0026] Embodiment 18. A method for analyzing a sequence of interest in a sample comprising a first and a second nucleic acid molecule, comprising: Preparing a library of nucleic acid molecules as defined in any one of embodiments 13 to 16; Optionally, amplifying the prepared nucleic acid molecule as defined in embodiment 17; sequencing, preferably deep sequencing, the library of nucleic acid molecules; A method comprising:

[0027] Embodiment 19. One or more adapters as defined in any one of embodiments 1 to 3; Optionally, a protelomerase, preferably a TeIN protelomerase Kit of parts including.

[0028] definition Various terms relating to the methods, compositions, uses, and other aspects of the present invention are used throughout the specification and claims. Such terms are to be given their ordinary meaning in the art to which the invention pertains, unless otherwise specified. Other specifically defined terms are to be construed in a manner consistent with the definitions provided herein. Although any methods and materials similar or equivalent to those described herein can be used in the practice for testing of the present invention, the preferred materials and methods are described herein.

[0029] Methods for carrying out conventional techniques used in the methods of the present invention will be apparent to those skilled in the art. The practice of conventional techniques in molecular biology, biochemistry, computational chemistry, cell culture, recombinant DNA, bioinformatics, genomics, sequencing, and related fields is well known to those skilled in the art and is discussed, for example, in the following references: Sambrook et al., Molecular Cloning. A Laboratory Manual, 2nd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989; Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1987 and periodic updates; and the series Methods in Enzymology, Academic Press, San Diego.

[0030] "A," "an," and "the": these singular terms include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a combination of two or more cells, and the like.

[0031] As used herein, the term "about" is used to describe and account for slight variations. For example, this term can mean ±10% or less, e.g., ±5% or less, ±4% or less, ±3% or less, ±2% or less, ±1% or less, ±0.5% or less, ±0.1%, or ±0.05% or less. Furthermore, amounts, ratios, and other numerical values ​​may be expressed herein in range format. Such range formats are used for convenience and brevity and should be interpreted flexibly to include the numerical values ​​explicitly set forth as range limits, but also to include all individual numerical values ​​or subranges subsumed within the range, as if each numerical value and subrange were expressly set forth. For example, a ratio in the range of about 1 to about 200 includes the explicitly recited range of about 1 to about 200, but should also be understood to include individual ratios such as about 2, about 3, and about 4, as well as subranges such as about 10 to about 50 and about 20 to about 100.

[0032] As used herein, the term "adapter" refers to a single-stranded, double-stranded, partially double-stranded, Y-shaped, or hairpin nucleic acid molecule, preferably chemically synthesized, that can be attached to, and preferably ligated to, one or both strands of a double-stranded DNA molecule, preferably of a limited length, e.g., about 10 to about 200, or about 10 to about 100 base pairs, or about 10 to about 80, or about 10 to about 50, or about 10 to about 30 base pairs, at the end of another nucleic acid. The double-stranded structure of the adapter can be formed by two separate oligonucleotide molecules that base pair with each other, or by a hairpin structure of a single oligonucleotide strand. As will be apparent, the attachable end of the adapter can be designed to be optionally ligatable so as to be compatible with overhangs generated upon cleavage by restriction enzymes and / or programmable nucleases, or can be designed to be compatible with overhangs generated after non-templated extension reactions (e.g., 3'-A addition), or can have blunt ends.

[0033] "And / or": The term "and / or" means that one or more of the stated cases may occur alone or in combination with at least one of the stated cases, up to all of the stated cases.

[0034] "Amplification," as used with respect to nucleic acids or nucleic acid reactions, refers to an in vitro method of generating copies of a specific nucleic acid, such as a target nucleic acid or a tagged nucleic acid. Numerous methods for amplifying nucleic acids are known in the art, including polymerase chain reaction, ligase chain reaction, strand displacement amplification reaction, rolling circle amplification reaction, transcription-mediated amplification, such as NASBA (e.g., U.S. Pat. No. 5,409,818), loop-mediated amplification (e.g., "LAMP" amplification using a loop-forming sequence, as described in U.S. Pat. No. 6,410,278), and isothermal amplification reactions. The nucleic acid being amplified can be DNA or RNA, or DNA comprising, consisting of, or derived from a mixture of DNA and RNA, including modified DNA and / or RNA. The products resulting from the amplification of one or more nucleic acid molecules (i.e., "amplification products"), regardless of whether the starting nucleic acid is DNA, RNA, or both, can be either DNA or RNA, or a mixture of both DNA and RNA nucleosides or nucleotides, or they can contain modified DNA or RNA nucleosides or nucleotides.

[0035] A "copy" can be, but is not limited to, a sequence that has perfect sequence complementarity or perfect sequence identity to a particular sequence. Alternatively, a copy need not necessarily have perfect sequence complementarity or identity to the particular sequence; for example, some degree of sequence variation may be tolerated. For example, a copy may contain nucleotide analogs, such as deoxyinosine or deoxyuridine, intentional sequence modifications (e.g., sequence modifications introduced via primers containing sequences that are hybridizable to, but not complementary to, a particular sequence), and / or sequence errors that occur during amplification.

[0036] The term "complementary" is defined herein as the sequence identity of a sequence relative to a fully complementary strand (e.g., a second or reverse strand). For example, a sequence that is 100% complementary (or fully complementary) is understood herein to have 100% sequence identity with the complementary strand, and for example, a sequence that is 80% complementary is understood herein to have 80% sequence identity with the (fully) complementary strand.

[0037] "Comprises": This term is to be interpreted as inclusive and open-ended, not exclusive. Specifically, this term and variations thereof mean that the specified features, steps, or components are included. These terms are not to be interpreted as excluding the presence of other features, steps, or components.

[0038] "Construct" or "nucleic acid construct" or "vector": This refers to an artificial nucleic acid molecule obtained by the use of recombinant DNA techniques, which can often be used to deliver exogenous DNA into a host cell for the purpose of expressing the DNA region contained in the construct in the host cell. The vector backbone of the construct can be, for example, a plasmid into which a (chimeric) gene has been incorporated, or, if appropriate transcriptional regulatory sequences are already present (e.g., an (inducible) promoter), only the desired nucleotide sequence (e.g., a coding sequence) is incorporated downstream of the transcriptional regulatory sequence. Vectors can contain additional genetic elements, such as selectable markers, multiple cloning sites, etc., to facilitate their use in molecular cloning.

[0039] As used herein, the terms "double stranded" and "duplex" refer to two complementary polynucleotides that base pair, i.e., hybridize together. Complementary nucleotide strands are also known in the art as reverse complements.

[0040] The term "effective amount," as used herein, refers to an amount of a biologically active agent that is sufficient to induce a desired biological effect. For example, in some embodiments, an effective amount of an exonuclease can refer to an amount of exonuclease sufficient to induce cleavage of unprotected nucleic acid. As will be apparent to one skilled in the art, the effective amount of an agent can vary depending on various factors, such as the agent used, the conditions under which the agent is used, and the desired biological effect, such as the degree of nuclease cleavage detected.

[0041] "Exemplary": This term means "serving as an example, instance, or illustration," and should not be construed as excluding other configurations disclosed herein.

[0042] "Expression": This refers to the process by which a DNA region that is operably linked to appropriate regulatory regions, in particular a promoter, is transcribed into RNA that can then be translated into proteins or peptides.

[0043] A "guide sequence" is herein understood as a sequence that directs an RNA- or DNA-guided endonuclease to a specific site in an RNA or DNA molecule. In the context of a gRNA-CAS complex, a "guide sequence" is herein further understood as the section of an sgRNA or crRNA that is required to target the gRNA-CAS complex to a specific site in duplex DNA.

[0044] The gRNA-CAS complex, also referred to herein as a CRISPR-endonuclease or CRISPR-nuclease, is understood to be a CAS protein complexed or hybridized with a guide RNA, which may be a crRNA and / or tracrRNA, or an sgRNA.

[0045] "Identity" and "similarity" can be easily calculated by known methods. "Sequence identity" and "sequence similarity" can be determined by aligning two peptide sequences or two nucleotide sequences using a global or local alignment algorithm, depending on the length of the two sequences. Sequences of similar length are preferably aligned using a global alignment algorithm (e.g., Needleman Wunsch) that optimally aligns the sequences over their entire length, while sequences of substantially different lengths are preferably aligned using a local alignment algorithm (e.g., Smith Waterman). Sequences can be said to be "substantially identical" or "essentially similar" if they share at least a certain percentage of sequence identity (defined below) (e.g., when optimally aligned using the programs GAP or BESTFIT using default parameters). GAP uses the Needleman and Wunsch global alignment algorithm to align two sequences over their entire length (full length), maximizing the number of matches and minimizing the number of gaps. Global alignment is appropriately used to determine sequence identity when two sequences have similar lengths. Generally, the GAP default parameters are used, with a gap creation penalty of 50 (nucleotides) / 8 (proteins) and a gap extension penalty of 3 (nucleotides) / 2 (proteins). For nucleotides, the default scoring matrix used is nwsgapdna, and for proteins, the default scoring matrix is ​​Blosum62 (Henikoff & Henikoff, 1992, PNAS 89, 915-919).Sequence alignment and sequence identity percentage scores can be determined using computer programs such as the GCG Wisconsin Package, version 10.3, available from Accelrys Inc., 9685 Scranton Road, San Diego, CA 92121-3752 USA, or open source software such as the EmbossWIN version 2.10.0 program "needle" (using the global Needleman Wunsch algorithm) or "water" (using the local Smith Waterman algorithm), using the same parameters as GAP described above, or using default settings (both "needle" and "water", and both protein and DNA alignments, the default gap creation penalty is 10.0, and the default gap extension penalty is 0.5; the default scoring matrix is ​​Blosum62 for proteins and DNAFull for DNA). When sequences have substantially different overall lengths, local alignment, such as using the Smith Waterman algorithm, is preferred.

[0046] Alternatively, the percentage of similarity or identity can be determined by searching public databases using algorithms such as FASTA, BLAST, etc. Thus, the nucleic acid and protein sequences of the present invention can further be used as "query sequences" to perform searches against public databases, for example, to identify other family members or related sequences. Such searches can be performed using the BLASTn and BLASTx programs (version 2.0) of Altschul et al. (1990) J. Mol. Biol. 215:403-10. BLAST nucleotide searches can be performed using the NBLAST program, score = 100, word length = 12, to obtain nucleotide sequences homologous to the nucleic acid molecules of the present invention. BLAST protein searches can be performed using the BLASTx program, score = 50, word length = 3, to obtain amino acid sequences homologous to the protein molecules of the present invention. To obtain gapped alignments for comparison purposes, Gapped BLAST can be used as described in Altschul et al. (1997) Nucleic Acids Res. 25(17): 3389-3402. When using BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., BLASTx and BLASTn) can be used. See the homepage of the National Center for Biotechnology Information at http: / / www.ncbi.nlm.nih.gov / .

[0047] The term "nucleotide" includes naturally occurring nucleotides, including, but not limited to, guanine, cytosine, adenine, and thymine (G, C, A, and T, respectively). The term "nucleotide" is further intended to include those portions that contain not only the known purine and pyrimidine bases, but also modified other heterocyclic bases. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, and alkylated riboses or other heterocycles. Additionally, the term "nucleotide" includes those portions that contain hapten or fluorescent labels and that may contain not only conventional ribose and deoxyribose sugars, but other sugars as well. Modified nucleosides or nucleotides also include modifications to the sugar moiety, e.g., one or more hydroxyl groups replaced with halogen atoms or aliphatic groups, or functionalized as ethers, amines, etc.

[0048] The terms "nucleic acid," "polynucleotide," and "nucleic acid molecule" are used interchangeably herein to describe polymers of any length composed of nucleotides, e.g., deoxyribonucleotides or ribonucleotides, e.g., polymers of greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1000 bases, up to about 10,000 or more bases, and can be produced enzymatically or synthetically (e.g., PNAs, as described in U.S. Pat. No. 5,948,902 and references cited therein). Nucleic acids can hybridize with naturally occurring nucleic acids in a sequence-specific manner similar to the sequences of two naturally occurring nucleic acids, e.g., participate in Watson-Crick base pairing interactions. Furthermore, nucleic acids and polynucleotides can be isolated (and optionally fragmented) from cells, tissues, and / or bodily fluids. Nucleic acids can be, for example, genomic DNA (gDNA), mitochondrial, cell-free DNA (cfDNA), library-derived DNA, and / or library-derived RNA.

[0049] As used herein, the term "nucleic acid sample" or "sample containing nucleic acid" refers to any sample containing nucleic acid, and a sample relates to a material or mixture of materials, typically, but not necessarily, in liquid form, containing one or more nucleic acid molecules of interest. The one or more nucleic acid molecules of interest preferably comprise a sequence of interest. The nucleic acid molecules of interest are preferably a first nucleic acid molecule or a second nucleic acid molecule as defined herein. The nucleic acid sample preferably comprises a sequence of interest. The nucleic acid sample used as starting material in the methods of the present invention can be from any source, e.g., a whole genome, a collection of chromosomes, a single chromosome, one or more regions derived from one or more chromosomes or transcribed genes, and can be directly purified from biological or experimental sources, e.g., a nucleic acid library. Nucleic acid samples can be obtained from the same individual, which may be human or other species (e.g., plants, bacteria, fungi, algae, archaea, etc.), or from different individuals of the same species, or from different individuals of different species. For example, the nucleic acid sample can be derived from a cell, a tissue, a biopsy, a body fluid, a genomic DNA library, a cDNA library, and / or an RNA library. The nucleic acid sample preferably comprises at least a first nucleic acid molecule and a second nucleic acid molecule.

[0050] The term "sequence of interest" includes, but is not limited to, any genetic sequence preferably present in a cell, such as a gene, a portion of a gene, or a non-coding sequence within or adjacent to a gene. The sequence of interest may be present in a chromosome, an episome, an organelle genome, such as the mitochondrial genome or chloroplast genome, or in genetic material that can exist independently of the body of genetic material, such as an infectious viral genome, a plasmid, an episome, a transposon, etc. The sequence of interest may be within the coding sequence of a gene, within a transcribed non-coding sequence, such as a leader sequence, a trailer sequence, or an intron. The nucleic acid sequence of interest may be present in double-stranded or single-stranded nucleic acid. Preferably, the sequence of interest is present in a first nucleic acid molecule or a second nucleic acid molecule.

[0051] A sequence of interest can be, but is not limited to, a sequence that has or is suspected of having a polymorphism, such as a SNP.

[0052] As used herein, the term "oligonucleotide" refers to a single-stranded multimer of nucleotides, preferably about 2 to 200 nucleotides in length, or up to 500 nucleotides in length. Oligonucleotides may be synthetic or enzymatically produced and, in some embodiments, are about 10 to 50 nucleotides in length. Oligonucleotides may contain ribonucleotide monomers (i.e., oligoribonucleotides) or deoxyribonucleotide monomers. Oligonucleotides may be, for example, about 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 100, 100 to 150, 150 to 200, or about 200 to 250 nucleotides in length.

[0053] "Plant": This includes plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant callus, plant mass, and intact plant cells of plants or plant parts such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, grains, ears, cobs, husks, stems, roots, root tips, anthers, grains, etc. Non-limiting examples of plants include crops and cultivated plants such as barley, cabbage, canola, cassava, cauliflower, chicory, cotton, cucumber, eggplant, grapes, peppers, lettuce, corn, melon, rapeseed, potato, pumpkin, rice, rye, sorghum, squash, sugarcane, sugar beet, sunflower, bell pepper, tomato, watermelon, wheat, and zucchini.

[0054] A "protospacer sequence" is a sequence that can be recognized or hybridized by a guide RNA, more particularly a crRNA, or in the case of an sgRNA, a guide sequence within the crRNA portion of the guide RNA. It is understood herein that a "protospacer sequence" in the context of the present invention is an example of a sequence that is present in a target sequence, i.e., a first or second nucleic acid molecule as defined herein.

[0055] An "endonuclease" is an enzyme that hydrolyzes at least one strand of a double-stranded DNA or RNA molecule when bound to its target or recognition site. An endonuclease is herein understood to be a site-specific endonuclease, and the terms "endonuclease" and "nuclease" are used interchangeably herein. A restriction endonuclease is herein understood to be an endonuclease that simultaneously hydrolyzes both strands of a duplex, introducing a double-strand break in DNA. A "nicking" endonuclease is an endonuclease that hydrolyzes only one strand of a duplex, generating a "nicked" rather than a cut DNA molecule.

[0056] An "exonuclease" is defined herein as any enzyme that cleaves one or more nucleotides from the end (exo) of a polynucleotide.

[0057] "Reducing complexity" or "complexity reduction" is herein understood to mean reducing the complexity of a nucleic acid sample, such as a sample derived from genomic DNA, cfDNA derived from liquid biopsy, isolated RNA sample, etc. Complexity reduction preferably results in the enrichment of one or more specific nucleic acids containing a sequence of interest contained within the complex starting material, and / or the generation of a subset of samples, wherein the subset preferably comprises or consists of one or more specific nucleic acids containing a sequence of interest contained within the complex starting material, and preferably the amount of non-specific nucleic acids not containing a sequence of interest is reduced by at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% compared to the amount of non-specific nucleic acids in the starting material, i.e., compared to before complexity reduction.

[0058] Complexity reduction is generally performed prior to further analytical or method steps, such as amplification, barcoding, sequencing, determination of epigenetic changes, etc. Preferably, the complexity reduction is a reproducible complexity reduction, meaning that when the complexity of the same sample is reduced using the same method, the same or at least comparable subsets are obtained, as opposed to random complexity reduction.

[0059] Examples of complexity reduction methods include, for example, AFLP® (Keygene NV, the Netherlands; see, e.g., EP 0 534 858), arbitrarily primed PCR amplification, capture probe hybridization, the method described by Dong (see, e.g., WO 03 / 012118, WO 00 / 24939), and indexed ligation (Unrau P. and Deugau KV (1994) Gene 145:163-169), methods described in WO 2006 / 137733; WO 2007 / 037678; WO 2007 / 073165; WO 2007 / 073171, U.S. Patent Application Publication Nos. 2005 / 260628, WO 03 / 010328, U.S. Patent Application Publication No. 2004 / 10153, genome partitioning (see, e.g., WO 2004 / 022758), serial analysis of gene expression (SAGE; see, e.g., Velculescu et al., 1995, supra, and Matsumura et al., 1999, The Plant Journal, vol. 20(6):719-726), and modifications of SAGE (see, e.g., Powell, 1998, Nucleic Acids Research, vol. 26(14):3445-3446; and Kenzelmann and Muhlemann, 1999, Nucleic Acids Research, vol. 27(3):917-918), MicroSAGE (see, e.g., Datson et al., 1999, Nucleic Acids Research, vol. 27(5):1300-1307), Massively Parallel Signature Sequencing (MPSS; see, e.g., Brenner et al., 2000, Nature Biotechnology, vol. 18:630-634, and Brenner et al., 2000, PNAS, vol. 97(4):1665-1670), self-subtracted cDNA libraries (Laveder et al., 2002, Nucleic Acids Research, vol.30(9):e38), real-time multiplex ligation-dependent probe amplification (RT-MLPA; see, for example, Eldering et al., 2003, vol. 31(23):el53), high-coverage expression profiling (HiCEP; see, for example, Fukumura et al., 2003, Nucleic Acids Research, vol. 31(16):e94), the universal microarray system disclosed in Roth et al. (Roth et al., 2004, Nature Biotechnology, vol. 22(4):418-426), transcriptome subtraction (see, for example, Li et al., Nucleic Acids Research, vol. 33(16):el36), and fragment display (see, for example, Metsis et al., 2004, Nucleic Acids Research, vol. 32(16):el27).

[0060] "Sequence" or "nucleotide sequence": This refers to the order of nucleotides of or within a nucleic acid. In other words, any order of nucleotides in a nucleic acid can be referred to as a sequence or nucleic acid sequence. For example, a target sequence is the order of nucleotides contained in one strand of a DNA duplex.

[0061] The term "sequencing," as used herein, refers to a method that obtains the identity of at least 10 consecutive nucleotides of a polynucleotide (e.g., the identity of at least 20, at least 50, at least 100, or at least 200 or more consecutive nucleotides). The term "next-generation sequencing" refers to sequencing by so-called parallelized sequencing-by-synthesis or ligation platforms, such as those currently employed by Illumina, Life Technologies, PacBio, and Roche. Next-generation sequencing methods may also include nanopore sequencing methods, such as those commercialized by Oxford Nanopore Technologies (ONT), or electronic detection-based methods, such as the Ion Torrent technology commercialized by Life Technologies. Preferably, the next-generation sequencing method is a nanopore sequencing method, preferably a nanopore selective sequencing method.

[0062] "Nanopore selective sequencing" should be understood as the use of nanopore sequencing technologies such as Oxford Nanopore or Ontera to selectively sequence single molecules in real time and map streaming nanopore current signals or base calls to a reference sequence in order to reject non-target sequences. Depending on the data being generated, the sequencer is operated to either pursue sequencing of the nucleic acid or terminate and remove the nucleic acid from the sequencing pore by reversing the polarity of the voltage of a specific pore for a specific short period of time sufficient to eject non-target molecules and make the nanopore available for a new sequencing read. Examples of nanopore selective sequencing methods are described in Payne et al., 2020 (Nanopore adaptive sequencing for mixed samples, whole exome capture and targeted panels, February 3, 2020; DOI: 10.1101 / 2020.02.03.926956) and Kovaka et al., 2020 (Targeted nanopore sequencing by real-time mapping of raw electrical signal with UNCALLED, February 3, 2020; doi: 10.1101 / 2020.02.03.931923), both of which are incorporated herein by reference.

[0063] A "first nucleic acid molecule" in the context of the present invention can be a smaller or longer stretch or selected portion of a single-stranded or double-stranded nucleic acid. Prior to carrying out the method of the present invention, the first nucleic acid molecule can be contained within a larger nucleic acid molecule, for example, within a larger nucleic acid molecule present in the sample to be analyzed. Preferably, the first nucleic acid molecule comprises a first target sequence.

[0064] A "second nucleic acid molecule" in the context of the present invention can be a smaller or longer stretch or selected portion of a single-stranded or double-stranded nucleic acid. Prior to performing the methods of the present invention, the second nucleic acid molecule can be contained within a larger nucleic acid molecule, for example, within a larger nucleic acid molecule present in the sample to be analyzed. The first nucleic acid molecule can be present in the same larger nucleic acid molecule. Alternatively, the first and second nucleic acid molecules can be present in separate larger nucleic acid molecules, which are present in the same sample. In some embodiments, the second nucleic acid molecule can comprise a second target sequence.

[0065] At least one of the first and second nucleic acid molecules may contain a sequence of interest. Preferably, the first nucleic acid molecule contains the sequence of interest. In an alternative embodiment, the second nucleic acid molecule contains the sequence of interest.

[0066] The sequence of interest can be any sequence within a nucleic acid sample, such as a gene, gene complex, locus, pseudogene, regulatory region, highly repetitive region, polymorphic region, or portion thereof. The sequence of interest can also be a region containing a genetic or epigenetic variation that is indicative of a phenotype or disease. The sequence of interest is preferably the subject of further analysis or action, such as, but not limited to, copying, amplification, sequencing, and / or other procedures for nucleic acid interrogation.

[0067] A "target sequence" is defined herein to be a sequence present in a first or second nucleic acid molecule as defined herein, which sequence is recognized by at least one of a nuclease and a nickase as defined herein.

[0068] In some embodiments, a plurality or "set" of nucleic acid molecules used in the methods of the invention comprises one or more sequences of interest selected for enrichment. Optionally, such a set consists of structurally or functionally related nucleic acid molecules. Nucleic acid molecules in the context of the present invention can include both natural and unnatural artificial or non-standard nucleotides, including, but not limited to, DNA, RNA, BNA (bridged nucleic acid), LNA (locked nucleic acid), PNA (peptide nucleic acid), morpholino nucleic acid, glycol nucleic acid, threos nucleic acid, epigenetically modified nucleotides such as methylated DNA, and mimetics and combinations thereof.

[0069] Preferably, the sequence of interest is a small or longer contiguous stretch of nucleotides (i.e., a polynucleotide) in a single strand of double-stranded DNA, said double-stranded DNA further comprising a complementary strand comprising a sequence complementary to the sequence of interest. Preferably, said double-stranded DNA is genomic DNA (gDNA) and / or cell-free DNA (cfDNA).

[0070] The present inventors have discovered that adapters containing protelomerase recognition sites can be used for library preparation. In particular, adapters containing a recognition site for the protelomerase enzyme can be ligated to nucleic acid molecules, which are either double-stranded or become double-stranded after adapter ligation. These adapters are then cleaved by the protelomerase enzyme, simultaneously covalently closing the ends of the nucleic acid molecule. When both ends of a nucleic acid molecule are closed in this manner, the molecule is protected from exonuclease degradation due to the lack of free "terminal" nucleotides.

[0071] The terminus of a double-stranded nucleic acid in which the 3'-terminal nucleotide of a corresponding upper strand is covalently linked to the 5'-terminal nucleotide of a corresponding lower strand is annotated herein as a "closed-ended end." Similarly, the terminus of a double-stranded nucleic acid in which the 5'-terminal nucleotide of a corresponding upper strand is covalently linked to the 3'-terminal nucleotide of a corresponding lower strand is also annotated herein as a "closed-ended end." Thus, a "closed-ended end" is understood herein to be the terminus of a double-stranded nucleic acid in which the terminal nucleic acids of opposite strands are covalently linked to each other, as opposed to an "open-ended end," which is understood herein to be the terminus of a double-stranded nucleic acid in which the terminal nucleic acids of opposite strands are not covalently linked to each other.

[0072] In the novel library preparation methods detailed herein, preferably, all nucleic acid molecules present in a particular nucleic acid sample are tagged on both sides with protelomerase adapters, which are therefore cleaved during protelomerase treatment to yield covalently closed nucleic acid molecules that are insensitive to 5'- or 3'-modifying enzymes. An optional step of exonuclease treatment of the protelomerase-treated sample can be added to remove any possible nucleic acid molecules that are not covalently closed at both ends. The (covalently closed) nucleic acid molecules can then be selectively opened, for example, by using targeted or programmable endonucleases. While all nucleic acid molecules are still present in the reaction mixture, only those cleaved in the final open-end reaction can be used in the subsequent (sequencing) process, for example, by ligating sequencing adapters to the open-ended ends, thereby selectively rendering these open-ended fragments ready for sequencing. Alternatively, exonuclease treatment can be used to degrade open-ended fragments, thereby enriching the unopened nucleic acid molecules for further processing. For example, such open-ended molecules can be opened in a second round of selective opening, for example, using a programmable endonuclease that targets such open-ended molecules.

[0073] The above mentioned approach has at least the following advantages:

[0074] Barcodes can be added to the ends of the nucleic acid molecules and the samples can then be pooled before further sample preparation steps are performed.

[0075] Only a single CRISPR enzyme / guide complex is required to target a locus, as opposed to the typical use of two gRNAs per target locus.

[0076] This approach is, in principle, independent of the sequencing platform.

[0077] This approach allows nucleic acid molecules to be targeted without an amplification step, thereby enabling the detection of natural base modifications.

[0078] The mentioned approach can be applied to nucleic acid molecules of any length, ie, short molecules (<1 Kbp) or long molecules (>5 Kbp).

[0079] Thus, in a first aspect, the present invention relates to an adapter comprising a protelomerase recognition sequence. Preferably, the adapter comprises a TeIN protelomerase recognition sequence. Preferably, the adapter is for use in a method of the present invention. Preferably, the adapter is capable of ligating to a nucleic acid molecule used in a method of the present invention.

[0080] The adapter may be single-stranded. The single-stranded adapter preferably comprises a section, preferably at its 3' end, capable of hybridizing to a nucleic acid molecule used in the method of the present invention. The single-stranded adapter can preferably hybridize to a single-stranded overhang of a nucleic acid molecule, preferably to a 3' overhang of a nucleic acid molecule. The single-stranded portion of the annealed single-stranded adapter can then be filled in, i.e., made double-stranded, using a polymerase such as, but not limited to, Klenow (known to those skilled in the art to have 5'->3' polymerase activity and 3'->5' exonuclease activity but lacking 5'->3' exonuclease activity) or Bst-polymerase (known to those skilled in the art to be a DNA polymerase derived from Bacillus stearothermophilus that has 5'->3' polymerase activity and strand displacement activity but lacking 3'->5' exonuclease activity). The filling step optionally results in the generation of a double-stranded protelomerase recognition sequence.

[0081] Preferably, the adapter is at least partially double-stranded. In the methods of the present invention defined herein, an at least partially double-stranded adapter can be ligated to a nucleic acid molecule. Preferably, at least 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the nucleotides of the adapter are double-stranded. Preferably, the protelomerase recognition sequence is double-stranded. The adapter may be 100% or "fully" double-stranded. After ligating the adapter to the nucleic acid molecule, the adapter can be made fully double-stranded by, for example, using a DNA polymerase to fill in the single-stranded portion of the adapter.

[0082] Preferably, an at least partially double-stranded adapter comprises two single-stranded molecules that are capable of at least partially annealing to one another, i.e., the double-stranded adapter preferably comprises two open ends prior to ligation of the adapter to a nucleic acid molecule as defined herein.

[0083] One end of an at least partially double-stranded adaptor can be ligated to a nucleic acid molecule. Therefore, preferably, at least one end ligated to a nucleic acid molecule is double-stranded. At least one of the double-stranded ends of the adaptor can be blunt, cohesive, or "sticky." Preferably, the adaptor comprises at least one cohesive end. Preferably, the end of the adaptor ligated to a nucleic acid molecule has an end that is compatible with the end of the nucleic acid molecule. For example, if the nucleic acid molecule comprises an end with an A overhang, the adaptor preferably comprises an end with a T overhang. Similarly, if the nucleic acid molecule is obtained by enzymatic digestion, leaving behind an overhang of 1, 2, 3, 4, 5, or more nucleotides, the adaptor preferably comprises an overhang of 1, 2, 3, 4, 5, or more nucleotides that is complementary to the overhang of the nucleic acid molecule.

[0084] The other end of the adaptor preferably cannot be ligated to a nucleic acid molecule or to an adaptor. Any means of blocking ligation of the adaptor end is suitable for use in the methods of the present invention. As a non-limiting example, the other end of the adaptor may be single-stranded or include a non-compatible overhang.

[0085] The adapters of the present invention comprise a protelomerase recognition sequence, preferably a TeIN protelomerase recognition sequence. A protelomerase recognition sequence is any DNA sequence whose presence in a DNA template allows for the enzymatic activity of protelomerase to convert the DNA into a closed-ended linear DNA. In other words, the protelomerase recognition sequence is required for the cleavage and religation of double-stranded DNA by protelomerase to form a covalently closed, linear DNA. Typically, the protelomerase recognition sequence comprises a perfect palindrome, i.e., a double-stranded DNA sequence with two-fold rotational symmetry.

[0086] The length of the perfect inverted repeat varies depending on the particular organism. In Borrelia burgdorferi, the perfect inverted repeat is 14 base pairs long. In various mesophilic bacteriophages, the perfect inverted repeat is 22 base pairs long or longer. Also, in some cases, such as E. coli N15, the central perfect inverted palindrome is adjacent to an inverted repeat sequence, thus forming part of a larger imperfect inverted palindrome.

[0087] The protelomerase recognition sequence used in the present invention preferably comprises a double-stranded palindrome (perfect inverted repeat) sequence at least 14 base pairs in length. Preferred perfect inverted repeat sequences include the sequences of SEQ ID NOS: 1-9 and variants thereof. SEQ ID NOS: 1 (NCATNNTANNCGNNTANNATGN) is a 22-base consensus sequence. For example, as disclosed in WO 2010 / 086626, the base pairs of a perfect inverted repeat are conserved at certain positions, but the sequence may be flexible at other positions. Thus, SEQ ID NOS: 1 is preferably the minimal consensus sequence for a perfect inverted repeat sequence for use with protelomerase in the methods of the present invention. The protelomerase recognition sequence may have the sequence described in WO 2010 / 086626, which is incorporated herein by reference.

[0088] Preferably, the protelomerase recognition sequence has at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 10. The sequence of SEQ ID NO: 10 is 5'-TATCAGCACACAATTGCCCATTATACGCGCGTATAATGGACTATTGTGTGCTGATA-3' is.

[0089] Preferably, the protelomerase cleaves the adapter sequence at positions 28 to 29 of the recognition sequence and closes the cleaved end.

[0090] The adapter may consist of a protelomerase recognition sequence. Alternatively, the adapter may comprise additional nucleotides. The adapter may comprise an identifier sequence or "barcode" or "tag." The identifier is preferably at least one of a sample identifier and a UMI. Preferably, the recognition sequence remains part of the nucleic acid molecule after cleavage and the cleaved ends are closed.

[0091] The UMI may be a separate sequence within the adapter, or if the protelomerase recognition sequence contains degenerate nucleotides, these degenerate nucleotides may be used to introduce the identifier. For example, in the case of degenerate nucleotides in the protelomerase recognition sequence of one sample, an adapter can be used with one or more specific nucleotides within this recognition sequence, while in a second and further sample, other specific nucleotides are used at this position, thereby creating an identifier sequence within the protelomerase recognition sequence. The adapter can include a sample identifier and a UMI.

[0092] A sample identifier can associate the sequence of a nucleic acid molecule with a particular sample. For example, the adapters used in the methods of the invention can include an identifier sequence specific to a particular sample. Each additional sample can be processed using an adapter with an identifier sequence specific to that additional sample. The processed samples can then be pooled, and the resulting sequence can be assigned to a particular sample using the sample identifier sequence.

[0093] A UMI is a nucleic acid molecule-specific, i.e., substantially unique, preferably completely unique, sequence or barcode unique to each nucleic acid molecule used in the methods of the present invention. A UMI can have a random, pseudorandom, partially random, or non-random nucleotide sequence. A UMI can be used to uniquely identify the source molecule from which a sequencing read originates. For example, reads of amplified nucleic acid molecules can converge to a single consensus sequence for each source nucleic acid molecule. As indicated above, a UMI can be completely or substantially unique. Completely unique is herein understood to mean that every adaptor-ligated nucleic acid molecule provided in the methods of the present invention contains a unique tag that is different from all other tags contained in additional adaptor-ligated nucleic acid molecules used in the methods of the present invention. Substantially unique is herein understood to mean that each adaptor-ligated nucleic acid molecule provided in the methods of the present invention contains a random UMI, but the percentage of such adaptor-ligated nucleic acid molecules that may contain the same UMI is low. Preferably, substantially unique molecular identifiers are used when the likelihood of tagging identical molecules containing the same sequence with the same UMI is negligible. Preferably, the UMI is completely unique with respect to the particular sequence of the nucleic acid molecule. The UMI preferably has a length sufficient to ensure this uniqueness. In some implementations, less unique molecular identifiers (i.e., the substantially unique identifiers set forth above) can be used in conjunction with other identification techniques to ensure that each nucleic acid molecule is uniquely identified during the sequencing process.

[0094] The identifier sequence may range in length from about 2 to 100 nucleotide bases or more, preferably about 4 to 16 nucleotide bases. The identifier sequence may be a continuous sequence or may be divided into several subunits. Each of these subunits may be present on a single adapter or on separate adapters. For example, if a nucleic acid molecule is flanked by two adapters, each of these two adapters may contain a subunit of the identifier sequence. To obtain a consensus sequence, the sequence reads obtained by the method of the present invention can be grouped based on the information of each of the two subunits.

[0095] Preferably, the identifier sequences do not contain two or more consecutive identical bases, and more preferably there are differences in at least two, preferably at least three bases between the identifier sequences.

[0096] Means for designing and constructing adapters for use in the present invention are well known to those of skill in the art, and the present invention is not limited to any particular adapter design and / or construction. As a non-limiting example, two oligonucleotides can be constructed and annealed to each other under controlled conditions to obtain an at least partially double-stranded adapter for use in the present invention. As a further non-limiting example, a long oligonucleotide and a short oligonucleotide can be constructed, with the short oligonucleotide annealing to the end of the long oligonucleotide. Preferably, at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the nucleotides of the short oligonucleotide can anneal to the long oligonucleotide. Preferably, the short oligonucleotide is 100% complementary to a section of the long oligonucleotide. Preferably, this complementary section is located 3' of the protelomerase recognition sequence, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides from 3' of the recognition sequence. The complementary section can be located between the protelomerase recognition sequence and the 3' end of the long oligonucleotide. The complementary section can be located at the 3' end of the long oligonucleotide. Alternatively, the complementary section can be located upstream of the 3' end of the long oligonucleotide, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or more nucleotides upstream of the 3' end of the long oligonucleotide. After annealing the short and long oligonucleotides, the portion of the long oligonucleotide located 5' of the complementary section can be filled in, thus producing a double-stranded adapter, which may have a 3' overhang, in which case the 3' overhang is the 3' end of the long oligonucleotide. Filling in the single-stranded sequence, i.e., generating a double-stranded sequence, can be performed using any conventional polymerase, such as, but not limited to, Klenov polymerase or BST polymerase. A preferred polymerase is BST polymerase.

[0097] Optionally, the adapters of the present invention further comprise a restriction enzyme recognition site between the protelomerase recognition sequence and the adapter portion for ligation to a nucleic acid molecule. In a further aspect, the present invention relates to a method for preparing a library of nucleic acid molecules, preferably comprising the following steps: a) providing a sample comprising at least a first and a second nucleic acid molecule, wherein the first nucleic acid molecule comprises a first target sequence that is not present in the second nucleic acid molecule; b) ligating an adaptor, as defined herein, i.e., comprising a protelomerase recognition sequence, to the end of the first nucleic acid molecule to provide an adaptor-ligated nucleic acid molecule; c) contacting the adaptor-ligated nucleic acid molecule with a protelomerase to cleave the adaptor-ligated nucleic acid molecule and covalently closing the cleaved ends to provide a first nucleic acid molecule comprising closed-ended ends; and d) cleaving the first nucleic acid molecule comprising the closed-ended end to provide a first nucleic acid comprising one open-ended end and one closed-ended end. It includes one or more of the following.

[0098] Optionally, the second nucleic acid molecule or its amplicon is not ligated to its end with an adapter containing a protelomerase recognition sequence. In such embodiments, the second nucleic acid molecule is eliminated, for example, by exonuclease treatment between steps c and d. Selective adapter ligation to a specific nucleic acid molecule can be achieved by creating a specific end on the first nucleic acid molecule suitable for selective adapter ligation in step b, without creating such a specific end on the second nucleic acid molecule. For example, a specific sticky end can be created by a specific endonuclease capable of creating such a sticky end, such as, but not limited to, a type V CRISPR endonuclease such as Cpf1, in combination with a first crRNA targeting a sequence upstream of the first target sequence and a second crRNA targeting a sequence downstream of the first target sequence. In such embodiments, the adapter used in step b should include an overhang on the adapter side that is compatible with ligation to the sticky end thus created for ligation to the first nucleic acid molecule. Within this embodiment, the closed-ended first nucleic acid molecule can be opened in step d by cleavage at a specific sequence within the adapter, for example, when using an adapter containing a specific restriction enzyme recognition site between the ligation side and the protelomerase recognition sequence. Alternatively, the closed-ended first nucleic acid molecule can be opened in step d by cleavage at a sequence within the first nucleic acid molecule, such as the first target sequence.

[0099] Alternatively, in step b of the method of the present invention, an adapter is ligated to both the first and second nucleic acid molecules. Within such an embodiment, the closed-ended second nucleic acid molecule obtained in step c can be specifically removed from the reaction mixture containing the closed-ended first nucleic acid molecule prior to step d. This can be achieved by cleaving the closed-ended second nucleic acid molecule at a specific sequence, i.e., a second target sequence, that is not present in the closed-ended first nucleic acid molecule. Within such an embodiment, the second nucleic acid molecule of the method defined herein comprises a second target sequence that is not present in the first nucleic acid molecule. The subsequent open-ended second nucleic acid molecule can be removed by exonuclease treatment. Since the second nucleic acid molecule is not present at this stage, the closed-ended first nucleic acid can be opened in a specific or non-specific manner, for example, by cleaving at a sequence in the adapter as indicated herein above or at a sequence present in the first nucleic acid molecule. In methods where the closed-ended second nucleic acid molecule is not removed prior to step d, the closed-ended second nucleic acid molecule is still present in the reaction mixture containing the closed-ended first nucleic acid molecule in step d. In such designs, the first nucleic acid is preferably selectively opened by cleaving at a first target sequence that is not present in the second nucleic acid molecule. Such methods preferably include the following steps: a) providing a sample comprising at least a first and a second nucleic acid molecule, wherein the first nucleic acid molecule comprises a first target sequence that is not present in the second nucleic acid molecule, and optionally the second nucleic acid molecule comprises a second target sequence; b) ligating an adapter, as defined herein, i.e., comprising a protelomerase recognition sequence, to the ends of the first and second nucleic acid molecules to provide an adapter-ligated nucleic acid molecule; c) contacting the adaptor-ligated nucleic acid molecules with a protelomerase to cleave the adaptor-ligated nucleic acid molecules and covalently close the cleaved ends, resulting in first and second nucleic acid molecules comprising closed-ended ends; d) cleaving a first nucleic acid molecule comprising a closed-ended end at a first target sequence to provide a first nucleic acid comprising one open-ended end and one closed-ended end; Includes.

[0100] A preferred protelomerase is TeIN protelomerase.

[0101] It is understood herein that an effective number of components are used in the methods of the present invention. The nucleic acid molecule library prepared by the methods of the present invention is preferably suitable for further processing of the nucleic acid molecules, including, but not limited to, cloning, amplification, and sequencing. Thus, in a further aspect, the present invention also relates to a method for cloning a nucleic acid molecule library, a method for amplifying a nucleic acid molecule library, or a method for sequencing a nucleic acid molecule library using the steps described herein.

[0102] Preferably, the prepared nucleic acid molecule library is enriched for nucleic acid molecules containing a sequence of interest. "Enriched" is understood herein to mean reducing or eliminating nucleic acid molecules that do not have a sequence of interest by either (i) selectively excluding nucleic acid molecules that do not have a sequence of interest from further processing steps, or (ii) selectively including nucleic acid molecules that have a sequence of interest for further processing steps. The selectively excluded nucleic acid molecules can be degraded, for example, by exonuclease treatment. The selectively included nucleic acid molecules can be cloned, amplified, and / or sequenced, for example.

[0103] The nucleic acid library prepared preferably comprises nucleic acid molecules having one closed end and one open end.

[0104] In one embodiment, the method defined herein comprises the step a) of providing a sample comprising at least a first and a second nucleic acid molecule. Preferably, the first nucleic acid molecule comprises a first target sequence that is not present in the second nucleic acid molecule. Preferably, the second nucleic acid molecule comprises a second target sequence. Optionally, the second target sequence is also present in the first nucleic acid molecule. Alternatively, the second target sequence is not present in the first nucleic acid molecule.

[0105] Preferably, the first nucleic acid molecule comprises a sequence of interest and the second nucleic acid molecule does not comprise said sequence of interest. In this embodiment, the first nucleic acid molecule will be present in the prepared library of nucleic acid molecules, which will preferably be further processed.

[0106] In an alternative embodiment, the first nucleic acid molecule does not contain a sequence of interest, but the second nucleic acid molecule does contain said sequence of interest, in which embodiment the second nucleic acid molecule will be present in the prepared library of nucleic acid molecules, and preferably will be further processed.

[0107] The sample containing at least the first and second nucleic acid molecules may be from any source, for example, human, animal, plant, or microorganism, and may be endogenous or exogenous to the cell, such as genomic DNA, chromosomal DNA, artificial chromosome, plasmid DNA, episomal DNA, cDNA, RNA, mitochondrial, or an artificial library such as BAC or YAC. The DNA may be nuclear DNA or organelle DNA. Preferably, the DNA is chromosomal DNA, preferably endogenous to the cell. Preferably, the first, second, and optionally further nucleic acid molecules present in the sample used as the starting material for the method of the present invention are any one of DNA, such as genomic DNA, chromosomal DNA, organelle DNA, mitochondrial DNA, artificial chromosome, plasmid DNA, episomal DNA, cDNA, and RNA.

[0108] The first and second nucleic acid molecules can be long nucleic acid molecules prepared, for example, by cell lysis and, optionally, organelle lysis. Nucleic acid molecules used in the methods of the invention can have a size of at least about 50 kb, 100 kb, 150 kb, 200 kb, 300 kb, 400 kb, 500 kb, 600 kb, 700 kb, 800 kb, 900 kb, or at least about 1000 kb (1 Mb). The first and / or second nucleic acids for use in the invention can be high molecular weight (HMW) or ultra-high molecular weight (uHMW) nucleic acids. uHMW nucleic acids can have a length of at least 1 Mb. Nucleic acid molecules used in the methods of the invention may have a size of at least 1.1 Mb, 1.3 Mb, 1.5 Mb, 1.7 Mb, 2 Mb, 2.5 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb, or at least about 10 Mb.

[0109] Alternatively, a long nucleic acid molecule may first be fragmented to provide the first and second nucleic acid molecules. Thus, in one embodiment, the first and second nucleic acid molecules of step a) are provided by fragmentation. The fragmentation is preferably fragmentation of a genomic nucleic acid molecule.

[0110] Those skilled in the art are familiar with the means for fragmenting longer nucleic acid molecules, and the present invention is not limited to any specific means for fragmenting longer nucleic acid molecules.The fragmented nucleic acid is preferably fragmented genomic DNA.DNA, particularly genomic DNA, can be fragmented using any suitable method known in the art.Methods for DNA fragmentation include, but are not limited to, enzymatic digestion and mechanical force.

[0111] Non-limiting examples of fragmenting nucleic acid molecules using mechanical forces include acoustic shearing, nebulization, sonication, point-sink shearing, needle shearing, and the use of a French pressure cell.

[0112] Enzymatic digestion for fragmenting nucleic acid molecules, including at least one of the first and second nucleic acid molecules defined herein, includes, but is not limited to, endonuclease restriction. For example, enzymatic digestion, such as that used in AFLP® technology, can further reduce the complexity of the nucleic acid sample. Those skilled in the art will know which enzymes to select for DNA fragmentation. As a non-limiting example, at least one high-frequency cleaving agent and at least one low-frequency cleaving agent can be used to fragment a nucleic acid sample. High-frequency cleaving agents, such as, but not limited to, MseI, preferably have a recognition site of about 3-5 bp. Low-frequency cleaving agents, such as, but not limited to, EcoRI, preferably have a recognition site of >5 bp.

[0113] In certain embodiments, it may be preferable to use a third enzyme that is a low or high frequency cutter to obtain a larger set of shorter sized restriction fragments, particularly when the sample contains or is derived from a relatively large genome.

[0114] The methods of the present invention are not limited to any particular restriction endonuclease. The endonuclease can be a type II endonuclease such as EcoRI, Msel, or Pstl. In certain embodiments, type IIS or type III endonucleases, i.e., endonucleases whose recognition sequence is distal to the restriction site, can be used, such as, but not limited to, Acell, AIwI, AIwXI, Alw26I, BbvI, BbvII, Bbsl, Bed, Bce83I, Bcefl, Bcgl, Binl, Bsal, Bsgl, BsmAI, BsmFl, BspMI, EarI, EciI, Eco31, Eco57I, Esp3I, Faul, Fokl, Gsul, Hgal, HinGUII, Hphl, Ksp632I, MboII, Mmel, MnII, NgoVIII, PIeI, RIeAI, Sapl, SfaNI, TaqJI, and Zthll III. Depending on the endonuclease used, restriction fragments may be blunt-ended or have overhanging ends.

[0115] In a preferred embodiment, the recognition site of at least one of the high-frequency cleaving agent and the low-frequency cleaving agent is located within or in close proximity to the target sequence; for example, the recognition site of the high-frequency cleaving agent or the low-frequency cleaving agent is located approximately 0 to 10,000, 10 to 5,000, 50 to 1,000, or approximately 100 to 500 bases from the target sequence.

[0116] The methods disclosed herein can also be used in AFLP® technology, for example, in polyploid cells. AFLP® technology is described in more detail, for example, in WO 2007 / 114693, WO 2006 / 137733, and WO 2007 / 073165, which are incorporated herein by reference. AFLP® technology described in the art can be modified by attaching an adapter containing a protelomerase recognition sequence, as described herein, to the restricted nucleic acid sample.

[0117] Additionally or alternatively, programmable nucleases may be used to digest nucleic acid samples, preferably using at least one of CRISPR nucleases, zinc finger nucleases, TALENs, and meganucleases.

[0118] Optionally, the first and / or second nucleic acid molecule may be modified to include an A-tail, preferably to facilitate ligation to a partially or fully double-stranded adapter that includes a protelomerase recognition sequence and further includes a T-overhang. Thus, prior to annealing the adapter to the fragmented nucleic acids, the methods of the present invention may optionally include a step of A-tailing the fragmented nucleic acid sample. A-tailing reactions are well known in the art, and one of skill in the art will clearly understand how to perform an A-tailing reaction, for example, using the Klenow fragment (exo-).

[0119] A nucleic acid sample containing at least one of a first and a second nucleic acid molecule may contain a plurality of additional nucleic acid molecules. Thus, in some embodiments, the nucleic acid sample contains only the first nucleic acid molecule and only the second nucleic acid molecule. In other embodiments, the nucleic acid sample contains the first nucleic acid molecule, the second nucleic acid molecule, in addition to a plurality of other nucleic acid molecules. Preferably, the additional nucleic acid molecule does not contain the first target sequence. Optionally, the additional nucleic acid molecule does not contain the second target sequence. The plurality of other nucleic acid molecules may be derived from at least one of the same organism, tissue, cell, organelle, and / or molecule from which the first and second nucleic acid molecules are derived.

[0120] It is understood herein that a nucleic acid sample containing a first nucleic acid molecule can also include a nucleic acid sample containing multiple first nucleic acid molecules. Similarly, it is understood herein that a nucleic acid sample containing a second nucleic acid molecule can also include a nucleic acid sample containing multiple second nucleic acid molecules. Preferably, the first nucleic acid molecule is derived from the same organism, tissue, cell, organelle, and / or molecule as the second nucleic acid molecule. The first and second nucleic acid molecules can have essentially the same sequence, except for one or more nucleotides. As a non-limiting example, the first and second nucleic acid molecules can be allelic variants. Alternatively, the first and second nucleic acid molecules can be highly different, for example, having less than 40%, 30%, 20%, 10%, or 5% sequence identity. The primary difference between the first and second nucleic acid molecules used in the present invention is that the first nucleic acid molecule contains a target sequence that is not present in the second nucleic acid molecule.

[0121] Optionally, the second nucleic acid molecule can include a second target sequence, which may or may not be present in the first nucleic acid molecule.

[0122] In one embodiment, the method includes the step of b) ligating adapters to the ends of the first and second nucleic acid molecules to prepare adapter-linked nucleic acid molecules. The adapters are preferably as defined herein, i.e., adapters comprising a protelomerase recognition sequence. The adapters are preferably ligated to both ends of the first nucleic acid molecule and both ends of the second nucleic acid molecule. Preferably, adapters are ligated to both ends of at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acids present in the sample. Preferably, after the ligation step, all nucleic acid molecules in the sample contain adapters at both ends. In other words, preferably, all or substantially all nucleic acids in the sample are flanked on both sides by covalently linked adapters. Adapter ligation can be performed using any conventional method known to those skilled in the art, and the present invention is not limited to any particular ligation method or ligase. Preferably, to facilitate ligation, the adaptors contain ends that are compatible with the ends of the nucleic acid molecules obtained by, for example, using restriction endonucleases and compatible cohesive ends on the adaptors.

[0123] In one embodiment, the fragmented nucleic acid molecules may be end-treated to create blunt ends, followed by the addition of 3'A cohesive overhangs. The end-treatment step may be performed using any conventional means known in the art. Similarly, the addition of 3'A overhangs may be achieved using any conventional method known to those skilled in the art. The nucleic acid molecules containing 3'A overhangs may then be ligated to compatible adapters containing 5'T overhangs.

[0124] In one embodiment, the fragmentation step and the adapter ligation step can be combined into a single step, for example, by tagmentation. In this embodiment, the adapters in step b) are preferably ligated by tagmentation using Tn5 transposase. The transposase randomly cleaves long DNA molecules into short nucleic acid molecules, and adapters can be ligated on either side of the cleavage site. Tagmentation, or "transposase-mediated fragmentation and tagging," is a process well known to those skilled in the art, as exemplified, for example, by the Nextera™ workflow. The adapters may contain sequences that make them suitable for use in a tagmentation reaction. Preferably, the adapters used in the tagmentation reaction further contain a transposase sequence. The transposase sequence is preferably compatible with the transposase used in the tagmentation reaction. The tagmentation reaction can be followed by a repair step to ensure that all or substantially all generated nucleic acid molecules contain adapters on both sides. Thus, the nucleic acid molecules containing ligated adapters obtained by tagmentation can optionally be repaired to remove any single-strand breaks. Preferably, the repair step is carried out before contacting the molecule with TeIN protelomerase in step c) Such a repair step can be carried out using any conventional means known in the art.

[0125] Optionally, the protelomerase recognition sequence is attached to the nucleic acid molecule by a primer instead of an adapter. Preferably, the primer comprises: i) a 3' end for annealing to a primer binding site present on at least a first and / or second nucleic acid molecule or to an optionally universal primer binding site of an adapter linked to said at least a first and / or second nucleic acid molecule; and ii) a protelomerase recognition site in the 5' tail of such a primer Includes.

[0126] Optionally, the primer binding site is a unique sequence, i.e., a sequence present only in the first and / or second nucleic acid molecule. One or more of these primers can be used to introduce a protelomerase sequence into amplicons produced by PCR using the first and / or second nucleic acid molecule as templates. Within such embodiments, instead of ligating adapters, as defined herein, to the ends of the first and optional second nucleic acid molecules in step b) to provide adapter-linked nucleic acid molecules, the first and optional second nucleic acid molecules are amplified using at least one primer containing a protelomerase recognition site, and subsequent steps are then performed on the resulting amplicons, which can be protelomerized to close the ends. Alternatively, the protelomerase sequence can be introduced by a single step of denaturation, primer annealing, and base filling in single-stranded overhangs.

[0127] Thus, in those embodiments where adapters are attached by primer or tagmentation instead of ligation by (partially) double-stranded adapters, the terms "ligate" or "ligation" can be replaced with the terms "attach" or "attachment" as used herein.

[0128] In one embodiment, the method of the invention comprises step c) contacting the adaptor-ligated nucleic acid molecules with a protelomerase to cleave the adaptor-ligated nucleic acid molecules and covalently close the cleaved ends to provide first and second nucleic acid molecules comprising closed ends. A preferred protelomerase is TeIN protelomerase.

[0129] Preferably, the first nucleic acid molecule comprises an adaptor at both ends of the molecule (i.e., the 5' and 3' ends), and the second nucleic acid molecule comprises an adaptor at both ends of the molecule, the adaptor having a protelomerase recognition sequence. Contacting the adaptor-containing first and second molecules with protelomerase under appropriate conditions results in cleavage or "restriction" of the adaptor. Concurrently, the protelomerase can covalently close the nucleic acid molecules to produce a closed-ended first nucleic acid and a closed-ended second nucleic acid. Closed-ended linear DNA molecules typically contain covalently closed ends, providing protection from loss or damage of terminal nucleotides.

[0130] A preferred protelomerase for use in the present invention is a bacteriophage protelomerase. The protelomerase may be selected from the group consisting of phiHAP-1 from Halomonas aquamarina, PY54 from Yersinia enterolytica, phiKO2 from Klebsiella oxytoca, VP882 from Vibrio sp., and N15 from Escherichia coli, or a variant thereof. The protelomerase may have the amino acid sequence described in WO 2010 / 086626, which is incorporated herein by reference. The use of bacteriophage N15 (TeIN) protelomerase or a variant thereof is particularly preferred. Preferred protelomerases have a sequence with at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 11. Variants include homologs or mutants thereof. Variants include truncations, substitutions, or deletions relative to the native sequence. Variants preferably produce closed-ended linear DNA from templates containing the protelomerase recognition sequence described above.

[0131] The method may optionally comprise a step c1) of exposing the sample to an exonuclease after obtaining the nucleic acid molecule comprising a closed-end in step c) but before cleaving the first nucleic acid molecule comprising a closed-end in step d). Thus, in one embodiment, the method of the invention comprises: a) providing a sample comprising at least a first and a second nucleic acid molecule, wherein the first nucleic acid molecule comprises a first target sequence that is not present in the second nucleic acid molecule, and optionally the second nucleic acid molecule comprises a second target sequence; b) ligating an adapter, as defined herein, i.e., comprising a protelomerase recognition sequence, to the ends of the first and second nucleic acid molecules to provide an adapter-ligated nucleic acid molecule; c) contacting the adaptor-ligated nucleic acid molecules with a protelomerase to cleave the adaptor-ligated nucleic acid molecules and covalently close the cleaved ends, resulting in first and second nucleic acid molecules comprising closed-ended ends; c1) exposing a sample comprising first and second nucleic acid molecules comprising closed-end ends to an exonuclease; d) cleaving a first nucleic acid molecule comprising a closed-ended end at a first target sequence to provide a first nucleic acid comprising one open-ended end and one closed-ended end; Includes.

[0132] The exonuclease can digest any nucleic acid molecule that does not contain two closed-ended ends, i.e., that contains one or two open-ended ends, such as, but not limited to, a nucleic acid molecule with no adaptors, a nucleic acid molecule with one or two adaptors that has an open-ended end, and / or a truncated nucleic acid molecule with one open-ended end and one closed-ended end.

[0133] Nucleic acid molecules with two closed-end ends are protected from degradation, while unprotected fragments are degraded, resulting in enrichment or reduction in the complexity of nucleic acid molecules containing the sequence of interest, i.e., the first nucleic acid molecule or, optionally, the second nucleic acid molecule. Thus, in one embodiment, the method of the present invention takes an approach to removing undesired (non-target) portions of a nucleic acid sample. As a non-limiting example, the adapters in step b) may be ligated to nucleic acid molecules with selectively attached overhangs, created, for example, by enzymatic digestion. The adapter-containing molecules are then closed-ended in step c), and exonuclease treatment in step c1) can digest any nucleic acid molecules that do not have two closed-end ends. Thus, exonuclease treatment in step c1) can result in enrichment of nucleic acid molecules containing closed-end ends.

[0134] The exonuclease can be Exonuclease I, III, V, VII, VIII, or related enzymes, or any combination thereof. Exonuclease III recognizes the nick and extends it into the gap until a piece of ssDNA is formed. Exonuclease VII can degrade this ssDNA. Exonuclease I also degrades ssDNA. ExoIII and ExoVII are a preferred combination of exonucleases for use in step c) of the method of the present invention.

[0135] Exonuclease V is capable of degrading ssDNA and dsDNA in both the 3' to 5' and 5' to 3' directions. Thus, in a preferred embodiment, the exonuclease in step c) of the method of the present invention is an exonuclease capable of degrading ssDNA and dsDNA in both the 3' to 5' and 5' to 3' directions, preferably Exonuclease V.

[0136] Further information regarding methods for degrading non-target sequences is provided in U.S. Patent Application Publication No. 2014 / 0134610, which is incorporated by reference in its entirety for all purposes.

[0137] Step c1) is preferably carried out under conditions (e.g., time, temperature, enzyme concentration) sufficient for the exonuclease to degrade substantially all of the unprotected fragments. Preferably, step c1) is carried out under conditions and for a time sufficient for the exonuclease to degrade all of the unprotected fragments. Step c1) is preferably carried out for about 1 minute to about 12 hours, preferably 30 minutes, at about 10-90°C, preferably about 37°C.

[0138] After step c1), the exonuclease may be inactivated by, for example, but not limited to, at least one of the following: treatment with a proteinase, e.g., proteinase K, or heat inactivation. Such techniques are standard in the art, and those skilled in the art will readily understand how to inactivate exonucleases. A preferred inactivation step is heating the sample to a temperature of about 50-90°C, preferably about 75°C, for about 1-120 minutes, preferably about 10 minutes.

[0139] In one embodiment, the method of the present invention includes a step d) of cleaving a first nucleic acid molecule comprising a closed-ended end at a first target sequence to provide a first nucleic acid comprising one open-ended end and one closed-ended end. "Cleaving" is herein understood to mean generating a double-stranded break. The double-stranded break can be created by using a nuclease or by using two nickases that cleave opposite strands. The double-stranded break can create blunt, open-ended ends in the first nucleic acid molecule and, optionally, in the second nucleic acid molecule. Thus, after cleavage, the cleaved nucleic acid molecule can have one open-ended blunt end and one closed-ended end. Alternatively, the double-stranded break can create a cohesive, open-ended end in the cleaved nucleic acid molecule. Thus, after cleavage, the cleaved nucleic acid molecule can have one open-ended cohesive end and one closed-ended end.

[0140] Preferably, the first nucleic acid molecule in step d) is cleaved by a programmable nuclease or restriction endonuclease. Thus, the first nucleic acid molecule comprises a target sequence that is not present in the second nucleic acid molecule. The first nucleic acid molecule may comprise more than one target sequence, for example, the first nucleic acid molecule may comprise 1, 2, 3, 4, 5, 6 or more target sequences. In one embodiment, the second nucleic acid molecule may comprise a target sequence that is not present in the first nucleic acid molecule. The second nucleic acid molecule may comprise more than one target sequence, for example, the second nucleic acid molecule may comprise 1, 2, 3, 4, 5, 6 or more target sequences.

[0141] Those skilled in the art will readily appreciate that this step can be extended to additional nucleic acid molecules, e.g., at least a third, fourth, or fifth, or further nucleic acid molecules, each of which can optionally contain a target sequence that is not present in any of the other nucleic acid molecules.

[0142] Therefore, a nucleic acid sample is herein understood to contain at least one nucleic acid molecule containing a sequence of interest, i.e., a first nucleic acid molecule as defined herein or optionally a second nucleic acid molecule as defined herein. Therefore, in other words, a nucleic acid sample may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more sequences of interest, such as at least about 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000, or more sequences of interest, preferably, each sequence of interest in the sample has a different target sequence. The method of the present invention can provide simultaneous enrichment of such sequences of interest from a nucleic acid sample. Therefore, optionally, in step d) of the method of the present invention, multiple gRNA-CAS complexes are added to enrich nucleic acid molecules from a nucleic acid sample. Preferably, these multiple gRNA-CAS complexes may contain the same CRISPR-nuclease, but their gRNAs may be different. For example, a different gRNA molecule can be used for each nucleic acid molecule containing a sequence of interest. For example, at least about 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000 or more nucleic acid molecules, preferably at least about 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 750, 1000 or more gRNA molecules can be used in the methods of the present invention.

[0143] The first nucleic acid molecule and optionally the second nucleic acid molecule comprising closed-end ends can be cleaved by a restriction endonuclease. In embodiments in which the first and second nucleic acid molecules are cleaved, the first and second nucleic acid molecules are cleaved by different endonucleases. Any sequence-specific endonuclease may be suitable for use in the present invention. The endonuclease may be a so-called "restriction endonuclease" or "restriction enzyme," for example, a type I, type II, type III, type IV, or type V restriction endonuclease. Preferred restriction endonucleases are type II restriction endonucleases, preferably type IIP or type IIS. When fragmentation in step a) is performed by cleaving DNA with a restriction enzyme, the enzyme used in step d) is preferably a different endonuclease.

[0144] The first nucleic acid molecule and optionally the second nucleic acid molecule can be cleaved by a programmable nuclease. In embodiments in which the first and second nucleic acid molecules are cleaved, the first and second nucleic acid molecules are cleaved by different programmable nucleases, i.e., programmable nucleases that recognize different target sequences. The programmable nuclease can be selected from the group consisting of zinc finger nucleases, meganucleases, TAL effector nucleases, and RNA-guided CRISPR nucleases. Preferably, the programmable nuclease is an RNA-guided CRISPR (clustered regularly interspaced short palindromic repeats) nuclease.

[0145] The RNA-guided CRISPR-nuclease is preferably part of a gRNA-Cas complex. The gRNA-CAS complex is herein understood to be a CRISPR-associated (CAS) protein or CRISPR-nuclease complexed with a guide RNA. The CRISPR-nuclease comprises a nuclease domain and at least one domain that interacts with the guide RNA. When complexed with the guide RNA, the CRISPR-nuclease is guided to the target sequence by the guide RNA. When the guide RNA is directed to a site containing a specific target sequence via the guide sequence by interacting with the CRISPR-nuclease and the target sequence, the CRISPR-nuclease can introduce a cleavage in the target sequence. Preferably, the CRISPR-nuclease can introduce a single-strand or double-strand break in the target sequence when one or both domains of the nuclease are catalytically active, respectively. Those skilled in the art are well aware how to design guide RNAs such that, when combined with a CRISPR-nuclease, the introduction of a single-strand or double-strand break at a predetermined target site in a first nucleic acid molecule and / or optionally a second nucleic acid molecule is achieved.

[0146] CRISPR-nucleases can generally be classified into six major types (types I to VI), which are further divided into subtypes based on the content and sequence of their core elements (Makarova et al., 2011, Nat Rev Microbiol 9:467-77, and Wright et al., 2016, Cell 164(1-2):29-44). Generally, the two key components of the CRISPR-CAS system complex are the CRISPR-nuclease and the crRNA. The crRNA consists of short repeat sequences interspersed with spacer sequences derived from invader DNA. The CAS protein has various activities, such as nuclease activity. Thus, the gRNA-CAS complex provides a mechanism for targeting specific sequences as well as specific enzymatic activity against those sequences.

[0147] Type I CRISPR-CAS systems typically contain the Cas3 protein, which has separate helicase and DNase activities. For example, in Type I-E systems, the crRNA is incorporated into a multi-subunit effector complex called Cascade (a CRISPR-associated complex for antiviral defense) (Brouns et al., 2008, Science 321:960-4), which specifically binds to double-stranded DNA and induces its degradation by the Cas3 protein (Sinkunas et al., 2011, EMSO J 30:1335-1342; Beloglazova et al., 2011, EMBO J 30:616-627).

[0148] Type II CRISPR-CAS systems contain the signature Cas9 protein, a single protein (approximately 160 kDa) capable of specifically cleaving double-stranded DNA. Cas9 proteins typically contain two nuclease domains: a RuvC-like nuclease domain near the amino terminus and an HNH (or McrA-like) nuclease domain near the center of the protein. Each nuclease domain of the Cas9 protein is specialized for cleaving one strand of the double helix (Jinek et al., 2012, Science 337(6096):816-821). The Cas9 protein is an example of a CAS protein in a type II CRISPR / CAS system. When combined with a crRNA and a second RNA called a trans-activating crRNA (tracrRNA), the Cas9 protein forms an endonuclease that targets invading pathogen DNA and degrades it by introducing a DNA double-strand break (DSB) at a location in the pathogen genome defined by the crRNA. Jinek et al. (2012, Science 337:816-820) showed that single-stranded chimeric guide RNAs (sgRNAs), generated by fusing essential portions of crRNA and tracrRNA, can combine with the Cas9 protein to form a functional endonuclease.

[0149] Type III CRISPR-CAS systems contain a polymerase and a RAMP module. Type III systems can be further divided into subtypes III-A and III-B. Type III-A CRISPR-CAS systems have been shown to target plasmids, and the polymerase-like protein in type III-A systems is responsible for specific cleavage of DNA (Marraffini and Sontheimer, 2008, Science 322:1843-1845). Type III-B CRISPR-CAS systems have also been shown to target RNA (Hale et al., 2009, Cell 139:945-956).

[0150] Type IV CRISPR-CAS systems include Csf1, an uncharacterized protein that has been shown to form part of a Cascade-like complex, but these systems are often identified as isolated cas genes without associated CRISPR arrays.

[0151] A V-type CRISPR-CAS system, Clustered Regularly Interspaced Short Palindromic Repeats (C2c1) from Prevotella and Francisella 1 (CRISPR / Cpf1), has recently been described. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that targets DNA using crRNA. Cpf1 is a smaller and simpler endonuclease than Cas9, potentially overcoming some of the limitations of the CRISPR-Cas9 system. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA and utilizing a T-rich protospacer adjacent motif. Cpf1 cleaves DNA via staggered DNA double-strand breaks (Zetsche et al. (2015) Cell 163(3):759-771). A V-type CRISPR-CAS system preferably includes at least one of Cpf1, C2c1, and C2c3.

[0152] The Type VI CRISPR-CAS system can include a Cas13a protein containing RNase A activity. When the target nucleic acid fragment is RNA, at least the first and second gRNA-CAS complexes of the present invention can include Cas13a, such as, but not limited to, Cas13a derived from Leptotrichia wadee (LwCas13a) or Cas13a derived from Leptotrichia shahii (LshCas13a), as described in Gootenberg et al., Science. 2017 Apr 28;356(6336):438-442.

[0153] The gRNA-CAS complex of the method of the present invention can comprise any CRISPR-nuclease as defined herein above.Preferably, the gRNA-CAS complex used in the method of the present invention comprises type II CRISPR-nuclease, such as Cas9 (for example, the protein of SEQ ID NO: 12, or the protein of SEQ ID NO: 14, encoded by SEQ ID NO: 13), or type V CRISPR-nuclease, such as Cpf1 (for example, the protein of SEQ ID NO: 15, encoded by SEQ ID NO: 16), or Mad7 (for example, the protein of SEQ ID NO: 17 or 18), or proteins derived therefrom, and preferably have at least about 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity with said protein over its entire length.

[0154] Preferably, the gRNA-CAS complex of the method of the present invention comprises a type II CRISPR-nuclease, preferably a Cas9 nuclease.

[0155] Those skilled in the art know how to prepare the various components of the CRISPR-CAS system, including CRISPR-nucleases. Numerous reports on their design and use are available in the prior art. See, for example, the reviews by Haeussler et al. (J Genet Genomics. (2016) 43(5):239-50. doi:10.1016 / j.jgg.2016.04.008.) or Lee et al. (Plant Biotechnology Journal (2016) 14(2) 448-462) on the design of guide RNAs and their use in combination with CAS proteins (originally obtained from S. pyogenes).

[0156] Generally, CRISPR nucleases, such as Cas9, contain two catalytically active nuclease domains. For example, a Cas9 protein can contain a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC and HNH domains function together to both cleave single strands and create double-strand breaks in DNA (Jinek et al., Science, 337:816-821). A dead CRISPR nuclease contains modifications such that neither nuclease domain exhibits cleavage activity. The CRISPR nuclease of the gRNA-CAS complex used in the methods of the present invention can be a variant of a CRISPR nuclease in which one of the nuclease domains has been mutated so that it is no longer functional (i.e., has no nuclease activity), thereby generating a nickase. An example is an SpCas9 variant with either a D10A mutation or an H840A mutation. Preferably, the nuclease of the gRNA-CAS complex is not a dead nuclease. Preferably, the CRISPR-nuclease of the gRNA-CAS complex is either a nickase or an (endo)nuclease.

[0157] The gRNA-CAS complexes that may be used in the methods of the invention may comprise or consist of the entire Cas9 protein or variant, or may comprise fragments thereof, preferably such fragments that bind to the crRNA and tracrRNA or sgRNA and maintain at least one of nuclease or nickase activity.

[0158] Preferably, the gRNA-CAS complex comprises a Cas9 protein. Cas9 proteins are derived from the bacteria Streptococcus pyogenes (SpCas9; NCBI Reference Sequence NC_017053.1; UniProtKB-Q99ZW2), Geobacillus thermodenitrificans (UniProtKB-A0A178TEJ9), Corynebacterium ulcerous (NCBI Refs: NC_015683.1, NC_017317.1), Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1), and Spiroplasma silphidicola. syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisl (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); from Listeria innocua (NCBI Ref: NP_472073.1); from Campylobacter jejuni (NCBI Ref: YP_002344900.1); or from Neisseria meningitidis (NCBI Ref: YP_002342100.1). Cas9 variants from these, such as SpCas9_D10A or SpCas9_H840A, that have inactivated HNH or RuvC domains homologous to SpCas9, or Cas9s with equivalent substitutions at positions corresponding to D10 or H840 in the SpCas9 protein to generate nickases, are also included.

[0159] The programmable nuclease can be derived from Cpf1, for example, Acidaminococcus sp. UniProtKB-U2UMQ6. A variant can be a Cpf1 nickase with an inactivated RuvC or NUC domain, in which the RuvC or NUC domain no longer has nuclease activity. Those skilled in the art are familiar with techniques available in the art, such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, which allow for inactivated nucleases, such as inactivated RuvC or NUC domains. An example of a Cpf1 nickase with an inactive NUC domain is Cpf1R1226A (see Gao et al., Cell Research (2016) 26:901-913; Yamano et al., Cell (2016) 165(4):949-962). In this variant, an arginine is converted to an alanine (R1226A) in the NUC domain, rendering the NUC domain inactive.

[0160] The gRNA-CAS complex further comprises a CRISPR-nuclease-associated guide RNA that directs the complex to a target sequence or "target site" in the nucleic acid molecule, also annotated as a protospacer sequence. The guide RNA is preferably located near, at, or within a sequence of interest within the nucleic acid molecule and comprises a guide sequence for targeting the gRNA-CAS complex to the protospacer sequence, which may be an sgRNA, or a combination of crRNA and tracrRNA (e.g., in the case of Cas9), or crRNA alone (e.g., in the case of Cpfl). Optionally, more than one type of guide RNA can be used in the same experiment, e.g., to target two or more different nucleic acid molecules of interest, or even the same nucleic acid molecule of interest.

[0161] In an optional embodiment, the method of the present invention is for polymorphism detection and / or genetic mutation detection by using an enzyme that recognizes and cleaves heteroduplexes at mismatched sites. Within such an embodiment, one or more nucleotide samples are fragmented and then subjected to at least one round of denaturation and annealing before or after step b) of the method of the present invention. Then, after step c) of the method of the present invention, the closed-ended nucleic acids can be treated with an enzyme that recognizes and cleaves heteroduplexes, such as CEL I, or with an enzyme described in Langhans MT and Palladino MJ (Curr Istumes Mol Biol. 2009;11(1):1-12), incorporated herein by reference. This results in open ends of only double-stranded DNA molecules containing heteroduplexes, which can then be selectively included for further processing (e.g., by ligating sequencing adapters to these open ends and subsequent sequencing) or selectively excluded for further processing (e.g., by degrading these fragments by exonuclease treatment).

[0162] In one embodiment, the method may include step e) of exposing the sample to an exonuclease after obtaining a first nucleic acid molecule comprising one open end and one closed end in step d). Thus, in this embodiment, the first nucleic acid comprises an open end and the second nucleic acid comprises two closed ends. Thus, the second nucleic acid molecule will be protected from exonuclease degradation, but the first nucleic acid molecule will not. Thus, exposure to exonuclease results in digestion of the first nucleic acid, but not the second nucleic acid. In this embodiment, the second nucleic acid molecule preferably comprises a sequence of interest.

[0163] The exonuclease may optionally be an exonuclease as defined herein in step c1), under the same or similar conditions as defined herein in step c1). Preferably, the exonuclease digestion results in digestion of all or substantially all nucleic acid molecules comprising at least one open end. Thus, in this embodiment, the method of the present invention comprises: a) providing a sample comprising at least a first and a second nucleic acid molecule, wherein the first nucleic acid molecule comprises a first target sequence that is not present in the second nucleic acid molecule, and optionally the second nucleic acid molecule comprises a second target sequence; b) ligating an adapter, as defined herein, i.e., comprising a protelomerase recognition sequence, to the ends of the first and second nucleic acid molecules to provide an adapter-ligated nucleic acid molecule; c) contacting the adaptor-ligated nucleic acid molecules with a protelomerase to cleave the adaptor-ligated nucleic acid molecules and covalently close the cleaved ends, resulting in first and second nucleic acid molecules comprising closed-ended ends; d) cleaving a first nucleic acid molecule comprising a closed-ended end at a first target sequence to provide a first nucleic acid comprising one open-ended end and one closed-ended end; e) exposing the sample to an exonuclease; may include:

[0164] Optionally, the method may further comprise, between steps c) and d), step c1) as described herein above.

[0165] Optionally, step e) may comprise step e1) of removing and / or inactivating restriction endonucleases and / or programmable nucleases, followed by step e2) of exposing the sample to an exonuclease.

[0166] Step e1) may include heating the sample to an appropriate temperature to remove and / or inactivate restriction endonucleases and / or programmable nucleases. By way of non-limiting example, the temperature may be raised to at least 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C or more. The temperature may be raised for a period of at least about 5', 10', 15', 20', 25', 30', 35', 40', 45', 50', 55', 60' (minutes) or more.

[0167] Alternatively, or in addition, step e1) may comprise purifying the cleaved first nucleic acid molecule. Purification of the cleaved first nucleic acid molecule can be carried out using any conventional means, such as the AMPure bead-based purification process and / or partial or complete digestion with a proteinase, such as, but not limited to, digestion with proteinase K, restriction endonuclease and / or programmable nuclease.

[0168] The second nucleic acid molecule comprising two closed-ended ends may then be cleaved at the target sequence. Thus, the method of the present invention may further comprise step f) of cleaving the second nucleic acid molecule comprising closed-ended ends at the second target sequence to result in a second nucleic acid comprising one open-ended end and one closed-ended end. The target sequence of the second nucleic acid molecule is preferably not present in the first nucleic acid molecule. However, within this embodiment, the first nucleic acid molecule has already been removed by the time cleavage of the second nucleic acid molecule is carried out, so optionally the target sequence of the second nucleic acid molecule is also present in the first nucleic acid molecule. Preferably, the method of the present invention comprises: a) providing a sample comprising at least a first and a second nucleic acid molecule, wherein the first nucleic acid molecule comprises a first target sequence that is not present in the second nucleic acid molecule, and optionally the second nucleic acid molecule comprises a second target sequence; b) ligating an adapter, as defined herein, i.e., comprising a protelomerase recognition sequence, to the ends of the first and second nucleic acid molecules to provide an adapter-ligated nucleic acid molecule; c) contacting the adaptor-ligated nucleic acid molecules with a protelomerase to cleave the adaptor-ligated nucleic acid molecules and covalently close the cleaved ends, resulting in first and second nucleic acid molecules comprising closed-ended ends; d) cleaving a first nucleic acid molecule comprising a closed-ended end at a first target sequence to provide a first nucleic acid comprising one open-ended end and one closed-ended end; e) exposing the sample to an exonuclease; f) cleaving the second nucleic acid molecule comprising the closed-ended end at the second target sequence to result in a second nucleic acid comprising one open-ended end and one closed-ended end; may include:

[0169] Optionally, the method may further comprise, between steps c) and d), step c1) as described herein above.

[0170] Preferably, the second nucleic acid molecule in step f) is cleaved by a programmable nuclease or restriction endonuclease, preferably the restriction endonuclease defined in step d) or the programmable nuclease defined in step d). Preferably, the second nucleic acid molecule in step f) can be digested using a programmable nuclease, preferably at least one of CRISPR nuclease, zinc finger nuclease, TALEN, and meganuclease. Preferably, the second nucleic acid molecule is digested by an RNA-guided CRISPR nuclease. The CRISPR nucleases used to cleave the first and second nucleic acid molecules can be the same or different. If the CRISPR nucleases used to cleave the first and second nucleic acid molecules are the same, the guide RNA sequences bound to the CRISPR nucleases are not the same. In other words, when a CRISPR nuclease is used to cleave a first and a second nucleic acid molecule, it is understood herein that the gRNA-Cas complex that recognizes and cleaves the first nucleic acid molecule is a different gRNA-Cas complex that recognizes and cleaves the second nucleic acid molecule.

[0171] The method may further include step g) ligating an additional (or "further") adaptor to the open end of at least one of the first and second nucleic acid molecules, wherein the first and second nucleic acid molecules comprise one open end and one closed end.

[0172] Thus, in one embodiment, the method may include steps a), b), c), d), and g). Optionally, the method may include steps a), b), c), c1), d), and g). In this embodiment, an additional adaptor is ligated to the open end of the first nucleic acid molecule. The first nucleic acid molecule preferably contains a sequence of interest.

[0173] In another embodiment, the method may include steps a), b), c), d), e), f), and g). Optionally, the method may include steps a), b), c), c1), d), e), and g). In this embodiment, an additional adaptor is ligated to the open end of a second nucleic acid molecule. The second nucleic acid molecule preferably contains a sequence of interest.

[0174] The additional adapter may be an adapter suitable for amplification and / or sequencing. The additional adapter may be a sequencing adapter, for example, including a functional domain that enables Roche 454A and 454B sequencing, ILLUMINA™ SOLEXA™ sequencing, Applied Biosystems' SOLID™ sequencing, Pacific Biosciences' SMRT™ sequencing, Pollonator Polony sequencing, Oxford Nanopore Technologies (ONT), Ontera sequencing, or Complete Genomics sequencing.

[0175] Thus, preferably, the additional adapter comprises at least one sequencing primer binding site, and / or the additional adapter comprises at least one amplification primer binding site. The additional adapter may comprise at least two sequencing primer binding sites, and / or the further adapter may comprise at least two amplification primer binding sites. The additional adapter may be a single-stranded, double-stranded, partially double-stranded, Y-shaped, or hairpin nucleic acid molecule. Preferably, the adapter is a hairpin adapter or a Y-shaped adapter.

[0176] Although stem-loop or hairpin adapters are single-stranded, their ends are complementary, resulting in the adapter folding back on itself, creating a double-stranded portion and a single-stranded loop. Stem-loop adapters can be ligated to the ends of linear double-stranded nucleic acid molecules. For example, when a stem-loop adapter is ligated to the open end of a corresponding first or second nucleic acid molecule in step g), there is no terminal nucleotide. Thus, the resulting molecule lacks a terminal nucleotide.

[0177] The first or second nucleic acid molecule of step g) can be ligated to a circularizable adaptor. In this regard, nucleic acid molecules containing open ends can be circularized by self-circularization of compatible structures on either side of the fragment (which can occur by adaptor ligation or as a result of restriction enzyme digestion of the ligated adaptor), or by hybridization to a selector probe complementary to the ends of the desired fragment. The final step of extension and ligation generates a covalently closed, circular, optionally double-stranded polynucleotide.

[0178] The additional adapter may be a protection adapter. In this context, a protection adapter is herein understood to be an adapter specifically designed to protect the nucleic acid molecule captured by the adapter from exonuclease digestion. Such adapters preferably protect against exonuclease degradation either by including a chemical moiety or a protecting group (e.g., phosphorothioate) or by the lack of a terminal nucleotide (hairpin or stem-loop adapters, or cyclizable adapters).

[0179] Optionally, the additional adaptor comprises an identifier sequence, preferably an identifier sequence as defined herein.

[0180] Preferably, the nucleic acid molecule library is prepared from multiple samples. Optionally, the method of the present invention is multiplexed, i.e., applied simultaneously to multiple nucleic acid samples, for example, at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 500, 1000 or more nucleic acid samples. Thus, the method can be performed in parallel on multiple samples, where "in parallel" should be understood herein to mean substantially simultaneously, but with each sample being processed in a separate reaction tube or vessel.

[0181] Additionally or alternatively, one or more steps of the method of the present invention can be performed on pooled samples. To trace the first and / or second nucleic acid molecules back to their original samples, the first and / or second nucleic acid molecules can be tagged with an identifier before pooling the samples. Such identifiers can be any detectable entity, including but not limited to radioactive or fluorescent labels, but are preferably specific nucleotide sequences or combinations of nucleotide sequences, preferably of defined length. Additionally or alternatively, samples can be pooled using inventive pooling strategies, including but not limited to 2D and 3D pooling strategies, whereby, after pooling, each sample is contained in at least two or three pools, respectively. A specific nucleic acid molecule can be traced back to its original sample by using the coordinates of the corresponding pool containing the first and / or second nucleic acid molecule. Multiple samples can be pooled before step b), step c), step d), step e), step f), or before step g), or after step g).

[0182] Between steps a) and b), between steps b) and c), between steps c) and d), and / or after step d) described herein, the nucleic acid sample may be purified and / or reaction enzymes may be inactivated.

[0183] In embodiments of the invention, the nucleic acid sample may be purified and / or the reaction enzymes may be inactivated between steps c) and c1) and / or between steps c1) and d) described herein.

[0184] In embodiments of the invention, the nucleic acid sample may be purified and / or reaction enzymes may be inactivated between steps d) and e), between e) and f), between f) and g), between d) and g), and / or after step g) described herein.

[0185] A purification step, e.g., an AMPure bead-based purification process, may be included to remove complexes, enzymes, free nucleotides, possible free adapters, and possible small unrelated nucleic acid molecules. After purification, the first nucleic acid molecule and / or optionally the second nucleic acid molecule can be recovered and subjected to further processing and / or analysis, such as single molecule sequencing.

[0186] An optional purification step is proteinase K treatment. Alternatively, or in addition, the purification may comprise the following steps: i. exposing the nucleic acid sample to one or more solid supports that specifically and effectively bind a first nucleic acid molecule and / or optionally a second nucleic acid molecule; and optionally, ii. washing the one or more solid supports and eluting the first nucleic acid molecule and / or optionally the second nucleic acid molecule from the one or more solid supports; may include:

[0187] The one or more solid supports can be, but are not limited to, Ampure beads. Since at least one isolated nucleic acid molecule is obtained after purification, the method defined herein can also be considered to be a method for isolating one or more nucleic acid molecules from a nucleic acid sample.

[0188] The method of the present invention may further comprise a size selection step, optionally performed before step b), between steps b) and c), between steps c) and d), and / or after step d) of the method of the present invention.

[0189] In one embodiment, a size selection step is performed between steps c) and c1) and / or between steps c1) and d) of the present invention.

[0190] In one embodiment, a size selection step is performed between steps d) and e), between steps e) and f), between steps f) and g), or after step g) of the present invention.

[0191] Alternatively, there are no further purification, inactivation, and / or size selection steps. Thus, in one embodiment, the method of the present invention does not require any purification steps between steps a), b), c), d), e), f), and g) or after step g). Additionally, or alternatively, the method of the present invention does not require any inactivation steps between steps a), b), c), d), e), f), and g) or after step g). Additionally, or alternatively, the method of the present invention does not require any size selection steps between steps a), b), c), d), e), f), and g) or after step g).

[0192] The method of the present invention may be followed by a step of sequencing one or more target nucleic acid molecules. The method defined herein may therefore also be considered as a method for sequencing one or more target nucleic acid molecules from a nucleic acid sample.

[0193] Preferably, the sequencing step is performed after addition of an adapter comprising a protelomerase recognition sequence. Preferably, the sequencing step is performed after step c), i.e., after sequencing of the circular nucleic acid molecule. Preferably, the sequencing step is performed after addition of a further adapter. Preferably, the sequencing step is performed after step g). Sequencing of at least one of the first and second nucleic acid molecules can be performed after step b), after step c1), after step d), after step e), or after step f).

[0194] Optionally, the method of the present invention further comprises an amplification step. The amplification step can be performed after closing the adaptor-containing nucleic acid molecule, where the adaptor comprises a protelomerase recognition sequence. Preferably, the amplification step is performed after step c), i.e., after amplification of the circular nucleic acid molecule. Optionally, the amplification step is performed after annealing an additional adaptor to the first or second nucleic acid molecule. Preferably, the amplification step is performed after step g). Amplification of at least one of the first and second nucleic acid molecules can be performed after step a), step b), step c1), step d), step e), and / or step f). Amplification can be performed by PCR or any amplification method known in the art.

[0195] In one embodiment, the method of the present invention is a sequencing method that does not include an amplification step and / or a cloning step. The reduction of the amplification step is beneficial because epigenetic information (e.g., 5-mC, 6-mA, etc.) is lost in the amplicon. Further amplification can introduce diversity into the amplicon (e.g., through errors during amplification), resulting in nucleotide sequences that do not reflect the original sample. Similarly, cloning a target region into another organism often does not maintain modifications present in the original sample nucleic acid, so in some embodiments, target sequences enriched for further analysis are typically not amplified and / or cloned in the methods herein.

[0196] In one aspect, the method of the present invention relates to a method for amplifying a library of nucleic acid molecules, which method preferably comprises the step of preparing a library of nucleic acid molecules as defined herein. The library of nucleic acid molecules preferably comprises: Steps a), b), and c) as defined herein; Steps a), b), c), and c1 as defined herein; steps a), b), c), and d) as defined herein; Steps a), b), c), c1), and d as defined herein; steps a), b), c), d), and g as defined herein; Steps a), b), c), c1), d), and g as defined herein; steps a), b), c), d), and e) as defined herein; Steps a), b), c), c1), d), and e as defined herein; Steps a), b), c), d), e), and f) as defined herein; Steps a), b), c), c1), d), e), and f) as defined herein; steps a), b), c), d), e), f), and g as defined herein; Steps a), b), c), c1), d), e), f), and g) as defined herein It is prepared using at least one of the following:

[0197] The method further comprises amplifying the library of nucleic acid molecules. Amplification can be carried out using a single primer, for example by "rolling circle" amplification. The single primer preferably comprises: i) a primer that anneals to the first nucleic acid molecule comprising one open end and one closed end obtained in step d); ii) a primer that anneals to the second nucleic acid molecule obtained in step f) that comprises one open end and one closed end; and iii) a primer annealing to the further adapter defined in step g). At least one of the following is true:

[0198] Alternatively, or in addition, amplification can be carried out using a primer pair, i.e., using a first and a second primer, preferably the first and second primers being capable of annealing to a first nucleic acid molecule and / or the first and second primers being capable of annealing to a second nucleic acid molecule, so as to allow amplification of the corresponding first and / or second nucleic acid molecule.

[0199] Preferably, the primer pair comprises a first primer and a second primer that can anneal to a first nucleic acid molecule, preferably to a first nucleic acid molecule obtained in step a), b), c), c1), d) or step g) as defined herein. Preferably, the primer pair comprises a first primer and a second primer that can anneal to a first nucleic acid molecule comprising one open end and one open end obtained in step d) or step g) as defined herein.

[0200] Alternatively, or in addition, the primer pair may comprise a first primer and a second primer capable of annealing to a second nucleic acid molecule, preferably to a second nucleic acid molecule obtained in step a), b), c), c1), d), e), f) or step g) as defined herein. Preferably, the primer pair comprises a first primer and a second primer capable of annealing to a second nucleic acid molecule comprising one open end and one open end obtained in step f) or step g) as defined herein.

[0201] Preferably, the first primer of the primer pair is not complementary or substantially not complementary to the second primer of the primer pair.

[0202] In one embodiment, at least one of the first and second primers may anneal to an adapter, preferably an adapter comprising a protelomerase recognition sequence as defined herein and / or to a sequence present in a further adapter as defined in step g).

[0203] The first and second primers can anneal to the same adapter, preferably to the first and second sequences present in the adapter of step g) defined herein. As a non-limiting example, the adapter can be a Y-shaped adapter, and the first primer binding site can be present in a first single-stranded arm of the Y-shaped adapter and the second primer binding site can be present in the other single-stranded arm of the Y-shaped adapter.

[0204] Alternatively, or in addition, a first amplification primer can anneal to a sequence present in the first nucleic acid molecule and a second amplification primer can anneal to an adapter, preferably an adapter comprising a protelomerase recognition sequence or to a sequence present in a further adapter of step g) as defined herein.

[0205] Alternatively, or in addition, the first amplification primer can anneal to a sequence present in the second nucleic acid molecule and the second amplification primer can anneal to an adapter, preferably to an adapter comprising a protelomerase recognition sequence or to a sequence present in a further adapter of step g) as defined herein.

[0206] Alternatively, or in addition, a first amplification primer can anneal to a sequence present in an adaptor comprising a protelomerase recognition sequence, and a second amplification primer can anneal to a sequence present in a further adaptor of step g) as defined herein.

[0207] In a further aspect, the present invention relates to a method for analyzing a sequence of interest in a sample comprising a first and a second nucleic acid molecule, the method preferably comprising the step of preparing a library of nucleic acid molecules as defined herein.

[0208] The sample may contain at least a first and a second nucleic acid molecule. The first and / or second nucleic acid molecule may be part of a longer nucleic acid molecule. The nucleic acid sample may contain multiple nucleic acid molecules, including the first and second nucleic acid molecules.

[0209] As described in detail herein, the prepared nucleic acid library preferably comprises at least one of a first and a second nucleic acid molecule. In one embodiment, the prepared nucleic acid library comprises the first nucleic acid molecule but does not comprise the second nucleic acid molecule. In an alternative embodiment, the prepared nucleic acid library comprises the second nucleic acid molecule but does not comprise the first nucleic acid molecule.

[0210] The first or second nucleic acid molecule preferably comprises a sequence of interest. The library of nucleic acid molecules preferably comprises: Steps a), b), and c) as defined herein; Steps a), b), c), and c1 as defined herein; steps a), b), c), d), and g as defined herein; Steps a), b), c), c1), d), and g as defined herein; Steps a), b), c), d), e), f), and g) as defined herein; and Steps a), b), c), c1), d), e), f), and g) as defined herein It is prepared using at least one of the following:

[0211] The method preferably further comprises the step of analyzing the prepared library of nucleic acid molecules. The analysis can be carried out using any conventional means known in the art. detecting the sequence using a label, such as a radioactive or fluorescent label; Analyzing the size of the prepared nucleic acid molecule library; cloning the library, optionally in part, into a vector, optionally followed by gene expression and / or restriction analysis; and Sequencing a library of nucleic acid molecules may include at least one of:

[0212] Preferably, the prepared nucleic acid molecule library is sequenced, preferably deep sequenced.Sequencing can include at least one of ILLUMINA™, SOLEXA™ sequencing, Ion Torrent sequencing, Pacific Biosciences' SMRT™ sequencing, Sanger sequencing, Genapsys, Pollonator Polony sequencing, Oxford Nanopore Technologies (ONT), Ontera sequencing and Complete Genomics sequencing.

[0213] In a preferred embodiment, the prepared nucleic acid molecule library is sequenced by nanopore selective sequencing. In nanopore selective sequencing, during real-time sequencing, the generated data (direct current signal or base call converted from such current signal) is compared with one or more reference sequences. If a set number of nucleotides or amount of signal of the target sequence is aligned with the reference sequence, sequencing proceeds; otherwise, the current is reversed, thereby removing nucleic acid from the pore, making the pore available for sequencing new nucleic acid. The set number of nucleotides can be at least the first 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of the read nucleic acid. The one or more reference sequences can be a number of different sequences. Preferably, each of these reference sequences is at least 50, 60, 70, 80, 90, 92, 93, 94, 95, 96, 97, 98, 99, or 100% identical to the sequences of the target nucleic acid fragments in the nucleic acid molecule library obtained by the method of the present invention. In one embodiment, each of the reference sequences is at least 50, 60, 70, 80, 90, 92, 93, 94, 95, 96, 97, 98, 99, or 100% identical to a specific subset of the sequences of one or more of the target nucleic acid fragments in the nucleic acid molecule library obtained by the method of the present invention. One advantage of selectively sequencing a specific subset by nanopore selective sequencing is that the prepared nucleic acid molecule library can be used to sequence different subsets in different sequencing runs.

[0214] In one embodiment, the adapter comprising a protelomerase recognition sequence comprises at least one binding site for a sequencing primer.

[0215] Alternatively, or in addition, the further adapter in step g) comprises at least one binding site for a sequencing primer. The further adapter in step g) may comprise two different binding sites for two sequencing primers. As a non-limiting example, the adapter in step g) may be a Y-shaped adapter, and the first sequencing primer binding site may be present on a first single-stranded arm of the Y-shaped adapter and the second sequencing primer binding site may be present on the other single-stranded arm of the Y-shaped adapter.

[0216] In one aspect, the present invention relates to a method for enriching a nucleic acid sample for nucleic acid molecules comprising a sequence of interest, which preferably uses at least method steps a) to d) detailed herein above, but may also use any of the additional steps detailed herein, such as step c1), step e), step f), and / or step g).

[0217] In one aspect, the present invention relates to a kit of parts for carrying out the methods of the invention as described herein. Preferably, the kit of parts is for use in the methods defined herein. Preferably, the kit of parts comprises at least one or more adapters comprising a protelomerase recognition sequence as defined herein.

[0218] An adaptor for use in the methods defined herein preferably does not comprise a recognition site for a restriction endonuclease or programmable nuclease used in step d) and / or step f) of the methods of the invention. More preferably, the portion of the adaptor located between the protelomerase recognition sequence and the end linked to the first and / or second nucleic acid molecule does not comprise a recognition site for a restriction endonuclease or programmable nuclease used in step d) and / or step f) of the methods of the invention.

[0219] The one or more adapters may be combined in one vial or may be present in separate vials, e.g., the adapters in one vial contain the same identifier sequence, preferably the same sample identifier sequence. The kit-of-parts may further comprise a vial containing a protelomerase as defined herein.

[0220] The kit of parts may include one or more reagents for carrying out the methods described herein. Thus, the kit of parts may include: one or more vials containing an adaptor comprising a protelomerase recognition sequence as defined herein; one or more vials containing a further adapter as defined herein for step g); one or more vials containing a protelomerase as defined herein; one or more vials containing a gRNA-CAS complex as defined herein; one or more vials containing a gRNA for complexing with a CRISPR-CAS protein to form a gRNA-CAS complex, and a further vial containing said CRISPR-CAS protein; Additional vials containing one or more exonucleases may include at least one of:

[0221] Preferably, the kit comprises at least 2, 4, 10, 20, 30, or 50 vials containing one or more gRNAs as defined herein. Preferably, the volume of any vial in the kit does not exceed 100 mL, 50 mL, 20 mL, 10 mL, 5 mL, 4 mL, 3 mL, 2 mL, or 1 mL.

[0222] The reagents may be present in lyophilized form or in a suitable buffer. The kit may also include any other components necessary to practice the invention, such as buffers, pipettes, microtiter plates, and written instructions. Such other components for kits of the invention are known to those of skill in the art.

[0223] In one aspect, the present invention provides a method for producing a pharmaceutical composition comprising: (i) Preparation of a nucleic acid molecule library; (ii) amplification of a library of nucleic acid molecules; and (iii) Analysis of the sequence of interest in the sample. The present invention relates to the use of an adapter comprising a protelomerase recognition sequence as defined herein for at least one of:

[0224] [Table 1] [Example]

[0225] Materials and Methods The adapter containing the TeIN recognition site was Oligo 19_04626 (100 μM): 2 μl Oligo 19_03053 (100 μM): 2 μl was prepared by combining

[0226] Oligo sequences: 19_04626 5'-AGGACCGGATCAACTTATCAGCACACAATTGCCCATTATACGCGCGTATAATGGACTATTGTGTGCTGATAAAGAAAGTTGTCGGTGTCTTTGTGAGATGTGTATAAGAGACAGT-3' (SEQ ID NO: 19) 19_03053 5'-CTGTCTCTTATACACATCTCACAAAGACACCGACAACTTTCTTTATCAGCACACAATAGTCCATTATACGCGCGTATAATGGGCAATTGTGTGCTGATAAGTTGATCCGGTCCT-3' (SEQ ID NO: 20). The 5' end is preferably phosphorylated.

[0227] To allow hybridization of the oligos, the following thermal profile was used: 95°C for 10 minutes 90°C for 1 minute The temperature is decreased 60 times at 1°C / cycle. Maintain 4℃ The resulting adaptor solution (50 μM) was diluted to a concentration of 15 μM.

[0228] The input material for this example was a 1 Kbp amplicon derived from lambda DNA. Amplification was performed using the following settings: Lambda DNA 5ng / μl 5μl MilliQ water 9.3 μl PCR buffer 4 μl 25mM dNTP (each) 0.2μl Herculase polymerase 0.5 μl Forward primer (10 μM) 0.5 μl Reverse primer (10 μM) 0.5 μl Forward primer: 18_03029: 5'-TCACGCTGATTTACAGCGGCA-3' (SEQ ID NO: 21) Reverse primer: 18_03032: 5'-CGATGCTGATTGCCGTTCCG-3' (SEQ ID NO: 22)

[0229] The thermal profile for amplification was as follows: 95°C for 2 minutes 95°C for 30 seconds 65℃ for 30 seconds -> Reduce temperature by 0.7℃ / cycle 72°C for 4 minutes 13 cycles 95°C for 30 seconds 56°C for 30 seconds 72°C for 5 minutes 25 cycles 72°C for 2 minutes Maintain at 12°C

[0230] The resulting amplicon was purified 0.8x and eluted in 20 ul MQ. The concentration was measured by QubitBR: 554 ng / ul. The purified amplicons were end-repaired and A-tailed.

[0231] End repair (two reactions performed): 2 μl of purified amplicon 7 μl NEBNext Ultra II End Prep Reaction Buffer (New England Biolabs Inc.) 3 μl of NEBNext Ultra II End Prep enzyme mix (New England Biolabs Inc.) 48 μl MilliQ water Total volume = 60 μl -> Incubate at 20°C for 30 min, 65°C for 30 min and keep at 4°C until further use.

[0232] Adapter Connection: 60 μl of NEBNext Ultra II End Prep reaction mixture (New England Biolabs Inc.) 30 μl of NEBNext Ultra II ligation master mix (New England Biolabs Inc.) 1 μl of NEBNext ligated enhancer (New England Biolabs Inc.) 2.5 μl adapter (50 μM) Total volume = 93.5 μl -> Incubate at 15°C for 20 minutes The resulting ligated sample was purified using 1:1 Ampure beads and eluted in 20 μl of MilliQ water.

[0233] An additional Ampure purification (0.75x) was performed to remove residual adapters.

[0234] The concentration of the adaptor ligated product is 40 ng / μl.

[0235] The adaptor ligation products were treated with TeIN to covalently close the ends. Adapter ligation product 4 μl ThermoPol reaction buffer (10x) (New England Biolabs Inc.) 2 μl TeIN protelomerase (New England Biolabs Inc.) 2 μl MilliQ water 12 μl

[0236] The reaction mixture was mixed gently by pipetting, briefly centrifuged, and incubated for 30 minutes at 30° C. The enzyme was inactivated by incubation at 75° C. for 5 minutes.

[0237] The resulting sample was purified using 1:1 Ampure beads and eluted with 15 μl of MilliQ water.

[0238] To verify exonuclease protection, TeIN-treated samples were incubated with exonuclease V. 10 μl sample NEB buffer 3.1 (10x) 2.0 μl ATP (100 mM) 1.0 μl Exonuclease V (10 units) 1.0 μl MilliQ water 6.0 μl The reaction mixture was incubated for 60 minutes at 37° C. The exonuclease was inactivated at 70° C. for 30 minutes.

[0239] Samples were purified using Ampure (1x) and eluted in 10 ul of MilliQ water.

[0240] result The results of the bioanalyzer analysis are shown in Figure 1. Briefly, Amplicons and adaptor-ligated amplicons are readily degraded using exonuclease V.

[0241] Adapter-ligated and TeIN-treated amplicons are resistant to exonuclease degradation.

[0242] conclusion Covalently closing the ends of the DNA fragments using TeIN results in exonuclease V resistant fragments.

Claims

1. 1. A method for preparing a library of nucleic acid molecules, comprising: a) providing a sample comprising at least a first and a second nucleic acid molecule, wherein the first nucleic acid molecule comprises a first target sequence that is not present in the second nucleic acid molecule; b) ligating an adaptor to the ends of the first and second nucleic acid molecules to provide an adaptor-ligated nucleic acid molecule, wherein the adaptor is at least partially double-stranded and comprises a double-stranded protelomerase recognition sequence; c) contacting the adaptor-ligated nucleic acid molecules with a protelomerase to cleave the adaptor-ligated nucleic acid molecules and covalently close the cleaved ends to provide first and second nucleic acid molecules comprising closed-ended ends; d) cleaving said first nucleic acid molecule comprising said closed-ended end at said first target sequence to provide a first nucleic acid comprising one open-ended end and one closed-ended end; A method comprising:

2. The method of claim 1 , wherein the second nucleic acid molecule comprises a second target sequence.

3. 3. The method of claim 1 or 2, wherein the protelomerase is TelN protelomerase.

4. The method of any one of claims 1 to 3, wherein the sample of step a) comprises the first and second nucleic acid molecules and a plurality of further nucleic acid molecules.

5. The method of any one of claims 1 to 4, wherein the first nucleic acid molecule of step d) is cleaved by a programmable nuclease or a restriction endonuclease.

6. 6. The method of claim 5, wherein the programmable nuclease is an RNA-guided CRISPR nuclease.

7. The method according to any one of claims 1 to 6, wherein the first and second nucleic acid molecules of step a) are prepared by fragmentation.

8. The method of claim 7, wherein the fragmentation is fragmentation of a genomic nucleic acid molecule.

9. The method of any one of claims 1 to 8, wherein the adapters in step b) are ligated by tagmentation.

10. 10. The method of any one of claims 1 to 9, comprising a step c1) of exposing the sample to an exonuclease after obtaining the nucleic acid molecule comprising closed-ended ends in step c) but before cleaving the first nucleic acid molecule comprising said closed-ended ends in step d).

11. 10. The method of any one of claims 1 to 9, comprising, after obtaining the first nucleic acid molecule comprising one open end and one closed end in step d), a step e) of exposing the sample to an exonuclease.

12. 12. The method of claim 11, comprising step f) cleaving the second nucleic acid molecule comprising the closed-ended end at the second target sequence to result in a second nucleic acid comprising one open-ended end and one closed-ended end.

13. 13. The method of any one of claims 1 to 12, comprising step g) ligating a further adaptor to the open-ended end of said first nucleic acid molecule comprising one open-ended end and one closed-ended end, wherein said further adaptor comprises at least one of an amplification primer binding site and a sequence primer binding site.

14. 13. The method of claim 12, comprising step g) ligating an additional adapter to the open-ended end of the second nucleic acid molecule comprising one open-ended end and one closed-ended end, wherein the additional adapter comprises at least one of an amplification primer binding site and a sequence primer binding site.

15. 15. The method of claim 13 or 14, wherein the further adapter comprises an identifier sequence.

16. The method according to any one of claims 1 to 15, wherein the library of nucleic acid molecules is prepared from a plurality of samples.

17. 17. The method of claim 16, wherein the plurality of samples is pooled.

18. 18. The method of claim 17, wherein the plurality of samples are pooled before step c), step d), step e), step f), or step g), or are pooled after step g).

19. 19. The method of any one of claims 1 to 18, wherein in step b) the adaptor-ligated nucleic acid molecule is repaired to remove single-strand breaks before contacting the molecule with TelN protelomerase in step c).

20. 1. A method for amplifying a library of nucleic acid molecules, comprising: Preparing a library of nucleic acid molecules as defined in any one of claims 1 to 19; i) a first primer that anneals to the first nucleic acid molecule comprising one open end and one closed end obtained in step d); ii) a first primer that anneals to the second nucleic acid molecule comprising one open end and one closed end obtained in step f); iii) a first primer annealing to said further adapter defined in step g); amplifying the library of nucleic acid molecules using at least one of A method comprising:

21. i) a first primer and a second primer that anneal to the first nucleic acid molecule comprising one open end and one closed end obtained in step d); ii) a first primer and a second primer that anneal to the second nucleic acid molecule comprising one open end and one closed end obtained in step f); iii) a first primer and a second primer that anneal to the further adapter defined in step g); and iv) a combination of a first primer defined in i) or ii) and a second primer defined in iii); 21. The method of claim 20, wherein the library of nucleic acid molecules is amplified using at least one of:

22. 1. A method for analyzing a sequence of interest in a sample comprising first and second nucleic acid molecules, comprising: Preparing a library of nucleic acid molecules as defined in any one of claims 1 to 19; sequencing the library of nucleic acid molecules; A method comprising:

23. 23. The method of claim 22, wherein the prepared nucleic acid molecules as defined in claim 20 or 21 are amplified before sequencing the library of nucleic acid molecules.

24. 24. The method of claim 22 or 23, wherein the sequencing is deep sequencing.

Citation Information

Patent Citations

  • Method for nucleic acid amplification

    EP2692870A1

  • Methods for Combining Single Cell Profiling with Combinatorial Nanoparticle Conjugate Library Screening and In Vivo Diagnostic System

    US20160289769A1

  • Method of DNA synthesis

    US20180037943A1

  • Detecting targets by unique identifier nucleotide tags

    WO2003031591A2

  • Methods and compositions for identifying or quantifying targets in a biological sample

    WO2018144813A1