Compositions and methods for nucleotide modification-based depletion
By leveraging nucleotide modification differences, the method efficiently enriches nucleic acids of interest using modification-sensitive restriction enzymes, addressing the inefficiencies of current depletion methods and enhancing sequencing applications.
Patent Information
- Application Number
- JP2021560052
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-09
- Filing Date
- 2020-04-08
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2040-04-08
AI Technical Summary
Existing methods for depleting unwanted sequences from nucleic acid libraries are time-consuming and expensive, and there is a need for a more efficient method to enrich sequences of interest.
Utilizing differences in nucleotide modifications between nucleic acids of interest and those targeted for depletion, employing modification-sensitive restriction enzymes to cleave and ligate adaptors to the ends of nucleic acids of interest, thereby enriching the sample.
Enriches the sample for nucleic acids of interest by at least 2-fold to 1000-fold without size selection or modification-sensitive target binding, reducing library complexity and enhancing downstream applications like PCR amplification and high-throughput sequencing.
Smart Images

Figure 0007780951000006 
Figure 0007780951000007 
Figure 0007780951000008
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to and the benefit of U.S. Provisional Application No. 62,831,302, filed April 9, 2019, the contents of which are incorporated herein by reference in their entirety.
[0002] Incorporating a sequence list The contents of the text file submitted electronically herewith are incorporated herein by reference in their entirety. A computer-readable copy of the Sequence Listing (Filename: ARCB_01301WO_SeqList, Date Recorded: April 6, 2020, File Size: 13 KB). [Background technology]
[0003] Sample libraries, such as cDNA libraries derived from human clinical DNA samples and RNA, contain sequences that have little useful value and increase the cost of sequencing. Methods have been developed to deplete these unwanted sequences (e.g., via hybridization capture) and enrich for sequences of interest, but these methods can often be time-consuming and expensive. Therefore, there is a need in the art for a method to deplete unwanted sequences from libraries. The present invention provides a method for depleting sequences from libraries and enriching for desired sequences using differences in nucleotide modifications between the sequences of interest and the sequences targeted for depletion. Summary of the Invention
[0004] The present disclosure provides methods for enriching a sample of a nucleic acid of interest by at least about two-fold compared to a nucleic acid targeted for depletion, comprising using differences in nucleotide modifications between the nucleic acid of interest and the nucleic acid targeted for depletion.
[0005] The present disclosure provides methods for enriching a sample of nucleic acids of interest by at least about two-fold compared to nucleic acids targeted for depletion, comprising using differences in nucleotide modifications between the nucleic acids of interest and the nucleic acids targeted for depletion, and not involving size selection or modification-sensitive target binding.
[0006] The present disclosure provides a method for enriching a sample of a nucleic acid of interest relative to a nucleic acid targeted for depletion by at least about two-fold, comprising using differences in nucleotide modifications between the nucleic acid of interest and the nucleic acid targeted for depletion to ligate the target nucleic acid to an adaptor, but not the nucleic acid targeted for depletion.
[0007] The present disclosure provides a method for enriching a sample for nucleic acids of interest, comprising: (a) providing a sample containing nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids of interest or at least a subset of the nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme; (b) terminally dephosphorylating the plurality of nucleic acids in the sample; (c) contacting the sample from (b) with the first modification-sensitive restriction enzyme under conditions that allow cleavage of at least some of the first modification-sensitive restriction sites in the nucleic acids in the sample; and (d) contacting the sample from (c) with adaptors under conditions that allow ligation of the adaptors to the 5' and 3' ends of the plurality of nucleic acids of interest, thereby producing a sample enriched for nucleic acids of interest that are adaptor-linked at their 5' and 3' ends.
[0008] In some embodiments of the disclosed methods, both the nucleic acid of interest and the nucleic acid targeted for depletion each comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme, and in some embodiments, the frequency of nucleotide modifications within or adjacent to the plurality of first recognition sites is not the same in the nucleic acid of interest as in the nucleic acid targeted for depletion.
[0009] In some embodiments of the disclosed methods, the activity of a first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to its cognate recognition site, hi some embodiments, the plurality of first recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of first recognition sites in the nucleic acid of interest.
[0010] In some embodiments of the disclosed methods, the first modification-sensitive restriction enzyme is active at recognition sites that include at least one modified nucleotide and is not active at recognition sites that do not include at least one modified nucleotide, hi some embodiments, the plurality of first recognition sites in the nucleic acid targeted for depletion are more frequently modified than the plurality of first recognition sites in the nucleic acid of interest.
[0011] In some embodiments of the disclosed methods, the method further comprises, prior to step (d), contacting the sample from (c) with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated terminus of the nucleic acid.
[0012] In some embodiments of the disclosed methods, the method further comprises (e) contacting the adaptor-linked nucleic acids from (d) with a second modification-sensitive restriction enzyme under conditions that allow the second modification-sensitive restriction enzyme to cleave the second recognition site, wherein at least a subset of the nucleic acids targeted for depletion comprise a plurality of second recognition sites for the second modification-sensitive restriction enzyme, and wherein the second modification-sensitive restriction enzyme targets the recognition sites that comprise at least one modified nucleotide and does not target the recognition sites that do not include at least one modified nucleotide, thereby producing a collection of nucleic acids targeted for depletion that are adaptor-linked at one end and a collection of nucleic acids of interest that are adaptor-linked at both ends.
[0013] In some embodiments of the disclosed methods, the method further comprises contacting the sample after step (d) with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes, where the gNAs are complementary to the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends. In some embodiments, the method comprises contacting the sample with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes after step (d), where the gNAs are complementary to the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends. 2 specific nucleic acid-guided nuclease-gNA complex, at least 10 3 specific nucleic acid-guided nuclease-gNA complex, 10 4 Specific nucleic acid-guided nuclease-gNA complex or 10 5 The method includes contacting the sample with a specific nucleic acid-guided nuclease-gNA complex of the following structure: In some embodiments, the nucleic acid-guided nuclease is Cas9, Cpf1, or a combination thereof.
[0014] The present disclosure provides a method for enriching a sample for nucleic acids of interest, comprising: (a) providing a sample containing nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids targeted for depletion comprise multiple recognition sites for a modification-sensitive restriction enzyme; (b) terminally dephosphorylating multiple nucleic acids in the sample; (c) contacting the sample from (b) with the modification-sensitive restriction enzyme under conditions that allow for cleavage of the modification-sensitive restriction sites of the nucleic acids in the sample, thereby generating nucleic acids with exposed terminal phosphates; and (d) contacting the sample with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated ends of the nucleic acids, thereby generating a sample enriched for the nucleic acids of interest.
[0015] In some embodiments of the disclosed methods, the nucleic acid of interest and the nucleic acid targeted for depletion each comprise multiple recognition sites for a modification-sensitive restriction enzyme, hi some embodiments, the multiple recognition sites in the nucleic acid targeted for depletion are modified more frequently than the multiple recognition sites in the nucleic acid of interest.
[0016] In some embodiments of the disclosed methods, the method includes (e) contacting the sample from (d) with adaptors under conditions that allow for ligation of the adaptors to the 5' and 3' ends of a plurality of nucleic acids of interest, thereby producing a sample enriched in nucleic acids of interest that are adaptor-linked at their 5' and 3' ends.
[0017] In some embodiments of the disclosed methods, the method further comprises contacting the sample after step (d) with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes, where the gNAs are complementary to the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends. In some embodiments, the method comprises contacting the sample with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes after step (d), where the gNAs are complementary to the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends. 2 specific nucleic acid-guided nuclease-gNA complex, at least 10 3 specific nucleic acid-guided nuclease-gNA complex, 10 4 Specific nucleic acid-guided nuclease-gNA complex or 10 5 The method includes contacting the sample with a specific nucleic acid-guided nuclease-gNA complex of the following structure: In some embodiments, the nucleic acid-guided nuclease is Cas9, Cpf1, or a combination thereof.
[0018] The present disclosure provides a method of enriching a sample for nucleic acids of interest, comprising: (a) providing a sample containing nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids targeted for depletion comprise a plurality of recognition sites for a modification-sensitive restriction enzyme; (b) contacting the sample with adapters under conditions that allow ligation of the adapters to the 5' and 3' ends of a plurality of nucleic acids in the sample; and (c) contacting the sample from (b) with the modification-sensitive restriction enzyme under conditions that allow cleavage of the modification-sensitive restriction sites of nucleic acids in the sample, thereby producing a sample enriched for nucleic acids of interest that are adapter-linked at their 5' and 3' ends.
[0019] In some embodiments of the disclosed methods, both the nucleic acid of interest and the nucleic acid targeted for depletion each comprise multiple recognition sites for a modification-sensitive restriction enzyme, hi some embodiments, the multiple recognition sites in the nucleic acid targeted for depletion are more frequently modified than the multiple recognition sites in the nucleic acid of interest.
[0020] In some embodiments of the disclosed methods, the method further comprises contacting the sample after step (d) with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes, where the gNAs are complementary to the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends. In some embodiments, the method comprises contacting the sample with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes after step (d), where the gNAs are complementary to the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends. 2 specific nucleic acid-guided nuclease-gNA complex, at least 10 3 specific nucleic acid-guided nuclease-gNA complex, 10 4 Specific nucleic acid-guided nuclease-gNA complex or 10 5 The method includes contacting the sample with a specific nucleic acid-guided nuclease-gNA complex of the following structure: In some embodiments, the nucleic acid-guided nuclease is Cas9, Cpf1, or a combination thereof.
[0021] The present disclosure provides a method of enriching a sample for nucleic acids of interest, comprising: (a) providing a sample comprising nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids of interest or at least a subset of the nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme, wherein activity of the first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to the cognate recognition site; (b) terminally dephosphorylating the plurality of nucleic acids in the sample; (c) contacting the sample from (b) with the first modification-sensitive restriction enzyme under conditions that allow cleavage of at least some of the first modification-sensitive restriction sites in the nucleic acids in the sample; and (d) contacting the sample from (c) with adapters under conditions that allow ligation of the adapters to the 5' and 3' ends of the plurality of nucleic acids of interest, thereby producing a sample enriched for nucleic acids of interest that are adapter-linked at their 5' and 3' ends.
[0022] In some embodiments of the disclosed methods, both the nucleic acid of interest and the nucleic acid targeted for depletion each comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme. In some embodiments, the frequency of nucleotide modifications within or adjacent to the plurality of first recognition sites is not the same in the nucleic acid of interest as in the nucleic acid targeted for depletion. In some embodiments, the plurality of first recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of first recognition sites in the nucleic acid of interest.
[0023] In some embodiments of the disclosed methods, the methods further comprise amplifying, sequencing, or cloning the nucleic acids of interest that are adapter-linked at their 5' and 3' ends using the adapters.
[0024] In some embodiments, the nucleotide modification comprises an adenine modification or a cytosine modification. In some embodiments, the adenine modification comprises adenine methylation. In some embodiments, the adenine methylation comprises Dam methylation or EcoKI methylation. In some embodiments, the cytosine modification comprises 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxylcytosine, 5-glucosylhydroxymethylcytosine, or 3-methylcytosine. In some embodiments, the cytosine modification comprises cytosine methylation. In some embodiments, the cytosine methylation comprises CpG methylation, CpA methylation, CpT methylation, CpC methylation, or a combination thereof. In some embodiments, the cytosine methylation comprises Dcm methylation, DNMT1 methylation, DNMT3A methylation, or DNMT3B methylation.
[0025] In some embodiments, the nucleic acid targeted for depletion comprises a host nucleic acid and the nucleic acid of interest comprises a non-host nucleic acid. [Brief explanation of the drawings]
[0026] [Figure 1] Figure 1 shows an exemplary method of the present disclosure. Nucleic acids in a sample are dephosphorylated and then digested with a restriction enzyme that is blocked by the presence of a modification in the restriction enzyme recognition site. The phosphates exposed by the digestion are then used to ligate adapters to the nucleic acids of interest.
[0027] [Figure 2] 2 illustrates an exemplary method of the present disclosure. Nucleic acids in a sample are dephosphorylated and then digested with a restriction enzyme that recognizes a restriction enzyme site containing one or more modified nucleotides. The cleaved nucleic acid is then digested with an exonuclease that uses the exposed terminal phosphates to ligate adapters to the remaining nucleic acid of interest.
[0028] [Figure 3]3 shows an exemplary method of the present disclosure: Nucleic acids in a sample are adaptor-ligated and then digested with a restriction enzyme that recognizes a restriction enzyme site containing one or more modified nucleotides, resulting in a nucleic acid of interest that is adaptor-ligated at both ends.
[0029] [Figure 4] 4 shows an exemplary method of the present disclosure. Nucleic acids in a sample are adaptor-linked and then cleaved with a nucleic acid-guided nuclease that cleaves the nucleic acid targeted for depletion, resulting in a nucleic acid of interest that is adaptor-linked at both ends. This method can be used in combination with the nucleotide modification-based methods of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0030] Epigenetic nucleotide modifications within genomes vary between species. For example, the frequency and type of nucleotide modifications differ between vertebrates and bacteria, fungi, and viruses. Furthermore, in some genomes, such as the human genome, modifications such as methylation occur more frequently at transcriptionally active sites (e.g., genes and / or gene promoters) and less frequently at other sites in the genome (e.g., repetitive regions). Some restriction enzymes are sensitive to nucleotide modifications at or adjacent to their cognate recognition site. Differences in nucleotide modifications between sequences can be exploited to enrich samples for nucleic acids of interest using modification-sensitive restriction enzymes.
[0031] The present disclosure provides methods for enriching a sample for a nucleic acid of interest relative to a nucleic acid targeted for depletion, comprising using differences in nucleotide modification frequency between the nucleic acid of interest and the nucleic acid targeted for depletion. The disclosed methods allow for the reduction of library complexity and the enrichment of sequences that can be used in a variety of downstream applications, including, but not limited to, PCR amplification, cloning, high-throughput sequencing, identification of rare sequences in mixed populations, and quantification of sequences within a library. In some embodiments, a sample is enriched for the nucleic acid of interest by at least about 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 50-fold, 100-fold, 200-fold, 500-fold, or 1000-fold. In some embodiments, the sample is at least about 2-fold enriched for the nucleic acid of interest. In some embodiments, the sample is at least about 3-fold enriched for the nucleic acid of interest. In some embodiments, the sample is at least about 2-fold to about 3-fold enriched for the nucleic acid of interest. In some embodiments, the sample is at least about 12-fold enriched for the nucleic acid of interest. In some embodiments, the sample is at least about 15-fold enriched for the nucleic acid of interest. In some embodiments, the sample is at least about 50% to about 70% depleted for the nucleic acid targeted for depletion. In some embodiments, the sample is at least about 95% depleted for the nucleic acid targeted for depletion.
[0032] The present disclosure provides a method for enriching a sample for nucleic acids of interest, comprising: (a) providing a sample comprising nucleic acids of interest and nucleic acids targeted for depletion, wherein at least at least a subset of the nucleic acids of interest or at least a subset of the nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme; (b) terminally dephosphorylating the plurality of nucleic acids in the sample; (c) contacting the sample from (b) with the first modification-sensitive restriction enzyme under conditions that allow cleavage of at least some of the first modification-sensitive restriction sites in the nucleic acids in the sample; and (d) contacting the sample from (c) with adaptors under conditions that allow ligation of the adaptors to the 5' and 3' ends of the plurality of nucleic acids of interest, thereby producing a sample enriched for nucleic acids of interest that are adaptor-linked at their 5' and 3' ends.
[0033] The present disclosure provides a method for enriching a sample for nucleic acids of interest, comprising: (a) providing a sample containing nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids targeted for depletion contain multiple recognition sites for a modification-sensitive restriction enzyme; (b) ultimately dephosphorylating multiple nucleic acids in the sample; (c) contacting the sample from (b) with the modification-sensitive restriction enzyme under conditions that allow for cleavage of the modification-sensitive restriction sites of the nucleic acids in the sample, thereby generating nucleic acids with exposed terminal phosphates; and (d) contacting the sample with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated ends of the nucleic acids, thereby generating a sample enriched for the nucleic acids of interest.
[0034] The present disclosure provides a method of enriching a sample for nucleic acids of interest, comprising: (a) providing a sample comprising nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids of interest or at least a subset of the nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme, wherein activity of the first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to the cognate recognition site; (b) terminally dephosphorylating the plurality of nucleic acids in the sample; (c) contacting the sample from (b) with the first modification-sensitive restriction enzyme under conditions that allow cleavage of at least some of the first modification-sensitive restriction sites in the nucleic acids in the sample; and (d) contacting the sample from (c) with adapters under conditions that allow ligation of the adapters to the 5' and 3' ends of the plurality of nucleic acids of interest, thereby producing a sample enriched for nucleic acids of interest that are adapter-linked at their 5' and 3' ends.
[0035] The present disclosure provides a method for enriching a sample for nucleic acids of interest, comprising: (a) providing a sample comprising nucleic acids of interest and nucleic acids targeted for depletion, wherein at least the nucleic acids targeted for depletion comprise a plurality of recognition sites for a modification-sensitive restriction enzyme; (b) contacting the sample with adapters under conditions that allow ligation of the adapters to the 5' and 3' ends of a plurality of nucleic acids in the sample; and (c) contacting the sample from (b) with the modification-sensitive restriction enzyme under conditions that allow cleavage of the modification-sensitive restriction sites of nucleic acids in the sample, thereby producing a sample enriched for nucleic acids of interest that are adapter-linked at their 5' and 3' ends.
[0036] The present disclosure provides methods for depleting nucleic acids targeted for depletion by digestion of the nucleic acids targeted for depletion, thereby enriching a sample for the nucleic acid of interest.
[0037] The present disclosure provides methods for depleting nucleic acids targeted for depletion by digestion of the targeted nucleic acids, thereby enriching a sample for nucleic acids of interest, by differential adapter attachment to the nucleic acids targeted for depletion and the nucleic acids of interest.
[0038] The present disclosure provides methods for depleting nucleic acids targeted for depletion without the use of size selection.
[0039] The present disclosure provides a method for depleting nucleic acids targeted for depletion without using modification-sensitive target binding, thereby enriching a sample for nucleic acids of interest. In some embodiments, the method for depleting nucleic acids targeted for depletion does not use CpG-sensitive target binding.
[0040] In some embodiments, the disclosed methods involving modification-sensitive restriction enzymes are used as an independent method to enrich a sample for nucleic acids of interest. In alternative embodiments, the disclosed methods based on differences in nucleotide modifications are combined with one or more additional methods of sample enrichment. In some embodiments, any of the enrichment methods disclosed herein are combined with any other additional enrichment method disclosed herein. In some embodiments, the additional method is a method based on nucleotide modifications. In some embodiments, the additional method uses a library of guide nucleic acids (gNAs) and nucleic acid-guided nucleases. In some embodiments, the additional method is a combination of an enrichment method based on nucleotide modifications and an enrichment method using a library of guide nucleic acids (gNAs) and nucleic acid-guided nucleases. In some embodiments, the additional method depletes nucleic acids targeted for depletion by digestion of the nucleic acids targeted for depletion. In some embodiments, the additional method uses the disclosed methods to deplete nucleic acids targeted for depletion by differential adapter attachment. In some embodiments, the additional method depletes nucleic acids targeted for depletion without using size selection. In some embodiments, the additional method depletes nucleic acids targeted for depletion without using a modification-sensitive targeting linkage, hi some embodiments, the additional method depletes nucleic acids targeted for depletion without using a CpG-sensitive targeting linkage.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, the preferred methods and materials are described.
[0042] Numeric ranges are inclusive of the numbers defining the range.
[0043] For the purposes of interpreting this specification, the following definitions shall apply, and whenever appropriate, terms used in the singular shall include the plural and vice versa. In the event that a definition set forth below conflicts with any document incorporated herein by reference, the set definition shall control.
[0044] As used herein, the singular forms "a," "an," and "the" include plural references unless specifically stated otherwise.
[0045] The term "about" as used herein refers to a normal error range for the respective value, readily known to one of ordinary skill in the art. Reference herein to "about" a value or parameter includes (and describes) embodiments directed to the value or parameter itself.
[0046] The term "nucleic acid" as used herein refers to a molecule comprising one or more nucleic acid subunits. A nucleic acid can comprise one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), and modified versions thereof. Nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and combinations or derivatives thereof. Nucleic acids can be single-stranded and / or double-stranded.
[0047] Nucleic acids comprise "nucleotides," which, as used herein, is intended to include moieties that contain purine and pyrimidine bases, as well as modified versions thereof.
[0048] The terms "nucleic acid" and "polynucleotide" are used interchangeably herein. Polynucleotide is used to describe nucleic acid polymers of any length, e.g., greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1000 bases, up to about 10,000 or more than 10,000 bases. Polynucleotides are composed of nucleotides, e.g., deoxyribonucleotides or ribonucleotides, and can be produced enzymatically or synthetically (e.g., PNAs, as described in U.S. Pat. No. 5,948,902 and references cited therein). They can hybridize with two naturally occurring nucleic acids in a similar sequence-specific manner, e.g., participate in Watson-Crick base pairing interactions. Naturally occurring nucleotides include guanine, cytosine, adenine, and thymine (G, C, A, and T, respectively). DNA and RNA have deoxyribose and ribose sugar backbones, respectively, while the PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. In PNAs, various purine and pyrimidine bases are linked to the backbone by methylene carbonyl bonds. Locked nucleic acids (LNAs), often referred to as inaccessible RNAs, are modified RNA nucleotides. The ribose moiety of LNA nucleotides is modified with an additional bridge connecting the 2' oxygen and 4' carbon. The bridge "locks" the ribose in a 3'-endo (north) conformation, which is commonly found in A-form duplexes. LNA nucleotides can be mixed with DNA or RNA residues in oligonucleotides as needed. The term "unstructured nucleic acid" or "UNA" refers to nucleic acids containing non-natural nucleotides that bind with reduced stability. For example, unstructured nucleic acids can contain G' and C' residues, which are non-naturally occurring analogs of G and C, which base pair with each other with reduced stability, but retain the base-pairing ability of naturally occurring C and G residues, respectively. Unstructured nucleic acids are described in U.S. Patent Application Publication No. 20050233340, which is incorporated herein by reference for its disclosure of UNAs.
[0049] "Modified nucleotides" include, but are not limited to, methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated riboses or other heterocycles. Exemplary modifications include, but are not limited to, cytosine modifications, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxylcytosine, 5-glucosylhydroxymethylcytosine, or 3-methylcytosine.
[0050] The term "cleavage," also known as "scission," as used herein refers to a reaction that breaks the phosphodiester bond between two adjacent nucleotides on both strands of a double-stranded DNA molecule, thereby resulting in a double-stranded break in the DNA molecule.
[0051] The term "nicking," as used herein, refers to a reaction that cleaves the phosphodiester bond between two adjacent nucleotides in only one strand of a double-stranded DNA molecule, thereby resulting in a single-strand break in the DNA molecule.
[0052] As used herein, the term "cleavage site" refers to the site at which a double-stranded DNA molecule is cleaved.
[0053] The terms "capture" and "enrichment" are used interchangeably herein to refer to the process of selectively isolating a nucleic acid region containing a sequence of interest, a target site of interest, a non-target sequence, or a non-target site of interest. In some embodiments, a sample is enriched for the sequence of interest or the captured sequence of interest by selectively depleting non-target sequences. Isolation of a nucleic acid region can, in some cases, be achieved by selectively modifying the nucleic acid region of interest in a manner suitable for downstream applications. For example, the isolated nucleic acid can selectively have adaptors linked to the 5' and 3' ends of the nucleic acid.
[0054] The term "next generation sequencing" refers to so-called parallelized sequencing-by-synthesis or sequencing-by-ligation platforms, such as those currently employed by Illumina, Life Technologies, and Roche. Next generation sequencing methods can also include nanopore sequencing methods or electronic detection-based methods, such as Oxford Nanopore or the Ion Torrent technology commercialized by Life Technologies.
[0055] sample Nucleic acids isolated or derived from any type of sample are considered within the scope of the disclosed methods.
[0056] In some embodiments of the disclosed methods, the sample is a biological sample, a clinical sample, a forensic sample, or an environmental sample. Clinical and forensic samples include, but are not limited to, whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bone, fingernails, feces, urine tissue, and biopsy samples.
[0057] In some embodiments, the sample is a metagenomic sample (a sample containing multiple biological species). In some embodiments, a metagenomic sample comprises a sample isolated or derived from an organism that is host to other non-host organisms (e.g., a mammal harboring one or more viral, bacterial, fungal, or eukaryotic parasites). In some embodiments, a metagenomic sample comprises a sample of a microbial community (e.g., a biofilm).
[0058] In some embodiments, the nucleic acids in the sample are fragmented. In some embodiments, the nucleic acids of interest and the nucleic acids targeted for depletion are fragmented.
[0059] In some embodiments, the nucleic acid in the sample is about 20 to about 5000 base pairs (bp) in length, about 20 to about 1000 bp in length, about 20 to about 500 bp in length, about 20 to about 400 bp in length, about 20 to about 300 bp in length, about 20 to about 200 bp in length, about 20 to 100 bp in length, about 50 to about 5000 bp in length, about 50 to about 1000 bp in length, about 50 to about 500 bp in length, about 50 to about 400 bp in length, about 50 to about 300 bp in length, about 50 to about 200 bp in length, about 50 to 100 bp in length, about 100 to about 5000 bp in length, about 100 to about 1000 bp in length, about 100 to about 500 bp in length, about 100 to about 400 bp in length, about 100 to about 300 bp in length, or about 100 to about 200 bp in length. In some embodiments, the nucleic acid in the sample is about 50 to about 1000 bp in length. In some embodiments, the nucleic acid in the sample is about 50 to about 500 bp in length. In some embodiments, the nucleic acid in the sample is about 100 to about 500 bp in length.
[0060] Nucleic acid of interest Provided herein are methods that can be used to enrich nucleic acids of interest in a sample for a variety of applications, including, but not limited to, amplification, cloning, high-throughput sequencing, detection and quantification of nucleic acids in the sample.
[0061] In some embodiments, the nucleic acid of interest comprises at least one recognition site for at least a first modification-sensitive restriction enzyme. In some embodiments, the nucleic acid of interest comprises multiple recognition sites for at least a first modification-sensitive restriction enzyme. In some embodiments, the nucleic acid of interest comprises multiple recognition sites for each of a first and a second modification-sensitive restriction enzyme. In some embodiments, the activity of the first and / or second modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to its cognate restriction site. In some embodiments, the first and / or second modification-sensitive restriction enzyme is active at a recognition site that comprises at least one modified nucleotide within or adjacent to the recognition site, but is not active at a recognition site that does not include at least one modified nucleotide within or adjacent to the recognition site. In some embodiments, only the nucleic acid of interest, and not the nucleic acid targeted for depletion, comprises one or more restriction sites for at least one first modification-sensitive restriction enzyme. In some embodiments, both the nucleic acid of interest and the nucleic acid targeted for depletion comprise multiple recognition sites for a first and optionally a second modification-sensitive restriction enzyme, but the recognition sites differ in the frequency with which they comprise modified nucleotides adjacent to or within the recognition site. In some embodiments, the nucleic acid of interest comprises multiple recognition sites for more than two (i.e., at least 3, 4, 5, 6, 7, 8, 9, or 10) modification-sensitive restriction enzymes. In some embodiments, the nucleic acid of interest and the nucleic acid targeted for depletion each comprise multiple recognition sites for more than two (i.e., at least 3, 4, 5, 6, 7, 8, 9, or 10) modification-sensitive restriction enzymes.
[0062] In some exemplary embodiments, the nucleic acid of interest is derived from a species lacking or having low levels of CpG methylation (e.g., a non-host species such as a virus, fungus, or bacterium). Conversely, in such embodiments, the nucleic acid targeted for depletion is derived from a species having higher levels of CpG methylation, such as a mammal (e.g., a human). One skilled in the art can select a modification-sensitive restriction enzyme with a recognition site that contains one or more CG dimers and whose activity is blocked by the presence of CpG methylation, and use the methods of the present disclosure to enrich for nucleic acids of interest.
[0063] In some exemplary embodiments, the nucleic acid of interest is derived from a species lacking or having low levels of CpG methylation (e.g., a non-host species such as a virus, fungus, or bacterium). Conversely, in such embodiments, the nucleic acid targeted for depletion is derived from a species having higher levels of CpG methylation, such as a mammal (e.g., a human). One skilled in the art can select a modification-sensitive restriction enzyme having a recognition site that includes one or more CG dimers and whose activity is specific for the presence of CpG methylation within or near the recognition site, and use the methods of the present disclosure to enrich for nucleic acids of interest.
[0064] In some embodiments, the nucleic acid of interest is a genomic sequence (genomic DNA). In some embodiments, the nucleic acid of interest is a mammalian genomic sequence. In some embodiments, the nucleic acid of interest is a eukaryotic genomic sequence. In some embodiments, the nucleic acid of interest is a prokaryotic genomic sequence. In some embodiments, the sequence of interest is a viral genomic sequence. In some embodiments, the nucleic acid of interest is a bacterial genomic sequence. In some embodiments, the nucleic acid of interest is a plant genomic sequence. In some embodiments, the nucleic acid of interest is a microbial genomic sequence. In some embodiments, the sequence of interest is a genomic sequence from a parasite, e.g., a eukaryotic parasite. In some embodiments, the nucleic acid of interest is a genomic sequence from a pathogen, e.g., a bacterium, virus, or fungus. In some embodiments, the nucleic acid of interest is genomic sequences from multiple bacterial, viral, or fungal species.
[0065] In some embodiments, the nucleic acid of interest can be a genome fragment comprising a region of the genome, or the entire genome. In one embodiment, the genome is a DNA genome. In another embodiment, the genome is an RNA genome.
[0066] In some embodiments, the nucleic acid of interest comprises a repetitive sequence. Exemplary, but non-limiting, repetitive sequences include, but are not limited to, mitochondrial sequences, ribosomal sequences, centromeric sequences, Alu sequences, long interspersed nuclear elements (LINEs), and short interspersed nuclear elements (SINEs).
[0067] In some embodiments, the nucleic acid of interest is from a eukaryote or a prokaryote, from a mammalian or non-mammalian organism, from an animal or a plant, from a bacterium or a virus, from an animal parasite, or from a pathogen.
[0068] In some embodiments, the nucleic acid of interest is derived from a species of bacteria, hi one embodiment, the bacteria is a bacterium that causes tuberculosis.
[0069] In some embodiments, the nucleic acid of interest is derived from a virus.
[0070] In some embodiments, the nucleic acid of interest is derived from a fungal species.
[0071] In some embodiments, the nucleic acid of interest is derived from an algal species.
[0072] In some embodiments, the nucleic acid of interest is derived from any mammalian parasite.
[0073] In some embodiments, the nucleic acid of interest is obtained from any mammalian parasite. In one embodiment, the parasite is a worm. In another embodiment, the parasite is a parasite that causes malaria. In another embodiment, the parasite is a parasite that causes leishmaniasis. In another embodiment, the parasite is an amoeba.
[0074] In some embodiments, the nucleic acid of interest is derived from a pathogen.
[0075] In some embodiments, the nucleic acid of interest is about 20 to about 5000 bp in length, about 20 to about 1000 bp in length, about 20 to about 500 bp in length, about 20 to about 400 bp in length, about 20 to about 300 bp in length, about 20 to about 200 bp in length, about 20 to about 100 bp in length, about 50 to about 5000 bp in length, about 50 to about 1000 bp in length, about 50 to about 500 bp in length, about 50 to about 400 bp in length, about 50 to about 300 bp in length, about 50 to about 200 bp in length, about 50 to about 100 bp in length, about 100 to about 5000 bp in length, about 100 to about 1000 bp in length, about 100 to about 500 bp in length, about 100 to about 400 bp in length, about 100 to about 300 bp in length, or about 100 to about 200 bp in length. In some embodiments, the nucleic acid of interest is about 50 to about 1000 bp in length. In some embodiments, the nucleic acid of interest is about 50 to about 500 bp in length. In some embodiments, the nucleic acid of interest is about 100 to about 500 bp in length.
[0076] In some embodiments, the nucleic acid of interest comprises less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1% of the total nucleic acids in the sample.
[0077] In some exemplary embodiments, the nucleic acid of interest comprises less than 50% of the total nucleic acid in the sample.
[0078] In some exemplary embodiments, the nucleic acid of interest comprises less than 30% of the total nucleic acid in the sample.
[0079] In some exemplary embodiments, the nucleic acid of interest comprises less than 5% of the total nucleic acid in the sample.
[0080] In some embodiments, the nucleic acid of interest comprises at least 0.5%, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% of the total nucleic acid in the sample.
[0081] Nucleic acids targeted for depletion Provided herein are methods that can be used to deplete nucleic acids from a sample to produce a sample enriched for a nucleic acid of interest that can be used for a variety of applications, including, but not limited to, amplification, cloning, high-throughput sequencing, and detection and quantification of nucleic acids in a sample.
[0082] In some embodiments, the nucleic acid targeted for depletion comprises at least one recognition site for at least one first modification-sensitive restriction enzyme. In some embodiments, the nucleic acid targeted for depletion comprises multiple recognition sites for at least one first modification-sensitive restriction enzyme. In some embodiments, the nucleic acid targeted for depletion comprises multiple recognition sites for each of a first and a second modification-sensitive restriction enzyme. In some embodiments, the activity of the first and / or second modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to its cognate restriction site. In some embodiments, the first and / or second modification-sensitive restriction enzyme is active at a recognition site that includes at least one modified nucleotide within or adjacent to the recognition site, but is not active at a recognition site that does not include at least one modified nucleotide within or near the recognition site. In some embodiments, only the nucleic acid targeted for depletion does not include the nucleic acid of interest and includes one or more restriction sites for at least the first modification-sensitive restriction enzyme. In some embodiments, both the nucleic acid of interest and the nucleic acid targeted for depletion comprise multiple recognition sites for a first and optionally a second modification-sensitive restriction enzyme, but the recognition sites differ in the frequency with which they contain modified nucleotides adjacent to or within the recognition site. In some embodiments, the nucleic acid targeted for depletion comprises multiple recognition sites for more than two (i.e., at least 3, 4, 5, 6, 7, 8, 9, or 10) modification-sensitive restriction enzymes. In some embodiments, the nucleic acid of interest and the nucleic acid targeted for depletion each comprise multiple recognition sites for more than two (i.e., at least 3, 4, 5, 6, 7, 8, 9, or 10) modification-sensitive restriction enzymes.
[0083] In some exemplary embodiments, the nucleic acids targeted for depletion include human RNA or DNA. In some cases, all human nucleic acids are targeted for depletion.
[0084] In some exemplary embodiments, the nucleic acid targeted for depletion is derived from a host species, such as a mammal (e.g., a human), that has a high level of CpG methylation relative to the nucleic acid of interest. One skilled in the art could select a modification-sensitive restriction enzyme with a recognition site that contains one or more CG dimers and whose activity is blocked by the presence of CpG methylation, and use the methods of the present disclosure to deplete the nucleic acid targeted for depletion and obtain a sample enriched for the nucleic acid of interest.
[0085] In some exemplary embodiments, the nucleic acid targeted for depletion is derived from a host species, such as a mammal (e.g., a human), that has a high level of CpG methylation relative to the nucleic acid of interest. One skilled in the art could select a modification-sensitive restriction enzyme with a recognition site that includes one or more CG dimers and whose activity is specific for the presence of CpG methylation within or near the recognition site, and use the methods of the present disclosure to deplete the nucleic acid targeted for depletion and obtain a sample enriched for the nucleic acid of interest.
[0086] In some embodiments, the nucleic acids targeted for depletion are abundant genomic sequences, such as sequences from one or more genomes of the most abundant species in the sample, hi some embodiments, the most abundant species in the sample is human.
[0087] In some embodiments, the nucleic acid targeted for depletion can be a genome fragment comprising a region of the genome, or the entire genome. In one embodiment, the genome is a DNA genome. In another embodiment, the genome is an RNA genome.
[0088] In some embodiments, the nucleic acid targeted for depletion is from any mammalian organism. In one embodiment, the mammal is a human. In another embodiment, the mammal is a livestock animal, such as a horse, sheep, cow, pig, or donkey. In another embodiment, the mammalian organism is a domesticated pet, such as a cat, dog, gerbil, mouse, or rat. In another embodiment, the mammal is a type of monkey.
[0089] In some embodiments, the nucleic acid targeted for depletion is from any bird or avian organism, including but not limited to chicken, turkey, duck, and goose.
[0090] In some embodiments, the nucleic acid targeted for depletion is derived from an insect, including but not limited to a honeybee, a solitary bee, an ant, a fly, a wasp, or a mosquito.
[0091] In some embodiments, the nucleic acid targeted for depletion is derived from a plant, hi one embodiment, the plant is rice, corn, wheat, rose, grape, coffee, fruit, tomato, potato, or cotton.
[0092] In some embodiments, the nucleic acid targeted for depletion comprises repetitive DNA. In some embodiments, the nucleic acid of interest comprises abundant DNA. In some embodiments, the nucleic acid targeted for depletion comprises mitochondrial DNA. In some embodiments, the nucleic acid targeted for depletion comprises ribosomal DNA. In some embodiments, the nucleic acid targeted for depletion comprises centromeric DNA. In some embodiments, the nucleic acid targeted for depletion comprises DNA containing Alu sequences (Alu DNA). In some embodiments, the nucleic acid targeted for depletion comprises long interspersed nucleotide sequences (LINE DNA). In some embodiments, the nucleic acid targeted for depletion comprises short interspersed nucleotide sequences (SINE DNA). In some embodiments, the abundant DNA comprises ribosomal DNA.
[0093] In some embodiments, the nucleic acid targeted for depletion comprises a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variation, an exon, a gene mutation, or a regulatory region.
[0094] In some embodiments, the nucleic acid targeted for depletion comprises a transcriptionally active sequence. For example, the transcriptionally active sequence comprises a promoter and a sequence of a transcriptionally active gene. According to some embodiments, the transcriptionally active region of the genome has a higher level of nucleotide modification than the transcriptionally silent region of the genome. According to some exemplary embodiments, the genome is a mammalian genome, and the nucleotide modification comprises CpG methylation. According to some exemplary embodiments, the genome is a human genome, and the nucleotide modification comprises CpG methylation.
[0095] In some embodiments, nucleic acids targeted for depletion include common or prevalent nucleic acids in a subject. For example, depleted nucleic acids can include nucleic acids common to all cell types or nucleic acids more abundant in typical or healthy cells. Following depletion, the remaining nucleic acids analyzed can then include less common or less prevalent nucleic acids, such as cell-type-specific nucleic acids. These less common nucleic acids may signal cell death, including cell death of one or more specific cell types. Such signals may indicate infection, cancer, and other diseases. In some cases, the signal is a signal of cancer-related apoptosis in a specific tissue. Nucleic acids in a sample isolated or derived from a mixed population of cells can be enriched for nucleic acids from specific cell types using differences in nucleotide modifications between the cell types and the methods disclosed herein.
[0096] In some embodiments, nucleic acids targeted for depletion are about 20 to about 5000 bp in length, about 20 to about 1000 bp in length, about 20 to about 500 bp in length, about 20 to about 400 bp in length, about 20 to about 300 bp in length, about 20 to about 200 bp in length, about 20 to about 100 bp in length, about 50 to about 5000 bp in length, about 50 to about 1000 bp in length, about 50 to about 500 bp in length, about 50 to about 400 bp in length, about 50 to about 300 bp in length, about 50 to about 200 bp in length, about 50 to about 100 bp in length, about 100 to about 5000 bp in length, about 100 to about 1000 bp in length, about 100 to about 500 bp in length, about 100 to about 400 bp in length, about 100 to about 300 bp in length, or about 100 to about 200 bp in length. In some embodiments, the nucleic acid targeted for depletion is about 50 to about 1000 bp in length. In some embodiments, the nucleic acid targeted for depletion is about 50 to about 500 bp in length. In some embodiments, the nucleic acid of interest is about 100 to about 500 bp in length.
[0097] In some embodiments, the nucleic acids targeted for depletion constitute at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the total nucleic acids in the sample.
[0098] host / non-host nucleic acids In some embodiments, the nucleic acid of interest comprises a non-host nucleic acid and the nucleic acid targeted for depletion comprises a host nucleic acid.
[0099] In some exemplary embodiments, the host is a vertebrate, and the non-host is a virus, bacterium, or fungus. In some embodiments, the vertebrate is a human. In some embodiments, the nucleotide modification comprises CpG, CpC, CpA, or CpT methylation, which occurs more frequently in the host genome than in the non-host genome. One skilled in the art can select a modification-sensitive restriction enzyme having a recognition site containing one or more CG, CC, CA, or CT dimers whose activity is blocked by the presence of methylation, and use the disclosed methods to deplete targeted host nucleic acids and obtain a sample enriched in non-host nucleic acids. In some embodiments, the host is a eukaryote. In some embodiments, the host is a mammal, bird, reptile, or insect. In some embodiments, the host is a plant. Exemplary mammals include, but are not limited to, humans, cows, horses, sheep, pigs, monkeys, dogs, cats, rabbits, rats, mice, or gerbils. In some embodiments, the host is a plant. Exemplary plants include, but are not limited to, agricultural plants such as corn, wheat, rice, tobacco, tomatoes, oranges, apples, and almonds.
[0100] In some embodiments, the host is a human.
[0101] In some embodiments, the non-host comprises multiple species of organisms. In some embodiments, the non-host is a single species of organism. In some embodiments, the non-host comprises a bacterium, a fungus, a virus, or a eukaryotic parasite. In some embodiments, the non-host is a pathogen.
[0102] Nucleotide Modifications The present disclosure provides a method for enriching a sample for nucleic acid of interest compared to the nucleic acid of interest that is targeted for depletion, comprising using the difference in nucleotide modification between the nucleic acid of interest and the nucleic acid of interest that is targeted for depletion.Any type of nucleotide modification is considered to be within the scope of the present disclosure.The exemplary but non-limiting examples of the nucleotide modification of the present disclosure are described below.
[0103] The nucleotide modifications used by the methods of the present disclosure can occur at any nucleotide (e.g., adenine, cytosine, guanine, thymine, or uracil). These nucleotide modifications can occur in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). These nucleotide modifications can occur in double- or single-stranded DNA molecules, or double- or single-stranded RNA molecules.
[0104] In some embodiments, the nucleotide modification comprises an adenine modification or a cytosine modification.
[0105] In some embodiments, the adenine modification comprises adenine methylation. In some embodiments, the adenine methylation comprises N 6 -Contains methyladenine (6mA). 6 -methyladenine (6mA) is present in both prokaryotic and eukaryotic genomes. The amount of 6mA methylation within a genome varies by species. For example, the abundance of 6mA is generally lower in mammalian and plant genomes than in prokaryotic genomes. In some cases, the abundance of 6mA is at least 1,000-fold higher in prokaryotic genomes compared to mammalian or plant genomes. In some embodiments, the location of 6mA methylation within a genome varies based on species. For example, the location of 6mA methylated nucleotides (e.g., within specific restriction enzyme recognition sites) depends on the activity of methyltransferases, the expression and activity of which differ between species. Thus, 6mA methylation can be used to distinguish between eukaryotic and prokaryotic genomes in a sample containing multiple genomes, and the methods of the present disclosure can be used to selectively enrich sequences from one genome to the other.
[0106] In some embodiments, adenine methylation includes Dam methylation. Dam methylation is a type of DNA nucleotide modification performed by deoxyadenosine methylase. Deoxyadenosine methylase (also called DNA adenine methyltransferase or Dam methylase) is an enzyme that transfers a methyl group from S-adenosylmethionine (SAM) to the N6 position of the adenine residue in the sequence 5'-GATC-3 to produce 6mA. Dam methylation and Dam methylase are found in prokaryotes and bacteriophages.
[0107] In some embodiments, the adenine methylation comprises EcoKI methylation. EcoKI methylation is a type of DNA nucleotide modification performed by EcoKI methylase. EcoKI methylase modifies adenine residues in the sequences AAC(N6)GTGC (SEQ ID NO: 1) and GCAC(N6)GTT (SEQ ID NO: 2). EcoKI methylase and EcoKI methylation are found in prokaryotes.
[0108] In some embodiments, the adenine modification is N by glycine. 6 It contains adenine modified with (momylation). The momylation modification is N6-(1-acetamido)-adenine. Momylation occurs in viruses such as bacteriophages.
[0109] In some embodiments, the modifications comprise cytosine modifications. In some embodiments, the abundance and type of cytosine modifications in the genome vary based on species. In some embodiments, the location of the cytosine modifications in the genome (e.g., within a particular restriction enzyme recognition site) varies based on species.
[0110] In some embodiments, the cytosine modification comprises 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), 5-glucosylhydroxymethylcytosine (5ghmC), or 3-methylcytosine (3mC).
[0111] In some embodiments, the cytosine modification comprises cytosine methylation, hi some embodiments, the cytosine methylation comprises 5-methylcytosine (5mC) or N4-methylcytosine (4mC).
[0112] In some embodiments, 4mC cytosine methylation is found in bacteria, hi some embodiments, the bacteria is a thermophilic bacterium, e.g., a thermophilic eubacterium or a thermophilic archaebacterium.
[0113] In some embodiments, cytosine methylation includes Dcm methylation. Dcm methylation is a type of methylation performed by Dcm methylase. In Dcm methylation, Dcm methylase (encoded by the DNA-cytosine methyltransferase or dcm gene) methylates internal (second) cytosine residues in the sequences CCAGG and CCTGG at the C5 position (5mC). Dcm methylase and Dcm methylation are found in bacteria, such as E. coli.
[0114] In some embodiments, cytosine methylation comprises DNMT1 methylation, DNMT3A methylation, or DNMT3B methylation. DNMT1 (DNA methyltransferase 1), DNMT3A (DNA methyltransferase 3 alpha), and DNMT3B (DNA methyltransferase 3 beta) are mammalian methyltransferases that mediate the methylation of CpG, CpA, CpT, and CpC cytosines.
[0115] In some embodiments, cytosine methylation includes CpG methylation, CpA methylation, CpT methylation, CpC methylation, or a combination thereof. CpG methylation, CpA methylation, CpT methylation, and CpC are found in mammals. While methylated cytosine is frequently found at CpG sites in mammals, non-CpG sites such as CpA, CpT, and CpC can also be methylated. In some embodiments, non-CpG methylation is restricted to specific cell types, including, but not limited to, pluripotent stem cells, oocytes, and cells of the nervous system. In some embodiments, non-CpG cytosine methylation is mediated by DNMT3A and DNTM3B methyltransferases. In some embodiments, cytosine is methylated at the C5 position (5mC). Thus, CpA, CpT, and CpC methylation can be used to distinguish nucleic acids isolated or derived from different cell types in a mixed cell type sample.
[0116] In some embodiments, cytosine methylation includes CpG methylation. CpG methylation in mammals is mediated by DNMT1, DNMT3A, and DNMT3B DNA methyltransferases. DNMT1 primarily binds to hemimethylated DNA at CpG sites. After DNA replication, newly synthesized strands lack methylation, while the parental strand retains methylated nucleotides. DNMT1 binds to hemimethylated CpG sites generated by DNA replication and methylates cytosines in the newly synthesized strands. DNMT3A and DNMT3B do not require hemimethylated DNA for binding and show equal affinity for both hemimethylated and unmethylated CpG sites. In some embodiments, DNMT1, DNMT3A, and DNMT3B mediate 5mC methylation. In mammals, CpG methylation occurs more frequently at transcriptionally active sites in the genome, such as promoters of active genes. Thus, CpG methylation can be used to selectively distinguish between active and inactive regions of a mammalian genome. For example, CpG methylation can be used to selectively target active regions of a mammalian genome for depletion using the methods of the present disclosure.
[0117] In some embodiments, the cytosine modification comprises 5-hydroxymethylcytosine (5hmC), which is an oxidized derivative of 5mC. 5hmC is found in viruses (e.g., bacteriophages) and some mammalian tissues (e.g., the brain).
[0118] In some embodiments, the cytosine modification comprises 5-formylcytosine (5fC), which is an oxidized derivative of 5mC. 5mC is oxidized to 5-hydroxymethylcytosine (5hmC), which is then oxidized to 5fC. In some embodiments, each of these oxidation steps is carried out by ten-eleven translocation (TET) enzymes. In some embodiments, 5fC is found in mammalian genomes.
[0119] In some embodiments, the cytosine modification comprises 5-carboxyl cytosine (5caC). 5caC is the final oxidized derivative of 5mC. 5mC is oxidized to 5hmC, which is then oxidized to 5fC and then to 5caC by enzymes of the TET family. In some embodiments, 5caC is found in mammalian genomes.
[0120] In some embodiments, the cytosine modification comprises 5-glucosylhydroxymethylcytosine. In some embodiments, 5-glucosylhydroxymethylcytosine is found in viruses. In some embodiments, the virus is a bacteriophage. In some embodiments, the virus is a non-host species, and the viral nucleic acid is the nucleic acid of interest in the sample.
[0121] In some embodiments, the cytosine modification comprises 3-methylcytosine.
[0122] Modification-sensitive restriction enzymes Provided herein are methods for enriching a sample for a nucleic acid of interest relative to a nucleic acid targeted for depletion, comprising using differences in nucleotide modifications between the nucleic acid of interest and the nucleic acid targeted for depletion, as recognized by one or more modification-sensitive restriction enzymes. Any type of restriction enzyme that is sensitive to any of the nucleotide modifications described herein is within the scope of this disclosure.
[0123] In some embodiments of the disclosed methods, the method uses at least a first modification-sensitive restriction enzyme and a second modification-sensitive restriction enzyme. In some embodiments, the first and second modification-sensitive restriction enzymes are the same. In some embodiments, the first and second modification-sensitive restriction enzymes are not the same. In some embodiments, the first or second modification-sensitive restriction enzyme is a single type of restriction enzyme (e.g., AluI or McrBC, but not both). In some embodiments, the first or second modification-sensitive restriction enzyme is a mixture of two or more types of modification-sensitive restriction enzymes (e.g., a mixture of FspEI and AbaSI). In some embodiments of the disclosed methods, the first or second modification-sensitive restriction enzyme comprises a mixture of at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 or more modification-sensitive restriction enzymes. In some embodiments of the disclosed methods, three or more different methods are combined, each using a different modification-sensitive restriction enzyme or cocktail of modification-sensitive restriction enzymes.
[0124] As used herein, the term "modification-sensitive restriction enzyme" refers to a restriction enzyme that is sensitive to the presence of modified nucleotides within or adjacent to its recognition site. Modification-sensitive restriction enzymes can be sensitive to modified nucleotides within the recognition site itself. Modification-sensitive restriction enzymes can be sensitive to modified nucleotides adjacent to the recognition site, e.g., 1 to 50 nucleotides within the 5' or 3' of the recognition site. Modification-sensitive restriction enzymes can be sensitive to modified nucleotides both within and adjacent to the recognition site. As used herein, the term "recognition site" refers to a site within a polynucleotide that contains a specific sequence recognized by a restriction enzyme. Restriction enzymes cleave within or near the recognition site of a polynucleotide. In some embodiments, restriction enzymes cleave within 1 to 105 nucleotides of the recognition site. In some embodiments, restriction enzymes recognize a pair of recognition half-sites that can be separated by 3 kilobases or more in a polynucleotide. In some embodiments, restriction enzymes recognize a specific sequence (recognition site) in a polynucleotide. In some embodiments, the recognition site is 3 to 20 bp in length. In some embodiments, the recognition site is palindromic.
[0125] Nucleotide modifications of the present disclosure can include nucleotides within the recognition site itself or adjacent to the recognition site (eg, within 1-50 nucleotides, 5' or 3', or both, of the recognition site).
[0126] In some embodiments, the modification-sensitive restriction enzyme is sensitive to a single modified nucleotide within or adjacent to the recognition site.
[0127] In some embodiments, the modification-sensitive restriction enzyme is sensitive to multiple modified nucleotides within or adjacent to the recognition site.
[0128] In some embodiments, a modification-sensitive restriction enzyme is sensitive to a particular type or types of modification (e.g., methylation, hydroxymethylation, or carboxylation) on one or more nucleotides within or adjacent to the recognition site.
[0129] In some embodiments, a modification-sensitive restriction enzyme is sensitive to a modification of a particular nucleotide or nucleotides within or adjacent to the recognition site.
[0130] In some embodiments, a modification-sensitive restriction enzyme is sensitive to a particular spatial arrangement of modified nucleotides within or adjacent to its recognition site. For example, a modification-sensitive restriction enzyme can be sensitive to a pair of modifications on opposite strands, separated by one or two nucleotides, within its recognition site of a DNA polynucleotide.
[0131] In some embodiments, a modification-sensitive restriction enzyme is blocked by the presence of one or more modified nucleotides within or adjacent to its recognition site. A modification-sensitive restriction enzyme that is blocked by the presence of modified nucleotides will cleave at recognition sites that do not contain the modified nucleotides and will not cleave, or will cleave at a reduced level, at recognition sites that do contain the modified nucleotides.
[0132] Modification-sensitive restriction enzymes whose activity is blocked by modified nucleotides include enzymes whose activity is blocked or reduced by any type of modified nucleotide, or any combination of modified nucleotides, within or adjacent to the recognition site. Exemplary modifications that can block or reduce the activity of modification-sensitive restriction enzymes include N 6Modifications that can block modification-sensitive restriction enzymes include, but are not limited to, 5-methyladenine, 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), 5-glucosylhydroxymethylcytosine, 3-methylcytosine (3mC), N4-methylcytosine (4mC), or combinations thereof. Exemplary modifications that can block modification-sensitive restriction enzymes include those mediated by Dam, Dcm, EcoKI, DNMT1, DNMT3A, DNMT3B, and TET enzymes.
[0133] In some embodiments, the modification comprises Dam methylation. Restriction enzymes that are blocked by Dam methylation include, but are not limited to, the enzymes in Table 1 below. [Table 1]
[0134] In some embodiments, the modification comprises Dcm methylation. Restriction enzymes that are blocked by Dcm methylation include, but are not limited to, those in Table 2 below. [Table 2]
[0135] In some embodiments, the modification comprises CpG methylation. Restriction enzymes that are blocked by CpG methylation include, but are not limited to, the enzymes in Table 3 below. [Table 3] TIFF0007780951000004.tif119151
[0136] In some embodiments, a modification-sensitive restriction enzyme is active at a recognition site that contains at least one modified nucleotide and is not active at a recognition site that does not contain at least one modified nucleotide, e.g., a modification-sensitive restriction enzyme cleaves at a recognition site that contains one or more modified nucleotides, but does not cleave a recognition site that does not contain one or more modified nucleotides.
[0137] Exemplary modifications recognized by modification-sensitive restriction enzymes that cleave at recognition sites containing one or more modified nucleotides include N 6 Modifications include, but are not limited to, 5-methyladenine, 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), 5-glucosylhydroxymethylcytosine, 3-methylcytosine (3mC), N4-methylcytosine (4mC), or combinations thereof. Exemplary modifications of recognized modification-sensitive restriction enzymes that specifically cleave recognition sites containing one or more modified nucleotides include those mediated by Dam, Dcm, EcoKI, DNMT1, DNMT3A, DNMT3B, and TET enzymes.
[0138] Exemplary, but non-limiting, modification-sensitive restriction enzymes that cleave at recognition sites that contain one or more modified nucleotides within or adjacent to the recognition site are listed in Table 4 below. [Table 4]
[0139] In some embodiments, the modification comprises 5-glucosylhydroxymethylcytosine and the modification-sensitive restriction enzyme comprises AbaSI, which cleaves AbaSI recognition sites that contain glucosylhydroxymethylcytosine but does not cleave AbaSI recognition sites that do not contain glucosylhydroxymethylcytosine.
[0140] In some embodiments, the nucleotide modification includes 5-hydroxymethylcytosine, and the modification-sensitive restriction enzymes include AbaSI and T4 phage β-glucosyltransferase. T4 phage β-glucosyltransferase specifically transfers the glucose moiety of uridine diphosphate glucose (UDP-Glc) to 5-hydroxymethylcytosine (5-hmC) residues in double-stranded DNA, for example, within the AbaSI recognition site, to create a glucosylhydroxymethylcytosine-modified AbaSI recognition site. AbaSI cleaves AbaSI recognition sites that contain glucosylhydroxymethylcytosine but does not cleave AbaSI recognition sites that do not contain glucosylhydroxymethylcytosine.
[0141] In some embodiments, the nucleotide modification includes methylcytosine and the modification-sensitive restriction enzyme includes McrBC. McrBC cleaves McrBC sites containing methylcytosine and does not cleave McrBC sites that do not contain methylcytosine. McrBC sites can be modified with methylcytosine on one or both DNA strands. In some embodiments, McrBC also cleaves McrBC sites containing hydroxymethylcytosine on one or both DNA strands. In some embodiments, McrBC half sites are separated by up to 3000 nucleotides. In some embodiments, McrBC half sites are separated by 55 to 103 nucleotides.
[0142] In some embodiments, the modification includes adenine methylation, and the method includes digestion with DpnI. DpnI cleaves the GATC recognition site when the adenines on both strands of the GATC recognition site are methylated. In some embodiments, DpnI GATC recognition sites containing both adenine methylation and cytosine modification occur in bacterial DNA but not in mammalian DNA. These recognition sites containing both methylated adenine and modified cytosine are selectively cleaved by DpnI in a sample (e.g., a mixed bacterial and mammalian DNA) and then treated with T4 polymerase, replacing the methylated adenine and modified cytosine with unmodified adenine and unmodified cytosine at the cleaved ends. T4 polymerase catalyzes DNA synthesis in the 5' to 3' direction in the presence of a template, primers, and nucleotides. T4 polymerase incorporates unmodified nucleotides into the newly synthesized DNA. This produces a sample in which the nucleic acid of interest contains unmodified cytosines and the nucleic acid targeted for depletion contains modified cytosines. These differences in modified cytosines can be used to enrich for nucleic acids of interest using the methods of the present disclosure.
[0143] phosphatase In some embodiments of the disclosed methods, nucleic acids in a sample are terminally dephosphorylated by contacting the nucleic acids in the sample with a modification-sensitive restriction enzyme to generate either nucleic acids of interest or nucleic acids targeted for depletion with exposed terminal phosphates that can be used in the disclosed methods to enrich a sample for nucleic acids of interest. For example, these exposed terminal phosphates can be used to target nucleic acids for depletion for exonucleolytic degradation (FIG. 2) or nucleic acids of interest for adaptor ligation (FIG. 1).
[0144] As used herein, the term "terminally dephosphorylated" refers to a nucleic acid that has had the terminal phosphate groups removed from the 5' and 3' ends of the nucleic acid molecule.
[0145] In some embodiments, nucleic acids in a sample are terminally dephosphorylated using a phosphatase. A phosphatase is an enzyme that nonspecifically catalyzes the dephosphorylation of the 5' and 3' ends of DNA and RNA molecules. In some embodiments, the phosphatase is alkaline phosphatase.
[0146] Exemplary phosphatases of the present disclosure include, but are not limited to, shrimp alkaline phosphatase (SAP), recombinant shrimp alkaline phosphatase (rSAP), calf intestinal alkaline phosphatase (CIP), and Antarctic phosphatase.
[0147] Exonuclease As used herein, the term "exonuclease" refers to a class of enzymes that sequentially remove nucleotides from the 3' or 5' end of a nucleic acid molecule. The nucleic acid molecule can be DNA or RNA. The DNA or RNA can be single-stranded or double-stranded. Exemplary exonucleases include, but are not limited to, lambda nuclease, exonuclease I, exonuclease III, and BAL-31. Exonucleases can be used to selectively degrade nucleic acids targeted for depletion using the methods of the present disclosure (e.g., Figure 2).
[0148] In some embodiments, exonuclease III is used to degrade cleaved DNA targeted for depletion while leaving the uncleaved DNA of interest intact. Exonuclease III initiates unidirectional 3'>5' degradation of one DNA strand by using blunt ends or 5' overhangs with terminal phosphates to generate single-stranded DNA and nucleotides. Because it has no activity against single-stranded DNA or DNA lacking a terminal phosphate, 3' overhangs, such as Y-shaped adapter ends, are resistant to degradation. As a result, intact double-stranded DNA fragments of interest that are not cleaved by modification-sensitive restriction enzymes and lack a terminal phosphate are not digested by exonuclease III, whereas DNA molecules targeted for depletion that are cleaved by modification-sensitive restriction enzymes are degraded by exonuclease III.
[0149] In some embodiments, exonuclease I is used to degrade cleaved DNA targeted for depletion while leaving the uncleaved DNA of interest intact. In some embodiments, a sample of nucleic acid fragments (e.g., single-stranded DNA) is dephosphorylated and cleaved with a modification-sensitive restriction enzyme that cleaves the nucleic acid targeted for depletion but not the nucleic acid of interest. Exonuclease I degrades single-stranded DNA in the 3' to 5' direction.
[0150] In some embodiments, lambda nuclease (lambda exonuclease) is used to degrade cleaved DNA targeted for depletion while leaving uncleaved DNA of interest intact. In some embodiments, a sample of nucleic acid fragments (e.g., DNA) is dephosphorylated and cleaved with a modification-sensitive restriction enzyme that cleaves the nucleic acid targeted for depletion but not the nucleic acid of interest. Lambda nuclease is a highly processive 5' to 3' exonuclease. Its preferred substrate is 5'-phosphorylated double-stranded DNA, and it degrades unphosphorylated DNA at a significantly reduced rate. Thus, the intact, dephosphorylated nucleic acid of interest is protected from lambda nuclease, while the cleaved nucleic acid targeted for depletion with its exposed 5' phosphate is degraded.
[0151] In some embodiments, exonuclease BAL-31 is used to degrade cleaved DNA targeted for depletion, while leaving the uncleaved DNA of interest intact. In some embodiments, a sample of nucleic acid fragments (e.g., DNA) is dephosphorylated and cleaved with a modification-sensitive restriction enzyme that cleaves the nucleic acid targeted for depletion but not the nucleic acid of interest. The sample is contacted with the modification-sensitive restriction enzyme, which cleaves the nucleic acid targeted for depletion and leaves the nucleic acid of interest intact. The resulting product is then contacted with exonuclease BAL-31. Exonuclease BAL-31 has two activities: double-stranded DNA exonuclease activity and single-stranded DNA / RNA endonuclease activity. Due to its double-stranded DNA exonuclease activity, BAL-31 can degrade DNA from the open ends of both strands, reducing the size of the double-stranded DNA. Longer incubation times result in a greater reduction in the size of double-stranded DNA, aiding in the depletion of medium- to large-sized DNA fragments (>200 bp). In some embodiments, the 3' ends of the nucleic acids are poly(dG)-tailed using terminal transferase. It has been noted that the single-stranded endonuclease activity of BAL-31 allows for very rapid digestion of poly-A, -C, or -T, but very slow digestion of poly-G. Because of this property, adding single-stranded poly-dG to the 3' ends of the library serves as protection from degradation by BAL-31. As a result, poly(dG)-tailed DNA molecules cleaved by modification-sensitive restriction enzymes can be degraded by BAL-31, whereas intact DNA libraries are not digested by BAL-31 due to the poly(dG) protection and / or lack of terminal phosphate at the 3' end.
[0152] In some embodiments of the disclosed methods, the method includes contacting a sample with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated ends of nucleic acids. In some embodiments, the nucleic acids in the sample are terminally dephosphorylated. In some embodiments, contacting the sample with the exonuclease includes cleaving the nucleic acids in the sample with a modification-sensitive restriction enzyme that exposes terminal phosphates on the ends of the cleaved nucleic acids in the sample, followed by contacting the sample with the exonuclease. In some embodiments, the nucleic acids in the sample with exposed terminal phosphates comprise nucleic acids targeted for depletion. In some embodiments, the exonuclease depletes targeted nucleic acids from the sample by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.
[0153] adapter The present disclosure provides adaptors ligated to the 5' and 3' ends of nucleic acids in a sample or nucleic acids of interest. In some embodiments of the disclosed methods, adaptors ligate to all nucleic acids in a sample, and then use differences in nucleotide modifications to selectively cleave nucleic acids targeted for depletion, generating nucleic acids of interest ligated to both ends with adaptors and nucleic acids targeted for depletion ligated to one end with an adaptor (Figures 3 and 4). In some embodiments, differences in nucleotide modifications are used to selectively deplete nucleic acids targeted for depletion, and then adaptors are ligated to the nucleic acids of interest (Figure 2). In some embodiments, differences in nucleotide modifications are used to generate nucleic acids of interest with exposed terminal phosphates, which are used to ligate adaptors to the nucleic acids of interest (Figure 1).
[0154] In some embodiments of the disclosed methods, adaptors are ligated to the 5' and 3' ends of nucleic acids in a sample. In some embodiments, the adaptors further comprise an intervening sequence between the 5' and / or 3' ends. For example, the adaptors can further comprise a barcode sequence.
[0155] In some embodiments, an adaptor is a nucleic acid that is capable of ligating to both strands of a double-stranded DNA molecule.
[0156] In some embodiments, the adaptors are ligated prior to depletion / enrichment, while in other embodiments, the adaptors are ligated at a later step.
[0157] In some embodiments, the adaptors are linear. In some embodiments, the adaptors are linear Y-shaped. In some embodiments, the adaptors are linear circular. In some embodiments, the adaptors are hairpin adaptors. In some embodiments, the adaptors comprise a poly-G sequence.
[0158] In various embodiments, the adaptor can be a hairpin adaptor, i.e., a molecule that base pairs with itself to form a structure having a double-stranded stem and loop, with the 3' and 5' ends of the molecule 5'-linking to the 5' and 3' ends of a fragment of a double-stranded DNA molecule.
[0159] Alternatively, the adaptor may be a Y adaptor ligated to one or both ends of the fragment, also referred to as a universal adaptor. Alternatively, the adaptor itself may be composed of two separate oligonucleotide molecules that base-pair to each other. Furthermore, the ligatable end of the adaptor may be designed to be compatible with the overhang created by restriction enzyme cleavage, or may have a blunt end or a 5' T overhang. In some embodiments, the restriction enzyme is a modification-sensitive restriction enzyme.
[0160] Adapters can include double-stranded and single-stranded molecules. Thus, adapters can be DNA or RNA, or a mixture of the two. RNA-containing adapters can be cleaved by RNase treatment or alkaline hydrolysis.
[0161] The adapters can be 10-100 bp in length, although adapters outside this range can be used without departing from the present disclosure. In certain embodiments, the adapters are at least 10 bp, at least 15 bp, at least 20 bp, at least 25 bp, at least 30 bp, at least 35 bp, at least 40 bp, at least 45 bp, at least 50 bp, at least 55 bp, at least 60 bp, at least 65 bp, at least 70 bp, at least 75 bp, at least 80 bp, at least 85 bp, at least 90 bp, or at least 95 bp in length.
[0162] In some embodiments, the nucleic acid of interest and the nucleic acid adapter junction targeted for depletion are about 20 to about 5000 bp in length, about 20 to about 1000 bp in length, about 20 to about 500 bp in length, about 20 to about 400 bp in length, about 20 to about 300 bp in length, about 20 to about 200 bp in length, about 20 to 100 bp in length, about 50 to about 5000 bp in length, about 50 to about 1000 bp in length p, about 50 to about 500 bp in length, about 50 to about 400 bp in length, about 50 to about 300 bp in length, about 50 to about 200 bp in length, about 50 to about 100 bp in length, about 100 to about 5,000 bp in length, about 100 to about 1,000 bp in length, about 100 to about 500 bp in length, about 100 to about 400 bp in length, about 100 to about 300 bp in length, about 100 to about 200 bp in length. In some embodiments, the adaptor-linked nucleic acids of interest and the nucleic acids targeted for depletion are in the range of about 50 to about 1,000 bp in length. In some embodiments, the adaptor-linked nucleic acids of interest and the nucleic acids targeted for depletion are in the range of about 50 to about 500 bp in length. In some embodiments, the adaptor-linked nucleic acids of interest and the nucleic acids targeted for depletion are in the range of about 100 to about 500 bp in length. In some embodiments, the adaptor-ligated nucleic acids of interest and nucleic acids targeted for depletion range in length from about 50 to about 300 bp.
[0163] In some embodiments, the adaptor may comprise an oligonucleotide designed to match the nucleotide sequence of a specific region of a host genome, for example, a chromosomal region whose sequence has been deposited in NCBI's GenBank database or other database. Such oligonucleotides can be used in assays using a sample containing a test genome, where the test genome contains a binding site for the oligonucleotide. In a further example, the fragmented nucleic acid sequences can be derived from one or more DNA sequencing libraries. The adaptor can be configured for next-generation sequencing platforms, for example, for use with Illumina sequencing platforms, IonTorrents platforms, or nanopore technology.
[0164] In some embodiments, the adapter comprises a sequencing adapter (e.g., an Illumina sequencing adapter). In some embodiments, the adapter comprises a unique molecular identifier (UMI) sequence. In some embodiments, the UMI sequence comprises a sequence unique to each original nucleic acid molecule (e.g., a random sequence). This allows for quantification of nucleic acid amounts without sequencing bias. In some embodiments, the adapter comprises a "barcode" sequence. In some embodiments, the barcode sequence comprises a barcode sequence shared among nucleic acid molecules from a particular source (e.g., subject, patient, environmental sample, partition (e.g., droplet, well, bead), etc.). This allows for pooling of sequencing information for subsequent analysis, enabling detection and elimination of cross-contamination. In some embodiments, the adapter comprises multiple distinct sequences, such as a UMI unique to each nucleic acid molecule, a barcode shared among nucleic acid molecules from a particular source, and sequencing adapters.
[0165] depletion Nucleic acids targeted for depletion can be depleted by a variety of techniques.
[0166] Nucleic acids targeted for depletion can be depleted by the attachment of different adapters. In some embodiments, adapters are attached to the nucleic acids of a sample, and then one or more adapters are removed from the nucleic acids targeted for depletion based on their modification state. For example, nucleic acids targeted for depletion with adapters attached to both ends are cleaved with a modification-sensitive restriction enzyme, thereby generating nucleic acids targeted for depletion with an adapter attached to only one end. A subsequent step (e.g., amplification) can be used to target only the nucleic acids with adapters attached to both ends, thereby depleting the nucleic acids targeted for depletion. In another example, the nucleic acids of a sample are treated (e.g., by dephosphorylation) so that only the cleaved nucleic acids can have adapters attached; subsequently, the nucleic acids of interest can be cleaved with a modification-sensitive restriction enzyme (e.g., thereby exposing phosphate groups) and adapters can be attached. A subsequent step (e.g., amplification) can be used to target only the nucleic acids with adapters attached, thereby depleting the nucleic acids targeted for depletion.
[0167] Nucleic acids targeted for depletion can be depleted by digestion. For example, the nucleic acids of a sample can be treated (e.g., by dephosphorylation) so that only the cleaved nucleic acids can be digested (e.g., by exonuclease). The nucleic acids targeted for depletion can be cleaved with a modification-sensitive restriction enzyme, thereby making them digestible. Subsequent digestion, such as with an exonuclease, can then be used to deplete the nucleic acids targeted for depletion.
[0168] Nucleic acids targeted for depletion can be depleted by size selection, for example, a modification-sensitive restriction enzyme can be used to cleave either the nucleic acid of interest or the nucleic acid targeted for depletion, and the nucleic acid of interest can then be separated from the nucleic acid targeted for depletion based on the size difference resulting from cleavage.
[0169] In some cases, nucleic acids targeted for depletion are depleted without the use of size selection.
[0170] Nucleic acids targeted for depletion can be depleted by targeted binding. For example, a modification-sensitive binding domain (e.g., a methylation-sensitive antibody or a DNA-binding domain) can be used to bind and separate either the nucleic acid targeted for depletion or the nucleic acid of interest based on their modification state. As used herein, a "modification-sensitive binding domain" refers to a protein, protein fragment, or fusion protein that binds to a nucleic acid in a modification-sensitive manner but does not cleave the nucleic acid, unlike the modification-sensitive restriction enzymes disclosed herein. "Modification-sensitive target binding" refers to the binding of a nucleic acid by a modification-sensitive binding domain. In some exemplary embodiments, the binding of the modification-sensitive binding domain to a nucleic acid is sufficiently stable to allow selective binding of either the nucleic acid targeted for depletion or the nucleic acid of interest, followed by purification, e.g., co-immunoprecipitation, or binding of the modification-sensitive binding domain to beads or a column.
[0171] In some cases, the nucleic acid targeted for depletion is depleted without using modification-sensitive target binding. In some cases, the nucleic acid targeted for depletion is depleted without using CpG-sensitive target binding.
[0172] method Protocol 1: An exemplary method for the applications described herein is shown in Figure 1. A nucleic acid sample containing a nucleic acid of interest (101) and a nucleic acid targeted for depletion (102) is terminally dephosphorylated (105) to produce an unphosphorylated nucleic acid of interest (106) and a nucleic acid targeted for depletion (107). In some embodiments, the nucleic acid is fragmented prior to dephosphorylation. In some embodiments, the nucleic acid in the sample is terminally dephosphorylated with a phosphatase, such as recombinant shrimp alkaline phosphatase (rSAP). In some embodiments, both the nucleic acid of interest and the nucleic acid targeted for depletion contain one or more recognition sites for a modification-sensitive restriction enzyme (103, 104, respectively). In the nucleic acid of interest, the recognition site for the modification-sensitive restriction enzyme either does not contain modified nucleotides (103) or contains modified nucleotides less frequently than the corresponding recognition site in the nucleic acid targeted for depletion. In the nucleic acid targeted for depletion, the recognition site of the modification-sensitive restriction enzyme contains modified nucleotides within or adjacent to the restriction site (104), or contains modified nucleotides more frequently than the corresponding recognition site in the nucleic acid of interest. The activity of the modification-sensitive restriction enzyme (109) is blocked by the presence of modified nucleotides within or adjacent to its cognate recognition site (108), thereby targeting the activity of the modification-sensitive restriction enzyme to the nucleic acid of interest (compare 110 and 111). In some embodiments, the modification-sensitive restriction enzyme (109) comprises AatII, AccII, Aor13HI, Aor51HI, BspT104I, BssHII, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, MluI, NaeI, NotI, NruI, NsbI, PmaCI, Pspl406I, PvuI, SacII, SalI, SmaI, SnaBI, AluI, or Sau3AI. In some embodiments, the modification-sensitive restriction enzyme (109) comprises AluI or Sau3AI. Digestion of the sample with the modification-sensitive restriction enzyme (113) produces nucleic acids of interest with terminal phosphates at the 5' and 3' ends of the terminal phosphate (114). These terminal phosphates are used to ligate adapters (115, ligation step; 116, adapter) to the ends of the nucleic acid of interest, generating a nucleic acid of interest that is adapter-ligated at both ends (117).In contrast, nucleic acids targeted for depletion are not adapter-linked (111). These adapters can be used for downstream applications, such as adapter-mediated PCR amplification, sequencing (e.g., high-throughput sequencing), quantification of nucleic acids of interest in a sample, and / or cloning. This depletion of nucleic acids targeted for depletion occurs by selectively ligating adapters to the nucleic acids of interest. This depletion can be performed without using size selection. Alternatively, adapter-linked nucleic acids of interest are used in one or more additional enrichment methods described herein. For example, adapter-linked nucleic acids are used in additional modification-dependent enrichment methods of the present disclosure (e.g., the method shown in Figure 3). Alternatively, or in addition, adapter-linked nucleic acids are used in nucleic acid-guided nuclease-based enrichment methods of the present disclosure (e.g., the method shown in Figure 4).
[0173] Protocol 2: An exemplary method for the applications described herein is shown in Figure 2. A nucleic acid sample containing a nucleic acid of interest (201) and a nucleic acid targeted for depletion (202) is terminally dephosphorylated (205) to produce an unphosphorylated nucleic acid of interest (206) and a nucleic acid targeted for depletion (207). In some embodiments, the nucleic acid is fragmented prior to dephosphorylation. In some embodiments, the nucleic acid in the sample is terminally dephosphorylated with a phosphatase, such as recombinant shrimp alkaline phosphatase (rSAP). In some embodiments, both the nucleic acid of interest and the nucleic acid targeted for depletion contain one or more recognition sites for a modification-sensitive restriction enzyme (203, 204, respectively). In the nucleic acid of interest, the recognition site for the modification-sensitive restriction enzyme either does not contain modified nucleotides (203) or contains modified nucleotides less frequently than the corresponding recognition site in the nucleic acid targeted for depletion. In the nucleic acid targeted for depletion, the recognition site of the modification-sensitive restriction enzyme contains modified nucleotides within or adjacent to the restriction site (204), or contains modified nucleotides more frequently than the corresponding recognition site in the nucleic acid of interest. The modification-sensitive restriction enzyme (209) cleaves its cognate recognition site (208) if one or more modified nucleotides are present within or adjacent to it, and does not cleave its cognate recognition site if the recognition site does not contain one or more modified nucleotides (208), thereby targeting the activity of the modification-sensitive restriction enzyme to the nucleic acid targeted for depletion (compare 210 and 211). In some embodiments, the modification-sensitive restriction enzyme comprises AbaSI, FspEI, LpnPI, MspJI, or McrBC. In some embodiments, the modification-sensitive restriction enzyme is FspEI. In some embodiments, the modification-sensitive restriction enzyme is MspJI. Digestion of a sample with a modification-sensitive restriction enzyme (212) produces nucleic acids targeted for depletion that have terminal phosphates at one end (213) or both the 5' and 3' ends of the nucleic acid (214). In contrast, nucleic acids of interest that were not cleaved by the modification-sensitive restriction enzyme do not have exposed terminal phosphates at the 5' and / or 3' ends of the nucleic acid (compare 210 with 213-214).The sample is then digested with an exonuclease (215, digestion step; 216, exonuclease). The exonuclease uses the terminal phosphate of the nucleic acid targeted for depletion to remove consecutive nucleotides from the end of the nucleic acid molecule, thus depleting the nucleic acid targeted for depletion from the sample. This depletion can be performed without size selection. Following exonuclease digestion, an adapter is ligated to the nucleic acid of interest (217) that lacks the terminal phosphate and has not been digested by the exonuclease. This generates a nucleic acid of interest that is adapter-linked at both ends (218). These adapters can be used for downstream applications, such as adapter-mediated PCR amplification, sequencing (e.g., high-throughput sequencing), quantification of the nucleic acid of interest in the sample, and / or cloning. Alternatively, the adapter-linked nucleic acid of interest can be used in one or more additional enrichment methods described herein. For example, the adapter-linked nucleic acid can be used in additional modification-dependent enrichment methods of the present disclosure (e.g., the method shown in FIG. 3). Alternatively, or in addition, the adaptor-ligated nucleic acids are used in the nucleic acid-guided, nuclease-based enrichment methods of the present disclosure (eg, the method shown in FIG. 4).
[0174] Protocol 3: An exemplary method for use as described herein is shown in Figure 3. A nucleic acid sample containing a nucleic acid of interest (301) and a nucleic acid targeted for depletion (302) is adapter-ligated (305) or used in an enrichment method (306) of the present disclosure (e.g., the method shown in Figure 1 or Figure 2) to generate an adapter-ligated nucleic acid of interest (307) and an adapter-ligated nucleic acid targeted for depletion (308). In some embodiments, both the nucleic acid of interest and the nucleic acid targeted for depletion contain one or more recognition sites (303, 304, respectively) for a modification-sensitive restriction enzyme. In the nucleic acid of interest, the recognition site for the modification-sensitive restriction enzyme either does not contain modified nucleotides (303) or contains modified nucleotides less frequently than the corresponding recognition site in the nucleic acid targeted for depletion. In the nucleic acid targeted for depletion, the recognition site for the modification-sensitive restriction enzyme either contains modified nucleotides within or adjacent to the restriction site (304) or contains modified nucleotides more frequently than the corresponding recognition site in the nucleic acid of interest. The modification-sensitive restriction enzyme (309) cleaves its cognate recognition site (308) if one or more modified nucleotides are present within or adjacent to the recognition site, and does not cleave its cognate recognition site if the recognition site does not contain one or more modified nucleotides (308), thereby targeting its activity to nucleic acids targeted for depletion (compare 310 and 311). In some embodiments, the modification-sensitive restriction enzyme comprises AbaSI, FspEI, LpnPI, MspJI, or McrBC. In some embodiments, the modification-sensitive restriction enzyme is FspEI. In some embodiments, the modification-sensitive restriction enzyme is MspJI. The sample is digested with the modification-sensitive restriction enzyme (311) to generate nucleic acids targeted for depletion that are not adapter-ligated (312) or are adapter-ligated at only one end (313). This selectively removes adapters from the nucleic acids targeted for depletion, thereby depleting the nucleic acids targeted for depletion. This depletion can be performed without using size selection. In contrast, nucleic acids of interest that were not cleaved by the modification-sensitive restriction enzyme are adaptor-ligated at both ends (controls 310 and 312-313).These adapters can be used in downstream applications, such as adapter-mediated PCR amplification, sequencing (e.g., high-throughput sequencing), quantification of nucleic acids of interest in a sample, and / or cloning.
[0175] Protocol 4: An exemplary method for the applications described herein is shown in Figure 4. Multiple gNAs (401) are used to target a nucleic acid-guided nuclease (402) to nucleic acids targeted for depletion (403) in a sample of adapter-ligated nucleic acids. The adapter-ligated nucleic acids are generated by any of the enrichment methods described herein, which use a modification-sensitive restriction enzyme to deplete the nucleic acids targeted for depletion from the sample, either before or after the initial adapter ligation. In this method, the gNAs specifically target the nucleic acids targeted for depletion (403) and not the nucleic acid of interest (404), and are therefore not cleaved by the nucleic acid-guided nuclease (402). Cleavage by the nucleic acid-guided nuclease results in a nucleic acid targeted for depletion (405) that is adapter-ligated at one end and a nucleic acid of interest (403) that is adapter-ligated at both ends. These adapters can be used in downstream applications, such as adapter-mediated PCR amplification, sequencing (e.g., high-throughput sequencing), quantification of nucleic acids of interest in a sample, and cloning.
[0176] Protocol 5: In some embodiments, the nucleic acid-guided nuclease is a nucleic acid-guided nickase. Multiple gNAs are used to target nucleic acids targeted for depletion in a sample of adapter-ligated nucleic acids. The adapter-ligated nucleic acids are generated by any of the enrichment methods described herein, in which a modification-sensitive restriction enzyme is used to deplete the nucleic acids targeted for depletion from the sample, either before or after the initial adapter ligation. In some embodiments, the multiple gNAs are designed so that all nucleic acids targeted for depletion have two gNA binding sites in close proximity (e.g., less than 15 bases apart) on opposite DNA strands of the double-stranded DNA targeted for depletion. In this embodiment, the nucleic acid-guided nickase recognizes its target site on the DNA to be removed and can cleave only one strand. To deplete DNA, two separate nucleic acid-guided nickases can cleave both strands of the DNA and deplete it in close proximity. Only the DNA to be depleted has two nucleic acid-guided nickase sites in close proximity, creating a double-strand break. If a nucleic acid-induced nickase, such as a CRISPR / Cas system protein nickase, recognizes a site on the target DNA nonspecifically or with low affinity, it can cleave only one strand and will not interfere with subsequent PCR amplification or downstream processing of the DNA molecule. In this embodiment, the possibility of two gNAs nonspecifically recognizing two sites in close enough proximity is negligible ( <lxl0 -14 This embodiment would be particularly useful when the normal CRISPR / Cas system protein-mediated cleavage cuts an excess of the DNA of interest.
[0177] Protocol 6: In some embodiments, the nucleic acid-guided nuclease is catalytically dead, and the method includes partitioning a nucleic acid targeted for depletion and a nucleic acid of interest in a sample. Multiple gNAs are used to target a catalytically dead nucleic acid-guided nuclease (e.g., dCas9 or dCpf1) to either the nucleic acid targeted for depletion or the nucleic acid of interest in a sample of adapter-ligated nucleic acids. The adapter-ligated nucleic acids are generated by any of the enrichment methods described herein, using a modification-sensitive restriction enzyme to deplete the nucleic acid targeted for depletion from the sample, either before or after the initial adapter ligation. The catalytically inactive nucleic acid-guided nuclease can bind to nucleic acids but cannot nick or cleave the nucleic acid. In some embodiments, the catalytically inactive nucleic acid-guided nuclease includes a tag, such as a biotin tag, that can be used to isolate the catalytically inactive nucleic acid-guided nuclease and any molecule to which it binds. In these embodiments, multiple gNAs are developed that hybridize to either the nucleic acid of interest or the nucleic acid targeted for depletion, but not both. The multiple gNAs and a catalytically inactive nucleic acid-guided nuclease are then contacted with the sample, allowing the catalytically inactive nucleic acid nuclease to bind to either the nucleic acid of interest or the nucleic acid targeted for depletion, depending on the design of the gNA. Instead of cleaving the target sequence, this method is used to divide the fragmented nucleic acid sample into two fractions that can be processed separately. Thus, the catalytically inactive nucleic acid-guided nuclease divides the mixture into unbound fragments (e.g., the nucleic acid of interest) and bound fragments (e.g., the nucleic acid targeted for depletion to which the gNA is targeted). The bound portion of the target nucleic acid sample is removed by binding of an affinity tag (e.g., biotin) previously attached to the catalytically inactive nucleic acid-guided nuclease protein. The bound nucleic acid sequence is eluted from the protein / gNA complex using denaturing conditions and amplified and sequenced. Similarly, unlinked nucleic acid sequences can be amplified and sequenced.
[0178] Any of the methods described herein can be used as an independent method to deplete nucleic acids targeted for depletion from a sample, thereby enriching for the nucleic acid of interest.
[0179] Alternatively, the methods described herein can be combined to achieve greater enrichment than any individual method alone. In some embodiments, a sample is first enriched using Protocol 1, followed by enrichment using Protocol 2. In some embodiments, a sample is first enriched using Protocol 1, followed by enrichment using Protocol 3. In some embodiments, a sample is first enriched using Protocol 1, followed by enrichment using Protocols 2 and 3. In some embodiments, a sample is first enriched using Protocol 1, followed by enrichment using any one of Protocols 4-6. In some embodiments, a sample is first enriched using Protocol 1, followed by enrichment using Protocols 2 and / or 3, and any one of Protocols 4-6.
[0180] Although particular combinations of methods and the order of combination of methods are described herein, these are in no way intended to limit the ways in which the methods of this disclosure can be combined. Any method of enriching a sample for a nucleic acid of interest of this disclosure that produces an adaptor-ligated nucleic acid of interest as the product of the method can be combined with any additional method of this disclosure that uses an adaptor-ligated nucleic acid as its starting substrate.
[0181] Nucleic acid-guided nuclease-based enrichment method In some embodiments of the disclosed methods, the modification-based enrichment method of the present disclosure is combined with a nucleic acid-guided nuclease-based enrichment method, which uses a nucleic acid-guided nuclease to enrich a sample for a sequence of interest. Nucleic acid-guided nuclease-based enrichment methods are described in WO / 2016 / 100955, WO / 2017 / 031360, WO / 2017 / 100343, WO / 2017 / 147345, and WO / 2018 / 227025, the contents of each of which are incorporated herein by reference in their entirety.
[0182] In some embodiments, the disclosed modification-based enrichment methods and nucleic acid-guided nuclease-based enrichment methods deplete different nucleic acids in a sample, thereby achieving a greater enrichment of the nucleic acid of interest than either method alone. For example, a sample contains a nucleic acid targeted for depletion from a mammalian host genome and a nucleic acid of interest from one or more non-host genomes (e.g., bacteria, viruses, or parasites). The disclosed methods are used to enrich the nucleic acid of interest in this sample, with modification-based enrichment methods that exploit differences in CpG methylation between the host and non-host nucleic acids being selected to deplete nucleic acids containing actively transcribed regions of the mammalian host genome, while nucleic acid-guided nuclease-based enrichment methods effectively target regions of repetitive sequences in the mammalian host genome using a library of guide nucleic acids (gNAs) that target those regions.
[0183] The term "nucleic acid-guided nuclease-gNA complex" refers to a complex comprising a nucleic acid-guided nuclease protein and a guide nucleic acid (gNA, e.g., gRNA or gDNA). For example, a "Cas9-gRNA complex" refers to a complex comprising a Cas9 protein and a guide RNA (gRNA). The nucleic acid-guided nuclease can be any type of nucleic acid-guided nuclease, including, but not limited to, a wild-type nucleic acid-guided nuclease, a catalytically inactive nucleic acid-guided nuclease, or a nucleic acid-guided nuclease-nickase.
[0184] Multiple gNAs Provided herein are a plurality of guide nucleic acids (gNAs) (interchangeably referred to as a library or collection).
[0185] The term "guide nucleic acid" refers to a guide nucleic acid (gNA) that can form a complex with a nucleic acid-guided nuclease and, optionally, additional nucleic acid(s). The gNA can be present as an isolated nucleic acid or as part of a nucleic acid-guided nuclease-gNA complex, e.g., a Cas9-gRNA complex.
[0186] As used herein, a plurality of gNAs is at least 10 2 In some embodiments, the plurality of gNAs comprises at least 10 unique gNAs. 2 unique gNAs, at least 10 3 unique gNAs, at least 10 4 unique gNAs, at least 10 5 unique gNAs, at least 10 6 unique gNAs, at least 10 7 unique gNAs, at least 10 8 unique gNAs, at least 10 9 unique gNAs, or at least 10 10 In some embodiments, the collection of gNAs comprises at least 10 unique gNAs in total. 2 unique gNAs, at least 10 3 unique gNAs, at least 10 4 unique gNAs, or at least 10 5 Contains unique gNAs.
[0187] In some embodiments, the collection of gNAs comprises a first NA segment comprising a target sequence and a second NA segment comprising a nucleic acid-guided nuclease system (e.g., CRISPR / Cas system) protein-binding sequence. In some embodiments, the first and second segments are in 5' to 3' order. In some embodiments, the first and second segments are in 3' to 5' order.
[0188] In some embodiments, the size of the first segment is 12 to 250 bp, or 12 to 100 bp, or 12 to 75 bp, or 12 to 50 bp, or 12 to 30 bp, or 12 to 25 bp, or 12 to 22 bp, or 12 to 20 bp, or 12 to 18 bp, or 12 to 16 bp, or 14 to 250 bp, or 14 to 100 bp, or 14 to 75 bp, or 14 to 50 bp across multiple gNAs. p, or 14-30bp, or 14-25bp, or 14-22bp, or 14-20bp, or 14-18bp, or 14-17bp, or 14-16bp, or 15-250bp, or 15-100bp, or 15-75bp, or 15-50bp, or 15-30bp, or 15-25bp, or 15-22bp, or 15-20bp, or 15-18bp, or 15-17bp, or 15 ~16bp, or 16-250bp, or 16-100bp, or 16-75bp, or 16-50bp, or 16-30bp, or 16-25bp, or 16-22bp, or 16-20bp, or 16-18bp, or 16-17bp, or 17-250bp, or 17-100bp, or 17-75bp, or 17-50bp, or 17-30bp, or 17-25bp, or 17-22bp p, or 17 to 20 bp, or 17 to 18 bp, or 18 to 250 bp, or 18 to 100 bp, or 18 to 75 bp, or 18 to 50 bp, or 18 to 30 bp, or 18 to 25 bp, or 18 to 22 bp, or 18 to 20 bp, or 19 to 250 bp, or 19 to 100 bp, or 19 to 75 bp, or 19 to 50 bp, or 19 to 30 bp, or 19 to 25 bp, or 19 to 22 bp. In some embodiments, the size of the first segment is 15-250 bp, or 30-100 bp, or 20-30 bp, or 22-30 bp, or 15-50 bp, or 15-75 bp, or 15-100 bp, or 15-125 bp, or 15-150 bp, or 15-175 bp, or 15-200 bp, or 15-225 bp, or 15-250 bp, or 22-50 bp, or 22-75 bp, or 22-100 bp, or 22-125 bp, or 22-150 bp, or 22-175 bp, or 22-200 bp, or 22-225 bp, or 22-250 bp across multiple gNAs.
[0189] In some embodiments, at least 10%, or at least 15%, or at least 20%, or at least 25%, or at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or 100% of the plurality of first segments are 15-50 bp.
[0190] In some embodiments, at least 10%, or at least 15%, or at least 20%, or at least 25%, or at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or 100% of the first segments of the collection are 15-20 bp.
[0191] In some specific embodiments, the size of the first segment is 15 bp. In some specific embodiments, the size of the first segment is 16 bp. In some specific embodiments, the size of the first segment is 17 bp. In some specific embodiments, the size of the first segment is 18 bp. In some specific embodiments, the size of the first segment is 19 bp. In some specific embodiments, the size of the first segment is 20 bp.
[0192] In some embodiments, the target sequences of the gNAs and / or gNAs in the plurality of gRNAs comprise unique 5' ends. In some embodiments, the plurality of gNAs exhibit variability in the sequence at the 5' end of the target sequence across the plurality of members. In some embodiments, the plurality of gNAs exhibits at least 5%, or at least 10%, or at least 15%, or at least 20%, or at least 25%, or at least 30%, or at least 35%, or at least 40%, or at least 45%, or at least 50%, or at least 55%, or at least 60%, or at least 65%, or at least 70%, or at least 75% variability in the sequence at the 5' end of the target sequence across the plurality of members.
[0193] In some embodiments, the 3' end of the gNA target sequence can be any purine or pyrimidine (and / or modified versions thereof). In some embodiments, the 3' end of the gNA target sequence is adenine. In some embodiments, the 3' end of the gNA target sequence is guanine. In some embodiments, the 3' end of the gNA target sequence is cytosine. In some embodiments, the 3' end of the gNA target sequence is uracil. In some embodiments, the 3' end of the gNA target sequence is thymine. In some embodiments, the 3' end of the gNA target sequence is not cytosine.
[0194] In some embodiments, the plurality of gNAs comprise target sequences capable of base pairing with target sequences in nucleic acids targeted for depletion, and the target sequences in the nucleic acids targeted for depletion are located at least every 1 bp, at least every 2 bp, at least every 3 bp, at least every 4 bp, at least every 5 bp, at least every 6 bp, at least every 7 bp, at least every 8 bp, at least every 9 bp, at least every 10 bp, at least every 11 bp, at least every 12 bp, at least every 13 bp, at least every 14 bp, at least every 15 bp, at least every 16 bp, at least every 17 bp, at least every 18 bp, at least every 19 bp, at least every 20 bp, at least every 25 bp, at least every 30 bp, at least every 31 bp, at least every 32 bp, at least every 33 bp, at least every 34 bp, at least every 35 bp, at least every 36 bp, at least every 37 bp, at least every 38 bp, at least every 39 bp, at least every 40 bp, at least every 41 bp, at least every 42 bp, at least every 43 bp, at least every 44 bp, at least every 45 bp, at least every 46 bp, at least every 47 bp, at least every 48 bp, at least every 49 bp, at least every 50 bp, at least every 51 bp, at least every 52 bp, at least every 53 bp, at least every 54 bp, at least every 55 bp, at least every 56 bp, at least every 57 bp, at least every 58 bp, at least every 59 bp, at least every 60 bp, at least every 61 spaced every 0 bp, at least every 40 bp, at least every 50 bp, at least every 100 bp, at least every 200 bp, at least every 300 bp, at least every 400 bp, at least every 500 bp, at least every 600 bp, at least every 700 bp, at least every 800 bp, at least every 900 bp, at least every 1000 bp, at least every 2500 bp, at least every 5000 bp, at least every 10,000 bp, at least every 15,000 bp, at least every 20,000 bp, at least every 25,000 bp, at least every 50,000 bp, at least every 100,000 bp, at least every 250,000 bp, at least every 500,000 bp, at least every 750,000 bp, or at least every 1,000,000 bp.
[0195] In some embodiments, the plurality of gNAs comprises a first NA segment comprising a target sequence and a second NA segment comprising a nucleic acid-guided nuclease system (e.g., a CRISPR / Cas system) protein binding sequence, and the plurality of gNAs can have different second NA segments with different specificities for protein members of the nucleic acid-guided nuclease system (e.g., a CRISPR / Cas system). For example, the collection of gNAs provided herein can include members whose second segment comprises a nucleic acid-guided nuclease system (e.g., a CRISPR / Cas system) protein binding sequence specific for a first nucleic acid-guided nuclease system (e.g., a CRISPR / Cas system) protein, and members whose second segment comprises a nucleic acid-guided nuclease system (e.g., a CRISPR / Cas system) protein binding sequence specific for a second nucleic acid-guided nuclease system (e.g., a CRISPR / Cas system) protein, wherein the first and second nucleic acid-guided nuclease system (e.g., a CRISPR / Cas system) proteins are not the same. In some embodiments, the collection of gNAs provided herein includes members that exhibit specificity for at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 proteins of a nucleic acid-guided nuclease system (e.g., a CRISPR / Cas system). In one particular embodiment, the plurality of gNAs provided herein includes members that exhibit specificity for Cas9 and other proteins selected from the group consisting of Cpfl, Cas3, Cas8a-c, CaslO, CasX, CasY, Casl3, Casl4, Cse1, Csy1, Csn2, Cas4, Csm2, and Cm5. In some embodiments, the nucleic acid-guided nuclease system protein binding sequences specific for the first and second nucleic acid-guided nuclease system proteins are both 5' of the first NA segment comprising the target sequence.In some embodiments, both the nucleic acid-guided nuclease system protein binding sequences specific to the first and second nucleic acid-guided nuclease system proteins are 3' of the first NA segment containing the target sequence. In some embodiments, the nucleic acid-guided nuclease system protein binding sequence specific to the first nucleic acid-guided nuclease system (e.g., CRISPR / Cas system) protein is 5' of the first NA segment containing the target sequence, and the second nucleic acid-guided nuclease system protein binding sequence specific to the second nucleic acid-guided nuclease system protein is 3' of the first NA segment containing the target sequence. The order of the first NA segment containing the target sequence and the second NA segment containing the nucleic acid-guided nuclease system protein binding sequence will depend on the nucleic acid-guided nuclease system protein. The appropriate 5' to 3' arrangement of the first and second NA segments and the selection of the nucleic acid-guided nuclease system protein will be apparent to those skilled in the art.
[0196] In some embodiments, gNAs comprise DNA and RNA. In some embodiments, gNAs are comprised of DNA (gDNA). In some embodiments, gNAs are comprised of RNA (gRNA).
[0197] In some embodiments, the gRNA comprises a gRNA, and the gRNA comprises two subsegments encoding the crRNA and tracrRNA. In some embodiments, the crRNA does not comprise an additional sequence in addition to the target sequence that can hybridize to the tracrRNA. In some embodiments, the crRNA comprises an additional sequence that can hybridize to the tracrRNA. In some embodiments, the two subsegments are transcribed independently. In some embodiments, the two subsegments are transcribed as a single unit. In some embodiments, the DNA encoding the crRNA comprises a target sequence 5' of the sequence GTTTTAGAGCTATGCTGTTTTG (SEQ ID NO: 26). In some embodiments, the DNA encoding the tracrRNA comprises the sequence GGAACCATTCAAAACAGCATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT (SEQ ID NO: 27).
[0198] Target sequence As used herein, target sequence refers to the sequence that directs gNA to the target sequence in the nucleic acid that is targeted for depletion in sample.For example, target sequence targets specific sequence, for example, target sequence targets the repeat sequence in the genome that is targeted for depletion in sample.
[0199] Provided herein are gNAs and multiple gNAs that comprise a segment that includes a target sequence.
[0200] In some embodiments, the target sequence comprises or consists of DNA.
[0201] In some embodiments, the target sequence comprises or consists of RNA.
[0202] In some embodiments, the target sequence comprises RNA and shares at least 70% sequence identity, at least 75% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, or 100% sequence identity to the 5' end of the PAM sequence of the sequence of interest, excluding RNA that contains uracil instead of thymine. In some embodiments, the target sequence comprises RNA and shares at least 70% sequence identity, at least 75% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, or 100% sequence identity to the 3' end of the PAM sequence of the sequence of interest, excluding RNA that contains uracil instead of thymine. In some embodiments, the PAM sequence is AGG, CGG, TGG, GGG, or NAG. In some embodiments, the PAM sequence is TTN, TCN, or TGN.
[0203] In some embodiments, the target sequence comprises DNA and shares at least 70% sequence identity, at least 75% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, or 100% sequence identity to the 5' of the PAM sequence of the sequence of interest. In some embodiments, the target sequence comprises DNA and shares at least 70% sequence identity, at least 75% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, or 100% sequence identity to the 3' of the PAM sequence of the sequence of interest.
[0204] In some embodiments, the target sequence comprises RNA and is complementary to the strand opposite the sequence of nucleotide 5' of the PAM sequence. In some embodiments, the target sequence is at least 70% complementary, at least 75% complementary, at least 80% complementary, at least 85% complementary, at least 90% complementary, at least 95% complementary, or 100% complementary to the strand opposite the sequence of nucleotide 5' of the PAM sequence. In some embodiments, the target sequence comprises RNA and is complementary to the strand opposite the sequence of nucleotide 3' of the PAM sequence. In some embodiments, the target sequence is at least 70% complementary, at least 75% complementary, at least 80% complementary, at least 85% complementary, at least 90% complementary, at least 95% complementary, or 100% complementary to the strand opposite the sequence of nucleotide 3' of the PAM sequence. In some embodiments, the PAM sequence is AGG, CGG, TGG, GGG, or NAG. In some embodiments, the PAM sequence is TTN, TCN, or TGN.
[0205] In some embodiments, the target sequence comprises DNA and is complementary to the strand opposite the sequence of nucleotide 5' of the PAM sequence. In some embodiments, the target sequence is at least 70% complementary, at least 75% complementary, at least 80% complementary, at least 85% complementary, at least 90% complementary, at least 95% complementary, or 100% complementary to the strand opposite the sequence of nucleotide 5' of the PAM sequence. In some embodiments, the target sequence comprises DNA and is complementary to the strand opposite the sequence of nucleotide 3' of the PAM sequence. In some embodiments, the target sequence is at least 70% complementary, at least 75% complementary, at least 80% complementary, at least 85% complementary, at least 90% complementary, at least 95% complementary, or 100% complementary to the strand opposite the sequence of nucleotide 3' of the PAM sequence. In some embodiments, the PAM sequence is AGG, CGG, TGG, GGG, or NAG. In some embodiments, the PAM sequence is TTN, TCN, or TGN.
[0206] Different CRISPR / Cas system proteins recognize different PAM sequences. The PAM sequence can be located 5' or 3' of the target sequence. For example, Cas9 can recognize an NGG PAM at the 3' end of the target sequence. Cpfl can recognize a TTN PAM at the 5' end of the target sequence. All PAM sequences recognized by all CRISPR / Cas system proteins are contemplated within the scope of this disclosure. It will be readily apparent to one of skill in the art which PAM sequences are compatible with a particular CRISPR / Cas system protein.
[0207] Nucleic acid-induced nuclease Provided herein are gNAs and gNAs comprising a segment containing a nucleic acid-guided nuclease protein-binding sequence. The nucleic acid-guided nuclease can be a nucleic acid-guided nuclease system protein (e.g., a CRISPR / Cas system). The nucleic acid-guided nuclease system can be an RNA-guided nuclease system. The nucleic acid-guided nuclease system can be a DNA-guided nuclease system.
[0208] The methods of the present disclosure can utilize nucleic acid-guided nucleases. As used herein, a "nucleic acid-guided nuclease" is any nuclease that cleaves DNA, RNA, or DNA / RNA hybrids and uses one or more guide nucleic acids (gNAs) to confer specificity. Nucleic acid-guided nucleases include CRISPR / Cas system proteins and non-CRISPR / Cas system proteins.
[0209] The nucleic acid-guided nucleases provided herein can be DNA-guided DNA nucleases, DNA-guided RNA nucleases, RNA-guided DNA nucleases, or RNA-guided RNA nucleases. The nucleases can be endonucleases. The nucleases can be exonucleases. In one embodiment, the nucleic acid-guided nuclease is a nucleic acid-guided DNA endonuclease. In one embodiment, the nucleic acid-guided nuclease is a nucleic acid-guided RNA endonuclease.
[0210] A nucleic acid-guided nuclease protein binding sequence is a nucleic acid sequence that binds to any protein member of a nucleic acid-guided nuclease system. For example, a CRISPR / Cas protein binding sequence is a nucleic acid sequence that binds to any protein member of a CRISPR / Cas system.
[0211] In some embodiments, the nucleic acid-guided nuclease is selected from the group consisting of CAS class I type I, CAS class I type III, CAS class I type IV, CAS class II type II, and CAS class II type V. In some embodiments, the CRISPR / Cas system proteins comprise proteins from a CRISPR type I system, a CRISPR type II system, and a CRISPR type III system. In some embodiments, the nucleic acid-guided nuclease is selected from the group consisting of Cas9, Cpfl, Cas3, Cas8a-c, CaslO, Casl3, Casl4, Cse1, Csy1, Csn2, Cas4, Csm2, Cm5, Csfl, C2c2, CasX, CasY, Casl4, and NgAgo.
[0212] In some embodiments, the nucleic acid-guided nuclease system protein (e.g., a CRISPR / Cas system protein) can be derived from any bacterial or archaeal species.
[0213] In some embodiments, the nucleic acid-guided nuclease system protein (e.g., a CRISPR / Cas system protein) is selected from the group consisting of Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Streptococcus thermophiles, Treponema denticola, Francisella tularensis, Pasteurella multocida, Campylobacter jejuni, Campylobacter lari, Mycoplasma gallisepticum, and the like. gallisepticum, Nitratifractor salsuginis, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria cinerea, Gluconacetobacter diazotrophicus, Azospirillum, Sphaerochaeta globus, Flavobacterium columnare, Fluviicola taffensis, Bacteroides coprophilus, Mycoplasma mobile, Lactobacillus farciminis farciminis, Streptococcus pasteurianus, Lactobacillus johnsoniia nucleic acid-guided nuclease system protein (e.g., a CRISPR / Cas system protein) from or derived from Staphylococcus johnsonii, Staphylococcus pseudintermedius, Filifactor alocis, Legionella pneumophila, Suterella wadsworthensis, Corynebacter diphtheria, Acidaminococcus, Lachnospiraceae bacterium, or Prevotella.
[0214] In some embodiments, examples of nucleic acid-guided nuclease system (e.g., CRISPR / Cas system) proteins can be naturally occurring or engineered versions.
[0215] In some embodiments, naturally occurring nucleic acid-guided nuclease system (e.g., CRISPR / Cas system) proteins include Cas9, Cpf1, Cas3, Cas8a-c, Cas10, CasX, CasY, Cas13, Cas14, Csel, Csy1, Csn2, Cas4, Csm2, and Cm5. Engineered versions of such proteins can also be used.
[0216] In some embodiments, engineered examples of nucleic acid-guided nuclease (e.g., CRISPR / Cas) system proteins also include nucleic acid-guided nickases (e.g., Cas nickases). Nucleic acid-guided nickases refer to modified versions of nucleic acid-guided nuclease system proteins that contain a single inactive catalytic domain. In one embodiment, the nucleic acid-guided nickase is a Cas nickase, such as Cas9 nickase. Cas9 nickases can contain a single inactive catalytic domain, such as either a RuvC domain or an HNH domain. With only one active nuclease domain, the Cas9 nickase cleaves only one strand of the target DNA, creating a single-strand break or "nick." Depending on the mutant used, the hybridized or non-hybridized strand of the guided NA can be cleaved. A nucleic acid-guided nickase bound to two NAs targeting opposite strands creates a double-strand break in the target double-stranded DNA. This "dual nickase" strategy can enhance cleavage specificity because both nucleic acid-guided nuclease / gRNA (e.g., Cas9 / gRNA) complexes must bind specifically at the site before a double-stranded break is formed. Naturally occurring nickase nucleic acid-guided nuclease system proteins can also be used.
[0217] In some embodiments, engineered examples of nucleic acid-guided nuclease system proteins also include nucleic acid-guided nuclease system fusion proteins, for example, a nucleic acid-guided nuclease (e.g., CRISPR / Cas) system protein can be fused to another protein, such as an activator, repressor, nuclease, fluorescent molecule, radioactive tag, or transposase.
[0218] In some embodiments, the nucleic acid-guided nuclease system protein-binding sequence comprises a gNA (e.g., gRNA) stem-loop sequence.
[0219] Different CRISPR / Cas system proteins are compatible with different nucleic acid-guided nuclease system protein-binding sequences, and it will be readily apparent to one skilled in the art which CRISPR / Cas system proteins are compatible with which nucleic acid-guided nuclease system protein-binding sequences.
[0220] In some embodiments, the double-stranded DNA sequence encoding the gNA (e.g., gRNA) stem-loop sequence comprises the following DNA sequence on one strand (5'>3', GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT (SEQ ID NO: 28)) and the reverse complementary DNA on the other strand (5'>3', AAAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTCTAGCTCTAAAAC (SEQ ID NO: 29)).
[0221] In some embodiments, the single-stranded DNA sequence encoding the gNA (e.g., gRNA) stem-loop sequence comprises the following DNA sequence (5'>3', AAAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTCTAGCTCTAAAAC (SEQ ID NO: 29)), wherein the single-stranded DNA serves as a transcription template.
[0222] In some embodiments, the gNA (e.g., gRNA) stem-loop sequence comprises the following RNA sequence (5'>3', GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU (SEQ ID NO: 30)).
[0223] In some embodiments, the double-stranded DNA sequence encoding the gNA (e.g., gRNA) stem-loop sequence comprises the following DNA sequence on one strand (5'>3', GTTTTAGAGCTATGCTGGAAACAGCATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTTC (SEQ ID NO: 31)) and the reverse complementary DNA on the other strand (5'>3', GAAAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATGCTGTTTCCAGCATAGCTCTAAAAC (SEQ ID NO: 32)).
[0224] In some embodiments, the single-stranded DNA sequence encoding the gNA (e.g., gRNA) stem-loop sequence comprises the following DNA sequence (5'>3', GAAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATGCTGTTTCCAGCATAGCTCTAAAAC (SEQ ID NO: 32)), wherein the single-stranded DNA serves as a transcription template.
[0225] In some embodiments, the gNA (e.g., gRNA) stem-loop sequence comprises the following RNA sequence (5'>3', GUUUUAGAGCUAUGCUGGAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUUC (SEQ ID NO: 33)).
[0226] In some embodiments, the CRISPR / Cas system protein is a Cpf1 protein. In some embodiments, the Cpf1 protein is isolated or derived from a Francisella species or an Acidaminococcus species. In some embodiments, the gNA (e.g., gRNA) CRISPR / Cas system protein-binding sequence comprises the following RNA sequence (5'>3', AAUUUCUACUGUUGUAGAU (SEQ ID NO: 34)).
[0227] In some embodiments, the CRISPR / Cas system protein is a Cpf1 protein. In some embodiments, the Cpf1 protein is isolated or derived from a Franciscella species or an Acidaminococcus species. In some embodiments, the DNA sequence encoding the gNA (e.g., gRNA) CRISPR / Cas system protein-binding sequence comprises the following DNA sequence (5'>3', AATTTCTACTGTTGTAGAT (SEQ ID NO: 35)). In some embodiments, the DNA is single-stranded. In some embodiments, the DNA is double-stranded.
[0228] In some embodiments, provided herein are gNAs (e.g., gRNAs) comprising a first NA segment comprising a target sequence and a second NA segment comprising a nucleic acid-guided nuclease (e.g., CRISPR / Cas) system protein-binding sequence. In some embodiments, the size of the first segment is 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, or 20 bp. In some embodiments, the second segment comprises a single segment comprising a gRNA stem-loop sequence. In some embodiments, the gRNA stem-loop sequence comprises the following RNA sequence (5'>3', GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU (SEQ ID NO: 30)). In some embodiments, the gRNA stem-loop sequence comprises the following RNA sequence (5'>3', GUUUUAGAGCUAUGCUGGAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUUC (SEQ ID NO: 33)). In some embodiments, the second segment comprises two subsegments. The first RNA subsegment (crRNA) hybridizes to the second RNA subsegment (tracrRNA), which together direct nucleic acid-guided nuclease (e.g., CRISPR / Cas) system protein binding. In some embodiments, the sequence of the second subsegment comprises GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 36). In some embodiments, the first RNA segment and the second RNA segment together form the crRNA sequence. In some embodiments, the other RNA that hybridizes to the second RNA segment is tracrRNA. In some embodiments, the tracrRNA comprises the 5'>3' sequence GGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU (SEQ ID NO: 37).
[0229] In some embodiments, provided herein are gNAs (e.g., gRNAs) comprising a first NA segment comprising a target sequence and a second NA segment comprising a nucleic acid-guided nuclease (e.g., CRISPR / Cas) system protein-binding sequence. In some embodiments, for example, in embodiments where the CRISPR / Cas system protein is a Cpf1 system protein, the second segment is 5' of the first segment. In some embodiments, the size of the first segment is 20 bp. In some embodiments, the size of the first segment is greater than 20 bp. In some embodiments, the size of the first segment is greater than 30 bp. In some embodiments, the second segment comprises a single segment comprising a gRNA stem-loop sequence. In some embodiments, the gRNA stem-loop sequence comprises the following RNA sequence (5'>3', AAUUUCUACUGUUGUAGAU (SEQ ID NO: 34)).
[0230] CRISPR / Cas system nucleic acid-guided nuclease In some embodiments, CRISPR / Cas system proteins are used in the embodiments provided herein, hi some embodiments, CRISPR / Cas system proteins include proteins from CRISPR type I systems, CRISPR type II systems, and CRISPR type III systems.
[0231] In some embodiments, the CRISPR / Cas system proteins may be derived from any bacterial or archaeal species.
[0232] In some embodiments, the CRISPR / Cas system proteins are isolated, recombinantly produced, or synthetic.
[0233] In some embodiments, the CRISPR / Cas system protein is selected from the group consisting of Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Streptococcus thermophiles, Treponema denticola, Francisella tularensis, Pasteurella multocida, Campylobacter jejuni, Campylobacter lari, Mycoplasma gallisepticum, Nitratifractor sarsuginis, and the like. salsuginis, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria cinerea, Gluconacetobacter diazotrophicus, Azospirillum, Sphaerochaeta globus, Flavobacterium columnare, Fluviicola taffensis, Bacteroides coprophilus, Mycoplasma mobile, Lactobacillus farciminis, Streptococcus Streptococcus pasteurianus, Lactobacillus johnsoniijohnsonii, Staphylococcus pseudintermedius, Filifactor alocis, Legionella pneumophila, Suterella wadsworthensis, Corynebacter diphtheria, Acidaminococcus, Lachnospiraceae bacterium, or Prevotella.
[0234] In some embodiments, examples of CRISPR / Cas system proteins can be naturally occurring or engineered versions.
[0235] In some embodiments, naturally occurring CRISPR / Cas system proteins can belong to CAS class I type I, III, or IV, or CAS class II type II or V, and can include Cas9, Cas3, Cas8a-c, Cas10, CasX, CasY, Cas13, Cas14, Cse1, Csy1, Csn2, Cas4, Csm2, Cmr5, Csf1, C2c2, and Cpf1.
[0236] In an exemplary embodiment, the CRISPR / Cas system protein comprises Cas9.
[0237] In an exemplary embodiment, the CRISPR / Cas system protein comprises Cpf1.
[0238] A "CRISPR / Cas system protein-gNA complex" refers to a complex comprising a CRISPR / Cas system protein and a guided NA (e.g., gRNA or gDNA). When the gNA is a gRNA, the gRNA may be composed of two molecules: one RNA ("crRNA") that hybridizes to the target and provides sequence specificity, and one RNA "tracrRNA" that can hybridize to the crRNA. Alternatively, the guide RNA may be a single molecule (i.e., gRNA) that comprises the crRNA and tracrRNA sequences. Alternatively, the guide RNA may be a single molecule (i.e., gRNA) that comprises the crRNA sequence.
[0239] The CRISPR / Cas system protein can be at least 60% identical (e.g., at least 70%, at least 80%, or 90% identical, at least 95% identical, or at least 98% identical, or at least 99% identical) to a wild-type CRISPR / Cas system protein. The CRISPR / Cas system protein can have all of the functions of the wild-type CRISPR / Cas system protein, or only one or a subset of the functions, including binding activity, nuclease activity, and nuclease activity.
[0240] The term "CRISPR / Cas system protein-associated induced NA" refers to an induced NA. A CRISPR / Cas system protein-associated induced NA can exist as an isolated NA or as part of a CRISPR / Cas system protein-gNA complex.
[0241] In some embodiments, the CRISPR / Cas system protein is an RNA-guided RNA nuclease (i.e., cleaves RNA). Exemplary CRISPR / Cas system proteins that cleave RNA include, but are not limited to, C2c2. C2c2 (also known as Cas13a) is a class 2 type VI RNA-guided RNA-targeted CRISPR / Cas system protein. In some embodiments, the C2c2 nuclease is isolated or derived from Leptotrichia shahii. In some embodiments, C2c2 is guided by a single crRNA that cleaves ssRNA carrying a complementary protospacer. Suitable C2c2 crRNA sequences will be readily apparent to those skilled in the art.
[0242] In some embodiments, the CRISPR / Cas system protein is a DNA-guided RNA nuclease. In some embodiments, the DNA cleaved by the CRISPR / Cas system protein is double-stranded. Exemplary RNA-guided DNA nucleases that cleave double-stranded DNA include, but are not limited to, Cas9, Cpfl, CasX, and CasY. Further exemplary RNA-guided DNA nucleases include Cas10, Csm2, Csm3, Csm4, and Csm5. In some embodiments, Cas10, Csm2, Csm3, Csm4, and Csm5 form a ribonucleoprotein complex with the gRNA.
[0243] In some embodiments, the RNA-guided DNA nuclease is CasX. In some embodiments, the CasX protein is dual-guided (i.e., the gNA comprises a crRNA and a tracrRNA). In some embodiments, CasX recognizes a TTCN PAM located immediately 5' to the sequence complementary to the target sequence. In some embodiments, the CasX protein is isolated or derived from Deltaproteobacteria or Planctomycetes. In some embodiments, the CasX protein is a CasX1, CasX2, or CasX3 protein. CasX proteins are described in WO / 2018 / 064371, the contents of which are incorporated herein by reference in their entirety. Suitable gNA sequences for CasX proteins will be readily apparent to those of skill in the art.
[0244] In some embodiments, the RNA-guided DNA nuclease is CasY. In some embodiments, the CasY protein is dual-guided (i.e., the gNA comprises a crRNA and a tracrRNA). In some embodiments, CasY recognizes a TA PAM located 5' of the target sequence. CasY proteins are described in WO / 2018 / 064352, the contents of which are incorporated herein by reference in their entirety. Suitable gNA sequences for CasY proteins will be readily apparent to those of skill in the art. In some embodiments, the CRISPR / Cas system protein is an RNA-guided DNA nuclease. In some embodiments, the DNA cleaved by the CRISPR / Cas system protein is single-stranded. Exemplary RNA-guided CRISPR / Cas system proteins that cleave single-stranded DNA include, but are not limited to, Cas3 and Cas14. In some embodiments, the Cas14 protein does not require a PAM site.
[0245] Cas9 In some embodiments, the CRISPR / Cas system protein nucleic acid-guided nuclease is or comprises Cas9. The Cas9 of the present disclosure can be isolated, recombinantly produced, or synthesized.
[0246] Examples of Cas9 proteins that can be used in embodiments herein can be found in F. A Ran, L. Cong, W. X. Yan, D. A. Scott, J. S. Gootenberg, A. J. Kriz, B. Zetsche, O. Shalem, X. Wu, K. S. Makarova, E. V. Koonin, P. A. Sharp, and F. Zhang; “In vivo genome editing using Staphylococcus aureus Cas9,” Nature 520, 186-191 (09 April 2015) doi:10.1038 / nature14299, incorporated herein by reference.
[0247] In some embodiments, Cas9 is capable of inhibiting Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Streptococcus thermophiles, Treponema denticola, Francisella tularensis, Pasteurella multocida, Campylobacter jejuni, Campylobacter lari, Mycoplasma gallisepticum, Nitratifractor sarsuginis, and / or Clostridium moniliforme. salsuginis, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria cinerea, Gluconacetobacter diazotrophicus, Azospirillum, Sphaerochaeta globus, Flavobacterium columnare, Fluviicola taffensis, Bacteroides coprophilus, Mycoplasma mobile, Lactobacillus farciminis, Streptococcus Streptococcus pasteurianus, Lactobacillus johnsonii, Staphylococcus pseudintermediusThese are type II CRISPR systems derived from Bacillus pseudintermedius, Filifactor alocis, Legionella pneumophila, Suterella wadsworthensis, or Corynebacter diphtheria.
[0248] In some embodiments, the Cas9 is a Type II CRISPR system derived from S. pyogenes, and the PAM sequence is NGG located immediately 3' of the target-specific guide sequence. Exemplary Type II CRISPR system PAM sequences from bacterial species include Streptococcus pyogenes (NGG), Staph aureus (NNGRRT), Neisseria meningitidis (NNNNGATT), Streptococcus thermophilus (NNAGAA), and Treponema denticola (NAAAAC), all of which may be used without departing from the present disclosure.
[0249] In one exemplary embodiment, the Cas9 sequence can be obtained, for example, from the pX330 plasmid (available from Addgene), reamplified by PCR, and then cloned into pET30 (from EMD biosciences) to express in bacteria and purify recombinant 6His-tagged protein.
[0250] "Cas9-gNA complex" refers to a complex containing a Cas9 protein and a derived NA. The Cas9 protein can be at least 60% identical (e.g., at least 70%, at least 80%, or 90% identical, at least 95% identical, at least 98% identical, or at least 99% identical) to a wild-type Cas9 protein, such as a Streptococcus pyogenes Cas9 protein. The Cas9 protein can have all of the functions of the wild-type Cas9 protein, or only one or some of the functions, including binding activity, nuclease activity, and nuclease activity.
[0251] The term "Cas9-associated induced NA" refers to the induced NA described above. Cas9-associated induced NA may exist independently or as part of a Cas9-gNA complex. Non-CRISPR / Cas system nucleic acid-guided nucleases
[0252] In some embodiments, non-CRISPR / Cas system proteins are used in the embodiments provided herein.
[0253] In some embodiments, the non-CRISPR / Cas system protein may be derived from any bacterial or archaeal species.
[0254] In some embodiments, the non-CRISPR / Cas system protein is isolated, recombinantly produced, or synthesized.
[0255] In some embodiments, the non-CRISPR / Cas system protein is selected from the group consisting of Aquifex aeolicus, Thermus thermophilus, Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Streptococcus thermophiles, Treponema denticola, Francisella tularensis, Pasteurella multocida, Campylobacter jejuni, Campylobacter lari, and the like. lari, Mycoplasma gallisepticum, Nitratifractor salsuginis, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria cinerea, Gluconacetobacter diazotrophicus, Azospirillum, Sphaerochaeta globus, Flavobacterium columnare, Fluviicola taffensis, Bacteroides coprophilus, Mycoplasma mobile mobile), Lactobacillus farciminis, Streptococcus pasteurianuspasteurianus, Lactobacillus johnsonii, Staphylococcus pseudintermedius, Filifactor alocis, Legionella pneumophila, Suterella wadsworthensis, Natronobacterium gregoryi, or Corynebacter diphtheria.
[0256] In some embodiments, the non-CRISPR / Cas system protein can be a naturally occurring or engineered version.
[0257] In some embodiments, the naturally occurring non-CRISPR / Cas system protein is NgAgo (Natronobacterium gregoryi).
[0258] A "non-CRISPR / Cas system protein-gNA complex" refers to a complex comprising a non-CRISPR / Cas system protein and a guided NA (e.g., gRNA or gDNA). When the gNA is a gRNA, the gRNA may be composed of two molecules: one RNA ("crRNA") that hybridizes to the target and provides sequence specificity, and one RNA ("tracrRNA") that can hybridize to the crRNA. Alternatively, the guide RNA may be a single molecule (i.e., gRNA) that comprises the crRNA and tracrRNA sequences.
[0259] The non-CRISPR / Cas system protein can be at least 60% identical (e.g., at least 70%, at least 80%, or 90% identical, at least 95% identical, or at least 98% identical, or at least 99% identical) to the wild-type non-CRISPR / Cas system protein. The non-CRISPR / Cas system protein can have all of the functions of the wild-type non-CRISPR / Cas system protein, or only one or some of the functions, including binding activity, nuclease activity, and nuclease activity.
[0260] The term "non-CRISPR / Cas system protein-associated induced NA" refers to an induced NA. Non-CRISPR / Cas system protein-associated induced NA can exist as an isolated NA or as part of a non-CRISPR / Cas system protein-gNA complex.
[0261] Cpf1 In some embodiments, the CRISPR / Cas system protein nucleic acid-guided nuclease is or comprises a Cpf1 system protein. The Cpf1 system proteins of the present disclosure can be isolated, recombinantly produced, or synthesized.
[0262] The Cpf1 system protein is a Class II, Type V CRISPR system protein. In some embodiments, the Cpf1 protein is isolated or derived from Francisella tularensis. In some embodiments, the Cpf1 protein is isolated or derived from Acidaminococcus, Lachnospiraceae bacterium, or Prevotella.
[0263] Cpf1 system proteins bind to a single guide RNA containing a nucleic acid-guided nuclease system protein-binding sequence (e.g., stem-loop) and a target sequence. The Cpf1 target sequence includes a sequence located immediately 3' of the Cpf1 PAM sequence in the target nucleic acid. Unlike Cas9, the Cpf1 nucleic acid-guided nuclease system protein-binding sequence is located 5' of the Cpf1 gRNA target sequence. Cpf1 can also generate a staggered cut in the target nucleic acid rather than a blunt-end cut. After targeting the Cpf1 protein-gRNA complex to the target nucleic acid, Francisella Cpf1, for example, staggers the target nucleic acid, creating a 5' overhang of approximately 5 nucleotides at the 3' end of the target sequence, 18-23 bases away from the PAM. In contrast, cleavage by wild-type Cas9 generates a blunt end three nucleotides upstream of the Cas9 PAM.
[0264] In the exemplary Cpf1, the gRNA stem-loop sequence comprises the following RNA sequence (5'>3', AAUUUCUACUGUUGUAGAU (SEQ ID NO: 34)).
[0265] "Cpf1 protein-gNA complex" refers to a complex comprising a Cpf1 protein and a derived NA (e.g., gRNA). When the gNA is a gRNA, the gRNA can be composed of a single molecule, i.e., one RNA ("crRNA") that hybridizes to the target and provides sequence specificity.
[0266] The Cpf1 protein can be at least 60% identical (e.g., at least 70%, at least 80%, or 90% identical, at least 95% identical, at least 98% identical, or at least 99% identical) to a wild-type Cpf1 protein. The Cpf1 protein can have all of the functions of the wild-type Cpf1 protein, or only one or some of the functions, including binding activity and nuclease activity.
[0267] Cpf1 system proteins recognize a variety of PAM sequences. Exemplary PAM sequences recognized by Cpf1 system proteins include, but are not limited to, TTN, TCN, and TGN. Additional Cpf1 PAM sequences include, but are not limited to, TTTN. One characteristic of Cpf1 PAM sequences is that they have a higher A / T content than the NGG or NAG PAM sequences used by Cas9 proteins. Target nucleic acids, such as different genomes, vary in their percent G / C content. For example, the genome of the human malaria parasite Plasmodium falciparum is known to be A / T-rich. Alternatively, protein-coding sequences within a genome often have a higher G / C content than the genome as a whole. The ratio of A / T to G / C nucleotides within a target genome affects the distribution and frequency of a given PAM sequence within that genome. For example, an A / T-rich genome may have fewer NGG or NAG sequences, while a G / C-rich genome may have fewer TTN sequences. Cpf1 system proteins expand the repertoire of PAM sequences available to those skilled in the art, providing greater flexibility and functionality for gRNA libraries.
[0268] Catalytically inactive nucleic acid-induced nuclease In some embodiments, engineered examples of nucleic acid-guided nuclease system (e.g., CRISPR / Cas system) proteins include catalytically inactive nucleic acid-guided nuclease system proteins. The term "catalytically inactive" generally refers to nucleic acid-guided nuclease system proteins with inactivated nucleases (e.g., HNH and RuvC nucleases). Such proteins can bind to target sites in any nucleic acid (the target site is determined by the induced NA), but the protein cannot cleave or nick the target nucleic acid (e.g., double-stranded DNA). In some embodiments, the catalytically inactive protein of a nucleic acid-guided nuclease system is a catalytically inactive CRISPR / Cas system protein, such as catalytically inactive Cas9 (dCas9). Thus, dCas9 allows the mixture to be separated into unbound nucleic acids and dCas9-bound fragments. In one embodiment, the dCas9 / gRNA complex binds to a target determined by the gRNA sequence. dCas9 binding can prevent cleavage by Cas9 while other manipulations are ongoing. In another embodiment, dCas9 can be fused to another enzyme, such as a transposase, to target the activity of that enzyme to a specific site. Naturally occurring catalytically inactive nucleic acid-guided nuclease system proteins can also be used.
[0269] In another embodiment, the catalytically inactive nucleic acid-guided nuclease can be fused to another enzyme, such as a transposase, to target the activity of that enzyme to a specific site.
[0270] In some embodiments, the catalytically inactive nucleic acid-guided nuclease is dCas9, dCpf1, dCas3, dCas8a-c, dCas10, dCsel, dCsy1, dCsn2, dCas4, dCsm2, dCm5, dCsf1, dC2C2, dCasX, dCasY, dCas13, dCas14, or dNgAgo.
[0271] In an exemplary embodiment, the catalytically inactive nucleic acid-guided nuclease protein is dCas9.
[0272] In an exemplary embodiment, the catalytically inactive nucleic acid-induced nuclease protein is dCpf1.
[0273] Nucleic acid-induced nuclease nickase In some embodiments, engineered examples of nucleic acid-guided nucleases include nucleic acid-guided nuclease nickases (interchangeably referred to as nickase nucleic acid-guided nucleases).
[0274] In some embodiments, engineered examples of nucleic acid-guided nucleases include CRISPR / Cas system nickases or non-CRISPR / Cas system nickases that contain a single, inactive catalytic domain.
[0275] In some embodiments, the nucleic acid-guided nuclease nickase is a Cas9 nickase, a Cpf1 nickase, a Cas3 nickase, a Cas8a-c nickase, a Cas10 nickase, a Csel nickase, a Csy1 nickase, a Csn2 nickase, a Cas4 nickase, a Csm2 nickase, a Cm5 nickase, a Csf1 nickase, a C2C2 nickase, a CasX nickase, a CasY nickase, a Cas13 nickase, a Cas14 nickase, or a NgAgo nickase.
[0276] In one embodiment, the nucleic acid-guided nuclease nickase is a Cas9 nickase.
[0277] In one embodiment, the nucleic acid-guided nuclease nickase is Cpf1 nickase.
[0278] In some embodiments, nucleic acid-guided nuclease nickases can be used to bind to target sequences. With only one active nuclease domain, the nucleic acid-guided nuclease nickases cleave only one strand of the target DNA, creating a single-strand break or "nick." Depending on the variant used, either the hybridized or non-hybridized strand of the guided NA may be cleaved. A nucleic acid-guided nuclease nickase bound to two gNAs targeting opposite strands can create a double-strand break in the nucleic acid. This "dual nickase" strategy increases the specificity of cleavage because both nucleic acid-guided nuclease / gNA complexes must bind specifically at the site before a double-strand break is formed.
[0279] In an exemplary embodiment, Cas9 nickase can be used to bind to a target sequence. The term "Cas9 nickase" refers to a modified version of the Cas9 protein that contains a single, inactive catalytic domain, i.e., either the RuvC domain or the HNH domain. With only one active nuclease domain, Cas9 nickase cleaves only one strand of the target DNA, creating a single-strand break or "nick." Depending on the variant used, the guide RNA-hybridized strand or the non-hybridized strand can be cleaved. Cas9 nickase bound to two gRNAs targeting opposite strands creates a double-strand break in the DNA. This "dual nickase" strategy can increase the specificity of cleavage because both Cas9 / gRNA complexes must specifically bind at the site before a double-strand break is formed.
[0280] Dissociative, thermostable nucleic acid-induced nucleases In some embodiments, thermostable nucleic acid-guided nucleases are used in the methods provided herein (thermostable CRISPR / Cas system nucleic acid-guided nucleases or thermostable non-CRISPR / Cas system nucleic acid-guided nucleases). In such embodiments, the reaction temperature is increased to induce protein dissociation. Lowering the reaction temperature allows for the generation of additional cleaved target sequences. In some embodiments, the thermostable nucleic acid-guided nuclease maintains at least 50% activity, at least 55% activity, at least 60% activity, at least 65% activity, at least 70% activity, at least 75% activity, at least 80% activity, at least 85% activity, at least 90% activity, at least 95% activity, at least 96% activity, at least 97% activity, at least 98% activity, at least 99% activity, or 100% activity when maintained at at least 75°C for at least 1 minute. In some embodiments, the thermostable nucleic acid-guided nuclease maintains at least 50% activity when maintained at at least 75°C, at least 80°C, at least 85°C, at least 90°C, at least 91°C, at least 92°C, at least 93°C, at least 94°C, at least 95°C, 96°C, at least 97°C, at least 98°C, at least 99°C, or at least 100°C for at least 1 minute. In some embodiments, the thermostable nucleic acid-guided nuclease maintains at least 50% activity when maintained at at least 75°C for at least 1 minute, 2 minutes, 3 minutes, 4 minutes, or 5 minutes. In some embodiments, the thermostable nucleic acid-guided nuclease maintains at least 50% activity when the temperature is increased and then decreased to between 25°C and 50°C. In some embodiments, the temperature is decreased to 25°C, 30°C, 35°C, 40°C, 45°C, or 50°C. In an exemplary embodiment, the thermostable enzyme retains at least 90% activity after 1 minute at 95°C.
[0281] In some embodiments, the thermostable nucleic acid-guided nuclease is thermostable Cas9, thermostable Cpf1, thermostable Cas3, thermostable Cas8a-c, thermostable Cas10, thermostable Cse1, thermostable Csy1, thermostable Csn2, thermostable Cas4, thermostable Csm2, thermostable Cm5, thermostable Csf1, thermostable C2C2, or thermostable NgAgo.
[0282] In some embodiments, the thermostable CRISPR / Cas system protein is a thermostable Cas9.
[0283] Thermostable nucleic acid-guided nucleases can be identified by sequence homology in the genomes of, for example, the thermophilic bacteria Streptococcus thermophilus and Pyrococcus furiosus. The nucleic acid-guided nuclease gene can then be cloned into an expression vector. In an exemplary embodiment, a thermostable Cas9 protein is isolated.
[0284] In another embodiment, a thermostable nucleic acid-guided nuclease can be obtained by in vitro evolution of a non-thermostable nucleic acid-guided nuclease. The sequence of the nucleic acid-guided nuclease can be mutagenized to improve its thermostability.
[0285] Kits and manufactured products The present disclosure provides kits comprising any one or more of the compositions described herein, including, but not limited to, adaptors, gNAs (e.g., gRNAs or gDNAs), gNA collections (e.g., multiple gRNAs or gDNAs), modification-sensitive restriction enzymes, controls, etc.
[0286] In an exemplary embodiment, the kit includes a gRNA, wherein the gRNA targets the human genome or other source of DNA sequence.
[0287] The present disclosure also provides all necessary reagents and instructions for carrying out the methods of enriching a sample for nucleic acids of interest using differences in nucleotide modifications as described herein.
[0288] Also provided herein is computer software that monitors information before and after enriching a sample using the methods provided herein. In an exemplary embodiment, the software can calculate and report the abundance of nucleic acid sequences targeted for depletion in a sample before and after applying a method described herein to assess the level of off-target depletion, and the software can test the effectiveness of target-depletion / enrichment / capture / partition / labeling / modulation / editing by comparing the abundance of sequences of interest before and after processing a sample using the enrichment methods provided herein.
[0289] All publications mentioned in the above specification are incorporated herein by reference. Various modifications and variations of the described products, systems, uses, processes, and methods of the present disclosure will be apparent to those skilled in the art without departing from the scope and spirit of the present disclosure. Although the present disclosure has been described in connection with specific preferred embodiments, it should be understood that the claimed disclosure should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the disclosure that are obvious to those skilled in molecular biology and biotechnology or related fields are intended to be within the scope of the following claims.
[0290] Enumerated Embodiments The invention can be defined by reference to the exemplary embodiments listed below.
[0291] 1. A method for enriching a sample of a nucleic acid of interest by at least about two-fold compared to a nucleic acid targeted for depletion, the method comprising using differences in nucleotide modifications between the nucleic acid of interest and the nucleic acid targeted for depletion.
[0292] 2. A method for enriching a sample of a nucleic acid of interest by at least about two-fold compared to a nucleic acid targeted for depletion, the method comprising using differences in nucleotide modifications between the nucleic acid of interest and the nucleic acid targeted for depletion, and the method does not involve size selection or modification-sensitive target binding.
[0293] 3. A method for enriching a sample of a nucleic acid of interest relative to a nucleic acid targeted for depletion by at least about two-fold, comprising using a difference in nucleotide modifications between the nucleic acid of interest and the nucleic acid targeted for depletion to ligate the target nucleic acid to an adapter, but not the nucleic acid targeted for depletion.
[0294] 4. A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample comprising nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids of interest or at least a subset of the nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme; b. terminally dephosphorylating a plurality of nucleic acids in the sample; c. contacting the sample from (b) with a first modification-sensitive restriction enzyme under conditions that allow for cleavage of at least some of the first modification-sensitive restriction sites in nucleic acids in the sample; d. contacting the sample from (c) with the adaptors under conditions that allow ligation of the adaptors to the 5' and 3' ends of a plurality of nucleic acids of interest; thereby producing a sample enriched for the nucleic acid of interest that is adaptor-linked at its 5' and 3' ends.
[0295] 5. The method of embodiment 4, wherein prior to (a), the nucleic acid of interest and the nucleic acid targeted for depletion are fragmented.
[0296] 6. The method of embodiment 4 or 5, wherein both the nucleic acid of interest and the nucleic acid targeted for depletion each comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme.
[0297] 7. The method of embodiment 6, wherein the frequency of nucleotide modifications within or adjacent to the plurality of first recognition sites is not the same in the nucleic acid of interest as in the nucleic acid targeted for depletion.
[0298] 8. The method of any one of embodiments 4 to 7, wherein the activity of the first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to its cognate recognition site.
[0299] 9. The method of embodiment 8, wherein the plurality of first recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of first recognition sites in the nucleic acid of interest.
[0300] 10. The method of embodiment 8 or 9, wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AatII, AccII, Aor13HI, Aor51HI, BspT104I, BssHII, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, MluI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, SnaBI, AluI, and Sau3AI.
[0301] 11. The method of embodiment 8 or 9, wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AluI and Sau3AI.
[0302] 12. The method of any one of embodiments 4 to 7, wherein the first modification-sensitive restriction enzyme is active at a recognition site that comprises at least one modified nucleotide and is not active at a recognition site that does not comprise at least one modified nucleotide.
[0303] 13. The method of embodiment 12, wherein the plurality of first recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of first recognition sites in the nucleic acid of interest.
[0304] 14. The method of embodiment 12 or 13, wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AbaSI, FspEI, LpnPI, MspJI, or McrBC.
[0305] 15. The method of any one of embodiments 12-13, wherein the modification comprises 5-hydroxymethylcytosine.
[0306] 16. The method of embodiment 15, wherein the first modification-sensitive restriction enzyme comprises AbaSI, and the method further comprises contacting the sample with T4 phage β-glucosyltransferase prior to step (c).
[0307] 17. The method of any one of embodiments 12-14, wherein the modification comprises glucosylhydroxymethylcytosine.
[0308] 18. The method of embodiment 17, wherein the first modification-sensitive restriction enzyme comprises AbaSI.
[0309] 19. The method of any one of embodiments 12-14, wherein the modification comprises methylcytosine.
[0310] 20. The method of embodiment 19, wherein the first modification-sensitive restriction enzyme comprises McrBC.
[0311] 21. The method of any one of embodiments 12 to 20, wherein the nucleic acid of interest comprises at least one DpnI recognition site, and the method further comprises, prior to step (c), contacting the sample with DpnI and T4 polymerase.
[0312] 22. The method of embodiment 21, wherein the T4 polymerase replaces methylated A and C nucleotides with unmethylated A and C nucleotides within or near at least one DpnI recognition site.
[0313] 23. The method of any one of embodiments 12 to 22, further comprising, prior to step (d), contacting the sample from (c) with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated ends of the nucleic acids.
[0314] 24. The method of embodiment 23, wherein the exonuclease comprises lambda nuclease, exonuclease III, or BAL-31.
[0315] 25. The method of any one of embodiments 4 to 24, wherein terminally dephosphorylating nucleic acids in the sample in step (b) comprises a phosphatase.
[0316] 26. The method of embodiment 25, wherein the phosphatase is alkaline phosphatase.
[0317] 27. The method of embodiment 26, wherein the alkaline phosphatase is shrimp alkaline phosphatase.
[0318] 28. e. contacting the adapter-ligated nucleic acid from (d) with a second modification-sensitive restriction enzyme under conditions that allow the second modification-sensitive restriction enzyme to cleave the second recognition site; at least a subset of the nucleic acids targeted for depletion comprise a plurality of second recognition sites for a second modification-sensitive restriction enzyme; a second modification-sensitive restriction enzyme that targets recognition sites that include at least one modified nucleotide and does not target recognition sites that do not include at least one modified nucleotide; The method of any one of embodiments 4 to 27, thereby generating a collection of nucleic acids targeted for depletion that are adapter-linked at one end and a collection of nucleic acids of interest that are adapter-linked at both ends.
[0319] 29. The method of embodiment 28, wherein the nucleic acid of interest and the nucleic acid targeted for depletion each comprise a plurality of second recognition sites for a second modification-sensitive restriction enzyme.
[0320] 30. The method of embodiment 29, wherein the plurality of second recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of second recognition sites in the nucleic acid of interest.
[0321] 31. The method of any one of embodiments 4 to 30, further comprising contacting the sample after step (d) with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes, wherein the gNAs are complementary to target sites of the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends.
[0322] 32. The method is at least 10 2 specific nucleic acid-guided nuclease-gNA complex, at least 10 3 specific nucleic acid-guided nuclease-gNA complex, 10 4 Specific nucleic acid-guided nuclease-gNA complex or 10 5 32. The method of embodiment 31, comprising contacting the sample with a specific nucleic acid-guided nuclease-gNA complex of
[0323] 33. The method of embodiment 31 or 32, wherein the nucleic acid-guided nuclease is Cas9, Cpf1, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn2, Cas4, Csm2, CasX, CasY, Cas13, Cas14, or Cm5.
[0324] 34. The method of embodiment 31 or 32, wherein the nucleic acid-guided nuclease is Cas9, Cpf1, or a combination thereof.
[0325] 35. The method of any one of embodiments 31-34, wherein the nucleic acid-guided nuclease is Cas9 or Cpf1 nickase.
[0326] 36. The method of any one of embodiments 31 to 35, wherein the nucleic acid-guided nuclease is thermostable.
[0327] 37. The method of any one of embodiments 31 to 36, wherein the gNA is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0328] 38. The method of any one of embodiments 4 to 37, further comprising using adapters to amplify, sequence, or clone the nucleic acid of interest that is adapter-linked at the 5' and 3' ends of the nucleic acid.
[0329] 39. The method of any one of embodiments 1-38, wherein the nucleotide modification comprises an adenine modification or a cytosine modification.
[0330] 40. The method of embodiment 39, wherein the adenine modification comprises adenine methylation.
[0331] 41. The method of embodiment 40, wherein the adenine methylation comprises Dam methylation or EcoKI methylation.
[0332] 42. The method of embodiment 39, wherein the cytosine modifications comprise 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxylcytosine, 5-glucosylhydroxymethylcytosine or 3-methylcytosine.
[0333] 43. The method of embodiment 39, wherein the cytosine modification comprises cytosine methylation.
[0334] 44. The method of embodiment 43, wherein the cytosine methylation comprises CpG methylation, CpA methylation, CpT methylation, CpC methylation, or a combination thereof.
[0335] 45. The method of embodiment 43, wherein the cytosine methylation comprises Dcm methylation, DNMT1 methylation, DNMT3A methylation, or DNMT3B methylation.
[0336] 46. The method of any one of embodiments 28-45, wherein the second modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AbaSI, FspEI, LpnPI, MspJI, or McrBC.
[0337] 47. The method of any one of embodiments 28-38, wherein the modification comprises 5-hydroxymethylcytosine.
[0338] 48. The method of embodiment 47, wherein the second modification-sensitive restriction enzyme comprises AbaSI, and the method further comprises contacting the sample with T4 phage β-glucosyltransferase prior to step (e).
[0339] 49. The method of any one of embodiments 28-38, wherein the modification comprises glucosylhydroxymethylcytosine.
[0340] 50. The method of embodiment 49, wherein the second modification-sensitive restriction enzyme comprises AbaSI.
[0341] 51. The method of any one of embodiments 28-38, wherein the modification comprises methylcytosine.
[0342] 52. The method of embodiment 51, wherein the second modification-sensitive restriction enzyme comprises McrBC.
[0343] 53. The method of any one of embodiments 28 to 52, wherein the nucleic acid of interest comprises at least one DpnI recognition site, and the method further comprises, prior to step (e), contacting the sample with DpnI and T4 polymerase.
[0344] 54. The method of embodiment 53, wherein the T4 polymerase replaces methylated A and C nucleotides with unmethylated A and C nucleotides within or near at least one DpnI recognition site.
[0345] 55. The method of any one of embodiments 1 to 54, wherein the nucleic acid targeted for depletion comprises a host nucleic acid and the nucleic acid of interest comprises a non-host nucleic acid.
[0346] 56. The method of embodiment 55, wherein the non-host comprises a bacterium, a fungus, or a virus.
[0347] 57. The method of embodiment 55, wherein the non-host comprises organisms of multiple species.
[0348] 58. The method of embodiment 55, wherein the host is a mammal, bird, reptile, or insect.
[0349] 59. The method of embodiment 58, wherein the mammal is a human, cow, horse, sheep, pig, monkey, dog, cat, rat, rabbit, mouse, or gerbil.
[0350] 60. The method of any one of embodiments 1 to 59, wherein the nucleic acid targeted for depletion comprises a transcriptionally active site and the nucleic acid of interest comprises a repetitive sequence.
[0351] 61. The method of any one of embodiments 4 to 60, wherein the adaptor-ligated nucleic acid of interest and the nucleic acid targeted for depletion are in the range of 50 to 1000 bp.
[0352] 62. The method of any one of embodiments 1-61, wherein the nucleic acid of interest constitutes less than 50% of the total nucleic acids in the sample.
[0353] 63. The method of any one of embodiments 1-61, wherein the nucleic acid of interest constitutes less than 30% of the total nucleic acids in the sample.
[0354] 64. The method of any one of embodiments 1-61, wherein the nucleic acid of interest constitutes less than 5% of the total nucleic acids in the sample.
[0355] 65. The method of any one of embodiments 1-64, wherein the sample is any one of a biological sample, a clinical sample, a forensic sample, or an environmental sample.
[0356] 66. The method of any one of embodiments 1-64, wherein the sample is selected from whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bone, fingernails, feces, urine, tissue, and biopsy.
[0357] 67. A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample containing nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids targeted for depletion contain multiple recognition sites for a modification-sensitive restriction enzyme; b. terminally dephosphorylating a plurality of nucleic acids in the sample; c. contacting the sample from (b) with a modification-sensitive restriction enzyme under conditions that allow for cleavage of the modification-sensitive restriction site in nucleic acids in the sample, thereby generating nucleic acids with exposed terminal phosphates; d. A method of enriching a sample for a nucleic acid of interest, comprising contacting the sample with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated end of the nucleic acid, thereby producing a sample enriched in the nucleic acid of interest.
[0358] 68. The method of embodiment 67, wherein prior to step (a), the nucleic acid of interest and the nucleic acid targeted for depletion are fragmented.
[0359] 69. The method of embodiment 67 or 68, wherein the nucleic acid of interest and the nucleic acid targeted for depletion each comprise multiple recognition sites for modification-sensitive restriction enzymes.
[0360] 70. The method of embodiment 69, wherein the multiple recognition sites in the nucleic acid targeted for depletion are modified more frequently than the multiple recognition sites in the nucleic acid of interest.
[0361] 71. The method of any one of embodiments 67 to 70, wherein the nucleic acid of interest comprises at least one DpnI recognition site, and the method further comprises, prior to step (c), contacting the sample with DpnI and T4 polymerase.
[0362] 72. The method of embodiment 71, wherein the T4 polymerase replaces methylated A and C nucleotides with unmethylated A and C nucleotides within or adjacent to at least one DpnI recognition site.
[0363] 73. The method of any one of embodiments 67-72, wherein the modification comprises an adenine modification or a cytosine modification.
[0364] 74. The method of embodiment 73, wherein the adenine modification comprises adenine methylation.
[0365] 74. The method of embodiment 73, wherein the 75 adenine methylation comprises Dam methylation or EcoKI methylation.
[0366] 76. The method of embodiment 73, wherein the cytosine modification comprises 5-glucosylhydroxymethylcytosine or 3-methylcytosine, including 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxylcytosine.
[0367] 77. The method of embodiment 73, wherein the cytosine modification comprises cytosine methylation.
[0368] 78. The method of embodiment 77, wherein the cytosine methylation comprises CpG methylation, CpA methylation, CpT methylation, CpC methylation, or a combination thereof.
[0369] 79. The method of embodiment 73, wherein the cytosine methylation comprises Dcm methylation, DNMT1 methylation, DNMT3A methylation, or DNMT3B methylation.
[0370] 80. The method of any one of embodiments 67-79, wherein the modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AbaSI, FspEI, LpnPI, MspJI or McrBC.
[0371] 81. The method of any one of embodiments 67-72, wherein the modification comprises 5-hydroxymethylcytosine.
[0372] 82. The method of embodiment 81, wherein the modification-sensitive restriction enzyme comprises AbaSI, and the method further comprises contacting the sample with T4 phage β-glucosyltransferase prior to step (c).
[0373] 83. The method of any one of embodiments 67-72, wherein the modification comprises glucosylhydroxymethylcytosine.
[0374] 84. The method of embodiment 83, wherein the modification-sensitive restriction enzyme comprises AbaSI.
[0375] 85. The method of any one of embodiments 67-72, wherein the modification comprises methylcytosine.
[0376] 86. The method of embodiment 85, wherein the modification-sensitive restriction enzyme comprises McrBC.
[0377] 87. The method of embodiments 67-86, wherein the exonuclease is lambda nuclease, exonuclease III, or BAL-31.
[0378] 88. The method of any one of embodiments 67-87, wherein terminally dephosphorylating nucleic acids in the sample in step (b) comprises a phosphatase.
[0379] 89. The method of embodiment 88, wherein the phosphatase is alkaline phosphatase.
[0380] 90. The method of embodiment 74, wherein the alkaline phosphatase is shrimp alkaline phosphatase.
[0381] contacting the sample from 91.e.(d) with an adapter under conditions that allow ligation of the adapter to the 5' and 3' ends of a plurality of nucleic acids of interest; The method according to any one of embodiments 67 to 90, thereby generating a sample enriched in target nucleic acids that are adaptor-linked at the 5' and 3' ends.
[0382] 92. The method of any one of embodiments 67 to 91, further comprising contacting the sample after step (d) with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes, wherein the gNAs are complementary to target sites of the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends.
[0383] 93. The method is at least 10 2 specific nucleic acid-guided nuclease-gNA complex, at least 10 3 specific nucleic acid-guided nuclease-gNA complex, 10 4 Specific nucleic acid-guided nuclease-gNA complex or 10 5 93. The method of embodiment 92, comprising contacting the sample with a specific nucleic acid-guided nuclease-gNA complex of
[0384] 94. The method of embodiment 92 or 93, wherein the nucleic acid-guided nuclease is Cas9, Cpf1, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn2, Cas4, Csm2, CasX, CasY, Cas13, Cas14 or Cm5.
[0385] 95. The method of embodiment 92 or 93, wherein the nucleic acid-guided nuclease is Cas9, Cpf1, or a combination thereof.
[0386] 96. The method of any one of embodiments 92 to 95, wherein the nucleic acid-guided nuclease is Cas9 or Cpf1 nickase.
[0387] 97. The method of any one of embodiments 92-96, wherein the nucleic acid-guided nuclease is thermostable.
[0388] 98. The method of any one of embodiments 92 to 97, wherein the gNA is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0389] 99. The method of any one of embodiments 67 to 98, further comprising using an adapter to amplify, sequence, or clone the nucleic acid of interest that is adapter-linked at the 5' and 3' ends of the nucleic acid.
[0390] 100. The method of any one of embodiments 67-99, wherein the nucleic acid targeted for depletion comprises a host nucleic acid and the nucleic acid of interest comprises a non-host nucleic acid.
[0391] 101. The method of embodiment 100, wherein the non-host comprises a bacterium, a fungus, or a virus.
[0392] 102. The method of embodiment 100, wherein the non-host comprises organisms of multiple species.
[0393] 103. The method of embodiment 100, wherein the host is a mammal, a bird, a reptile, or an insect.
[0394] 104. The method of embodiment 103, wherein the mammal is a human, cow, horse, sheep, pig, monkey, dog, cat, rat, rabbit, mouse, or gerbil.
[0395] 105. The method of any one of embodiments 67 to 104, wherein the nucleic acid targeted for depletion comprises a transcriptional activation site and the nucleic acid of interest comprises a repetitive sequence.
[0396] 106. The method of any one of embodiments 67 to 105, wherein the adaptor-ligated nucleic acid of interest and the nucleic acid targeted for depletion are in the range of 50 to 1000 bp.
[0397] 107. The method of any one of embodiments 67 to 106, wherein the nucleic acid of interest constitutes less than 50% of the total nucleic acids in the sample.
[0398] 108. The method of any one of embodiments 67 to 106, wherein the nucleic acid of interest constitutes less than 30% of the total nucleic acids in the sample.
[0399] 109. The method of any one of embodiments 67-106, wherein the nucleic acid of interest constitutes less than 5% of the total nucleic acids in the sample.
[0400] 110. The method of any one of embodiments 67 to 106, wherein the sample is any one of a biological sample, a clinical sample, a forensic sample, or an environmental sample.
[0401] 111. The method of any one of embodiments 67 to 106, wherein the sample is selected from whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bone, fingernails, feces, urine, tissue, and biopsy.
[0402] 112. A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample containing nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids targeted for depletion contain multiple recognition sites for a modification-sensitive restriction enzyme; b. contacting the sample with the adaptor under conditions that allow ligation of the adaptor to the 5' and 3' ends of a plurality of nucleic acids in the sample; c. contacting the sample from (b) with a modification-sensitive restriction enzyme under conditions that allow for cleavage of the modification-sensitive restriction site in nucleic acid in the sample; thereby producing a sample enriched for the nucleic acid of interest that is adaptor-linked at its 5' and 3' ends.
[0403] 113. The method of embodiment 112, wherein prior to step (a), the nucleic acid of interest and the nucleic acid targeted for depletion are fragmented.
[0404] 114. The method of embodiment 112 or 113, wherein both the nucleic acid of interest and the nucleic acid targeted for depletion each contain multiple recognition sites for modification-sensitive restriction enzymes.
[0405] 115. The method of any one of embodiments 112-114, wherein the multiple recognition sites in the nucleic acid targeted for depletion are modified more frequently than the multiple recognition sites in the nucleic acid of interest.
[0406] 116. The method of any one of embodiments 112 to 115, wherein the nucleic acid of interest comprises at least one DpnI recognition site, and the method further comprises, prior to step (c), contacting the sample with DpnI and T4 polymerase.
[0407] 117. The method of embodiment 116, wherein the T4 polymerase replaces methylated A and C nucleotides with unmethylated A and C nucleotides within or adjacent to at least one DpnI recognition site.
[0408] 118. The method of any one of embodiments 112-117, wherein the modification comprises an adenine modification or a cytosine modification.
[0409] 119. The method of embodiment 118, wherein the adenine modification comprises adenine methylation.
[0410] 120. The method of embodiment 119, wherein the adenine methylation comprises Dam methylation or EcoKI methylation.
[0411] 121. The method of embodiment 118, wherein the cytosine modification comprises 5-glucosylhydroxymethylcytosine or 3-methylcytosine, including 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxylcytosine.
[0412] 122. The method of embodiment 118, wherein the cytosine modification comprises cytosine methylation.
[0413] 123. The method of embodiment 122, wherein the cytosine methylation comprises CpG methylation, CpA methylation, CpT methylation, CpC methylation, or a combination thereof.
[0414] 124. The method of embodiment 122, wherein the cytosine methylation comprises Dcm methylation, DNMT1 methylation, DNMT3A methylation, or DNMT3B methylation.
[0415] 125. The method of any one of embodiments 112-124, wherein the modification-sensitive restriction enzyme comprises AbaSI, FspEI, LpnPI, MspJI, or McrBC.
[0416] 126. The method of any one of embodiments 112-117, wherein the modification comprises 5-hydroxymethylcytosine.
[0417] 127. The method of embodiment 126, wherein the modification-sensitive restriction enzyme comprises AbaSI, and the method further comprises contacting the sample with T4 phage β-glucosyltransferase prior to (c).
[0418] 128. The method of any one of embodiments 112-117, wherein the modification comprises glucosylhydroxymethylcytosine.
[0419] 129. The method of embodiment 128, wherein the modification-sensitive restriction enzyme comprises AbaSI.
[0420] 130. The method of any one of embodiments 112-117, wherein the modification comprises methylcytosine.
[0421] 131. The method of embodiment 130, wherein the modification-sensitive restriction enzyme comprises McrBC.
[0422] 132. The method of any one of embodiments 112 to 131, further comprising contacting the sample after step (c) with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes, wherein the gNAs are complementary to target sites of the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends.
[0423] 133. The method is at least 10 2 specific nucleic acid-guided nuclease-gNA complex, at least 10 3 specific nucleic acid-guided nuclease-gNA complex, 10 4 Specific nucleic acid-guided nuclease-gNA complex or 10 5 133. The method of embodiment 132, comprising contacting the sample with a specific nucleic acid-guided nuclease-gNA complex.
[0424] 134. The method of embodiment 132 or 133, wherein the nucleic acid-guided nuclease is Cas9, Cpf1, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn2, Cas4, Csm2, CasX, CasY, Cas13, Cas14 or Cm5.
[0425] 135. The method of embodiment 132 or 133, wherein the nucleic acid-guided nuclease is Cas9, Cpf1, or a combination thereof.
[0426] 136. The method of any one of embodiments 132 to 135, wherein the nucleic acid-guided nuclease is Cas9 or Cpf1 nickase.
[0427] 137. The method of any one of embodiments 132-136, wherein the nucleic acid-guided nuclease is thermostable.
[0428] 138. The method of any one of embodiments 112 to 137, wherein the gNA is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0429] 139. The method of any one of embodiments 112 to 138, further comprising using an adapter to amplify, sequence, or clone the nucleic acid of interest that is adapter-linked at the 5' and 3' ends of the nucleic acid.
[0430] 140. The method of any one of embodiments 112-139, wherein the nucleic acid targeted for depletion comprises a host nucleic acid and the nucleic acid of interest comprises a non-host nucleic acid.
[0431] 141. The method of embodiment 140, wherein the non-host comprises a bacterium, a fungus, or a virus.
[0432] 142. The method of embodiment 140, wherein the non-host comprises organisms of multiple species.
[0433] 143. The method of embodiment 140, wherein the host is a mammal, bird, reptile, or insect.
[0434] 144. The method of embodiment 143, wherein the mammal is a human, cow, horse, sheep, pig, monkey, dog, cat, rat, rabbit, mouse, or gerbil.
[0435] 145. The method of any one of embodiments 112 to 144, wherein the nucleic acid targeted for depletion comprises a transcriptional activation site and the nucleic acid of interest comprises a repetitive sequence.
[0436] 146. The method of any one of embodiments 112 to 145, wherein the adaptor-ligated nucleic acid of interest and the nucleic acid targeted for depletion are in the range of 50 to 1000 bp.
[0437] 147. The method of any one of embodiments 112-146, wherein the nucleic acid of interest constitutes less than 50% of the total nucleic acids in the sample.
[0438] 148. The method of any one of embodiments 112-146, wherein the nucleic acid of interest constitutes less than 30% of the total nucleic acids in the sample.
[0439] 149. The method of any one of embodiments 112-146, wherein the nucleic acid of interest constitutes less than 5% of the total nucleic acids in the sample.
[0440] 150. The method of any one of embodiments 112-149, wherein the sample is any one of a biological sample, a clinical sample, a forensic sample, or an environmental sample.
[0441] 151. The method of any one of embodiments 112 to 149, wherein the sample is selected from whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bone, fingernails, feces, urine, tissue, and biopsy.
[0442] 152. A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample containing a nucleic acid of interest and a nucleic acid targeted for depletion; at least a subset of the nucleic acids of interest or at least a subset of the nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme; the activity of a first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to its cognate recognition site; b. terminally dephosphorylating a plurality of nucleic acids in the sample; c. contacting the sample from (b) with a first modification-sensitive restriction enzyme under conditions that allow for cleavage of at least some of the first modification-sensitive restriction sites in nucleic acids in the sample; d. contacting the sample from (c) with the adaptors under conditions that allow ligation of the adaptors to the 5' and 3' ends of a plurality of nucleic acids of interest; thereby producing a sample enriched for the nucleic acid of interest that is adaptor-linked at its 5' and 3' ends.
[0443] 153. The method of embodiment 152, wherein prior to (a), the nucleic acid of interest and the nucleic acid targeted for depletion are fragmented.
[0444] 154. The method of embodiment 152 or 153, wherein both the nucleic acid of interest and the nucleic acid targeted for depletion each comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme.
[0445] 155. The method of embodiment 154, wherein the frequency of nucleotide modifications within or adjacent to the plurality of first recognition sites is not the same in the nucleic acid of interest as in the nucleic acid targeted for depletion.
[0446] 156. The method of embodiment 155, wherein the plurality of first recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of first recognition sites in the nucleic acid of interest.
[0447] 157. The method of embodiment 155 or 156, wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AatII, AccII, Aor13HI, Aor51HI, BspT104I, BssHII, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, MluI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, SnaBI, AluI, and Sau3AI.
[0448] 158. The method of embodiment 155 or 156, wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AluI and Sau3AI. Another aspect of the present invention may be as follows. [1] A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample comprising nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids of interest or at least a subset of the nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme; b. terminally dephosphorylating a plurality of said nucleic acids in said sample; c. contacting the sample from (b) with the first modification-sensitive restriction enzyme under conditions that allow for cleavage of at least some of the first modification-sensitive restriction sites in the nucleic acids in the sample; d. contacting the sample from (c) with the adaptors under conditions that allow ligation of the adaptors to the 5' and 3' ends of a plurality of the nucleic acids of interest; thereby generating a sample enriched in target nucleic acids that are adapter-linked at the 5' and 3' ends. [2] The method according to [1], wherein both the target nucleic acid and the nucleic acid targeted for depletion contain multiple first recognition sites for the first modification-sensitive restriction enzyme. [3] The method according to [2], wherein the frequency of nucleotide modifications within or adjacent to the plurality of first recognition sites is not the same in the target nucleic acid as in the nucleic acid targeted for depletion. [4] The method according to any one of [1] to [3], wherein the activity of the first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to its cognate recognition site. [5] The method according to [4], wherein the plurality of first recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of first recognition sites in the target nucleic acid. [6] The method according to [4] or [5], wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AatII, AccII, Aor13HI, Aor51HI, BspT104I, BssHII, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, MluI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, SnaBI, AluI, and Sau3AI. [7] The method according to [4] or [5], wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AluI and Sau3AI. [8] The method according to any one of [1] to [3] above, wherein the first modification-sensitive restriction enzyme is active at a recognition site containing at least one modified nucleotide, but is not active at a recognition site that does not contain at least one modified nucleotide. [9] The method according to [8], wherein the plurality of first recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of first recognition sites in the target nucleic acid.
[10] The method according to [8] or [9], wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AbaSI, FspEI, LpnPI, MspJI, or McrBC.
[11] The modification includes 5-hydroxymethylcytosine; The method of [8] or [9], wherein the first modification-sensitive restriction enzyme comprises AbaSI, and the method further comprises contacting the sample with T4 phage β-glucosyltransferase prior to step (c).
[12] The method according to [8] or [9], wherein the modification comprises glucosylhydroxymethylcytosine and the first modification-sensitive restriction enzyme comprises AbaSI.
[13] The method according to [8] or [9], wherein the modification comprises methylcytosine and the first modification-sensitive restriction enzyme comprises McrBC.
[14] The method according to any one of [8] to
[13] , wherein the target nucleic acid contains at least one DpnI recognition site, and the method further comprises, prior to step (c), contacting the sample with DpnI and T4 polymerase, thereby replacing methylated A and C nucleotides with unmethylated A and C nucleotides within or adjacent to the at least one DpnI recognition site.
[15] The method according to any one of [8] to
[14] above, further comprising, prior to step (d), contacting the sample from (c) with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated end of the nucleic acid.
[16] e. contacting the adapter-ligated nucleic acid from (d) with a second modification-sensitive restriction enzyme under conditions that allow the second modification-sensitive restriction enzyme to cleave the second recognition site; at least a subset of the nucleic acids targeted for depletion comprises a plurality of second recognition sites for a second modification-sensitive restriction enzyme; the second modification-sensitive restriction enzyme targets recognition sites that include at least one modified nucleotide and does not target recognition sites that do not include at least one modified nucleotide; This results in the generation of a collection of nucleic acids targeted for depletion that are adapter-linked at one end, and a collection of target nucleic acids that are adapter-linked at both ends.
[17] The method according to any one of [1] to
[16] , further comprising contacting the sample with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes after step (d), wherein the gNAs are complementary to target sites of the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adapter-linked at one end, and nucleic acids of interest that are adapter-linked at both the 5' and 3' ends.
[18] The method according to any one of [1] to
[17] , further comprising amplifying, sequencing, or cloning the target nucleic acid that is adapter-linked at the 5' and 3' ends using the adapter.
[19] The method according to any one of [1] to
[18] above, wherein the nucleotide modification comprises an adenine modification or a cytosine modification.
[20] The method according to
[19] , wherein the adenine modification or cytosine modification comprises methylation.
[21] The method according to
[19] , wherein the cytosine modification includes 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxylcytosine, 5-glucosylhydroxymethylcytosine, or 3-methylcytosine.
[22] The method according to any one of
[16] to
[21] , wherein the second modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AbaSI, FspEI, LpnPI, MspJI, or McrBC.
[23] The method according to any one of [1] to
[22] above, wherein the nucleic acid targeted for depletion comprises a host nucleic acid and the target nucleic acid comprises a non-host nucleic acid.
[24] The method according to
[23] , wherein the non-host comprises a bacterium, a fungus, or a virus.
[25] The method described in
[23] , wherein the non-host comprises organisms of multiple species.
[26] The method according to
[23] , wherein the host is a mammal, a bird, a reptile or an insect.
[27] The method according to
[26] , wherein the mammal is a human.
[28] The method according to any one of [1] to
[27] above, wherein the nucleic acid targeted for depletion comprises a transcriptional activation site and the target nucleic acid comprises a repetitive sequence.
[29] The method according to any one of [1] to
[28] above, wherein the target nucleic acid to which the adapter is linked and the nucleic acid targeted for depletion are in the range of 50 to 1000 bp.
[30] The method according to any one of [1] to
[29] , wherein the sample is any one of a biological sample, a clinical sample, a forensic sample, and an environmental sample.
[31] A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample containing nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids targeted for depletion contain multiple recognition sites for a modification-sensitive restriction enzyme; b. terminally dephosphorylating a plurality of said nucleic acids in said sample; c. contacting the sample from (b) with the modification-sensitive restriction enzyme under conditions that allow for cleavage of the modification-sensitive restriction site of the nucleic acid in the sample, thereby producing nucleic acid with an exposed terminal phosphate; d. contacting the sample with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated ends of the nucleic acids, thereby producing a sample enriched in the nucleic acid of interest.
[32] The method of
[31] , wherein both the target nucleic acid and the nucleic acid targeted for depletion contain multiple recognition sites for the modification-sensitive restriction enzyme.
[33] The method according to
[32] , wherein the plurality of recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of recognition sites in the nucleic acid of interest.
[34] The method of any one of
[31] to
[33] , wherein the target nucleic acid contains at least one DpnI recognition site, and the method further comprises, prior to step (c), contacting the sample with DpnI and T4 polymerase, thereby replacing methylated A and C nucleotides with unmethylated A and C nucleotides within or adjacent to the at least one DpnI recognition site.
[35] The method according to any one of
[31] to
[34] above, wherein the modification comprises an adenine modification or a cytosine modification.
[36] The method according to
[35] , wherein the adenine modification or cytosine modification comprises methylation.
[37] The method according to
[35] , wherein the cytosine modification includes 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxylcytosine, 5-glucosylhydroxymethylcytosine, or 3-methylcytosine.
[38] The method according to any one of
[31] to
[37] , wherein the modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AbaSI, FspEI, LpnPI, MspJI, or McrBC.
[39] The method according to any one of
[31] to
[34] , wherein the modification comprises 5-hydroxymethylcytosine, the modification-sensitive restriction enzyme comprises AbaSI, and the method further comprises contacting the sample with T4 phage β-glucosyltransferase prior to step (c).
[40] The method according to any one of
[31] to
[34] , wherein the modification comprises glucosylhydroxymethylcytosine and the modification-sensitive restriction enzyme comprises AbaSI.
[41] The method according to any one of
[31] to
[34] , wherein the modification comprises methylcytosine and the modification-sensitive restriction enzyme comprises McrBC.
[42] e. contacting the sample from (d) with the adaptors under conditions that allow ligation of the adaptors to the 5' and 3' ends of a plurality of the nucleic acids of interest; The method according to any one of
[31] to
[41] above, further comprising generating a sample enriched in target nucleic acids that are adapter-linked at the 5' and 3' ends.
[43] The method of any one of
[31] to
[42] , further comprising contacting the sample with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes after step (d), wherein the gNAs are complementary to target sites of the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adapter-linked at one end, and nucleic acids of interest that are adapter-linked at both the 5' and 3' ends.
[44] The method according to any one of
[31] to
[43] , further comprising amplifying, sequencing, or cloning the target nucleic acid that is adapter-linked at the 5' and 3' ends using the adapter.
[45] The method according to any one of
[31] to
[44] , wherein the nucleic acid targeted for depletion comprises a host nucleic acid and the target nucleic acid comprises a non-host nucleic acid.
[46] The method described in
[45] , wherein the non-host comprises a bacterium, a fungus, or a virus.
[47] The method described in
[45] above, wherein the host is a human.
[48] The method according to any one of
[31] to
[47] , wherein the nucleic acid targeted for depletion comprises a transcriptional activation site and the target nucleic acid comprises a repetitive sequence.
[49] The method according to any one of
[31] to
[48] , wherein the target nucleic acid to which the adapter is linked and the nucleic acid targeted for depletion are in the range of 50 to 1000 bp.
[50] The method according to any one of
[31] to
[49] , wherein the sample is any one of a biological sample, a clinical sample, a forensic sample, and an environmental sample.
[51] A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample containing nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids targeted for depletion contain multiple recognition sites for a modification-sensitive restriction enzyme; b. contacting the sample with the adaptor under conditions that allow ligation of the adaptor to the 5' and 3' ends of a plurality of the nucleic acids in the sample; c. contacting the sample from (b) with the modification-sensitive restriction enzyme under conditions that allow cleavage of the modification-sensitive restriction site of the nucleic acid in the sample; thereby generating a sample enriched in target nucleic acids that are adapter-linked at the 5' and 3' ends.
[52] The method of
[51] , wherein both the target nucleic acid and the nucleic acid targeted for depletion contain multiple recognition sites for the modification-sensitive restriction enzyme.
[53] The method of
[51] or
[52] , wherein the plurality of recognition sites in the nucleic acid targeted for depletion are modified more frequently than the plurality of recognition sites in the nucleic acid of interest.
[54] The method of any one of
[51] to
[53] , wherein the target nucleic acid contains at least one DpnI recognition site, and the method further comprises, prior to step (c), contacting the sample with DpnI and T4 polymerase, thereby replacing methylated A and C nucleotides with unmethylated A and C nucleotides within or adjacent to at least one DpnI recognition site.
[55] The method according to any one of
[51] to
[54] above, wherein the modification comprises an adenine modification or a cytosine modification.
[56] The method according to
[55] , wherein the adenine modification or cytosine modification comprises methylation.
[57] The method according to
[55] , wherein the cytosine modification includes 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxylcytosine, 5-glucosylhydroxymethylcytosine, or 3-methylcytosine.
[58] The method according to any one of
[51] to
[57] above, wherein the modification-sensitive restriction enzyme comprises AbaSI, FspEI, LpnPI, MspJI, or McrBC.
[59] The method of any one of
[51] to
[53] , wherein the modification comprises 5-hydroxymethylcytosine, the modification-sensitive restriction enzyme comprises AbaSI, and the method further comprises, before (c), contacting the sample with T4 phage β-glucosyltransferase.
[60] The method according to any one of
[51] to
[53] , wherein the modification comprises glucosylhydroxymethylcytosine and the modification-sensitive restriction enzyme comprises AbaSI.
[61] The method according to any one of
[51] to
[53] , wherein the modification comprises methylcytosine and the modification-sensitive restriction enzyme comprises McrBC.
[62] The method of any one of
[51] to
[61] , further comprising contacting the sample with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes after step (c), wherein the gNAs are complementary to target sites of the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adapter-linked at one end, and nucleic acids of interest that are adapter-linked at both the 5' and 3' ends.
[63] The method according to any one of
[51] to
[62] , further comprising amplifying, sequencing or cloning the target nucleic acid that is adapter-linked at the 5' and 3' ends using the adapter.
[64] The method according to any one of
[51] to
[63] , wherein the nucleic acid targeted for depletion comprises a host nucleic acid and the target nucleic acid comprises a non-host nucleic acid.
[65] The method described in
[64] , wherein the non-host comprises a bacterium, a fungus, or a virus.
[66] The method described in
[65] above, wherein the host is a human.
[67] The method according to any one of
[51] to
[66] , wherein the nucleic acid targeted for depletion comprises a transcriptional activation site and the target nucleic acid comprises a repetitive sequence.
[68] The method according to any one of
[51] to
[67] , wherein the target nucleic acid to which the adapter is linked and the nucleic acid targeted for depletion are in the range of 50 to 1000 bp.
[69] The method according to any one of
[51] to
[68] , wherein the sample is any one of a biological sample, a clinical sample, a forensic sample, and an environmental sample.
[70] A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample containing a nucleic acid of interest and a nucleic acid targeted for depletion; at least one subset of the nucleic acids of interest or at least one subset of the nucleic acids targeted for depletion comprises a plurality of first recognition sites for a first modification-sensitive restriction enzyme; the activity of the first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to its cognate recognition site; b. terminally dephosphorylating a plurality of said nucleic acids in said sample; c. contacting the sample from (b) with the first modification-sensitive restriction enzyme under conditions that allow for cleavage of at least some of the first modification-sensitive restriction sites of the nucleic acids in the sample; d. contacting the sample from (c) with the adaptors under conditions that allow ligation of the adaptors to the 5' and 3' ends of a plurality of the nucleic acids of interest; thereby generating a sample enriched in target nucleic acids that are adapter-linked at the 5' and 3' ends.
Claims
1. 1. A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample comprising nucleic acids of interest and nucleic acids targeted for depletion, wherein at least a subset of the nucleic acids of interest or the subset of nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme; b. terminally dephosphorylating a plurality of said nucleic acids in said sample; c. contacting the sample from (b) with the first modification-sensitive restriction enzyme under conditions that allow for cleavage of at least some of the first modification-sensitive restriction sites in the nucleic acids in the sample; d. contacting the sample from (c) with the adaptors under conditions that allow for ligation of the adaptors to the 5' and 3' ends of a plurality of the nucleic acids of interest; thereby producing a sample enriched in 5' and 3' adaptor-ligated nucleic acids of interest; e. amplifying the 5'- and 3'-end adapter-linked nucleic acid of interest using the adapter; f. contacting the adaptor-ligated nucleic acid from (d) and prior to step (e) with a second modification-sensitive restriction enzyme under conditions that allow the second modification-sensitive restriction enzyme to cleave the second recognition site; at least a subset of the nucleic acids targeted for depletion comprise a plurality of second recognition sites for a second modification-sensitive restriction enzyme; the second modification-sensitive restriction enzyme targets recognition sites that include at least one modified nucleotide and does not target recognition sites that do not include at least one modified nucleotide; This results in the generation of a collection of nucleic acids targeted for depletion that are adapter-ligated at one end and a collection of nucleic acids of interest that are adapter-ligated at both ends.
2. 2. The method of claim 1, wherein both the nucleic acid of interest and the nucleic acid targeted for depletion comprise a plurality of first recognition sites for the first modification-sensitive restriction enzyme.
3. 3. The method of claim 2, wherein the frequency of nucleotide modifications within or adjacent to the plurality of first recognition sites is not the same in the nucleic acid of interest as in the nucleic acid targeted for depletion.
4. 4. The method of claim 1, wherein the activity of the first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to the first recognition site.
5. 5. The method of claim 4, wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AatII, AccII, Aor13HI, Aor51HI, BspT104I, BssHII, Cfr10I, ClaI, CpoI, Eco52I, HaeII, HapII, HhaI, MluI, NaeI, NotI, NruI, NsbI, PmaCI, Psp1406I, PvuI, SacII, SalI, SmaI, SnaBI, AluI, and Sau3AI.
6. 4. The method of claim 1, wherein the first modification-sensitive restriction enzyme is active at a recognition site that comprises at least one modified nucleotide and is not active at a recognition site that does not contain at least one modified nucleotide.
7. 7. The method of claim 6, wherein the first modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AbaSI, FspEI, LpnPI, MspJI, or McrBC.
8. the modification comprises 5-hydroxymethylcytosine; 7. The method of claim 6, wherein the first modification-sensitive restriction enzyme comprises AbaSI, and the method further comprises contacting the sample with T4 phage β-glucosyltransferase prior to step (c).
9. 7. The method of claim 6, wherein the modification comprises glucosylhydroxymethylcytosine and the first modification-sensitive restriction enzyme comprises AbaSI, or wherein the modification comprises methylcytosine and the first modification-sensitive restriction enzyme comprises McrBC.
10. 7. The method of claim 6, wherein the nucleic acid of interest comprises at least one DpnI recognition site, and the method further comprises, prior to step (c), contacting the sample with DpnI and T4 polymerase, thereby replacing methylated A and C nucleotides with unmethylated A and C nucleotides within or near the at least one DpnI recognition site.
11. 7. The method of claim 6, further comprising, prior to step (d), contacting the sample from (c) with an exonuclease under conditions that allow for the sequential removal of nucleotides from the phosphorylated ends of nucleic acids.
12. 4. The method of any one of claims 1 to 3, further comprising contacting the sample with a plurality of nucleic acid-guided nuclease-guide nucleic acid (gNA) complexes after step (d) and before step (e), wherein the gNAs are complementary to target sites of the nucleic acids targeted for depletion, thereby generating cleaved nucleic acids targeted for depletion that are adaptor-linked at one end, and nucleic acids of interest that are adaptor-linked at both the 5' and 3' ends.
13. 4. The method of claim 1, further comprising, after step (e), using the adapters to sequence or clone the nucleic acid of interest that is adapter-linked at the 5' and 3' ends.
14. The method of any one of claims 1 to 3, wherein the nucleotide modification comprises methylation of adenine or cytosine.
15. 13. The method of claim 12, wherein the second modification-sensitive restriction enzyme comprises a restriction enzyme selected from the group consisting of AbaSI, FspEI, LpnPI, MspJI, or McrBC.
16. 4. The method of any one of claims 1 to 3, wherein the nucleic acid targeted for depletion comprises a host nucleic acid and the nucleic acid of interest comprises a non-host nucleic acid.
17. 17. The method of claim 16, wherein the non-host comprises a bacterium, a fungus, or a virus.
18. 17. The method of claim 16, wherein the host is a mammal, a bird, a reptile, or an insect.
19. The method of any one of claims 1 to 3, wherein the sample is one of a biological sample, a clinical sample, a forensic sample, or an environmental sample.
20. 1. A method for enriching a sample for a nucleic acid of interest, comprising: a. providing a sample containing a nucleic acid of interest and a nucleic acid targeted for depletion; at least a subset of the nucleic acids of interest or the subset of nucleic acids targeted for depletion comprise a plurality of first recognition sites for a first modification-sensitive restriction enzyme; the activity of the first modification-sensitive restriction enzyme is blocked by modification of a nucleotide within or adjacent to the first recognition site; said providing; b. terminally dephosphorylating a plurality of said nucleic acids in said sample; c. contacting the sample from (b) with the first modification-sensitive restriction enzyme under conditions that allow for cleavage of at least some of the first modification-sensitive restriction sites of the nucleic acids in the sample; d. contacting the sample from (c) with the adaptors under conditions that allow for ligation of the adaptors to the 5' and 3' ends of a plurality of the nucleic acids of interest; thereby producing a sample enriched in 5' and 3' adaptor-ligated nucleic acids of interest; e. amplifying the 5'- and 3'-end adapter-linked nucleic acid of interest using the adapter; contacting the adaptor-ligated nucleic acid from (d) and prior to step (e) with a second modification-sensitive restriction enzyme under conditions that allow the second modification-sensitive restriction enzyme to cleave a second recognition site; at least a subset of the nucleic acids targeted for depletion comprise a plurality of second recognition sites for a second modification-sensitive restriction enzyme; the second modification-sensitive restriction enzyme targets recognition sites that include at least one modified nucleotide and does not target recognition sites that do not include at least one modified nucleotide; This results in the generation of a collection of nucleic acids targeted for depletion that are adapter-ligated at one end and a collection of nucleic acids of interest that are adapter-ligated at both ends.
Citation Information
Patent Citations
DNA size selection for chromatin analysis
JP2013541941A
Isolation of target nucleic acids from mixed nucleic acid samples
JP2014520530A
Detection of DNA from Special Cell Types and Related Methods
JP2017514499A
Differential enzymatic fragmentation by whole genome amplification
US20110076726A1
Methods and systems for evaluating DNA methylation in cell-free DNA
WO2019006269A1