Capture of nucleic acids using nucleic acid-guided nuclease-based systems

Nucleic acid-guided nuclease-based systems efficiently capture and enrich target sequences by using CRISPR/Cas proteins to ligate and cleave specific sites, addressing the inefficiencies of existing methods and improving sequencing accuracy.

JP7792241B2Active Publication Date: 2025-12-25ARC BIO LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021193224
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-08-19
Filing Date
2021-11-29
Publication Date
2025-12-25
Estimated Expiration
2036-08-18

AI Technical Summary

Technical Problem

Current methods for targeted sequencing, such as hybridization-based enrichment and multiplex PCR, are time-consuming, expensive, and inefficient in capturing specific nucleic acid regions, often resulting in off-target sequences.

Method used

A method using nucleic acid-guided nuclease-based systems, including CRISPR/Cas proteins and their variants, to selectively capture and enrich target nucleic acid sequences by ligating adaptors and cleaving or nicking specific sites, followed by amplification and labeling.

Benefits of technology

Enables efficient, rapid, and cost-effective capture of target nucleic acid sequences, reducing off-target sequences and improving sequencing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007792241000003
    Figure 0007792241000003
  • Figure 0007792241000004
    Figure 0007792241000004
  • Figure 0007792241000005
    Figure 0007792241000005
Patent Text Reader

Abstract

Methods and compositions are provided for the capture of nucleic acids using nucleic acid-guided nuclease-based systems. [Solution] Methods and compositions enable selective capture of nucleic acid sequences of interest. The nucleic acids can contain DNA or RNA, and the methods and compositions are particularly useful for working with complex nucleic acid samples. In one aspect, a method for capturing a target nucleic acid sequence includes the steps of: (a) providing a sample containing a plurality of adaptor-ligated nucleic acids, wherein the nucleic acids are ligated at one end to a first adaptor and at the other end to a second adaptor.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 62 / 207,359, filed August 19, 2015, which is incorporated herein by reference in its entirety. [Background technology]

[0002] Targeted sequencing of specific regions of the genome continues to be of interest to researchers, especially in clinical settings. Clinical diagnosis of genetic diseases, cancer, and many research projects rely on targeted sequencing to enable broad-coverage sequencing of targeted sites while reducing sequencing costs. Currently, the main methods used for this purpose are 1) hybridization-based enrichment and 2) multiplex PCR. In the former approach, short oligonucleotide probes labeled with biotin are used to "extract" sequences of interest from a library. This process can be time-consuming and expensive, requiring many hands-on steps. Furthermore, some "off-target" sequences often remain in the resulting product. While multiplex PCR-based approaches may be faster, they can be limited in the number of targets and expensive. There is a need for a method for efficient capture of nucleic acid regions of interest that is easy, specific, rapid, and inexpensive. Methods and compositions that address this need are provided herein.

[0003] All patents, patent applications, publications, documents, web links and articles cited herein are hereby incorporated by reference in their entirety. Summary of the Invention [Means for solving the problem]

[0004] Provided herein is a method and composition that allows selective capture of target nucleic acid sequence.Nucleic acid can contain DNA or RNA.Provided herein is a method and composition that is particularly useful for handling complex nucleic acid samples.

[0005] In one aspect, the present invention provides a method for capturing a target nucleic acid sequence, the method comprising: (a) providing a sample containing a plurality of adaptor-ligated nucleic acids, wherein the nucleic acids are ligated at one end to a first adaptor and at the other end to a second adaptor; (b) contacting the sample with a plurality of nucleic acid-guided nuclease-gNA complexes, thereby generating a plurality of nucleic acid fragments ligated at one end to a first or second adaptor and at the other end to no adaptor, wherein the gNAs are complementary to a targeting site of interest contained in a subset of the nucleic acids; and (c) contacting the plurality of nucleic acid fragments with a third adaptor, thereby generating a plurality of nucleic acid fragments ligated at one end to the first or second adaptor and at the other end to the third adaptor. In one embodiment, the nucleic acid-guided nuclease is a CRISPR / Cas protein. In one embodiment, the nucleic acid-guided nuclease is a non-CRISPR / Cas protein. In one embodiment, the nucleic acid-guided nuclease is selected from the group consisting of CAS class I type I, CAS class I type III, CAS class I type IV, CAS class II type II, and CAS class II type V. In one embodiment, the nucleic acid-guided nuclease is selected from the group consisting of Cas9, Cpf1, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn2, Cas4, Csm2, Cm5, Csf1, C2c2, and NgAgo. In one embodiment, the gNA is a gRNA. In one embodiment, the gNA is a gDNA. In one embodiment, contacting with a plurality of nucleic acid-guided nuclease-gNA complexes cleaves the targeting site of interest contained in a subset of the nucleic acids, thereby generating a plurality of nucleic acid fragments comprising a first or second adaptor at one end and no adaptor at the other end. In one embodiment, the method further comprises amplifying the product of step (c) using PCR specific for the first or second and third adapters.In one embodiment, the nucleic acid is selected from the group consisting of single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA, and DNA / RNA hybrids. In one embodiment, the nucleic acid is double-stranded DNA. In one embodiment, the nucleic acid is derived from genomic DNA. In one embodiment, the genomic DNA is human. In one embodiment, the adaptor-ligated end of the nucleic acid is 20 bp to 5000 bp in length. In one embodiment, the targeting site of interest is a single nucleotide polymorphism (SNP), short tandem repeat (STR), oncogene, insertion, deletion, structural variant, exon, gene mutation, or regulatory region. In one embodiment, the amplified product is used for cloning, sequencing, or genotyping. In one embodiment, the adaptor is 20 bp to 100 bp in length. In one embodiment, the adaptor comprises a primer binding site. In one embodiment, the adaptor comprises a sequencing adaptor or a restriction site. In one embodiment, the targeting site of interest accounts for less than 50% of the total nucleic acids in the sample. In one embodiment, the sample is obtained from a biological sample, a clinical sample, a forensic sample, or an environmental sample. In one embodiment, the first and second adapters are the same. In one embodiment, the first and second adapters are different. In one embodiment, the sample comprises a sequencing library.

[0006] In one aspect, the present invention provides a method for introducing labeled nucleotides into a target site of interest, the method comprising: (a) providing a sample containing a plurality of nucleic acid fragments; (b) contacting the sample with a plurality of nucleic acid-guided nuclease nickase-gNA complexes, thereby generating a plurality of nicked nucleic acid fragments at the target site of interest, wherein the gNAs are complementary to the target site of interest in the nucleic acid fragments; and (c) contacting the plurality of nicked nucleic acid fragments with an enzyme capable of initiating nucleic acid synthesis at the nick site and labeled nucleotides, thereby generating a plurality of nucleic acid fragments containing labeled nucleotides at the target site of interest. In one embodiment, the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. In one embodiment, the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Cse1 nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cm5 nickase, Csf1 nickase, C2C2 nickase, and NgAgo nickase. In one embodiment, the gNA is gRNA. In one embodiment, the gNA is gDNA. In one embodiment, the nucleic acid fragment is selected from the group consisting of a single-stranded DNA fragment, a double-stranded DNA fragment, a single-stranded RNA fragment, a double-stranded RNA fragment, and a DNA / RNA hybrid fragment. In one embodiment, the nucleic acid fragment is a double-stranded DNA fragment. In one embodiment, the double-stranded DNA fragment is derived from genomic DNA. In one embodiment, the genomic DNA is human. In one embodiment, the nucleic acid fragment is 20 bp to 5000 bp in length. In one embodiment, the target site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region.In one embodiment, the targeted site of interest accounts for less than 50% of the total nucleic acids in the sample. In one embodiment, the sample is obtained from a biological sample, a clinical sample, a forensic sample, or an environmental sample. In one embodiment, the labeled nucleotide is a biotinylated nucleotide. In one embodiment, the labeled nucleotide is part of an antibody-conjugate pair. In one embodiment, the method further comprises contacting the nucleic acid fragment comprising the biotinylated nucleotide with avidin or streptavidin, thereby capturing the targeted nucleic acid site of interest. In one embodiment, the enzyme capable of initiating nucleic acid synthesis at a nicked site is DNA polymerase I, Klenow fragment, TAQ polymerase, or Bst DNA polymerase. In one embodiment, the Cas9 nickase nicks the 5' end of the nucleic acid fragment. In one embodiment, the nucleic acid fragment is 20 bp to 5000 bp in length. In one embodiment, the sample is obtained from a biological sample, a clinical sample, a forensic sample, or an environmental sample.

[0007] In one aspect, the present invention provides a method for capturing a target nucleic acid sequence of interest, comprising: (a) providing a sample containing a plurality of adaptor-ligated nucleic acids, the nucleic acids being ligated at one end to a first adaptor and at the other end to a second adaptor; and (b) contacting the sample with a plurality of catalytically inactive nucleic acid-guided nuclease-gNA complexes, the catalytically inactive nucleic acid-guided nucleases being fused to a transposase, and the gRNAs being complementary to targeting sites of interest contained in a subset of the nucleic acids, and loading the complexes with a plurality of third adaptors to generate a plurality of nucleic acid fragments comprising the first or second adaptor at one end and the third adaptor at the other end. In one embodiment, the method further comprises amplifying the products of step (b) using PCR specific for the first or second and third adaptors. In one embodiment, the catalytically inactive nucleic acid-guided nucleases are derived from CRISPR / Cas system proteins. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is derived from a non-CRISPR / Cas system protein. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of inactive CAS class I type I, inactive CAS class I type III, inactive CAS class I type IV, inactive CAS class II type II, and inactive CAS class II type V. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of dCas9, dCpf1, dCas3, dCas8a-c, dCas10, dCsel, dCsy1, dCsn2, dCas4, dCsm2, dCm5, dCsf1, dC2C2, and dNgAgo. In one embodiment, the gNA is a gRNA. In one embodiment, the gNA is a gDNA. In one embodiment, the nucleic acid sequence is from genomic DNA. In one embodiment, the genomic DNA is human. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is fused to the N-terminus of the transposase.In one embodiment, the catalytically inactive nucleic acid-guided nuclease is fused to the C-terminus of the transposase. In one embodiment, the adapter-ligated nucleic acid is 20 bp to 5000 bp in length. In one embodiment, the contacting step (b) allows for insertion of the second adapter into the target nucleic acid sequence. In one embodiment, the targeting site of interest is a single nucleotide polymorphism (SNP), short tandem repeat (STR), oncogene, insertion, deletion, structural variant, exon, gene mutation, or regulatory region. In one embodiment, the amplified product is used for cloning, sequencing, or genotyping. In one embodiment, the adapter is 20 bp to 100 bp in length. In one embodiment, the adapter comprises a primer binding site. In one embodiment, the adapter comprises a sequencing adapter or a restriction site. In one embodiment, the targeting site of interest accounts for less than 50% of the total nucleic acids in the sample. In one embodiment, the sample is obtained from a biological sample, a clinical sample, a forensic sample, or an environmental sample.

[0008] In one aspect, the present invention provides a method for capturing a target nucleic acid sequence of interest, the method comprising the steps of: (a) providing a sample comprising a plurality of adaptor-ligated nucleic acids, wherein the nucleic acids are ligated to the adaptors at their 5' and 3' ends; (b) contacting the sample with a plurality of catalytically inactive nucleic acid-guided nuclease-gNA complexes, thereby generating a plurality of nucleic acids bound to the catalytically inactive nucleic acid-guided nuclease-gNA complexes and adapted at their 5' and 3' ends, wherein the gNAs are complementary to targeting sites of interest contained in a subset of the nucleic acids; and (c) contacting the sample with a plurality of catalytically inactive nucleic acid-guided nuclease-gNA complexes, thereby generating a plurality of nucleic acid fragments comprising a non-target nucleic acid sequence, adapted at only one of their 5' or 3' ends, wherein the gNAs are complementary to both targeting sites of interest and non-targeting sites in the nucleic acids. In one embodiment, the contacting step (c) does not displace the plurality of adaptor-ligated nucleic acids at their 5' and 3' ends that are bound to the CAS9-gRNA complexes of step (b). In one embodiment, the contacting step (d) cleaves the unintended targeting sites contained in a subset of the nucleic acids, thereby generating a plurality of nucleic acid fragments containing unintended nucleic acid sequences that are adaptor-ligated at only one of their 5' or 3' ends. In one embodiment, the method further comprises removing the bound CAS9-gRNA complexes and amplifying the product of (b) using adaptor-specific PCR. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is a CRISPR / Cas protein. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is a non-CRISPR / Cas protein. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of inactive CAS class I type I, inactive CAS class I type III, inactive CAS class I type IV, inactive CAS class II type II, and inactive CAS class II type V.In one embodiment, the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of dCas9, dCpf1, dCas3, dCas8a-c, dCas10, dCsel, dCsy1, dCsn2, dCas4, dCsm2, dCm5, dCsf1, dC2C2, dCPF1, and dNgAgo. In one embodiment, the gNA is a gRNA. In one embodiment, the gNA is a gDNA. In one embodiment, the nucleic acid is selected from the group consisting of single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA, and a DNA / RNA hybrid. In one embodiment, the nucleic acid is double-stranded DNA. In one embodiment, the nucleic acid is derived from genomic DNA. In one embodiment, the genomic DNA is human. In one embodiment, the nucleic acid ligated to adapters at the 5' and 3' ends is 20 bp to 5000 bp in length. In one embodiment, the target site of interest is a single nucleotide polymorphism (SNP), short tandem repeat (STR), oncogene, insertion, deletion, structural variant, exon, gene mutation, or regulatory region. In one embodiment, the amplified product is used for cloning, sequencing, or genotyping. In one embodiment, the adapter is 20 bp to 100 bp in length. In one embodiment, the adapter comprises a primer binding site. In one embodiment, the adapter comprises a sequencing adapter or a restriction site. In one embodiment, the target site of interest accounts for less than 50% of the total nucleic acids in the sample. In one embodiment, the sample is obtained from a biological sample, a clinical sample, a forensic sample, or an environmental sample.

[0009] In one aspect, the present invention provides a method for capturing target DNA sequences of interest, comprising the steps of: (a) providing a sample comprising a plurality of nucleic acid sequences, wherein the nucleic acid sequences comprise methylated nucleotides, and wherein the nucleic acid sequences are ligated to adaptors at their 5' and 3' ends; (b) contacting the sample with a plurality of nucleic acid-guided nucleases, nickase-gNA complexes, thereby generating a plurality of nicked sites of interest in a subset of the nucleic acid sequences, wherein the gNAs are complementary to targeting sites of interest in the subset of nucleic acid sequences, and wherein the target nucleic acid sequences are ligated to adaptors at their 5' and 3' ends; and (c) (d) contacting the sample with an enzyme capable of cleaving methylated nucleic acids, thereby generating a plurality of nucleic acid fragments comprising methylated nucleic acids, wherein the plurality of nucleic acid fragments comprising methylated nucleic acids are ligated to an adaptor at at most one of the 5' and 3' ends. In one embodiment, the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. In one embodiment, the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cm5 nickase, Csf1 nickase, C2c2 nickase, and NgAgo nickase. In one embodiment, the gNA is gRNA. In one embodiment, the gNA is gDNA. In one embodiment, the DNA is double-stranded DNA.In one embodiment, the double-stranded DNA is derived from genomic DNA. In one embodiment, the genomic DNA is human. In one embodiment, the DNA sequence is 20 bp to 5000 bp in length. In one embodiment, the target site of interest is a single nucleotide polymorphism (SNP), short tandem repeat (STR), oncogene, insertion, deletion, structural variant, exon, gene mutation, or regulatory region. In one embodiment, the target site of interest accounts for less than 50% of the total DNA in the sample. In one embodiment, the sample is obtained from a biological sample, a clinical sample, a forensic sample, or an environmental sample. In one embodiment, the enzyme capable of initiating nucleic acid synthesis at a nicked site is DNA polymerase I, Klenow fragment, TAQ polymerase, or Bst DNA polymerase. In one embodiment, the nucleic acid-guided nuclease, nickase, nicks the 5' end of the DNA sequence. In one embodiment, the enzyme capable of cleaving methylated DNA is DpnI.

[0010] In one aspect, the present invention provides a method for capturing a target DNA sequence of interest, comprising: (a) contacting the sample with a plurality of nucleic acid-guided nucleases, nickase-gNA complexes, thereby generating a plurality of nicked DNAs at sites adjacent to the region of interest, wherein the gNAs are complementary to targeting sites of interest adjacent to the region of interest in a subset of the DNA sequences; (b) heating to 65°C to generate double-stranded breaks at the adjacent nicks; (c) contacting these double-stranded breaks with a thermostable ligase, thereby allowing adapter sequences to be ligated only to these sites; and (d) repeating these three steps to place a second adapter on the opposite side of the region of interest, thereby allowing enrichment of the region of interest. In one embodiment, the gNA is a gRNA. In one embodiment, the gNA is a gDNA. In one embodiment, the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. In one embodiment, the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cm5 nickase, Csf1 nickase, C2C2 nickase, and NgAgo nickase. In one embodiment, the DNA is double-stranded DNA. In one embodiment, the double-stranded DNA is derived from genomic DNA. In one embodiment, the genomic DNA is human. In one embodiment, the DNA sequence is 20 bp to 5000 bp in length. In one embodiment, the targeting site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region. In one embodiment, the targeting site of interest accounts for less than 50% of the total DNA in the sample.In one embodiment, the sample is obtained from a biological sample, a clinical sample, a forensic sample, or an environmental sample. In one embodiment, the ligase capable of catalyzing the double-stranded break is a thermostable 5'App DNA / RNA ligase or a T4 RNA ligase. In one embodiment, the nucleic acid-guided nuclease, a nickase, nicks the 5' end of the DNA sequence. In one embodiment, the enzyme capable of cleaving methylated DNA is DpnI.

[0011] In one aspect, the present invention provides a method for enriching a sample for a sequence of interest, the method comprising: (a) providing a sample containing a sequence of interest and a targeting sequence for depletion, wherein the sequence of interest comprises less than 50% of the sample; and (b) contacting the sample with a plurality of nucleic acid-guided RNA endonuclease-gRNA complexes or a plurality of nucleic acid-guided DNA endonuclease-gDNA complexes, wherein the gRNAs and gDNAs are complementary to the targeting sequences, thereby cleaving the targeting sequences. In one embodiment, the method further comprises extracting the sequence of interest and the targeting sequence for depletion from the sample. In one embodiment, the method further comprises fragmenting the extracted sequence. In one embodiment, the cleaved targeting sequence is removed by size exclusion. In one embodiment, the sample is any one of a biological sample, a clinical sample, a forensic sample, and an environmental sample. In one embodiment, the sample contains a host nucleic acid sequence targeted for depletion and a non-host nucleic acid sequence of interest. In one embodiment, the non-host nucleic acid sequence comprises a microbial nucleic acid sequence. In one embodiment, the microbial nucleic acid sequence is a nucleic acid sequence from a bacterium, virus, or eukaryotic parasite. In one embodiment, the gRNA and gDNA are complementary to a ribosomal RNA sequence, a spliced ​​transcript, an unspliced ​​transcript, an intron, an exon, or a non-coding RNA. In one embodiment, the extracted nucleic acid comprises single-stranded or double-stranded RNA. In one embodiment, the extracted nucleic acid comprises single-stranded or double-stranded DNA. In one embodiment, the sequence of interest comprises less than 10% of the extracted nucleic acid. In one embodiment, the nucleic acid-guided RNA endonuclease comprises C2c2. In one embodiment, the C2c2 is catalytically inactive. In one embodiment, the nucleic acid-guided DNA endonuclease comprises NgAgo. In one embodiment, the NgAgo is catalytically inactive. In one embodiment, the sample is selected from whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bone, fingernails, feces, urine, tissue, and biopsy material.

[0012] In another aspect, the present invention provides a method for enriching a sample, the method comprising: (a) providing a sample containing host nucleic acid and non-host nucleic acid; (b) contacting the sample with a plurality of nucleic acid-guided RNA endonuclease-gRNA complexes or a plurality of nucleic acid-guided DNA endonuclease-gDNA complexes, wherein the gRNAs and gDNAs are complementary to target sites of the host nucleic acid; and (c) enriching the sample for non-host nucleic acid. In one embodiment, the nucleic acid-guided RNA endonuclease comprises C2c2. In one embodiment, the nucleic acid-guided RNA endonuclease comprises catalytically inactive C2c2. In one embodiment, the nucleic acid-guided DNA endonuclease comprises NgAgo. In one embodiment, the nucleic acid-guided DNA endonuclease comprises catalytically inactive NgAgo. In one embodiment, the host is selected from the group consisting of human, bovine, equine, ovine, porcine, monkey, canine, feline, gerbil, avian, mouse, and rat. In one embodiment, the non-host organism is a prokaryote. In one embodiment, the non-host organism is selected from the group consisting of eukaryotes, viruses, bacteria, fungi, and protozoa. In one embodiment, the adaptor-ligated host nucleic acid and non-host nucleic acid are within the length range of 50 bp to 1000 bp. In one embodiment, the non-host nucleic acid constitutes less than 50% of the total nucleic acid in the sample. In one embodiment, the sample is one of a biological sample, a clinical sample, a forensic sample, or an environmental sample. In one embodiment, step (c) comprises reverse transcribing the product of step (b) into cDNA. In one embodiment, step (c) comprises removing the host nucleic acid by size exclusion. In one embodiment, step (c) comprises removing the host nucleic acid by the use of biotin. In one embodiment, the sample is selected from whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bone, fingernail, feces, urine, tissue, and biopsy.

[0013] In another aspect, the present invention provides a method for using a nucleic acid-guided RNA endonuclease to enrich for a target in an RNA sample using a labeled catalytically inactive nucleic acid-guided RNA endonuclease protein. In some embodiments, the nucleic acid-guided RNA endonuclease protein targets HIV RNA in a blood RNA sample, and host RNA is washed away. In some embodiments, the nucleic acid-guided RNA endonuclease is C2c2.

[0014] In another aspect, the present invention provides a composition comprising a nucleic acid fragment, a nucleic acid-guided nuclease nickase-gNA complex, and a labeled nucleotide. In one embodiment, the nucleic acid comprises DNA. In one embodiment, the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. In one embodiment, the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cm5 nickase, Csf1 nickase, C2C2 nickase, and NgAgo nickase. In one embodiment, the gNA is a gRNA. In one embodiment, the gNA is a gDNA. In one embodiment, the nucleic acid fragment comprises DNA. In one embodiment, the nucleic acid fragment comprises RNA. In one embodiment, the nucleotide is labeled with biotin. In one embodiment, the nucleotide is part of an antibody conjugate pair.

[0015] In another aspect, the present invention provides a composition comprising a nucleic acid fragment and a catalytically inactive nucleic acid-guided nuclease-gNA complex, wherein the catalytically inactive nucleic acid-guided nuclease is fused to a transposase. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of inactive CAS class I type I, inactive CAS class I type III, inactive CAS class I type IV, inactive CAS class II type II, and inactive CAS class II type V. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of dCas9, dCpf1, dCas3, dCas8a-c, dCas10, dCsel, dCsy1, dCsn2, dCas4, dCsm2, dCm5, dCsf1, dC2C2, and dNgAgo. In one embodiment, the gNA is a gRNA. In one embodiment, the gNA is a gDNA. In one embodiment, the nucleic acid fragment comprises DNA. In one embodiment, the nucleic acid fragment comprises RNA. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is fused to the N-terminus of the transposase. In one embodiment, the catalytically inactive nucleic acid-guided nuclease is fused to the C-terminus of the transposase. In one embodiment, the composition comprises a DNA fragment and a dCas9-gRNA complex, wherein dCas9 is fused to a transposon.

[0016] In one aspect, the present invention provides a composition comprising a nucleic acid fragment containing a methylated nucleotide, a nucleic acid-guided nuclease (nickase-gNA) complex, and unmethylated nucleotides. In one embodiment, the nucleic acid-guided nuclease (nickase) is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. In one embodiment, the nucleic acid-guided nuclease (nickase) is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cm5 nickase, Csf1 nickase, C2C2 nickase, and NgAgo nickase. In one embodiment, the gNA is a gRNA. In one embodiment, the gNA is gDNA. In one embodiment, the nucleic acid fragment comprises DNA. In one embodiment, the nucleic acid fragment comprises RNA. In one embodiment, the nucleotide is labeled with biotin. In one embodiment, the nucleotide is part of an antibody conjugate pair. In one embodiment, the composition comprises a DNA fragment containing a methylated nucleotide, a Cas9-gRNA complex that is a nickase, and an unmethylated nucleotide. In an embodiment of the present invention, for example, the following items are provided: (Item 1) 1. A method for capturing a target nucleic acid sequence, comprising: (a) providing a sample comprising a plurality of adaptor-ligated nucleic acids, wherein the nucleic acids are ligated at one end to a first adaptor and at the other end to a second adaptor; (b) contacting the sample with a plurality of nucleic acid-guided nuclease-gNA complexes, thereby generating a plurality of nucleic acid fragments ligated at one end to a first or second adaptor and lacking an adaptor at the other end, wherein the gNAs are complementary to targeting sites of interest contained in the subset of nucleic acids; and (c) contacting the plurality of nucleic acid fragments with a third adaptor, thereby generating a plurality of nucleic acid fragments ligated at one end to the first or second adaptor and at the other end to the third adaptor. A method comprising: (Item 2) 2. The method of claim 1, wherein the nucleic acid-guided nuclease is a CRISPR / Cas system protein. (Item 3) 2. The method of claim 1, wherein the nucleic acid-guided nuclease is a non-CRISPR / Cas system protein. (Item 4) 2. The method of claim 1, wherein the nucleic acid-guided nuclease is selected from the group consisting of CAS class I type I, CAS class I type III, CAS class I type IV, CAS class II type II, and CAS class II type V. (Item 5) 2. The method of claim 1, wherein the nucleic acid-guided nuclease is selected from the group consisting of Cas9, Cpf1, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn2, Cas4, Csm2, Cmr5, Csf1, C2c2, and NgAgo. (Item 6) 6. The method of any one of items 1 to 5, wherein the gNA is a gRNA. (Item 7) 6. The method of any one of items 1 to 5, wherein the gNA is gDNA. (Item 8) 8. The method of any one of items 1 to 7, wherein the step of contacting with the plurality of nucleic acid-guided nuclease-gNA complexes cleaves the targeting sites of interest contained in a subset of the nucleic acids, thereby generating a plurality of nucleic acid fragments comprising a first or second adaptor at one end and no adaptor at the other end. (Item 9) 9. The method of any one of items 1 to 8, further comprising amplifying the product of step (c) using PCR specific for the first or the second and third adapters. (Item 10) 10. The method of any one of items 1 to 9, wherein the nucleic acid is selected from the group consisting of single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA and DNA / RNA hybrids. (Item 11) 11. The method of claim 10, wherein the nucleic acid is double-stranded DNA. (Item 12) 12. The method of any one of items 1 to 11, wherein the nucleic acid is from genomic DNA. (Item 13) Item 13. The method of item 12, wherein the genomic DNA is human. (Item 14) 14. The method of any one of items 1 to 13, wherein the nucleic acid that is ligated to an adaptor is between 20 bp and 5000 bp in length. (Item 15) 15. The method of any one of items 1 to 14, wherein the target site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region. (Item 16) 10. The method of item 9, wherein the amplified product is used for cloning, sequencing, or genotyping. (Item 17) 17. The method of any one of items 1 to 16, wherein the adapter is 20 bp to 100 bp in length. (Item 18) 18. The method of any one of items 1 to 17, wherein the adapter comprises a primer binding site. (Item 19) 19. The method of any one of items 1 to 18, wherein the adapter comprises a sequencing adapter or a restriction site. (Item 20) 20. The method of any one of items 1 to 19, wherein the targeted site of interest accounts for less than 50% of the total nucleic acid in the sample. (Item 21) 21. The method of any one of items 1 to 20, wherein the sample is obtained from a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 22) 22. The method of any one of items 1 to 21, wherein the first and second adapters are identical. (Item 23) 22. The method of any one of items 1 to 21, wherein the first and second adapters are different. (Item 24) 24. The method of any one of items 1 to 23, wherein the sample comprises a sequencing library. (Item 25) 1. A method for introducing a labeled nucleotide into a target site of interest, comprising: (a) providing a sample containing a plurality of nucleic acid fragments; (b) contacting the sample with a plurality of nucleic acid-guided nucleases, nickase-gNA complexes, thereby generating a plurality of nicked nucleic acid fragments at the target sites of interest, wherein the gNAs are complementary to the target sites of interest in the nucleic acid fragments; and (c) contacting the plurality of nicked nucleic acid fragments with an enzyme capable of initiating nucleic acid synthesis at the sites of nicking and with labeled nucleotides, thereby generating a plurality of nucleic acid fragments comprising labeled nucleotides at the targeted sites of interest. A method comprising: (Item 26) 26. The method of claim 25, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. (Item 27) 26. The method of claim 25, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cm5 nickase, Csf1 nickase, C2C2 nickase, CPF1 nickase, and NgAgo nickase. (Item 28) 28. The method of any one of items 25 to 27, wherein the gNA is a gRNA. (Item 29) 28. The method of any one of items 25 to 27, wherein the gNA is gDNA. (Item 30) 30. The method of any one of items 25 to 29, wherein the nucleic acid fragment is selected from the group consisting of a single-stranded DNA fragment, a double-stranded DNA fragment, a single-stranded RNA fragment, a double-stranded RNA fragment and a DNA / RNA hybrid fragment. (Item 31) 31. The method of item 30, wherein the nucleic acid fragment is a double-stranded DNA fragment. (Item 32) 32. The method of item 31, wherein the double-stranded DNA fragment is from genomic DNA. (Item 33) 33. The method of claim 32, wherein the genomic DNA is human. (Item 34) 34. The method of any one of items 25 to 33, wherein the nucleic acid fragment is 20 bp to 5000 bp in length. (Item 35) 35. The method of any one of items 25 to 34, wherein the target site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region. (Item 36) 36. The method of any one of items 25 to 35, wherein the targeted site of interest accounts for less than 50% of the total nucleic acid in the sample. (Item 37) 37. The method of any one of items 25 to 36, wherein the sample is obtained from a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 38) 38. The method of any one of items 25 to 37, wherein the labeled nucleotide is a biotinylated nucleotide. (Item 39) 38. The method of any one of items 25 to 37, wherein the labeled nucleotide is part of an antibody conjugate pair. (Item 40) 39. The method of claim 38, further comprising contacting the nucleic acid fragment containing biotinylated nucleotides with avidin or streptavidin, thereby capturing the targeting site of interest. (Item 41) 41. The method of any one of items 25 to 40, wherein the enzyme capable of initiating nucleic acid synthesis at the nicked site is DNA polymerase I, Klenow fragment, TAQ polymerase or Bst DNA polymerase. (Item 42) 42. The method of any one of items 25 to 41, wherein the nucleic acid-guided nuclease, nickase, nicks the 5' end of the nucleic acid fragment. (Item 43) 43. The method of any one of items 25 to 42, wherein the nucleic acid fragment is between 20 bp and 5000 bp in length. (Item 44) 44. The method of any one of items 25 to 43, wherein the sample is obtained from a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 45) 1. A method for capturing a target nucleic acid sequence of interest, comprising: (a) providing a sample comprising a plurality of adaptor-ligated nucleic acids, the nucleic acids being ligated at one end to a first adaptor and at the other end to a second adaptor; and (b) contacting the sample with a plurality of catalytically inactive nucleic acid-guided nuclease-gNA complexes, wherein the catalytically inactive nucleic acid-guided nucleases are fused to a transposase and the gRNAs are complementary to targeting sites of interest contained in the subset of nucleic acids, and loading the complexes with a plurality of third adapters to generate a plurality of nucleic acid fragments comprising a first or second adapter at one end and a third adapter at the other end. A method comprising: (Item 46) 46. ​​The method of claim 45, further comprising amplifying the product of step (b) using PCR specific for the first or second and third adapters. (Item 47) 46. ​​The method of claim 45, wherein the catalytically inactive nucleic acid-guided nuclease is derived from a CRISPR / Cas system protein. (Item 48) 46. ​​The method of claim 45, wherein the catalytically inactive nucleic acid-guided nuclease is derived from a non-CRISPR / Cas system protein. (Item 49) Item 46. The method of item 45, wherein the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of inactive CAS class I type I, inactive CAS class I type III, inactive CAS class I type IV, inactive CAS class II type II, and inactive CAS class II type V. (Item 50) 46. ​​The method of claim 45, wherein the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of dCas9, dCpf1, dCas3, dCas8a-c, dCas10, dCsel, dCsy1, dCsn2, dCas4, dCsm2, dCmr5, dCsf1, dC2C2, and dNgAgo. (Item 51) 51. The method of any one of items 45 to 50, wherein the gNA is a gRNA. (Item 52) 51. The method of any one of items 45 to 50, wherein the gNA is gDNA. (Item 53) 53. The method of any one of items 45 to 52, wherein the nucleic acid sequence is from genomic DNA. (Item 54) 54. The method of claim 53, wherein the genomic DNA is human. (Item 55) 55. The method of any one of items 45 to 54, wherein the catalytically inactive nucleic acid-guided nuclease is fused to the N-terminus of the transposase. (Item 56) 55. The method of any one of items 45 to 54, wherein the catalytically inactive nucleic acid-guided nuclease is fused to the C-terminus of the transposase. (Item 57) 57. The method of any one of items 45 to 56, wherein the nucleic acid that is ligated to an adaptor is 20 bp to 5000 bp in length. (Item 58) 58. The method of any one of items 45 to 57, wherein the contacting step of (b) allows insertion of the second adaptor into the targeted nucleic acid sequence. (Item 59) 59. The method of any one of items 45 to 58, wherein the target site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region. (Item 60) 47. The method of item 46, wherein the amplified product is used for cloning, sequencing or genotyping. (Item 61) 60. The method of any one of items 45 to 59, wherein the adapter is 20 bp to 100 bp in length. (Item 62) 62. The method of any one of items 45 to 61, wherein the adapter comprises a primer binding site. (Item 63) 63. The method of any one of items 45 to 62, wherein the adapter comprises a sequencing adapter or a restriction site. (Item 64) 64. The method of any one of items 45 to 63, wherein the targeted site of interest accounts for less than 50% of the total nucleic acid in the sample. (Item 65) 65. The method of any one of items 45 to 64, wherein the sample is obtained from a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 66) 1. A method for capturing a target nucleic acid sequence of interest, comprising: (a) providing a sample comprising a plurality of adaptor-ligated nucleic acids, wherein the nucleic acids are ligated to the adaptors at their 5' and 3' ends; (b) contacting the sample with a plurality of catalytically inactive nucleic acid-guided nuclease-gNA complexes, thereby generating a plurality of nucleic acids bound to the catalytically inactive nucleic acid-guided nuclease-gNA complexes and ligated to adapters at their 5' and 3' ends, wherein the gNAs are complementary to targeting sites of interest contained in the subset of nucleic acids; and (c) contacting the sample with a plurality of catalytically inactive nucleic acid-guided nuclease-gNA complexes, thereby generating a plurality of nucleic acid fragments comprising undesired nucleic acid sequences ligated to adapters at only one of the 5' or 3' ends, wherein the gNAs are complementary to both the desired targeting site and the undesired targeting site in the nucleic acid. A method comprising: (Item 67) 67. The method of claim 66, wherein the catalytically inactive nucleic acid-guided nuclease is a CRISPR / Cas system protein. (Item 68) 67. The method of claim 66, wherein the catalytically inactive nucleic acid-guided nuclease is a non-CRISPR / Cas system protein. (Item 69) 67. The method of item 66, wherein the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of inactive CAS class I type I, inactive CAS class I type III, inactive CAS class I type IV, inactive CAS class II type II, and inactive CAS class II type V. (Item 70) 67. The method of claim 66, wherein the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of dCas9, dCpf1, dCas3, dCas8a-c, dCas10, dCsel, dCsy1, dCsn2, dCas4, dCsm2, dCm5, dCsf1, dC2C2, dCPF1, and dNgAgo. (Item 71) 71. The method of any one of items 66 to 70, wherein the gNA is a gRNA. (Item 72) 71. The method of any one of items 66 to 70, wherein the gNA is gDNA. (Item 73) 73. The method of any one of items 66 to 72, wherein the contacting step (c) does not displace the plurality of adaptor-ligated nucleic acids at their 5' and 3' ends bound to the catalytically inactive nucleic acid-guided nuclease-gNA complex of step (b). (Item 74) 74. The method of any one of items 66 to 73, wherein the contacting step of (d) cleaves the undesired targeting sites contained in the subset of nucleic acids, thereby generating a plurality of nucleic acid fragments comprising undesired nucleic acid sequences ligated to adapters at only one of the 5' or 3' ends. (Item 75) 75. The method of any one of items 66 to 74, further comprising removing the bound catalytically inactive nucleic acid-guided nuclease-gNA complex and amplifying the product of (b) using adapter-specific PCR. (Item 76) 76. The method of any one of items 66 to 75, wherein the nucleic acid is selected from the group consisting of single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA and DNA / RNA hybrids. (Item 77) 77. The method of claim 76, wherein the nucleic acid is double-stranded DNA. (Item 78) 78. The method of any one of items 66 to 77, wherein the nucleic acid is from genomic DNA. (Item 79) 79. The method of claim 78, wherein the genomic DNA is human. (Item 80) 80. The method of any one of items 66 to 79, wherein the nucleic acid ligated to adapters at the 5' and 3' ends is between 20 bp and 5000 bp. (Item 81) 81. The method of any one of items 66 to 80, wherein the target site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region. (Item 82) 76. The method of item 75, wherein the amplified product is used for cloning, sequencing or genotyping. (Item 83) 83. The method of any one of items 66 to 82, wherein the adapter is 20 bp to 100 bp in length. (Item 84) 83. The method of any one of items 66 to 82, wherein the adapter comprises a primer binding site. (Item 85) 83. The method of any one of items 66 to 82, wherein the adapter comprises a sequencing adapter or a restriction site. (Item 86) 83. The method of any one of paragraphs 66 to 82, wherein the targeted site of interest accounts for less than 50% of the total nucleic acids in the sample. (Item 87) 83. The method of any one of items 66 to 82, wherein the sample is obtained from a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 88) 1. A method for capturing a target nucleic acid sequence of interest, comprising: (a) providing a sample comprising a plurality of nucleic acid sequences, wherein the nucleic acid sequences comprise methylated nucleotides, and the nucleic acid sequences are ligated to adaptors at their 5' and 3' ends; (b) contacting the sample with a plurality of nucleic acid-guided nucleases, nickase-gNA complexes, thereby generating a plurality of nicked sites of interest in a subset of the nucleic acid sequences, wherein the gNAs are complementary to target sites of interest in the subset of nucleic acid sequences, and the target nucleic acid sequences are ligated to adapters at their 5' and 3' ends; (c) contacting the sample with an enzyme capable of initiating DNA synthesis at the nicked site and an unmethylated nucleotide, thereby generating a plurality of nucleic acid sequences comprising an unmethylated nucleotide at the targeted site of interest, wherein the nucleic acid sequences are ligated to adapters at their 5' and 3' ends; and (d) contacting the sample with an enzyme capable of cleaving methylated nucleic acids, thereby generating a plurality of nucleic acid fragments comprising methylated nucleic acids, wherein the plurality of nucleic acid fragments comprising methylated nucleic acids are ligated to an adaptor at at most one of the 5' and 3' ends. A method comprising: (Item 89) Item 89. The method of item 88, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. (Item 90) 89. The method of claim 88, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cmr5 nickase, Csf1 nickase, C2C2 nickase, and NgAgo nickase. (Item 91) 89. The method of claim 88, wherein the gNA is a gRNA. (Item 92) 89. The method of claim 88, wherein the gNA is gDNA. (Item 93) 93. The method of any one of items 88 to 92, wherein the DNA is double-stranded DNA. (Item 94) 94. The method of item 93, wherein the double-stranded DNA is from genomic DNA. (Item 95) 95. The method of item 94, wherein the genomic DNA is human. (Item 96) 96. The method of any one of items 88 to 95, wherein the DNA sequence is 20 bp to 5000 bp in length. (Item 97) 97. The method of any one of items 88 to 96, wherein the target site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region. (Item 98) 98. The method of any one of items 88 to 97, wherein the targeted site of interest accounts for less than 50% of the total DNA in the sample. (Item 99) 99. The method of any one of items 88 to 98, wherein the sample is obtained from a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 100) 90. The method of any one of items 88 to 99, wherein the enzyme capable of initiating nucleic acid synthesis at the nicked site is DNA polymerase I, Klenow fragment, TAQ polymerase or Bst DNA polymerase. (Item 101) 101. The method of any one of items 88 to 100, wherein the nucleic acid-guided nuclease, nickase, nicks the 5' end of the DNA sequence. (Item 102) 102. The method of any one of items 88 to 101, wherein the enzyme capable of cleaving methylated DNA is DpnI. (Item 103) 1. A method for capturing a target DNA sequence of interest, comprising: (a) contacting the sample with a plurality of nucleic acid-guided nuclease nickase-gNA complexes, thereby generating a plurality of nicked DNAs at sites flanking regions of interest, wherein the gNAs are complementary to targeting sites of interest flanking the regions of interest in the subset of DNA sequences; (b) heating the sample to 65°C, thereby generating double-strand breaks at adjacent nicks; (c) contacting the double-stranded breaks with a thermostable ligase, thereby allowing ligation of adapter sequences only to these sites; and (d) repeating steps a-c to place a second adaptor on the opposite side of the region of interest, thus allowing enrichment of the region of interest. A method comprising: (Item 104) The method of claim 103, wherein the gNA is a gRNA. (Item 105) Item 104. The method of item 103, wherein the gNA is gDNA. (Item 106) 106. The method of any one of items 103 to 105, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. (Item 107) 106. The method of any one of paragraphs 103 to 105, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cmr5 nickase, Csf1 nickase, C2C2 nickase, and NgAgo nickase. (Item 108) Item 104. The method of item 103, wherein the DNA is double-stranded DNA. (Item 109) 109. The method of claim 108, wherein the double-stranded DNA is from genomic DNA. (Item 110) 110. The method of claim 109, wherein the genomic DNA is human. (Item 111) 111. The method of any one of items 103 to 110, wherein the DNA sequence is 20 bp to 5000 bp in length. (Item 112) 112. The method of any one of items 103 to 111, wherein the target site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region. (Item 113) 113. The method of any one of items 103 to 112, wherein the targeted site of interest accounts for less than 50% of the total DNA in the sample. (Item 114) 114. The method of any one of items 103 to 113, wherein the sample is obtained from a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 115) the thermostable ligase capable of ligating the double strand break is a thermostable 5'App 114. The method of any one of items 103 to 113, wherein the enzyme is a DNA / RNA ligase or a T4 RNA ligase. (Item 116) 116. The method of any one of items 103 to 115, wherein the nucleic acid-guided nuclease, nickase, nicks the 5' end of the DNA sequence. (Item 117) 1. A method for enriching a sample for a sequence of interest, comprising: (a) providing a sample containing a sequence of interest and a target sequence for depletion, wherein the sequence of interest comprises less than 50% of the sample; and (b) contacting the sample with a plurality of nucleic acid-guided RNA endonuclease-gRNA complexes or a plurality of nucleic acid-guided DNA endonuclease-gDNA complexes, wherein the gRNAs and gDNAs are complementary to the targeting sequences, thereby cleaving the targeting sequences. A method comprising: (Item 118) 118. The method of claim 117, further comprising extracting the sequence of interest and the targeting sequence for depletion from the sample. (Item 119) 119. The method of claim 118, further comprising the step of fragmenting the extracted sequences. (Item 120) 120. The method of any one of items 117 to 119, wherein the cleaved targeting sequence is removed by size exclusion. (Item 121) 121. The method of any one of items 117 to 120, wherein the sample is any one of a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 122) 122. The method of any one of items 117 to 121, wherein the sample comprises host nucleic acid sequences targeted for depletion and non-host nucleic acid sequences of interest. (Item 123) 123. The method of claim 122, wherein the non-host nucleic acid sequence comprises a microbial nucleic acid sequence. (Item 124) 124. The method of claim 123, wherein the microbial nucleic acid sequence is a nucleic acid sequence of a bacterium, a virus, or a eukaryotic parasite. (Item 125) 125. The method of any one of items 117 to 124, wherein the gRNA and gDNA are complementary to a ribosomal RNA sequence, a spliced ​​transcript, an unspliced ​​transcript, an intron, an exon, or a non-coding RNA. (Item 126) 119. The method of claim 118, wherein the extracted nucleic acid comprises single-stranded or double-stranded RNA. (Item 127) 119. The method of claim 118, wherein the extracted nucleic acid comprises single-stranded or double-stranded DNA. (Item 128) 128. The method of any one of items 117 to 127, wherein the sequence of interest constitutes less than 10% of the nucleic acid extracted. (Item 129) 129. The method of any one of items 117 to 128, wherein the nucleic acid-guided RNA endonuclease comprises C2c2. (Item 130) 130. The method of claim 129, wherein the C2c2 is catalytically inactive. (Item 131) 129. The method of any one of items 117 to 128, wherein the nucleic acid-guided DNA endonuclease comprises NgAgo. (Item 132) Item 132. The method of item 131, wherein the NgAgo is catalytically inactive. (Item 133) 133. The method of any one of items 117 to 132, wherein the sample is selected from whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bone, fingernails, faeces, urine, tissue and biopsy. (Item 134) 1. A method for concentrating a sample, comprising: (a) providing a sample containing host nucleic acid and non-host nucleic acid; (b) contacting the sample with a plurality of nucleic acid-guided RNA endonuclease-gRNA complexes or a plurality of nucleic acid-guided DNA endonuclease-gDNA complexes, wherein the gRNAs and gDNAs are complementary to target sites of the host nucleic acid; and (c) enriching the sample for non-host nucleic acids. A method comprising: (Item 135) 135. The method of item 134, wherein the nucleic acid-guided RNA endonuclease comprises C2c2. (Item 136) 135. The method of item 134, wherein the nucleic acid-guided RNA endonuclease comprises catalytically inactive C2c2. (Item 137) 135. The method of claim 134, wherein the nucleic acid-guided DNA endonuclease comprises NgAgo. (Item 138) 135. The method of claim 134, wherein the nucleic acid-guided DNA endonuclease comprises catalytically inactive NgAgo. (Item 139) 139. The method of any one of items 134 to 138, wherein the host is selected from the group consisting of humans, cows, horses, sheep, pigs, monkeys, dogs, cats, gerbils, birds, mice and rats. (Item 140) 139. The method of any one of items 134 to 138, wherein the non-host organism is a prokaryote. (Item 141) 139. The method of any one of items 134 to 138, wherein the non-host is selected from the group consisting of a eukaryote, a virus, a bacterium, a fungus, and a protozoan. (Item 142) 142. The method of any one of items 134 to 141, wherein the adaptor-ligated host nucleic acid and non-host nucleic acid are within the range of 50 bp to 1000 bp. (Item 143) 143. The method of any one of paragraphs 134 to 142, wherein the non-host nucleic acids constitute less than 50% of the total nucleic acids in the sample. (Item 144) 144. The method of any one of items 134 to 143, wherein the sample is any one of a biological sample, a clinical sample, a forensic sample or an environmental sample. (Item 145) 145. The method of any one of items 134 to 144, wherein step (c) comprises reverse transcribing the product of step (b) into cDNA. (Item 146) 146. The method of any one of items 134 to 145, wherein step (c) comprises removing the host nucleic acid by size exclusion. (Item 147) 146. The method of any one of items 134 to 145, wherein step (c) comprises removing the host nucleic acid by use of biotin. (Item 148) 148. The method of any one of items 134 to 147, wherein the sample is selected from whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bone, fingernails, faeces, urine, tissue and biopsy. (Item 149) A composition comprising a nucleic acid fragment, a nickase-gNA complex which is a nucleic acid-guided nuclease, and a labeled nucleotide. (Item 150) Item 149. The composition of item 149, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. (Item 151) 150. The composition of claim 149, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cmr5 nickase, Csf1 nickase, C2C2 nickase, and NgAgo nickase. (Item 152) 152. The composition of any one of items 149 to 151, wherein the gNA is a gRNA. (Item 153) 152. The composition of any one of items 149 to 151, wherein the gNA is gDNA. (Item 154) 154. The composition of any one of items 149 to 153, wherein the nucleic acid fragment comprises DNA. (Item 155) 154. The composition of any one of items 149 to 153, wherein the nucleic acid fragment comprises RNA. (Item 156) 156. The composition of any one of items 149 to 155, wherein the nucleotide is labeled with biotin. (Item 157) 157. The composition of any one of items 149 to 156, wherein the nucleotide is part of an antibody conjugate pair. (Item 158) A composition comprising a nucleic acid fragment and a catalytically inactive nucleic acid-guided nuclease-gNA complex, wherein the catalytically inactive nucleic acid-guided nuclease is fused to a transposase. (Item 159) Item 159. The composition of item 158, wherein the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of inactive CAS class I type I, inactive CAS class I type III, inactive CAS class I type IV, inactive CAS class II type II, and inactive CAS class II type V. (Item 160) 159. The composition of claim 158, wherein the catalytically inactive nucleic acid-guided nuclease is selected from the group consisting of dCas9, dCpf1, dCas3, dCas8a-c, dCaslO, dCsel, dCsy1, dCsn2, dCas4, dCsm2, dCmr5, dCsfl, dC2C2, and dNgAgo. (Item 161) 161. The composition of any one of items 158, 159, or 160, wherein the gNA is a gRNA. (Item 162) 162. The composition of any one of items 158 to 161, wherein the gNA is gDNA. (Item 163) 163. The composition of any one of items 158 to 162, wherein the nucleic acid fragment comprises DNA. (Item 164) 164. The composition of any one of items 158 to 163, wherein the nucleic acid fragment comprises RNA. (Item 165) 165. The composition of any one of items 158 to 164, wherein the catalytically inactive nucleic acid-guided nuclease is fused to the N-terminus of the transposase. (Item 166) 166. The composition of any one of items 158 to 165, wherein the catalytically inactive nucleic acid-guided nuclease is fused to the C-terminus of the transposase. (Item 167) A composition comprising a nucleic acid fragment containing a methylated nucleotide, a nickase-gNA complex which is a nucleic acid-guided nuclease, and an unmethylated nucleotide. (Item 168) Item 168. The composition of item 167, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of CAS class I type I nickase, CAS class I type III nickase, CAS class I type IV nickase, CAS class II type II nickase, and CAS class II type V nickase. (Item 169) 168. The composition of claim 167, wherein the nucleic acid-guided nuclease nickase is selected from the group consisting of Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cmr5 nickase, Csf1 nickase, C2C2 nickase, and NgAgo nickase. (Item 170) 169. The composition of any one of items 167, 168, or 169, wherein the gNA is a gRNA. (Item 171) 171. The composition of any one of items 167 to 170, wherein the gNA is gDNA. (Item 172) 172. The composition of any one of items 167 to 171, wherein the nucleic acid fragment comprises DNA. (Item 173) 172. The composition of any one of items 167 to 171, wherein the nucleic acid fragment comprises RNA. (Item 174) 174. The composition of any one of items 167 to 173, wherein the nucleotide is labeled with biotin. (Item 175) 175. The composition of any one of items 167 to 174, wherein the nucleotide is part of an antibody conjugate pair. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 1 illustrates a first protocol for capture of target nucleic acids from a library of human genomic DNA.

[0018] [Figure 2] Figure 2 further illustrates a protocol for capture of target nucleic acids (e.g., DNA) from a nucleic acid mixture. The target nucleic acid is cleaved with a nucleic acid-guided nuclease, after which adapters are ligated to the newly available blunt ends.

[0019] [Figure 3] Figure 3 illustrates that Cas9 cleavage followed by adapter ligation allows for specific amplification of target DNA.

[0020] [Figure 4] Figure 4 illustrates that after sequencing of the amplified DNA, adapter ligation occurred only at positions specified by the guide RNA.

[0021] [Figure 5]FIG. 5 illustrates that the method of FIG. 1 efficiently amplifies DNA that is under-represented in any given library.

[0022] [Figure 6] Figure 6 illustrates a second protocol for capture: the use of a nucleic acid-guided nuclease, nickase, to label target nucleic acids (e.g., DNA), allowing for further capture and purification.

[0023] [Figure 7] FIG. 7 illustrates a proof of principle experiment using a restriction nickase as a substitute for a nucleic acid-guided nuclease, nickase.

[0024] [Figure 8] Figure 8 illustrates approximately a 50-fold enrichment of test DNA for the experiment depicted in Figure 7 (using Cas9-nickase).

[0025] [Figure 9] Figure 9 illustrates a third protocol for capture: the use of catalytically inactive nucleic acid-guided nuclease-transposase fusions to insert adapters into a human genomic library, allowing enrichment of specific SNPs.

[0026] [Figure 10] Figure 10 illustrates a fourth protocol for capture: the use of an inactive nucleic acid-guided nuclease to protect the targeted site from subsequent fragmentation by the nucleic acid-guided nuclease, allowing enrichment of the region of interest.

[0027] [Figure 11] Figure 11 illustrates a fifth protocol for capture: the use of a nucleic acid-guided nuclease, nickase, to protect and then enrich any targeted region, e.g., a SNP or STR, from, e.g., human genomic DNA, by replacing methylated DNA with unmethylated DNA.

[0028] [Figure 12] FIG. 12 illustrates that methylation of the test DNA in the fifth protocol makes it susceptible to DpnI-mediated cleavage.

[0029] [Figure 13] Figure 13 illustrates a sixth protocol for capture: the use of a nucleic acid-guided nuclease, nickase, to introduce two double-stranded breaks that pinpoint the region of interest, allowing 3' single-stranded ligation of adapters and subsequent enrichment. DETAILED DESCRIPTION OF THE INVENTION

[0030] definition Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are described.

[0031] The headings provided herein are not limitations of the various aspects or embodiments of the present invention. Accordingly, the terms defined hereinafter are more fully defined by reference to the specification as a whole.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs. Singleton et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY, 2nd ed., John Wiley and Sons, New York (1994) and Hale and Markham, THE HARPER COLLINS DICTIONARY OF BIOLOGY, Harper Perennial, NY (1991) provide those skilled in the art with the general meaning of many of the terms used herein. However, for clarity and ease of reference, certain terms are defined below.

[0033] Numeric ranges are inclusive of the numbers defining the range.

[0034] As used herein, the term "sample" generally, but not exclusively, refers to a liquid material or mixture of materials containing one or more analytes of interest.

[0035] The term "nucleic acid sample" as used herein refers to a sample containing nucleic acids. Nucleic acid samples as used herein may be complex in that they contain multiple different molecules with sequences. Genomic DNA from mammals is a type of complex sample. A complex sample may be 10 4 , 10 5 , 10 6 or 10 7 The DNA target can have more than 10 different nucleic acid molecules.The DNA target can originate from any source, for example, genomic DNA, cDNA or artificial DNA construct.Any sample that contains nucleic acid can be used herein, for example, genomic DNA or tissue sample that is produced from tissue culture cells.

[0036] The term "nucleotide" includes moieties containing not only the known purine and pyrimidine bases but also modified other heterocyclic bases. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated riboses or other heterocycles. Additionally, the term "nucleotide" includes moieties containing haptens or fluorescent labels, and can contain not only traditional ribose and deoxyribose sugars but other sugars as well. Modified nucleosides or nucleotides also include modifications to the sugar moiety, for example, in which one or more of the hydroxyl groups are replaced with halogen atoms or aliphatic groups, or functionalized as ethers, amines, etc.

[0037] The terms "nucleic acid" and "polynucleotide" are used interchangeably herein. Polynucleotide is used to describe a nucleic acid polymer of any length, for example, greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1000 bases, up to about 10,000 or more bases, composed of nucleotides, such as deoxyribonucleotides or ribonucleotides, and can be produced enzymatically or synthetically (e.g., PNAs described in U.S. Pat. No. 5,948,902 and references therein), which can hybridize with naturally occurring nucleic acids in a sequence-specific manner similar to that of two naturally occurring nucleic acids, for example, can participate in Watson-Crick base pairing interactions. Naturally occurring nucleotides include guanine, cytosine, adenine, and thymine (G, C, A, and T, respectively). While DNA and RNA have deoxyribose and ribose sugar backbones, respectively, the backbone of PNAs is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. In PNAs, various purine and pyrimidine bases are linked to the backbone by methylene carbonyl bonds. Locked nucleic acids (LNAs), often referred to as inaccessible RNAs, are modified RNA nucleotides. The ribose moiety of LNA nucleotides is modified with an extra bridge connecting the 2' oxygen and 4' carbon. The bridge "locks" the ribose in the 3'-endo (north) conformation, which is often found in A-form duplexes. Whenever desired, LNA nucleotides can be mixed with DNA or RNA residues in oligonucleotides. The term "unstructured nucleic acid" or "UNA" refers to nucleic acids containing non-natural nucleotides that bind to each other with low stability. For example, unstructured nucleic acids can contain G' and C' residues, where these residues correspond to non-naturally occurring forms, or analogs, of G and C, which base pair with each other with low stability but retain the ability to base pair with naturally occurring C and G residues, respectively. Unstructured nucleic acids are described in US Patent Application Publication No. 20050233340, which is incorporated herein by reference for its disclosure of UNAs.

[0038] As used herein, the term "oligonucleotide" refers to a single-stranded polymer of nucleotides.

[0039] Unless otherwise specified, nucleic acids are written left to right in 5' to 3' orientation. Amino acid sequences are written left to right in amino to carboxy orientation, respectively.

[0040] As used herein, the term "cleaving" refers to a reaction that cleaves the phosphodiester bond between two adjacent nucleotides in both strands of a double-stranded DNA molecule, thereby producing a double-stranded break in the DNA molecule.

[0041] As used herein, the term "cleavage site" refers to the site where a double-stranded DNA molecule is cleaved.

[0042] A "nucleic acid-guided nuclease-gNA complex" refers to a complex comprising a nucleic acid-guided nuclease protein and a guide nucleic acid (gNA, e.g., gRNA or gDNA). For example, a "Cas9-gRNA complex" refers to a complex comprising a Cas9 protein and a guide RNA (gRNA). The nucleic acid-guided nuclease can be any type of nucleic acid-guided nuclease, including, but not limited to, a wild-type nucleic acid-guided nuclease, a catalytically inactive nucleic acid-guided nuclease, or a nickase that is a nucleic acid-guided nuclease.

[0043] The term "nucleic acid-guided nuclease-associated guide NA" refers to a guide nucleic acid (guide NA). The nucleic acid-guided nuclease-associated guide NA can exist as an isolated nucleic acid or as part of a nucleic acid-guided nuclease-gNA complex, such as a Cas9-gRNA complex.

[0044] The terms "capture" and "enrichment" are used interchangeably herein to refer to the process of selectively isolating a nucleic acid region containing a sequence of interest, a targeting site of interest, a sequence of interest, or a targeting site of interest that is not of interest.

[0045] The term "hybridization," as known in the art, refers to the process by which a nucleic acid strand joins with a complementary strand through base pairing. A nucleic acid is considered to be "selectively hybridizable" to a reference nucleic acid sequence if the two sequences specifically hybridize to each other under moderate to high stringency hybridization and wash conditions. Moderate and high stringency hybridization conditions are known (e.g., Ausubel et al., Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons 1995 and Sambrook et al., Molecular Cloning: A Laboratory (See, The High Stringency Analysis Manual, 3rd ed., 2001, Cold Spring Harbor, NY). One example of high stringency conditions includes hybridization at about 42°C in 50% formamide, 5x SSC, 5x Denhardt's solution, 0.5% SDS, and 100 µg / ml denatured carrier DNA, followed by two washes in 2x SSC and 0.5% SDS at room temperature and two additional washes in 0.1x SSC and 0.5% SDS at 42°C.

[0046] As used herein, the terms "duplex" or "duplexed" describe two complementary polynucleotides that are base-paired, i.e., hybridized together.

[0047] As used herein, the term "amplifying" refers to producing one or more copies of a target nucleic acid using the target nucleic acid as a template.

[0048] As used herein, the term "genomic region" refers to a region of a genome, for example, an animal or plant genome, for example, a human, monkey, rat, fish, or insect or plant genome.In certain cases, the oligonucleotides used in the methods described herein can be designed using a reference genome region, i.e., a genome region of known nucleotide sequence, for example, a chromosome region whose sequence is deposited in, for example, the Genbank database of NCBI or other databases.

[0049] As used herein, the term "genomic sequence" refers to a sequence present in a genome. Because RNA is transcribed from a genome, this term encompasses sequences present in the nuclear genome of an organism, as well as sequences present in a cDNA copy of an RNA (e.g., mRNA) transcribed from such a genome.

[0050] As used herein, the term "genomic fragment" refers to a region of a genome, for example, an animal or plant genome, for example, a human, monkey, rat, fish, or insect or plant genome. A genomic fragment may be an entire chromosome or a fragment of a chromosome. A genomic fragment may be adaptor-ligated (in which case it has adaptors ligated to one or both ends of the fragment, or at least to the 5' end of the molecule), or may not be adaptor-ligated.

[0051] In certain cases, oligonucleotides used in the methods described herein can be designed using reference genomic regions, i.e., genomic regions of known nucleotide sequence, e.g., chromosomal regions whose sequences are deposited, for example, in NCBI's Genbank database or other databases. Such oligonucleotides can be used in assays that use samples containing test genomes, where the test genome contains a binding site for the oligonucleotide.

[0052] As used herein, the term "ligate" refers to the enzyme-catalyzed joining of the terminal nucleotide at the 5' end of a first DNA molecule with the terminal nucleotide at the 3' end of a second DNA molecule.

[0053] When two nucleic acids are "complementary," each base of one of the nucleic acids forms a base pair with a corresponding nucleotide in the other nucleic acid. The terms "complementary" and "fully complementary" are used interchangeably herein.

[0054] As used herein, the term "separating" refers to the physical separation of two elements (e.g., by size or affinity, etc.), as well as the degradation of one element while the other remains intact. For example, size exclusion can be used to separate nucleic acids containing cleaved targeting sequences.

[0055] In cells, DNA usually exists in double-stranded form, thus having two complementary nucleic acid strands, referred to herein as "top" and "bottom" strands. In certain cases, the complementary strands of a chromosomal region can be referred to as "plus" and "minus" strands, "first" and "second" strands, "coding" and "non-coding" strands, "Watson" and "Crick" strands, or "sense" and "antisense" strands. The assignment of strands to top or bottom strands is arbitrary and does not imply any particular orientation, function, or structure. Until they are covalently linked, the first and second strands are separate molecules. For ease of description, the "top" and "bottom" strands of a double-stranded nucleic acid in which the top and bottom strands are covalently linked will still be referred to as the "top" and "bottom" strands. In other words, for the purposes of this disclosure, the top and bottom strands of double-stranded DNA do not need to be separate molecules. The first strand nucleotide sequences of several exemplary mammalian chromosomal regions (eg, BACs, assemblies, chromosomes, etc.) are known and can be found, for example, in NCBI's Genbank database.

[0056] As used herein, the term "top strand" refers to either strand of a nucleic acid, but not both strands of the nucleic acid. When an oligonucleotide or primer binds or anneals "only to the top strand," it binds to only one strand and not the other. As used herein, the term "bottom strand" refers to the strand that is complementary to the "top strand." When an oligonucleotide binds or anneals "only to one strand," it binds to only one strand, e.g., the first or second strand, and not the other strand. When an oligonucleotide binds or anneals to both strands of a double-stranded DNA, the oligonucleotide can have two regions, one region that hybridizes to the top strand of the double-stranded DNA and a second region that hybridizes to the bottom strand of the double-stranded DNA.

[0057] The term "double-stranded DNA molecule" refers to both double-stranded DNA molecules in which the top and bottom strands are not covalently linked, and double-stranded DNA molecules in which the top and bottom strands are covalently linked. The top and bottom strands of double-stranded DNA form base pairs with each other through Watson-Crick interactions.

[0058] As used herein, the term "denaturing" refers to the separation of at least a portion of the base pairs of a nucleic acid duplex by subjecting the duplex to suitable denaturing conditions. Denaturing conditions are well known in the art. In one embodiment, to denature a nucleic acid duplex, the duplex is subjected to a T mThe nucleic acid can be denatured by exposing it to a temperature above 90°C, thereby releasing one strand of the duplex from the other. In certain embodiments, the nucleic acid can be denatured by exposing it to a temperature of at least 90°C for a suitable time (e.g., at least 30 seconds to 30 minutes). In certain embodiments, fully denaturing conditions can be used to completely separate the base pairs of the duplex. In other embodiments, partial denaturing conditions (e.g., at a temperature lower than fully denaturing conditions) can be used to separate the base pairs of a certain portion of the duplex (e.g., regions enriched in AT base pairs can be separated, while regions enriched in GC base pairs can remain paired). Nucleic acids can also be chemically denatured (e.g., using urea or NaOH).

[0059] As used herein, the term "genotyping" refers to any type of analysis of a nucleic acid sequence, including sequencing, polymorphism (SNP) analysis, and analysis to identify rearrangements.

[0060] As used herein, the term "sequencing" refers to a method by which the identity of consecutive nucleotides of a polynucleotide is obtained.

[0061] The term "next generation sequencing" refers to so-called parallelized sequencing-by-synthesis or sequencing-by-ligation platforms, such as those currently used by Illumina, Life Technologies, and Roche, etc. Next generation sequencing methods can also include nanopore sequencing methods or methods based on electronic detection, such as the Ion Torrent technology commercialized by Life Technologies.

[0062] The term "complementary DNA" or cDNA refers to a double-stranded DNA sample generated from an RNA sample by reverse transcription of RNA (using a primer such as a random hexamer or an oligo-dT primer), followed by digestion of the RNA with RNase H and synthesis of the second strand with a DNA polymerase.

[0063] The term "RNA promoter adapter" refers to an adapter that contains a promoter for a bacteriophage RNA polymerase, for example, an RNA polymerase from bacteriophage T3, T7, SP6, etc.

[0064] Other definitions of terms may appear throughout the specification.

[0065] Exemplary Methods of the Invention As described herein, the present invention provides exemplary protocols for capturing nucleic acids and compositions for use in these protocols. Exemplary protocols are illustrated in Figures 1, 6, 9, 10, 11, and 13, respectively, and in the Examples section. Overall, a variety of uses are contemplated. Specific terms referred to in this section are described in more detail in subsequent sections.

[0066] In one embodiment, the present invention provides a capture method (represented as Protocol 1) provided in Figures 1-5. In this embodiment, the method is used to capture target nucleic acid sequences. With reference to Figure 1, the method includes providing a sample or library 100 that is subjected to an extraction protocol 101 (e.g., a DNA extraction protocol) to result in a sample 102 containing >99% off-target sequences and <1% target sequences. The sample is subjected to a library construction protocol 103 to result in a nucleic acid library that includes sequencing indexing adaptors 104, resulting in a plurality of adaptor-ligated nucleic acids, where the nucleic acids are ligated at one end to a first adaptor and at the other end to a second adaptor. To generate a library of target-specific guide NAs (e.g., gRNAs), a library of target-specific gNA precursors 110, each containing an RNA polymerase promoter 111, a specific base-pair region 112 (e.g., a 20-base-pair region), and a stem-loop binding site 113 for a nucleic acid-guided nuclease 113, was subjected to in vitro transcription 114 to obtain a library of target-specific guide RNAs 115. To obtain a library of nucleic acid-guided nuclease-gNA complexes, the library of target-specific guide NAs was then combined with a nucleic acid-guided nuclease protein 116. The nucleic acid-guided nuclease-gNA complex was then combined with the nucleic acid library so that the nucleic acid-guided nuclease cleaved the corresponding target nucleic acid sequence and left other nucleic acids uncleaved 117. A second adaptor 118 was added and specifically ligated to the 5'-phosphorylated blunt ends of the cleaved nucleic acids 119. This allows for downstream applications, such as using an adapter-specific PCR120 to amplify a nucleic acid fragment containing the first or second adapter at one end and the third adapter at the other end.

[0067] In an exemplary depiction of Protocol 1, with reference to Figure 1, the method is used to capture a target nucleic acid sequence. The method includes providing a sample or library containing a plurality of adaptor-ligated nucleic acids, where the nucleic acids are ligated to a first adaptor at one end and a second adaptor at the other end. The sample is then contacted with a plurality of Cas9-gRNA complexes, where the gRNAs are complementary to targeting sites of interest contained in a subset of the nucleic acids. The contacting step cleaves the targeting sites of interest, thereby generating a plurality of nucleic acid fragments ligated to the first or second adaptor at one end and adapter-free at the other end. Following this step, the resulting plurality of nucleic acid fragments are ligated to a third adaptor, thereby generating a plurality of nucleic acid fragments ligated to the first or second adaptor at one end and the third adaptor at the other end. This allows for downstream applications, for example, using adapter-specific PCR to amplify nucleic acid fragments containing the first or second adapter at one end and the third adapter at the other end.

[0068] In one embodiment, the present invention provides a capture method (represented as Protocol 2) as provided in Figures 6-8. In this embodiment, the method is used to introduce labeled nucleotides at a target site of interest. The method includes providing a sample containing a plurality of double-stranded nucleic acid fragments 601 (e.g., double-stranded DNA); and contacting the sample with a plurality of nucleic acid-guided nucleases, namely, nickase-gNA complexes. The nickase-nucleic acid-guided nucleases 603 are guided by target-specific guide NAs 604, where the gNAs are complementary to the target sites of interest in the nucleic acid fragments, thereby generating a plurality of nicked nucleic acid fragments at the target sites of interest. The nickase is used to nick a target sequence 605. The nickase-nucleic acid-guided nuclease cleaves at the target sequence. The single-stranded break (nick) is a substrate for DNA polymerase I, which can be used to replace the DNA downstream of the nick with biotin-labeled DNA 606.

[0069] In an exemplary depiction of Protocol 2, with reference to Figure 6, the method is used to introduce labeled nucleotides at target sites of interest. The method includes providing a sample containing a plurality of nucleic acid fragments; contacting the sample with a plurality of Cas9 nickase-gRNA complexes, where the gRNAs are complementary to target sites of interest in the nucleic acid fragments, thereby generating a plurality of nicked nucleic acid fragments at the target sites of interest; and subsequently contacting the plurality of nicked nucleic acid fragments with an enzyme capable of initiating nucleic acid synthesis at the nicked sites and labeled nucleotides, thereby generating a plurality of nucleic acid fragments comprising labeled nucleotides at the target sites of interest.

[0070] In one embodiment, the present invention provides a capture method (represented as Protocol 3) provided in Figure 9. In this embodiment, the method is used to capture a target nucleic acid sequence 901 of interest. The method includes first providing a sample containing a plurality of adaptor-ligated nucleic acids 902, where the nucleic acids are ligated to a first adaptor at one end and a second adaptor at the other end. This is followed by contacting the sample with a plurality of catalytically inactive nucleic acid-guided nuclease-gNA complexes, where the catalytically inactive nucleic acid-guided nuclease is fused to a transposase 903, and the gNA 904 is complementary to a target site of interest contained in a subset of the nucleic acids, where the catalytically inactive nucleic acid-guided nuclease-gNA transposase complexes are loaded with a plurality of third adaptors 905 to generate a plurality of nucleic acid fragments 906 comprising the first or second adaptor at one end and a third adaptor at the other end. These fragments can then be amplified 907 using adapter sequences and then sequenced 908 .

[0071] In an exemplary depiction of Protocol 3, with reference to Figure 9, the method is used to capture a target nucleic acid sequence of interest. The method includes first providing a sample containing a plurality of adaptor-ligated nucleic acids, where the nucleic acids are ligated at one end to a first adaptor and at the other end to a second adaptor. The sample is then contacted with a plurality of dCas9-gRNA complexes, where the dCas9 is fused to a transposase and the gRNA is complementary to a targeting site of interest contained in a subset of the nucleic acids, and where the dCas9-gRNA transposase complexes are loaded with a plurality of third adaptors to generate a plurality of nucleic acid fragments comprising the first or second adaptor at one end and a third adaptor at the other end.

[0072] In one embodiment, the present invention provides a capture method (represented as Protocol 4), for example, as provided in FIG. 10. In this embodiment, the method is used to capture a target nucleic acid sequence 1001 of interest. The method includes first providing a sample containing a plurality of adaptor-ligated nucleic acids 1002, where the nucleic acids are adaptor-ligated at their 5' and 3' ends. The method then includes contacting the sample with a plurality of catalytically inactive nucleic acid-guided nuclease-gNA complexes 1003, where gNAs 1004 are complementary to targeting sites of interest contained in a subset of the nucleic acids, thereby generating a plurality of nucleic acids ligated to adaptors at their 5' and 3' ends that are bound to catalytically inactive nucleic acid-guided nuclease-gNA complexes 1005. This is followed by contacting the sample with a plurality of nucleic acid-guided nuclease-gNA complexes 1006, where gNAs 1007 are complementary to both the desired targeting site and the undesired targeting site of the nucleic acid, thereby generating a plurality of nucleic acid fragments 1008 comprising undesired nucleic acid sequences ligated to adapters at only one of the 5' or 3' ends. In this method, for the second contacting step, the contacting with the plurality of nucleic acid-guided nuclease-gNA complexes does not displace the plurality of nucleic acids ligated to adapters at the 5' and 3' ends that are bound to the catalytically inactive nucleic acid-guided nuclease-gNA complexes of step (b).

[0073] In an exemplary depiction of Protocol 4, with reference to Figure 10, the method is used to capture a target nucleic acid sequence of interest. The method includes first providing a sample containing a plurality of adaptor-ligated nucleic acids, where the nucleic acids are adaptor-ligated at their 5' and 3' ends. The method then includes contacting the sample with a plurality of inactive nucleic acid-guided nuclease-gNA complexes (e.g., dCas9-gRNAs), where the gNAs are complementary to targeting sites of interest contained in a subset of the nucleic acids, thereby generating a plurality of adaptor-ligated nucleic acids at their 5' and 3' ends that are bound to the inactive nucleic acid-guided nuclease-gNA complexes (e.g., dCAS9-gRNA complexes). This is followed by contacting the sample with a plurality of nucleic acid-guided nuclease-gNA complexes (e.g., Cas9-gRNA complexes), where the gNAs are complementary to both the intended targeting site and the unintended targeting site of the nucleic acid, thereby generating a plurality of nucleic acid fragments comprising the unintended nucleic acid sequence ligated to an adaptor at only one of the 5' or 3' ends. In this method, the second contacting step, contacting with the plurality of nucleic acid-guided nuclease-gNA complexes (e.g., Cas9-gRNA complexes) does not displace the plurality of adaptor-ligated nucleic acids at their 5' and 3' ends that are bound to the inactive nucleic acid-guided nuclease-gNA complexes (e.g., dCAS9-gRNA complexes) of step (b).

[0074] In one embodiment, the present invention provides a capture method (represented as Protocol 5), e.g., as provided in Figures 11 and 12. In this embodiment, the method is used to capture a target nucleic acid sequence 1101 of interest. The method includes first providing a sample containing a plurality of sequences 1102, where the sequences include methylated nucleotides (e.g., treated with Dam methyltransferase), and the sequences are ligated to adapters at their 5' and 3' ends. The method then includes first contacting the sample with a plurality of nucleic acid-guided nucleases, nickase-gNA complexes 1103, where gNAs 1104 are complementary to target sites of interest in a subset of the sequences, thereby generating a plurality of nicked nucleic acid sequences 1105 at the target sites of interest, where the nucleic acid sequences are ligated to adapters at their 5' and 3' ends. When the nucleic acid is DNA, the single-strand break (nick) can be a substrate for, for example, DNA polymerase I, which replaces the DNA downstream of the nick with unmethylated DNA. After this, the sample can then be contacted with an enzyme capable of initiating DNA synthesis at the nick site and an unmethylated nucleotide, thereby generating a plurality of DNA fragments containing unmethylated nucleotides 1106 at the target site of interest, where the DNA sequence is ligated to an adaptor at the 5' and 3' ends. After this, the sample can be contacted with an enzyme capable of cleaving methylated DNA (e.g., DpnI) 1107, thereby generating a plurality of DNA fragments containing methylated DNA, where the plurality of DNA fragments containing methylated DNA is ligated to an adaptor at only one of the 5' and 3' ends. The remaining intact nucleic acid can be amplified and sequenced 1108. Figure 12, for example, shows results from an experiment performed according to this protocol. The first column on the gel shows a 1 kb ladder, the second column shows test DNA treated with Dam methyltransferase and then digested with DpnI, and the third column shows test DNA digested with DpnI. The second column shows the bands corresponding to DpnI-digested DNA, and the third column shows the bands corresponding to uncut test DNA.

[0075] In an exemplary depiction of Protocol 5, with reference to Figures 11-12, the method is used to capture a target DNA sequence of interest. The method includes first providing a sample containing a plurality of DNA sequences, where the DNA sequences comprise methylated nucleotides and the DNA sequences are ligated to adaptors at their 5' and 3' ends. The method then includes first contacting the sample with a plurality of Cas9 nickase-gRNA complexes, where the gRNAs are complementary to target sites of interest in a subset of the DNA sequences, thereby generating a plurality of nicked DNAs at the target sites of interest, where the DNAs are ligated to adaptors at their 5' and 3' ends. This is followed by contacting the sample with an enzyme capable of initiating DNA synthesis at the nick site and an unmethylated nucleotide, thereby generating a plurality of DNAs comprising unmethylated nucleotides at the target sites of interest, where the DNA sequences are ligated to adaptors at their 5' and 3' ends. Thereafter, the sample is contacted with an enzyme capable of cleaving methylated DNA, thereby generating a plurality of DNA fragments comprising methylated DNA, wherein the plurality of DNA fragments comprising methylated DNA are ligated to an adaptor at only one of the 5' and 3' ends.

[0076] In one embodiment, the present invention provides a capture method (represented as Protocol 6) as provided in Figure 13. The purpose of this method is to enrich a region of nucleic acid 1301 from any source (e.g., library 1302, genome, or PCR), as depicted in Figure 13. A nucleic acid-guided nuclease, nickase 1303, can be targeted to proximal sites using two guide NAs 1304 and 1305, resulting in nicking of the nucleic acid at each position 1306. Alternatively, adapters can be ligated on only one side, then filled in, and then adapters ligated on the other side. The two nicks can be close to each other (e.g., within 10-15 bp). A single nick can be generated for off-target molecules. Because the two nick sites are close to each other, heating the reaction to, for example, 65°C 1307 can result in a double-stranded break, resulting in a long (e.g., 10-15 bp) 3' overhang. These overhangs can be recognized by a thermostable single-stranded DNA / RNA ligase to allow site-specific ligation of single-stranded adapters 1308. The ligase can, for example, recognize only the long 3' overhangs, thus ensuring that adapters are not ligated at other sites. This process can be repeated using a nucleic acid-guided nuclease, a nickase, and a guide NA targeting the opposite side of the region of interest, followed by ligation as above using a second single-stranded adapter. Once two adapters have been ligated to either side of the region of interest, the region can be amplified or directly sequenced 1309.

[0077] In an exemplary depiction of Protocol 6, with reference to Figure 13, the method enriches regions of DNA from any DNA source (e.g., library, genome, or PCR). Cas9 nickase can be targeted to proximal sites using two guide RNAs, resulting in nicking of the DNA at each position. Alternatively, adapters can be ligated on only one side, then filled in, and then adapters ligated on the other side. The two nicks can be close to each other (e.g., within 10-15 bp). A single nick can be generated on the off-target molecule. Due to the proximity of the two nick sites, when the reaction is heated, for example, to 65°C, a double-stranded break can occur, resulting in a long (e.g., 10-15 bp) 3' overhang. These overhangs can be recognized by a thermostable single-stranded DNA / RNA ligase, such as a thermostable 5' App DNA / RNA ligase, to enable site-specific ligation of single-stranded adapters. The ligase can, for example, recognize only the long 3' overhang, thus ensuring that the adapter is not ligated at other sites. This process can be repeated using a Cas9 nickase and a guide RNA targeting the opposite side of the region of interest, followed by ligation as above using a second single-stranded adapter. Once two adapters are ligated on either side of the region of interest, the region can be amplified or directly sequenced.

[0078] In one embodiment, provided herein is a method for enriching a sample for a sequence of interest, the method comprising: (a) providing a sample containing the sequence of interest and a targeting sequence for depletion, wherein the sequence of interest comprises less than 50% of the sample; and (b) contacting the sample with a plurality of nucleic acid-guided RNA endonuclease-gRNA complexes or a plurality of nucleic acid-guided DNA endonuclease-gDNA complexes, wherein the gRNA and gDNA are complementary to the targeting sequence. In some embodiments, the targeting sequence is cleaved thereby. In one embodiment, the nucleic acid-guided RNA endonuclease is C2c2. In one embodiment, C2c2 is catalytically inactive. In one embodiment, the nucleic acid-guided DNA endonuclease is NgAgo (Argonaute from Natronobacterium gregoryi). In one embodiment, NgAgo is catalytically inactive.

[0079] In one embodiment, provided herein is a method of enriching a sample, the method comprising: (a) providing a sample containing host nucleic acid and non-host nucleic acid; (b) contacting the sample with a plurality of nucleic acid-guided RNA endonuclease-gRNA complexes or a plurality of nucleic acid-guided DNA endonuclease-gDNA complexes, wherein the gRNAs are complementary to target sites of the host nucleic acid; and (c) enriching the sample for non-host nucleic acid. In one embodiment, the nucleic acid-guided RNA endonuclease is C2c2. In one embodiment, C2c2 is catalytically inactive. In one embodiment, the nucleic acid-guided DNA endonuclease is NgAgo. In one embodiment, NgAgo is catalytically inactive.

[0080] Nucleic acids, samples The nucleic acids of the present invention (targeted for capture) can be any DNA, any RNA, single-stranded DNA, single-stranded RNA, double-stranded DNA, double-stranded RNA, artificial DNA, artificial RNA, synthetic DNA, synthetic RNA and RNA / DNA hybrids.

[0081] The nucleic acids of the present invention may be genome fragments comprising regions of the genome, or the entire genome itself. In one embodiment, the genome is a DNA genome. In another embodiment, the genome is an RNA genome.

[0082] Nucleic acids of the invention can be obtained from eukaryotic or prokaryotic organisms; from mammalian or non-mammalian organisms; from animals or plants; from bacteria or viruses; from animal parasites; or from pathogens.

[0083] The nucleic acids of the present invention can be obtained from any mammalian organism. In one embodiment, the mammal is a human. In another embodiment, the mammal is a livestock animal, such as a horse, sheep, cow, pig, or donkey. In another embodiment, the mammalian organism is a domestic pet, such as a cat, dog, gerbil, mouse, or rat. In another embodiment, the mammal is a monkey.

[0084] The nucleic acids of the invention can be obtained from any bird or avian organism, including, but not limited to, chickens, turkeys, ducks, and geese.

[0085] The nucleic acids of the invention can be obtained from plants, hi one embodiment, the plant is rice, corn, wheat, rose, grape, coffee, berry, tomato, potato or cotton.

[0086] In some embodiments, the nucleic acids of the invention are obtained from a bacterial species. In one embodiment, the bacterium is a bacterium that causes tuberculosis.

[0087] In some embodiments, the nucleic acids of the invention are obtained from a virus.

[0088] In some embodiments, the nucleic acids of the invention are obtained from a fungal species.

[0089] In some embodiments, the nucleic acids of the invention are obtained from algal species.

[0090] In some embodiments, the nucleic acids of the invention are obtained from any mammalian parasite.

[0091] In some embodiments, the nucleic acids of the invention are obtained from any mammalian parasite. In one embodiment, the parasite is a worm. In another embodiment, the parasite is a parasite that causes malaria. In another embodiment, the parasite is a parasite that causes leishmaniasis. In another embodiment, the parasite is an amoeba.

[0092] In some embodiments, the pathogen is a non-mammalian pathogen (pathogenic in non-mammalian organisms).

[0093] In one embodiment, the nucleic acids of the present invention include nucleic acids that are targets of gNAs and nucleic acids that are not targets of gNAs in the same sample.

[0094] In one embodiment, the nucleic acids of the invention include nucleic acids that are targets of a gRNA and nucleic acids that are not targets of a gRNA in the same sample.

[0095] In one embodiment, nucleic acids of the invention include nucleic acids that are targets of gDNA and nucleic acids that are not targets of gDNA in the same sample.

[0096] In one embodiment, the nucleic acids of the invention include target nucleic acids (targets of the gNAs) and nucleic acids of interest (not targeted by the gNAs) from a sample.

[0097] In one embodiment, the nucleic acids of the invention include a target nucleic acid (target of the gRNA) and a nucleic acid of interest (not targeted by the gRNA) from a sample.

[0098] In one embodiment, the nucleic acids of the invention include target nucleic acids (targeted by gDNA) and nucleic acids of interest (not targeted by gDNA) from a sample.

[0099] In one embodiment, the target DNA (target of the gNA, gRNA, gDNA) may be human non-mitochondrial DNA (e.g., genomic DNA), the DNA of interest (for capture) may be human mitochondrial DNA, and human mitochondrial DNA is enriched by targeting non-mitochondrial human DNA.

[0100] In one embodiment, the nucleic acid that is captured may be a non-mappable region of the genome; the nucleic acid that is retained for further analysis / sequencing / cloning may be a mappable region of the genome. In one embodiment, the nucleic acid that is captured out may be a mappable region of the genome; the nucleic acid that is retained for further analysis / sequencing / cloning may be a non-mappable region of the genome. Examples of non-mappable regions include telomeres, centromeres, or other genomic regions with features that are more difficult to map.

[0101] In one embodiment, the nucleic acid of the present invention is obtained from a biological sample. Biological samples from which nucleic acids can be obtained include, but are not limited to, whole blood, plasma, serum, tears, saliva, mucus, cerebrospinal fluid, teeth, bones, fingernails, feces, urine, tissues, and biopsies. Biological samples may include forensic samples such as teeth, bones, and fingernails. Biological samples may include tissues, tissue biopsies, such as resected lung tissues. Biological samples may include clinical samples, which refer to samples obtained in clinical settings, such as hospitals or clinics.

[0102] In one embodiment, the nucleic acids of the invention are obtained from an environmental sample, for example, from water, soil, air, or rock.

[0103] In one embodiment, the nucleic acids of the invention are obtained from forensic samples, eg, samples obtained from individuals at a crime scene, from evidence, post-mortem, as part of an ongoing investigation, etc.

[0104] In one embodiment, the nucleic acids of the invention are provided in a library.

[0105] The nucleic acids of the present invention may be provided or extracted from a sample. Extraction can extract substantially all nucleic acid sequences from a specimen.

[0106] The methods of the present invention can produce ratios of captured nucleic acids to non-captured nucleic acids anywhere between 99.999:0.001 and 0.001:99.999. The methods of the present invention can produce ratios of targeted nucleic acids and nucleic acids of interest anywhere between 99.999:0.001 and 0.001:99.999. The methods of the present invention can produce ratios of captured nucleic acids to retained / analyzed / sequenced nucleic acids anywhere between 99.999:0.001 and 0.001:99.999. In these embodiments, the ratio may be equal to or fall anywhere within a range between 99.999:0.001 and 0.001:99.999, for example, the ratio may be 99:1, 95:5, 90:10, 85:15, 80:20, 75:25, 70:30, 65:35, 60:40, 55:45, 50:50, 45:55, 40:60, 35:65, 30:70, 25:75, 20:80, 15:85, 10:90, 5:95, and 1:99.

[0107] After capture, the captured or retained nucleic acid sequences can be fragmented to reduce the length of each extracted nucleic acid to a more manageable length for amplification, sequencing, etc.

[0108] As provided herein, at least 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90% of the starting nucleic acid material can be captured, which can be achieved in 10 minutes or less, 15 minutes or less, 20 minutes or less, 30 minutes or less, 45 minutes or less, 60 minutes or less, 75 minutes or less, 90 minutes or less, 105 minutes or less, 120 minutes or less, 150 minutes or less, 180 minutes or less, or 240 minutes or less.

[0109] In some cases, the targeted site of interest accounts for less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, or less than 10% of the total DNA in the sample.

[0110] adapter As provided herein, to facilitate the performance of the methods provided herein, the nucleic acids (interchangeably referred to as nucleic acids or nucleic acid fragments) of the invention are ligated to adaptors.

[0111] The size of the adaptor-ligated nucleic acid of the present invention may be within the range of 20 bp to 5000 bp. For example, the adaptor-ligated nucleic acid may be at least 20, 25, 50, 75, 100, 125, 150, 175, 200, 25, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, or 5000 bp. In a specific embodiment, the adaptor-ligated nucleic acid is 100 bp. In a specific embodiment, the adaptor-ligated nucleic acid is 200 bp. In a specific embodiment, the adaptor-ligated nucleic acid is 300 bp. In a specific embodiment, the adaptor-ligated nucleic acid is 400 bp. In a specific embodiment, the adaptor-ligated nucleic acid is 500 bp.

[0112] Adapters can be ligated to each end of each nucleic acid or nucleic acid fragment at the 5' and 3' ends. In other embodiments, adapters can be ligated to only one end of each fragment, or in other cases, adapters can be ligated in a later step. In one example, the adapter is a nucleic acid that can be ligated to both strands of a double-stranded DNA molecule. In various embodiments, the adapter can be a hairpin adapter, e.g., a single molecule that base-pairs with itself to form a double-stranded stem and loop structure, where the 3' and 5' ends of the molecule ligate to the 5' and 3' ends of the fragment's double-stranded DNA molecule, respectively. Alternatively, the adapter can be a Y adapter, also called a universal adapter, that is ligated to one or both ends of the fragment. Alternatively, the adapter itself can be composed of two different oligonucleotide molecules that base-pair with each other. Furthermore, the ligatable end of the adapter can be designed to fit overhangs created by restriction enzyme cleavage, or it can be blunt-ended or have a 5' T overhang. Generally, adapters can include double-stranded and single-stranded molecules. Thus, adapters can be DNA or RNA, or a mixture of the two. RNA-containing adapters can be cleavable by RNase treatment or alkaline hydrolysis.

[0113] The adapters can be 10 to 100 bp in length, although adapters outside this range can be used without departing from the invention. In specific embodiments, the adapters are at least 10 bp, at least 15 bp, at least 20 bp, at least 25 bp, at least 30 bp, at least 35 bp, at least 40 bp, at least 45 bp, at least 50 bp, at least 55 bp, at least 60 bp, at least 65 bp, at least 70 bp, at least 75 bp, at least 80 bp, at least 85 bp, at least 90 bp, or at least 95 bp in length.

[0114] In a further example, the captured nucleic acid sequences can be derived from one or more DNA sequencing libraries. The adapters can be configured for use on next-generation sequencing platforms, such as Illumina sequencing platforms or IonTorrent platforms.

[0115] The adapters may contain restriction sites or primer binding sites of interest.

[0116] Exemplary adaptors include P5 and P7 adaptors.

[0117] Guide nucleic acid (gNA) Provided herein is a guide nucleic acid (gNA), wherein the gNA is complementary to (selective for, can hybridize to) a target site or sequence of interest in a nucleic acid, for example, in genomic DNA from a host, or a sequence of non-interest. The gNA directs a nucleic acid-guided nuclease to a specific site in the nucleic acid.

[0118] In some embodiments, the gNA is a guide RNA (gRNA). In other embodiments, the gNA is a guide DNA (gDNA). In some embodiments, the gNA comprises a mixture of gRNA and gDNA.

[0119] The host into which the gNA is derived may be an animal, such as a human, cow, horse, sheep, pig, monkey, dog, cat, gerbil, bird, mouse, or rat. The host may be a plant. The non-host may be a prokaryote, eukaryote, virus, bacterium, fungus, or protozoan.

[0120] In one embodiment, the invention provides a guide nucleic acid (gNA) library comprising a collection of gNAs configured to hybridize with nucleic acid sequences targeted for capture. In another embodiment, the invention provides a guide NA library comprising a collection of gNAs configured to hybridize with nucleic acid sequences not targeted for capture.

[0121] In one embodiment, the present invention provides a guide RNA library comprising a collection of gRNAs configured to hybridize with nucleic acid sequences targeted for capture. In another embodiment, the present invention provides a guide RNA library comprising a collection of gRNAs configured to hybridize with nucleic acid sequences not targeted for capture.

[0122] In one embodiment, the invention provides a guide DNA library comprising a collection of gDNA configured to hybridize with nucleic acid sequences targeted for capture, hi another embodiment, the invention provides a guide DNA library comprising a collection of gDNA configured to hybridize with nucleic acid sequences not targeted for capture.

[0123] In one embodiment, the gNA is selective for the target nucleic acid in the sample, but not for the sequence of interest from the sample.

[0124] In one embodiment, gNAs are used to sequentially capture nucleic acid sequences.

[0125] In some embodiments, the gNA is selective for the target nucleic acid sequence followed by a protospacer adjacent motif (PAM) sequence to which the nucleic acid-guided nuclease can bind. In some embodiments, the sequence of the gNA is determined by the type of nucleic acid-guided nuclease. In various embodiments, the PAM sequence may vary depending on the species of organism from which the nucleic acid-guided nuclease is derived, so the gNA can be tailored to different nucleic acid-guided nuclease types.

[0126] The gNAs (gRNAs or gDNAs) of the present invention can vary in size within a range of 50 to 200 base pairs. For example, gNAs of the present invention can be at least 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 110 bp, 120 bp, 125 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 175 bp, 180 bp, 190 bp, or 195 bp. In specific embodiments, the gNAs are 80 bp, 90 bp, 100 bp, or 110 bp. In some embodiments, the target-specific gNAs comprise a base pair sequence that can be complementary to a predetermined site in the target nucleic acid followed by a Protospacer adjacent motif or (PAM) sequence to which a nucleic acid-guided nuclease protein (e.g., Cas9) derived from a bacterial species can bind. In specific embodiments, the base pair sequence of the gNA that is complementary to a predetermined site in the target nucleic acid is 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45 or 50 base pairs.

[0127] The present invention also provides gNA libraries (e.g., gRNA libraries or gDNA libraries). A gNA library can contain several different species-specific guide NA (e.g., gRNA or gDNA) elements, each configured to hybridize (selectively) with a nucleic acid sequence targeted for capture, a nucleic acid sequence of interest, or a non-target nucleic acid sequence. Each gNA contains a target-specific guide sequence and a stem-loop binding site configured to bind to a nucleic acid-guided nuclease protein. In some embodiments, the library can contain multiple different guide NAs, each with a different 15-30 base pair sequence complementary to a different predetermined site in the targeted nucleic acid, followed by an appropriate PAM sequence to which the nucleic acid-guided nuclease protein can bind. For each guide NA, the PAM sequence is present in the predetermined DNA or RNA target sequence of the nucleic acid of interest, but is absent from the corresponding target-specific guide sequence.

[0128] Generally, according to the present invention, any nucleic acid sequence in a genome of interest, having a predetermined target sequence and a suitable subsequent PAM sequence, can be provided in a guided NA library, and the corresponding guided RNA that the nucleic acid guided nuclease binds can hybridize with.In various embodiments, the PAM sequence may vary depending on the bacterial species from which the nucleic acid guided nuclease is derived, so the gNA library can be tailored to different nucleic acid guided nuclease types.However, in some variations, the predetermined target sequence is not followed by a PAM sequence.

[0129] Different target-specific sequences in gNA can be generated. This can be done using promoters for bacteriophage RNA polymerases, such as RNA polymerases from bacteriophages T3, T7, SP6, etc. Thus, each different T7 RNA polymerase promoter provides a different target-specific sequence suitable for hybridizing with a different target nucleic acid sequence. Non-limiting exemplary sets of forward primers that can be used in annealing and subsequent PCR reactions are listed in Table 1 below.

[0130] A gNA library (e.g., a gRNA or gDNA library) can be amplified to contain multiple copies of each different guide NA element and multiple different guide NA elements, as appropriate for the desired capture results. The number of unique guide NA elements in a given guide NA library can range from one unique guide NA element to as many as 300,000,000 unique guide NA elements, or approximately one unique guide NA sequence per 10 base pairs in the human genome. The number of unique gNAs (e.g., gRNAs or gDNAs) can be at least about 10, 10, 10, 10, 10, 10, or 10 unique gNAs. The number of unique gNAs can result in the number of unique nucleic acid-guided nuclease-gNA complexes (e.g., CRISPR / Cas system protein-gRNA complexes).

[0131] Without being bound by theory, if gNAs exhibit approximately 100% efficacy, one can calculate the distance between gNAs to achieve >95% cleavage of the target nucleic acid: this can be calculated by measuring the size distribution of the library and determining the mean, N, and standard deviation, SD; N-2SD = minimum size for >95% of the library, ensuring there is one guide NA per fragment of this size to ensure >95% capture. This can also be written as maximum distance between guide NAs = mean library size - 2 x (standard deviation of library size).

[0132] In various embodiments of the present invention, gNAs may be specific for various targeting sites of interest, including, but not limited to, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), oncogenes, insertions, deletions, structural variants, exons, genetic mutations, and regulatory regions.

[0133] Nucleic Acid-Guided Nucleases Provided herein are compositions and methods for capturing nucleic acid from sample.These compositions and methods utilize nucleic acid guided nuclease.As used herein, " nucleic acid guided nuclease " refers to any endonuclease that uses one or more nucleic acid guided nucleic acids (gNA) to cleave DNA, RNA or DNA / RNA hybrids and provide specificity.Nucleic acid guided nuclease includes CRISPR / Cas system proteins and non-CRISPR / Cas system proteins.

[0134] The nucleic acid-guided nucleases provided herein can be DNA-guided DNA endonucleases; DNA-guided RNA endonucleases; RNA-guided DNA endonucleases; or RNA-guided RNA endonucleases.

[0135] In one embodiment, the nucleic acid-guided nuclease is a nucleic acid-guided DNA endonuclease.

[0136] In one embodiment, the nucleic acid-guided nuclease is a nucleic acid-guided RNA endonuclease.

[0137] Nucleic acid-guided nucleases of the CRISPR / Cas system In some embodiments, CRISPR / Cas system proteins are used in the embodiments provided herein. In some embodiments, CRISPR / Cas system proteins include proteins from CRISPR type I systems, CRISPR type II systems, and CRISPR type III systems.

[0138] In some embodiments, the CRISPR / Cas system protein may be from any bacterial or archaeal species.

[0139] In some embodiments, the CRISPR / Cas system proteins are isolated, recombinantly produced, or synthetic.

[0140] In some embodiments, the CRIPR / Cas family protein is Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Streptococcus thermophiles, Treponema denticola, Francisella tularensis, Pasteurella multocida, Campylobacter jejuni, Campylobacter lari, Mycoplasma gallisepticum, Nitratifractor salsuginis, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria cinerea, Gluconacetobacter diazotrophicus, Azospirillum, Sphaerochaeta globus, Flavobacterium columnare, Fluviicola taffensis, Bacteroides coprophilus, Mycoplasma mobile, Lactobacillus farciminis, Streptococcus pasteurianus, Lactobacillus johnsonii, Staphylococcus pseudintermedius, Filifactor alocis, Legionella pneumophila, Suterella wadsworthensis, or Corynebacter diphtheria.

[0141] In some embodiments, examples of CRISPR / Cas system proteins may be naturally occurring or engineered versions.

[0142] In some embodiments, naturally occurring CRISPR / Cas system proteins can belong to CAS class I type I, III or IV, or CAS class II type II or V, and can include Cas9, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn2, Cas4, Csm2, Cmr5, Csf1, C2c2, and Cpf1.

[0143] In an exemplary embodiment, the CRISPR / Cas system protein comprises Cas9.

[0144] A "CRISPR / Cas system protein-gNA complex" refers to a complex containing a CRISPR / Cas system protein and a guide NA (e.g., gRNA or gDNA). When the gNA is a gRNA, the gRNA can be composed of two molecules: one RNA ("crRNA") that hybridizes with the target and provides sequence specificity, and one RNA "tracrRNA" that can hybridize with the crRNA. Alternatively, the guide RNA can be a single molecule (i.e., gRNA) that contains the crRNA and tracrRNA sequences.

[0145] The CRISPR / Cas system protein may be at least 60% identical (e.g., at least 70%, at least 80%, or 90% identical, at least 95% identical, or at least 98% identical, or at least 99% identical) to a wild-type CRISPR / Cas system protein. The CRISPR / Cas system protein may have all of the functions of the wild-type CRISPR / Cas system protein, or only one or a portion of the functions, including binding activity, nuclease activity, and nuclease activity.

[0146] The term "CRISPR / Cas system protein-associated guide NA" refers to a guide NA. A CRISPR / Cas system protein-associated guide NA can exist as an isolated NA or as part of a CRISPR / Cas system protein-gNA complex.

[0147] Cas9 In some embodiments, the CRISPR / Cas system protein, nucleic acid-guided nuclease, is or comprises Cas9. The Cas9 of the present invention may be isolated, recombinantly produced, or synthetic.

[0148] Examples of Cas9 proteins that can be used in embodiments herein can be found in FA Ran, L. Cong, WX Yan, DA Scott, JS Gootenberg, AJ Kriz, B. Zetsche, O. Shalem, X. Wu, KS Makarova, EV Koonin, PA Sharp and F. Zhang; "In vivo genome editing using Staphylococcus aureus Cas9," Nature 520, 186-191 (April 9, 2015) doi:10.1038 / nature14299, which is incorporated herein by reference.

[0149] In some embodiments, Cas9 is Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Streptococcus thermophiles, Treponema denticola, Francisella tularensis, Pasteurella multocida, Campylobacter jejuni, Campylobacter lari, Mycoplasma gallisepticum, Nitratifractor salsuginis, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria cinerea, Gluconacetobacter diazotrophicus, Azospirillum, Sphaerochaeta globus, Flavobacterium columnare, Fluviicola taffensis, Bacteroides coprophilus, Mycoplasma mobile, Lactobacillus farciminis, Streptococcus pasteurianus, Lactobacillus johnsonii, Staphylococcus pseudintermedius, Filifactor alocis, Legionella pneumophila, Suterella wadsworthensis or Corynebacter diphtheria.

[0150] In some embodiments, the Cas9 is a Type II CRISPR system derived from S. pyogenes, and the PAM sequence is NGG located immediately at the 3' end of the target-specific guide sequence. PAM sequences for Type II CRISPR systems from exemplary bacterial species can also include: Streptococcus pyogenes (NGG), Staph aureus (NNGRRT), Neisseria meningitidis (NNNNGA). TT), Streptococcus thermophilus (NNAGAA) and Treponema denticola (NAAAAC), all of which may be used without departing from the invention.

[0151] In one exemplary embodiment, the Cas9 sequence can be obtained from the pX330 plasmid (available from Addgene), which is reamplified by PCR and then cloned into pET30 (from EMD biosciences) for bacterial expression and purification of the recombinant 6His-tagged protein, for example.

[0152] "Cas9-gNA complex" refers to a complex containing a Cas9 protein and a guide NA. The Cas9 protein can be at least 60% identical (e.g., at least 70%, at least 80%, or 90% identical, at least 95% identical, at least 98% identical, or at least 99% identical) to a wild-type Cas9 protein, for example, to the Cas9 protein of Streptococcus pyogenes. The Cas9 protein can have all of the functions of the wild-type Cas9 protein, or only one or a portion of the functions, including binding activity, nuclease activity, and nuclease activity.

[0153] The term "Cas9-associated guide NA" refers to such a guide NA. The Cas9-associated guide NA can be present in isolation or as part of a Cas9-gNA complex.

[0154] Non-CRISPR / Cas nucleic acid-guided nucleases In some embodiments, non-CRISPR / Cas system proteins are used in the embodiments provided herein.

[0155] In some embodiments, the non-CRISPR / Cas system protein may be from any bacterial or archaeal species.

[0156] In some embodiments, the non-CRISPR / Cas system protein is isolated, recombinantly produced, or synthetic.

[0157] In some embodiments, the non-CRISPR / Cas-based protein is Aquifex aeolicus, Thermus thermophilus, Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Streptococcus thermophiles, Treponema denticola, Francisella tularensis, Pasteurella multocida, Campylobacter jejuni, Campylobacter lari, Mycoplasma gallisepticum, Nitratifractor salsuginis, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria cinerea, Gluconacetobacter diazotrophicus, Azospirillum, Sphaerochaeta globus, Flavobacterium columnare, Fluviicola taffensis, Bacteroides coprophilus, Mycoplasma mobile, Lactobacillus farciminis, Streptococcus pasteurianus, Lactobacillus johnsonii, Staphylococcus pseudintermedius, Filifactor alocis, Legionella pneumophila, Suterella wadsworthensis, Natronobacterium gregoryi or Corynebacter diphtheria.

[0158] In some embodiments, the non-CRISPR / Cas system protein may be a naturally occurring or an engineered version.

[0159] In some embodiments, the naturally occurring non-CRISPR / Cas system protein is NgAgo (Argonaute from Natronobacterium gregoryi).

[0160] A "non-CRISPR / Cas system protein-gNA complex" refers to a complex containing a non-CRISPR / Cas system protein and a guide NA (e.g., gRNA or gDNA). When the gNA is a gRNA, the gRNA can be composed of two molecules: one RNA ("crRNA") that hybridizes with the target and provides sequence specificity, and one RNA "tracrRNA" that can hybridize with the crRNA. Alternatively, the guide RNA can be a single molecule (i.e., gRNA) that contains the crRNA and tracrRNA sequences.

[0161] The non-CRISPR / Cas system protein may be at least 60% identical (e.g., at least 70%, at least 80%, or 90% identical, at least 95% identical, or at least 98% identical, or at least 99% identical) to the wild-type non-CRISPR / Cas system protein. The non-CRISPR / Cas system protein may have all of the functions of the wild-type non-CRISPR / Cas system protein, or only one or some of the functions, including binding activity, nuclease activity, and nuclease activity.

[0162] The term "non-CRISPR / Cas system protein-associated guide NA" refers to a guide NA. Non-CRISPR / Cas system protein-associated guide NA can exist as an isolated NA or as part of a non-CRISPR / Cas system protein-gNA complex.

[0163] Catalytically inactive nucleic acid-guided nucleases In some embodiments, engineered examples of nucleic acid-guided nucleases include catalytically inactive nucleic acid-guided nucleases (CRISPR / Cas-based nucleic acid-guided nucleases or non-CRISPR / Cas-based nucleic acid-guided nucleases). The term "catalytically inactive" generally refers to nucleic acid-guided nucleases with inactivated nucleases, such as inactivated HNH and RuvC nucleases. Such proteins can bind to target sites in any nucleic acid (the target site is determined by the guide NA), but the proteins cannot cleave or nick the nucleic acid.

[0164] Thus, catalytically inactive nucleic acid-guided nucleases allow the separation of the mixture into unbound nucleic acids and fragments bound to the catalytically inactive nucleic acid-guided nuclease. The use of inactive nucleic acid-guided nucleases is, for example, depicted in Protocols 3 and 4 of Figures 9 and 10, respectively. In an exemplary embodiment, the dCas9 / gRNA complex binds to a target determined by the gRNA sequence. As depicted in Figure 10, dCas9 binding can prevent cleavage by Cas9, while other operations proceed.

[0165] In another embodiment, a catalytically inactive nucleic acid-guided nuclease can be fused to another enzyme, such as a transposase, to target the activity of that enzyme to a specific site.

[0166] In some embodiments, the catalytically inactive nucleic acid-guided nuclease is dCas9, dCpf1, dCas3, dCas8a-c, dCas10, dCsel, dCsy1, dCsn2, dCas4, dCsm2, dCm5, dCsf1, dC2C2, or dNgAgo.

[0167] In one exemplary embodiment, the catalytically inactive nucleic acid-guided nuclease protein is dCas9.

[0168] Nickase, a nucleic acid-guided nuclease In some embodiments, engineered examples of nucleic acid-guided nucleases include nucleic acid-guided nucleases, nickases (interchangeably referred to as nickase-nucleic acid-guided nucleases).

[0169] In some embodiments, engineered examples of nucleic acid-guided nucleases include CRISPR / Cas-based nickases or non-CRISPR / Cas-based nickases that contain a single, inactive catalytic domain.

[0170] In some embodiments, the nucleic acid-guided nuclease nickase is a Cas9 nickase, Cpf1 nickase, Cas3 nickase, Cas8a-c nickase, Cas10 nickase, Csel nickase, Csy1 nickase, Csn2 nickase, Cas4 nickase, Csm2 nickase, Cm5 nickase, Csf1 nickase, C2C2 nickase, or NgAgo nickase.

[0171] In one embodiment, the nucleic acid-guided nuclease nickase is a Cas9 nickase.

[0172] In some embodiments, a nucleic acid-guided nuclease, or nickase, can be used to bind to a target sequence. Because of its single active nuclease domain, the nucleic acid-guided nuclease, or nickase, cleaves only one strand of the target DNA, generating a single-strand break or "nick." Depending on which variant is used, it can cleave either the hybridized strand or the unhybridized strand with the guide NA. A nucleic acid-guided nuclease, or nickase, bound to two gNAs targeting opposite strands can create a double-strand break in the nucleic acid. This "dual nickase" strategy increases cleavage specificity because both nucleic acid-guided nuclease / gNA complexes must specifically bind to the site before the double-strand break is formed.

[0173] In an exemplary embodiment, Cas9 nickase can be used to bind to a target sequence. The term "Cas9 nickase" refers to a modified version of the Cas9 protein that contains a single inactive catalytic domain, i.e., a RuvC or HNH domain. Due to only one active nuclease domain, Cas9 nickase cleaves only one strand of the target DNA, generating a single-strand break or "nick." Depending on which variant is used, it can cleave the strand hybridized to the guide RNA or the unhybridized strand. Cas9 nickase bound to two gRNAs targeting opposite strands creates a double-strand break in the DNA. This "dual nickase" strategy can increase cleavage specificity because it requires both Cas9 / gRNA complexes to specifically bind at the site before the double-strand break is formed.

[0174] Capture of DNA can be performed using a nucleic acid-guided nuclease, nickase, as illustrated in protocols 2 and 5 in Figures 6 and 11, respectively. In an exemplary embodiment, as depicted in Figures 6 and 11, the nucleic acid-guided nuclease nickase cleaves one strand of a double-stranded nucleic acid, and the double-stranded region contains methylated nucleotides.

[0175] Releasable and thermostable nucleic acid-guided nucleases In some embodiments, a thermostable nucleic acid-guided nuclease is used in the methods provided herein (a thermostable CRISPR / Cas-based nucleic acid-guided nuclease or a thermostable non-CRISPR / Cas-based nucleic acid-guided nuclease). In such embodiments, the reaction temperature is increased to induce protein dissociation. The reaction temperature is decreased to allow for the generation of additional cleaved target sequences. In some embodiments, when maintained at at least 75°C for at least 1 minute, the thermostable nucleic acid-guided nuclease maintains at least 50% activity, at least 55% activity, at least 60% activity, at least 65% activity, at least 70% activity, at least 75% activity, at least 80% activity, at least 85% activity, at least 90% activity, at least 95% activity, at least 96% activity, at least 97% activity, at least 98% activity, at least 99% activity, or 100% activity. In some embodiments, a thermostable nucleic acid-guided nuclease maintains at least 50% activity when maintained at at least 75°C, at least 80°C, at least 85°C, at least 90°C, at least 91°C, at least 92°C, at least 93°C, at least 94°C, at least 95°C, 96°C, at least 97°C, at least 98°C, at least 99°C, or at least 100°C for at least 1 minute. In some embodiments, a thermostable nucleic acid-guided nuclease maintains at least 50% activity when maintained at at least 75°C for at least 1 minute, 2 minutes, 3 minutes, 4 minutes, or 5 minutes. In some embodiments, a thermostable nucleic acid-guided nuclease maintains at least 50% activity when the temperature is increased and then decreased to between 25°C and 50°C. In some embodiments, the temperature is decreased to 25°C, 30°C, 35°C, 40°C, 45°C, or 50°C. In an exemplary embodiment, the thermostable enzyme retains at least 90% activity after 1 minute at 95°C.

[0176] In some embodiments, the thermostable nucleic acid-guided nuclease is thermostable Cas9, thermostable Cpf1, thermostable Cas3, thermostable Cas8a-c, thermostable Cas10, thermostable Cse1, thermostable Csy1, thermostable Csn2, thermostable Cas4, thermostable Csm2, thermostable Cm5, thermostable Csf1, thermostable C2C2, or thermostable NgAgo.

[0177] In some embodiments, the thermostable CRISPR / Cas system protein is a thermostable Cas9.

[0178] Thermostable nucleic acid-guided nucleases can be isolated and identified, for example, by sequence homology in the genomes of thermophilic bacteria Streptococcus thermophilus and Pyrococcus furiosus. The nucleic acid-guided nuclease gene can then be cloned into an expression vector. In an exemplary embodiment, a thermostable Cas9 protein is isolated.

[0179] In another embodiment, a thermostable nucleic acid-guided nuclease can be obtained by in vitro evolution of a non-thermostable nucleic acid-guided nuclease. The sequence of the nucleic acid-guided nuclease can be mutagenized to improve its thermostability.

[0180] Exemplary Compositions of the Invention In one embodiment, a composition is provided herein that includes a nucleic acid fragment, a nickase nucleic acid-guided nuclease-gNA complex, and a labeled nucleotide. In an exemplary embodiment, a composition is provided herein that includes a nucleic acid fragment, a nickase Cas9-gRNA complex, and a labeled nucleotide. In such an embodiment, the nucleic acid can include DNA. The nucleotide can be labeled, for example, with biotin. The nucleotide can be part of an antibody-conjugate pair.

[0181] In one embodiment, provided herein is a composition comprising a nucleic acid fragment and a catalytically inactive nucleic acid-guided nuclease-gNA complex, wherein the catalytically inactive nucleic acid-guided nuclease is fused to a transposase. In one exemplary embodiment, provided herein is a composition comprising a DNA fragment and a dCas9-gRNA complex, wherein the dCas9 is fused to a transposase.

[0182] In one embodiment, provided herein is a composition comprising a nucleic acid fragment comprising a methylated nucleotide, a nickase nucleic acid-guided nuclease-gRNA complex, and an unmethylated nucleotide. In an exemplary embodiment, provided herein is a composition comprising a DNA fragment comprising a methylated nucleotide, a nickase Cas9-gRNA complex, and an unmethylated nucleotide.

[0183] In one embodiment, provided herein is gDNA complexed with a nucleic acid-guided DNA endonuclease. In an exemplary embodiment, the nucleic acid-guided DNA endonuclease is NgAgo.

[0184] In one embodiment, provided herein is gDNA complexed with a nucleic acid-guided RNA endonuclease.

[0185] In one embodiment, provided herein is a gRNA complexed with a nucleic acid-guided DNA endonuclease.

[0186] In one embodiment, provided herein is a gRNA complexed with a nucleic acid-guided RNA endonuclease. In one embodiment, the nucleic acid-guided RNA endonuclease comprises C2c2.

[0187] Kits and Manufactured Products The application provides kits comprising any one or more of the compositions described herein, including, but not limited to, adaptors, gNAs, gDNAs, gRNAs, gNA libraries, gRNA libraries, gDNA libraries, nucleic acid-guided nucleases, catalytically inactive nucleic acid-guided nucleases, nickase nucleic acid-guided nucleases, CRISPR / Cas system proteins, nickase CRISPR / Cas system proteins, catalytically inactive CRISPR / Cas system proteins, Cas9, dCas9, Cas9 nickase, methylated nucleotides, labeled nucleotides, biotinylated nucleotides, avidin, streptavidin, enzymes capable of initiating nucleic acid synthesis at nick sites, DNA polymerase I, TAQ polymerase, bst DNA polymerase, enzymes capable of cleaving methylated nucleotides, Dpn1 enzyme, and enzymes capable of methylating DNA (e.g., Dam / Dcm1 ​​methyltransferase).

[0188] In one embodiment, the kit includes a collection or library of gNAs, wherein the gNAs target human genomic DNA sequences, such as specific genes of interest (e.g., cancer genes), SNPs, or STRs. In another exemplary embodiment, the kit includes a collection or library of gNAs, wherein the gNAs target DNA sequences of mammals other than humans. In another exemplary embodiment, the kit includes a collection or library of gNAs, wherein the gNAs target human ribosomal RNA sequences. In another exemplary embodiment, the kit includes a collection or library of gNAs, wherein the gNAs target human mitochondrial DNA sequences.

[0189] In one exemplary embodiment, the kit includes a collection or library of gRNAs, wherein the gRNAs target human genomic DNA sequences, e.g., specific genes of interest (e.g., cancer genes), SNPs, STRs. In another exemplary embodiment, the kit includes a collection or library of gRNAs, wherein the gRNAs target DNA sequences of mammals other than humans. In another exemplary embodiment, the kit includes a collection or library of gRNAs, wherein the gRNAs target human ribosomal RNA sequences. In another exemplary embodiment, the kit includes a collection or library of gRNAs, wherein the gRNAs target human mitochondrial DNA sequences.

[0190] The present application also provides an article of manufacture comprising any one of the kits described herein. Examples of articles of manufacture include vials (including sealed vials).

[0191] The following examples are included for illustrative purposes and are not intended to limit the scope of the invention. [Example]

[0192] Example 1: Capture of mitochondrial DNA from total human genomic DNA (Protocol 1 for DNA capture) overview The goal of this method, as depicted in Figure 1, was to capture mitochondrial DNA from a library of human genomic DNA. Human tissue specimens were subjected to a DNA extraction protocol, resulting in a DNA sample containing >99% human DNA and <1% target sequences. The DNA samples were then subjected to a sequencing library construction protocol, resulting in a nucleic acid library containing sequencing indexing adapters. To generate a library of target-specific guide RNAs (gRNAs), a library of target-specific gRNA precursors, each containing a T7 RNA polymerase promoter, a human-specific 20-base pair region, and a stem-loop binding site for Cas9, was subjected to in vitro transcription to obtain a library of target-specific guide RNAs. The library of target-specific guide RNAs was then combined with Cas9 protein to obtain a library of Cas9-gRNA complexes. The Cas9-gRNA complexes were then combined with the nucleic acid library so that Cas9 cleaved the corresponding target DNA sequences and left other DNA uncut. A second adapter was added and specifically ligated to the 5' phosphorylated blunt ends of the cleaved DNA. PCR was then used to specifically amplify using the first and second adapters.

[0193] Mitochondrial DNA comprises approximately 0.1-0.2% of total human genomic DNA. To test the precise site-specific cleavage of DNA by Cas9 and subsequent adapter ligation, as shown in Figure 2, test DNA (e.g., a plasmid) 201 was cleaved by Cas9 at either the first position 202 or the second position 203, resulting in a first product 204 or a second product 205 (see, e.g., Figure 2). Adapters were ligated to the newly available blunt DNA ends 206. PCR amplification was performed for AUK-F / P7 or MB1OriR / P7, yielding two distinct products per reaction. When the first position was cleaved, the product was 212 and 1.9kb in size. When the second position was cleaved, the product was 359 and 1.75kb in size. Verification was performed by sequencing.

[0194] The results showed that Cas9 cleavage and subsequent adapter ligation enabled specific amplification of target DNA (see, e.g., Figure 3). Sequencing of the amplified DNA showed that adapter ligation occurred only at the position specified by the guide RNA (see, e.g., Figure 4).

[0195] Next, mitochondrial DNA was enriched from a mixture containing primarily human nuclear DNA. 25 guide RNAs specific for mitochondrial DNA were used, followed by cleavage with Cas9 and adapter ligation, followed by amplification. As can be seen in Figure 5, two separate reactions allowed for the amplification of mitochondrial DNA, while the reaction without guide RNA did not amplify any DNA, thus demonstrating that this method efficiently amplifies DNA that is underrepresented in any given library.

[0196] DNA library preparation A human genomic DNA library was generated by end-repairing 500 ng of fragmented human genomic DNA (treated with dsDNA fragmentase, NEB, at 37°C for 1 hour) using a blunt-end repair kit (NEB) for 20 minutes at 25°C. The reaction was then heat-inactivated at 75°C for 20 minutes, cooled to 25°C, and ligated to 15 pmoles of P5 / Myc adapter using T4 DNA ligase (NEB) for 2 hours at 25°C. Adapter dimers were removed using an NGS cleanup kit (Life Technologies).

[0197] Cas9 expression To insert a hexahistidine tag immediately upstream of the Cas9 start codon, Cas9 (from S. pyogenes) was cloned into the pET30 expression vector (EMD biosciences). The resulting plasmid was transformed into the Rosetta (DE3) BL21 bacterial strain (EMD biosciences) and grown in 1 L of LB medium with vigorous aeration until the culture reached an optical density (OD at 600 nm) of 0.4. The temperature was reduced to 25°C, 0.2 mM IPTG was added, and the culture was grown for an additional 4 hours. Cells were then harvested by centrifugation (1,000 × g for 20 minutes at 4°C), resuspended in 10 ml binding buffer (20 mM Tris pH 8, 0.5 M NaCl, 5 mM imidazole, 0.05% NP40), and lysed by sonication (7 × 10-second bursts at 30% power, Sonifier 250, Branson). Insoluble cell debris was removed by centrifugation at 10,000 × g for 20 minutes. The supernatant containing the soluble protein was then mixed with 0.4 ml of NTA beads (Qiagen) and loaded onto the column. The beads were washed three times with 4 ml of binding buffer and then eluted with 3 × 0.5 ml of binding buffer supplemented with 250 mM imidazole. The eluted fraction was then concentrated using a 30,000 MWCO protein concentrator (Life Technologies), buffer exchanged with storage buffer (10 mM Tris pH 8, 0.3 M NaCl, 0.1 mM EDTA, 1 mM DTT, 50% glycerol), and examined by SDS-PAGE followed by Colloidal Blue staining (Life Technologies), quantified, and then stored at -20 °C for later use.

[0198] A mutant Cas9 nickase, the D10A mutant of S. pyogenes Cas9, can be generated and purified using the same techniques used to generate Cas9 as described above.

[0199] Preparation of gRNA1 and gRNA2 Three oligonucleotides, T7-guide RNA 1 and 2 (5' to 3' sequences: GCCTCGAGCTAATACGACTCACTATAGGGATTTATACAGCACTTTAA and GCCTCGAGCTAATACGACTCACTATAGGGTCTTTTTGGTCCTCGAAG) and stlgR (sequence: GT TTT AGA GCT AGA AAT AGC AAG TTA AAA TAA GGC TAG TCC GTT ATC AAC TTG AAA AAG TGG CAC CGA GTC GGT GCT TTT TTT GGA TCC GAT GC) were ordered and synthesized (IDT). The stlgR oligonucleotide (300 pmoles) was sequentially 5' phosphorylated using T4 PNK (New England Biolabs) according to the manufacturer's instructions, followed by 5' adenylation using a 5' adenylation kit (New England Biolabs). T7-guide RNA oligonucleotides (5 pmoles) and 5' adenylated stlgR (10 pmoles) were then ligated using thermostable 5' App DNA / RNA ligase (New England Biolabs) at 65°C for 1 hour. The ligation reaction was heat-inactivated at 90°C for 5 minutes and then amplified by PCR (OneTaq, New England Biolabs, 30 cycles: 95°C for 30 seconds, 57°C for 20 seconds, and 72°C for 20 seconds) with primers ForT7 (sequence GCC TCG AGC TAA TAC GAC TCA C) and gRU (sequence AAAAAAAGCACCGACTCGGTG). The PCR product was purified using a PCR cleanup kit (Life Technologies) and verified by agarose gel electrophoresis and sequencing. The verified product was then used as a template for in vitro transcription.

[0200] Guide RNA library preparation T7-guide RNA oligonucleotides (Table 1) and another oligonucleotide, stlgR (sequence, GT TTT AGA GCT AGA AAT AGC AAG TTA AAA TAA GGC TAG TCC GTT ATC AAC TTG AAA AAG TGG CAC CGA GTC GGT GCT TTT TTT GGA TCC GAT GC) were ordered and synthesized (IDT).

[0201] stlgR oligonucleotide (300 pmoles) was added to T4 Sequential 5' phosphorylation was performed using PNK (New England Biolabs), followed by 5' adenylation using a 5' adenylation kit (New England Biolabs). The T7-guide RNA oligonucleotide (5 pmoles) and 5' adenylated stlgR (10 pmoles) were then ligated using thermostable 5' App DNA / RNA ligase (New England Biolabs) at 65°C for 1 hour. The ligation reaction was heat-inactivated at 90°C for 5 minutes and then amplified by PCR (OneTaq, New England Biolabs, 30 cycles: 95°C for 30 seconds, 57°C for 20 seconds, and 72°C for 20 seconds) with primers ForT7 (sequence GCC TCG AGC TAA TAC GAC TCA C) and gRU (sequence AAAAAAAGCACCGACTCGGTG). The PCR product was purified using a PCR cleanup kit (Life Technologies) and verified by agarose gel electrophoresis and sequencing. The validated product was then used as a template for in vitro transcription. [Table 1-1] [Table 1-2]

[0202] In vitro transcription The validated product was then used as a template for an in vitro transcription reaction using the HiScribe T7 Transcription Kit (New England Biolabs). Following the manufacturer's instructions, 500–1000 ng of template was incubated overnight at 37°C. To transcribe the guide library into guide RNA, the following in vitro transcription reaction mixture was assembled: 10 μl of purified library (approximately 500 ng), 6.5 μl of HO, 2.25 μl of ATP, 2.25 μl of CTP, 2.25 μl of GTP, 2.25 μl of UTP, 2.25 μl of 10x Reaction Buffer (NEB), and 2.25 μl of T7 RNA Polymerase Mix. The reaction was incubated at 37°C for 24 hours, then purified using an RNA Cleanup Kit (Life Technologies), eluted in 100 μl of RNase-free water, quantified, and stored at -20°C until use.

[0203] DNA-specific Cas9-mediated fragmentation Test guide RNAs 1 and 2 were used separately to cleave test DNA. Guide RNA series 1 and 2 were used to cleave and enrich mitochondrial DNA from a human genomic DNA library. Diluted guide RNA (1 μl, equivalent to 2 pmoles) was combined with 3 μl of 10x Cas9 reaction buffer (NEB), 20 μl of HO, and 1 μl of recombinant Cas9 enzyme (NEB, 1 pmole / μl). A control reaction using a control guide RNA targeting the following sequence (5'-GGATTTATACAGCACTTTAA-3') was run separately using the same parameters. This sequence is absent from both human chromosomal and mitochondrial DNA. The reaction was incubated at 37°C for 15 minutes, then supplemented with 5 μl of diluted DNA library (50 pg / μl), and incubation at 37°C continued for 90 minutes. The reaction was terminated by adding RNase A (Thermo Fisher Scientific) at a dilution of 1:100, and the DNA was then purified using a PCR clean-up kit (Life Technologies) and eluted in 30 μl of 10 mM Tris-Cl pH 8. The reaction was then stored at −20°C until use.

[0204] Adapter ligation and PCR analysis For test DNA, the reaction after Cas9 digestion was incubated with 15 pmoles of P5 / P7 adapter and T4 DNA ligase (NEB) for 1 hour at 25°C. The ligation was then used as a template for PCR using the test DNA-F primer (sequence ATGCCGCAGCACTTGG) and the P5 primer (sequence AATGATACGGCGACCACCGA). Successful PCR products were confirmed by agarose gel electrophoresis followed by sequencing using the test DNA-F primer (ElimBio) to demonstrate that cleavage and ligation occurred at the target DNA sequence.

[0205] Elimination of adapters ligated at old ends For mitochondrial DNA enrichment, after Cas9 digestion, the reaction was incubated with 15 pmoles of Flag5 / P7 adapter and T4 DNA ligase (NEB) at 25°C for 1 hour. Several molecular biology methods can be used to eliminate adapter ligation at the ends of the previous P5 and P7 adapters from the original library, for example, using enzyme treatment. The reaction was then used as a template for PCR using P7 (sequence CAAGCAGAAGACGGCATACGA) and P5 primers (sequence AATGATACGGCGACCACCGA). Successful PCR products were confirmed by agarose gel electrophoresis.

[0206] Example 2: Use of Cas9 Nickase to Label and Subsequent Purification of Test DNA from a DNA Mixture (Protocol 2 for Capture of DNA) overview The goal of this method was to capture regions of interest (e.g., SNPs, STRs, etc.) from a library of human genomic DNA, as depicted in Figure 6. To apply this protocol, test DNA containing a site for a nicking enzyme (NtAlwI (NEB)) guided by a target-specific guide RNA was mixed at 1% or 5% with another pool of DNA that did not contain this site. The nickase Cas9 cleaves only at the target sequence and only cleaves one strand of DNA. The nickase was used to nick the target sequence. The single-strand break (nick) is a substrate for DNA polymerase I, which can be used to replace the DNA downstream of the nick with biotin-labeled DNA. The biotinylated DNA of interest was then purified, amplified, and sequenced.

[0207] A mixture of DNA containing a nicking target site (e.g., GGATC) 703 (701) and a mixture of DNA not containing the target site (702) were nicked using NtAlwI 704 (see, e.g., Figure 7). DNA lacking the nickase site was present in 100-fold excess compared to the target DNA containing the nickase site. The single-strand break (nick) is a substrate for DNA polymerase I 705, which is used to replace the DNA downstream of the nick with biotin-labeled DNA. The target biotin-labeled DNA was isolated by streptavidin binding and washing 706. PCR amplification was performed with specific primers 707 to obtain the amplified region of interest 708 and the amplified unlabeled sequence 709, which may have been present due to nonspecific labeling or capture, for example.

[0208] Two specific PCR reactions (one for the test DNA with the region of interest and one for the other DNA without the region of interest) showed that the test DNA was enriched approximately 50-fold (see, e.g., Figure 8).

[0209] Expression and purification of Cas9 nickase A mutant Cas9 nickase, the D10A mutant of S. pyogenes Cas9, can be generated and purified using the same techniques used to generate Cas9 as described above.

[0210] Sequence-specific Cas9 nickase-mediated nicking Diluted guide RNA (1 μl, equivalent to 2 pmoles) targeting the following sequence was combined with 3 μl of 10x Cas9 reaction buffer (NEB), 20 μl of HO, and 1 μl of recombinant Cas9 nickase enzyme (10 pmoles / μl). This target sequence is present only on the target DNA (accounting for 5% or 1% of the total DNA). The reaction was incubated at 37°C for 15 minutes, then replenished with 5 μl of diluted DNA (100 ng total), and incubation at 37°C continued for 90 minutes. The reaction was terminated by adding RNase A (Thermo Fisher Scientific) at a 1:100 dilution. The DNA was then purified using a PCR cleanup kit (Life Technologies) and eluted in 30 μl of 10 mM Tris-Cl pH 8. The reaction was then stored at -20°C until use.

[0211] Biotin Nick Translation The nicked DNA was incubated with E. coli DNA polymerase I, which has 5'>3' exonuclease activity and is capable of initiating DNA synthesis at the nick, thereby replacing the nucleotide downstream of the nick with a labeled nucleotide (in this case, a biotin-labeled nucleotide). The nick-labeling reaction was carried out at 25°C for 30 minutes in 20 μl of DNA polymerase buffer (NEB) containing 1 unit of E. coli DNA polymerase I (NEB) and 0.02 mM each of dCTP, dGTP, and dTTP, 0.01 mM dATP, and 0.01 mM biotin-C14-labeled dATP (Life Technologies). The reaction was terminated by adding 1 mM EDTA.

[0212] Concentration of biotin-labeled DNA Streptavidin C1 beads (5 μl per reaction, Life Technologies) were resuspended in 1 ml of binding buffer (50 mM Tris-Cl pH 8, 1 mM EDTA, 0.1% Tween 20), attached to a magnetic rack, and washed twice with binding buffer. The beads were then resuspended in 30 μl of binding buffer, mixed with the nick-translation reaction, and then incubated at 25°C for 30 minutes. The beads were captured using a magnetic rack and washed four times with 0.5 ml of binding buffer, followed by three times with 0.5 ml of 10 mM Tris-Cl pH 8. The beads were then resuspended in 20 μl of 10 mM Tris-Cl pH 8 and subsequently used as a template for PCR to determine the ratio of test DNA to other DNAs.

[0213] Example 3: Use of catalytically inactive nucleic acid-guided nuclease-transposase fusions to insert adapters into a human genomic library and subsequently enrich for specific SNPs (Protocol 3 for DNA capture) In this example, a catalytically inactive nucleic acid-guided nuclease-transposase fusion protein (e.g., dCas9-transposase fusion protein) is expressed and purified from E. coli as described for Cas9 purification. The fusion protein forms a complex with an adapter (Nextera) and then a guide RNA (e.g., gRNA) targeting a region of interest (e.g., a human SNP). The complex is then added to human genomic DNA. The region of interest is targeted by the catalytically inactive nucleic acid-guided nuclease, and the transposase-adapter complex is positioned near the region of interest, allowing for adapter insertion. Human SNPs can then be amplified by PCR and subsequently sequenced using MiSeq, thus enriching human SNPs from human genomic DNA.

[0214] Example 4: Use of catalytically inactive and nucleic acid-guided nucleases to protect mitochondrial DNA and digest remaining human nuclear DNA from a human genomic DNA library (Protocol 4 for DNA capture) A human genomic DNA library containing P5 / P7 adaptors is obtained from clinical, forensic, or environmental samples. To enrich for specific regions (e.g., SNPs), guide RNAs targeting these regions are generated and incubated with a catalytically inactive nucleic acid-guided nuclease (e.g., dCas9) and then added to the human genomic DNA library for 20 minutes at 37°C. Next, a library of guide RNAs covering the human genome, complexed with an active nucleic acid-guided nuclease (e.g., Cas9), is added. The catalytically inactive nucleic acid-guided nuclease remains bound to the target location, protecting the region of interest (e.g., SNPs) from cleavage, which would render them unamplifiable by PCR and unsequencing. Thus, the DNA of interest remains intact, while all other DNA is cleaved and eliminated. The DNA of interest is recovered from the reaction using a PCR cleanup kit, PCR amplified, and sequenced using a MiSeq.

[0215] For proof-of-concept testing, we enriched mitochondrial DNA from a total human genomic DNA library. A mitochondrial-specific guide RNA was added to dCas9, which then added it to the library. The random guide RNA complexed with Cas9 then degraded any sequences except those that were inaccessible because they were protected by the already bound dCas9. This allowed for enrichment of mitochondrial DNA to levels much higher than the original 0.1–0.2%.

[0216] Example 5: Use of a nucleic acid-guided nuclease, nickase (e.g., Cas9 nickase), to protect and then enrich SNPs from human genomic DNA by replacing methylated DNA with unmethylated DNA (Protocol 5 for DNA capture) overview The goal of this method is to capture a region of interest from a library of human genomic DNA, as depicted in Figure 11. As a proof of principle, test DNA containing a site for the nicking enzyme, NtBbvCI (NEB), was mixed at 1% or 5% with another pool of DNA that did not contain this site. The entire mixture was then treated with Dam methyltransferase to add methyl groups to all GATC sequences. Both the test DNA and the other DNA contained GATC motifs in their sequences. The mixture was then nicked using NtBbvCI, then incubated with DNA polymerase I and unlabeled nucleotides (dATP, dCTP, dGTP, dTTP), and then heat-inactivated at 75°C for 20 minutes. The mixture was then digested with DpnI, which digests only methylated DNA but not unmethylated or semi-methylated DNA. Figure 12 shows that methylation of the test DNA renders it susceptible to DpnI-mediated cleavage.

[0217] Expression and purification of Cas9 nickase A mutant Cas9 nickase, the D10A mutant of S. pyogenes Cas9, can be generated and purified using the same techniques used to generate Cas9 as described above.

[0218] Sequence-specific Cas9 nickase-mediated nicking Diluted guide RNA (1 μl, equivalent to 2 pmoles) targeting the following sequence was combined with 3 μl of 10x Cas9 reaction buffer (NEB), 20 μl of HO, and 1 μl of recombinant Cas9 nickase enzyme (10 pmoles / μl). This target sequence is present only on the target DNA (accounting for 5% or 1% of the total DNA). The reaction was incubated at 37°C for 15 minutes, then replenished with 5 μl of diluted DNA (100 ng total), and incubation at 37°C continued for 90 minutes. The reaction was terminated by adding RNase A (Thermo Fisher Scientific) at a 1:100 dilution. The DNA was then purified using a PCR cleanup kit (Life Technologies) and eluted in 30 μl of 10 mM Tris-Cl pH 8. The reaction was then stored at -20°C until use.

[0219] Unlabeled DNA nick translation The nicked DNA was incubated with E. coli DNA polymerase I, which has 5'>3' exonuclease activity and is capable of initiating DNA synthesis at the nick, thereby replacing the nucleotide downstream of the nick with a labeled nucleotide (in this procedure, a biotin-labeled nucleotide). The nick-labeling reaction was carried out at 25°C for 30 minutes in 20 μl of DNA polymerase buffer (NEB) with 1 unit of E. coli DNA polymerase I (NEB) and 0.02 mM each of dATP, dCTP, dGTP, and dTTP. The reaction was terminated by heat inactivation at 75°C for 20 minutes.

[0220] Digestion with DpnI The reaction is then incubated with DpnI (NEB) for 1 hour at 37° C. The DNA is recovered using a PCR cleanup kit (Life Technologies) and then used as a template for PCR using test DNA-specific primers.

[0221] Example 6: Use of a nucleic acid-guided nuclease, a nickase (e.g., Cas9 nickase), to generate long 3' overhangs flanking a region of interest and then ligate adapters to these overhangs to enrich target DNA (Protocol 6 for DNA capture) This method, as depicted in Figure 13, can be used to enrich regions of DNA from any DNA source (e.g., library, genome, or PCR). A nucleic acid-guided nuclease, a nickase (e.g., Cas9 nickase), can be targeted to proximal sites using two guide NAs (e.g., gRNAs), resulting in nicking of the DNA at each site. Alternatively, adapters can be ligated on only one side, then filled in, and then adapters ligated on the other side. The two nicks can be close to each other (e.g., within 10–15 bp). A single nick can be generated on an off-target molecule. The proximity of the two nick sites can result in a double-stranded break when the reaction is heated, for example, to 65°C, resulting in a long (e.g., 10–15 bp) 3' overhang. These overhangs can be recognized by a thermostable single-stranded DNA / RNA ligase, such as a thermostable 5' App DNA / RNA ligase, to enable site-specific ligation of single-stranded adapters. The ligase, for example, can recognize only the long 3' overhang, thus ensuring that the adapter is not ligated at other sites. This process can be repeated using a nucleic acid-guided nuclease, a nickase, and a guide NA targeting the opposite side of the region of interest, followed by ligation as above using a second single-stranded adapter. Once two adapters are ligated to either side of the region of interest, the region can be amplified or directly sequenced.

Claims

1. 1. A method for capturing a target nucleic acid sequence, comprising: (a) providing a sample comprising a plurality of adaptor-ligated nucleic acids, wherein the nucleic acids are ligated at one end to a first adaptor and at the other end to a second adaptor; (b) contacting the sample with a plurality of nucleic acid-guided nuclease-gNA complexes, thereby generating a plurality of nucleic acid fragments ligated at one end to a first or second adaptor and lacking an adaptor at the other end, wherein the gNAs are complementary to target sites of interest contained in the subset of nucleic acids, and wherein contacting with the plurality of nucleic acid-guided nuclease-gNA complexes cleaves the target sites of interest contained in the subset of nucleic acids; and (c) contacting the plurality of nucleic acid fragments with a third adaptor, thereby generating a plurality of nucleic acid fragments ligated at one end to the first or second adaptor and at the other end to the third adaptor. wherein the ratio of uncaptured nucleic acid to target nucleic acid sequence is between 99.999:0.001 and 99:

1.

2. 2. The method of claim 1, wherein the nucleic acid-guided nuclease is a CRISPR / Cas system protein.

3. 2. The method of claim 1, wherein the nucleic acid-guided nuclease is selected from the group consisting of CAS class I type I, CAS class I type III, CAS class I type IV, CAS class II type II, and CAS class II type V.

4. 2. The method of claim 1, wherein the nucleic acid-guided nuclease is selected from the group consisting of Cas9, Cpf1, Cas3, Cas8a-c, Cas10, Cse1, Csy1, Csn2, Cas4, Csm2, Cmr5, Csf1, C2c2, and NgAgo.

5. The method of any one of claims 1 to 4, wherein the gNA is a gRNA.

6. 6. The method of claim 1, further comprising amplifying the product of step (c) using PCR specific for the first or second and third adapters.

7. The method of claim 6, wherein the amplified product is used for cloning, sequencing or genotyping.

8. The method of claim 1 , wherein the nucleic acid is double-stranded DNA.

9. 9. The method of claim 1, wherein the nucleic acid is from genomic DNA.

10. 10. The method of claim 1, wherein the nucleic acid ligated to the adaptor is between 20 bp and 5000 bp in length.

11. 11. The method of any one of claims 1 to 10, wherein the target site of interest is a single nucleotide polymorphism (SNP), a short tandem repeat (STR), an oncogene, an insertion, a deletion, a structural variant, an exon, a gene mutation, or a regulatory region.

12. 12. The method of any one of claims 1 to 11, wherein the adapter is 20 bp to 100 bp in length.

13. 13. The method of any one of claims 1 to 12, wherein the targeted site of interest accounts for less than 50% of the total nucleic acid in the sample.

14. 14. The method of any one of claims 1 to 13, wherein the sample is obtained from a biological sample, a clinical sample, a forensic sample or an environmental sample.

15. 15. The method of any one of claims 1 to 14, wherein the first and second adapters are identical.

16. 15. The method of any one of claims 1 to 14, wherein the first and second adapters are different.

Citation Information

Patent Citations

  • Method for fragmenting genomic DNA using cas9

    US20140357523A1

  • In vitro DNA immortalization and whole genome amplification using libraries generated from randomly fragmented DNA

    WO2004081183A2