Methods for nucleic acid enrichment using site-specific nucleases and subsequent capture
Patent Information
- Application Number
- CN202211226016.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-12
- Filing Date
- 2019-12-12
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2039-12-12
AI Technical Summary
然而,根据AT:GC比率和被扩增片段的二级结构,扩增会产生偏差,并且随着扩增片段长度的增加,效率会降低
Smart Images

Figure HDA0003879853470000011 
Figure HDA0003879853470000012 
Figure HDA0003879853470000021
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application 2019800921549, filed on December 12, 2019, entitled “Method for enriching nucleic acids using site-specific nucleases and subsequent capture”. Technical Field
[0002] This invention relates to methods for the isolation and enrichment of nucleic acids. In fact, the isolation and enrichment of target nucleic acids represents a crucial first step in nucleic acid research, affecting both the quantity and quality of nucleic acids, and consequently directly impacting the quality of data obtained in downstream applications (e.g., sensitivity, coverage, robustness, and reproducibility). This is particularly important in applications analyzing only certain target nucleic acids from more complex mixtures, or in the presence of small amounts of target nucleic acids. For example, the human exome (protein-coding regions) comprises only about 1% of the total genome, yet contains 85% of DNA variations known to be associated with genetic diseases. Therefore, isolation and enrichment are particularly important in exome-related clinical applications, such as in diagnosis and genetic risk assessment. Background Technology
[0003] While whole-genome analysis can be used even when there are few nucleic acid targets of interest, sequencing the entire genome is generally not feasible due to technical, economic, and / or time constraints. Furthermore, whole-genome sequencing requires significantly increased computing power and storage space to analyze the massive amounts of data generated. Therefore, nucleic acid isolation is necessary to limit analysis to specific subsets of nucleic acid molecules.
[0004] To date, the main methods for isolating specific subsets of nucleic acid fragments are based on hybridization capture and / or targeted amplification techniques (see, for example, Mertes et al., Brief Funct Genomics, 2011, 10(6): 374-86 and WO 2016 / 014409). However, current hybridization capture methods are inefficient in enrichment, with 15-25% off-target capture (Garcia-Garcia, Sci Rep., 2016, 6: 20948), and typically require at least two rounds of selection. Nucleic acids are also denatured before capture, thereby removing any information encoded by the complementarity of the two strands or by the complementary strands themselves. When using hybridization capture, amplification is usually also required to increase the amount of nucleic acid material. However, amplification can be biased depending on the AT:GC ratio and the secondary structure of the amplified fragment, and efficiency decreases with increasing fragment length. Furthermore, the number of target regions that can be amplified in multiple ways is limited due to primer cross-reactions. In addition, all chemical modifications present in the original sequence (e.g., base modifications) are lost during amplification. Finally, amplification may introduce artifacts (i.e. unwanted or non-specific nucleic acid sequences) or errors into the nucleic acid.
[0005] Given these limitations, there is a need for new methods to isolate and / or enrich target nucleic acids, particularly methods that preserve the original features of the nucleic acid molecules of interest (e.g., chemical modifications, such as base modifications, and nucleic acid sequence information, such as mismatches or SNPs), do not require multiple rounds of selection, and are compatible with downstream analytical techniques (e.g., nucleic acid sequencing). Summary of the Invention
[0006] Before describing the invention in detail, it should be understood that the invention is not limited to the aspects of the specific examples and, of course, variations are possible. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting.
[0007] All publications, patents, and patent applications cited herein, both above and below, are incorporated herein by reference in their entirety. Furthermore, unless otherwise indicated, the practice of this invention utilizes conventional techniques of protein chemistry, molecular biology, microbiology, recombinant DNA technology, and pharmacology, techniques well within the scope of the art. These techniques are fully explained in the literature. See, for example, Ausubel et al., Current Protocols in Molecular Biology, Eds., John Wiley & Sons, Inc., New York, 1995, Remington's Pharmaceutical Sciences, 17th edition, Mack Publishing Co., Easton, Pa., 1985, and Sambrook et al., Molecular cloning: A laboratory manual, 2nd edition, Cold Spring Harbor Laboratory Press-Cold Spring Harbor, NY, USA, 1989.
[0008] In the following claims and the foregoing description, the words “comprising,” “including,” “containing,” and other variations are used in the sense of inclusion, that is, specifying the presence of the stated feature but not excluding the presence or addition of further features in various embodiments of the invention, unless the context otherwise requires due to the language of expression or necessary implication. Furthermore, unless expressly stated otherwise in the content of this application, the terms “a,” “an,” and “the” as used herein include the plural forms. For example, “target region” therefore also includes two or more target regions.
[0009] In a first aspect, the present invention relates to a novel method for isolating nucleic acid target regions from a group of nucleic acid molecules, comprising contacting the group with a class 2 type V Cas protein-gRNA complex, followed by contacting with an enzyme having single-stranded 3' to 5' exonuclease activity. Indeed, the inventors have surprisingly discovered that these steps produce nucleic acid molecules with a 5' single-stranded overhang of at least 9 nucleotides in length. Consequently, the target region can then be specifically enriched by a capture method, such as hybridization capture, by hybridizing oligonucleotides to the overhang. Advantageously, only the target region of interest is isolated because the class 2 type V Cas protein-gRNA complex targets and cleaves highly specific sites, which may occur only once throughout the genome. In stark contrast, restriction enzymes recognize and cleave shorter sites, which therefore appear multiple times in a given sequence, further producing relatively short overhangs (e.g., 3 bases).
[0010] Because the original nucleic acid molecule remains intact throughout all steps of the method of the present invention, this method is highly advantageous over current methods, as it preserves all characteristics of the target nucleic acid (e.g., chemical modifications, mismatches). Unlike existing hybridization capture methods, nucleic acids isolated by the method of the present invention do not require denaturation. Bias is also reduced because no amplification step is required in the method of the present invention. Furthermore, multiplex detection can be easily designed without the risk of primer interactions or cross-recognition. Small sample sizes and samples with low levels of target nucleic acids can also be used in the method of the present invention without target amplification, as nucleic acid target separation is highly efficient and has good specificity. Advantageously, separation is sufficient when only a single round is performed. Moreover, this method is simpler and less prone to error compared to prior art methods, as all steps can be performed in the same container. Sample loss is further reduced due to the absence of material transfer between containers. Finally, the method of the present invention is superior to existing methods because it is rapid and inexpensive, can be performed directly on the sample, involves few processing steps, and is compatible with existing downstream nucleic acid analysis platforms, including “third-generation” sequencing technologies, where individual nucleic acid molecules are analyzed in microstructures such as nanopores, zero-mode waveguides, or micropores. It is worth noting that the method of the present invention provides isolated specific nucleic acid target regions that may contain specific single-stranded nucleic acid protrusions at either end or both ends, to which various aptamers or adapters may be specifically attached, thereby providing flexibility for the use of the target regions in a variety of downstream analyses and applications.
[0011] More specifically, the method for isolating nucleic acid target regions from a population of nucleic acid molecules includes the following steps:
[0012] a) Contact the nucleic acid molecule group with a type V Cas protein-gRNA complex, wherein the gRNA contains a guide segment complementary to a first site adjacent to the target region, thereby forming a type V Cas protein-gRNA nucleic acid complex.
[0013] b) Contact the nucleic acid molecule group containing the type 2 V Cas protein-gRNA nucleic acid complex with at least one enzyme having single-stranded 3' to 5' exonuclease activity, thereby forming a 5' single-stranded overhang at the first site.
[0014] c) Remove the type V Cas protein-gRNA complex from the group in step b).
[0015] d) Contact the group from step c) with an oligonucleotide probe, the probe containing a sequence at least partially complementary to the overhang, thereby forming a double strand between the probe and the overhang.
[0016] e) Isolate the double strand from the nucleic acid molecule population of step d) to isolate the nucleic acid target region.
[0017] In the above method, the steps are performed in the provided order: first step a), then step b), then step c), and then step d). Alternatively, especially when the target nucleic acid region or molecule is double-stranded, steps a and b) can be performed simultaneously.
[0018] In some cases, the method may include additional steps. As a non-limiting example, at any stage of the method described above, such as before step a), simultaneously with step a), between steps a) and b), or simultaneously with steps b), c), or e), an additional step of fragmenting one or more nucleic acid molecules to obtain a population of nucleic acid molecules may be included. As a non-limiting example, an additional incubation step may be further included before, during, or after any step of the method described above. As a non-limiting example, a storage step may be further included after step c), after step d), or after step e). These optional additional steps will be further detailed below.
[0019] As used herein, the term "contact" refers to placing two or more molecules and / or products in the same solution such that the molecules and / or products can interact with each other. For example, contacting a group of nucleic acid molecules with a type V Cas protein-gRNA complex allows these molecules to interact and form a complex in which the type V Cas protein-gRNA complex has already bound to the nucleic acid molecule at a specific site. Similarly, contacting a group of nucleic acid molecules with an oligonucleotide probe will result in at least a single-stranded region of the probe hybridizing with at least a partially complementary single-stranded region contained within the nucleic acid molecule group. In step d) of the method provided herein, this more specifically corresponds to the hybridization of the single-stranded region of the probe with at least a partially complementary 5' overhang. As a further example, when the molecule or product is or contains a substrate of the enzyme, contacting the molecule or product with an enzyme such as an exonuclease will result in an enzymatic reaction. For example, contacting a group of nucleic acid molecules with an enzyme having single-stranded 3' to 5' exonuclease activity will result in the interaction of these molecules and the degradation of enzyme substrates (e.g., single-stranded nucleic acid molecules or regions with a 3' free end) accessible to the enzyme.
[0020] As used herein, the term "separation" refers to an increase in the ratio of one or more nucleic acid target regions in a sample relative to one or more other regions or molecules. As a non-limiting example, these other molecules may include proteins, lipids, carbohydrates, metabolites, nucleic acids, or combinations thereof. As a non-limiting example, these other regions may correspond to nucleic acid regions that are present on the same molecules as the target region but are not included in the target region (i.e., "non-target regions"). As used herein, "separation" of a target nucleic acid region may more specifically refer to an increase in the ratio of one or more target nucleic acid regions in a sample by at least 2-fold (e.g., 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 250, 500, 750, 1000, or 10,000 or more times) compared to one or more other molecules in the sample, or compared to the total number of molecules in the initial sample (i.e., prior to performing the method for separating the target region of the present invention). Separation of target nucleic acid regions can also refer to an increase in the proportion of target nucleic acid regions in a sample by at least 5% (e.g., 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) compared to the levels of one or more other molecules in the sample. When the proportion of target nucleic acid regions is 100%, the sample contains no other molecules. As used herein, the term "enrichment" more specifically refers to the separation of one or more target nucleic acid regions relative to other nucleic acid molecules in the sample. For example, target region enrichment refers to an increase in the ratio of isolated target regions compared to the initial total nucleic acid amount, wherein the increase in the ratio of isolated target regions is at least 10%, 20%, 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. According to a preferred embodiment, the increase in the ratio of isolated target regions compared to the initial total nucleic acid amount is at least 10%, more preferably at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, and even more preferably at least 99% or 100%.
[0021] According to one implementation, the isolated nucleic acid target region is enriched by at least 10-fold, at least 20-fold, at least 50-fold, at least 100-fold, at least 250-fold, at least 500-fold, at least 750-fold, preferably at least 1000-fold, at least 10,000-fold, at least 100,000-fold, and even more preferably at least 1,000,000-fold, at least 2,000,000-fold, or at least 3,000,000-fold. As a specific example, 100% enrichment of a single 1kb fragment from a population of nucleic acid molecules equivalent to approximately 3.2 billion bp of the human genome represents a 3,000,000-fold increase.
[0022] According to an alternative embodiment, the isolated target region is substantially pure. "Substantially pure" means that after isolating the target region according to the method of the invention, the isolated target region contains at least 99%, preferably at least 99.5%, of the total nucleic acids in the sample.
[0023] According to a preferred embodiment, the target region comprises less than 10%, preferably less than 5%, more preferably less than 2%, less than 0.05%, less than 0.02%, and even more preferably less than 0.01%, less than 0.005%, less than 0.001%, less than 0.0005%, less than 0.0001%, less than 0.00005%, less than 0.0001%, less than 0.000005%, less than 0.00001%, or less than 0.0000005% of the total nucleic acids in the sample. Those skilled in the art will recognize that the amount or percentage of the target region in the total nucleic acids of the sample will vary depending on the number and length of the target regions to be separated. As a non-limiting example, in a human genome of approximately 3.2 billion bp, a 1 kb target region of interest represents less than 0.0000005% of the total genome.
[0024] Nucleic acid target regions are isolated from a group of nucleic acid molecules, typically contained in a sample. As used herein, the term "sample" refers to any material or substance containing a group of nucleic acid molecules, including, for example, biological, environmental, or synthetic samples. A "biological sample" can be any sample that may contain a biological organism, such as bacteria, viruses, archaea, animals, plants, and / or fungi. According to the invention, a "biological sample" also refers to a sample that can be obtained from a biological organism, such as cell extracts obtained from, for example, bacteria, viruses, archaea, plants, fungi, animals, and / or other eukaryotes. Nucleic acid molecules of interest can be obtained directly from an organism or from a biological sample obtained from an organism, such as from blood, urine, cerebrospinal fluid, semen, saliva, sputum, feces, and tissues (e.g., cell tissues or plant tissues). In the context of the invention, any cell, tissue, or body fluid can be used as a source of nucleic acids. Nucleic acid molecules can also be recovered or purified from cultured cells, for example from primary cell cultures or cell lines. Cells or tissues from which nucleic acids of interest can be obtained by infecting with viruses or other intracellular pathogens. A sample can also be total nucleic acids extracted from a biological sample. "Environmental samples" can be any sample containing nucleic acids not directly derived from a biological organism (e.g., soil, seawater, air, etc.), and may contain nucleic acids no longer present in the biological organism. "Synthetic samples" include artificially created or engineered nucleic acids. Alternatively, samples can originate from any source suspected of containing the target nucleic acid region.
[0025] In some embodiments, the method of the present invention may include one or more steps of processing a sample to facilitate the isolation of nucleic acids containing the target region according to the method of the present invention. As a non-limiting example, the sample may be concentrated, diluted, or destroyed (e.g., by mechanical or enzymatic lysis). Prior to step a) of the method of the present invention, the nucleic acid may be fully or partially purified, or may be in an unpurified form.
[0026] As used herein, the terms “nucleic acid,” “nucleic acid region,” and “nucleic acid molecule” refer to polymers of nucleotide monomers, including deoxyribonucleotides (DNA), ribonucleotides (RNA), or analogs thereof, and combinations thereof (e.g., DNA / RNA chimeras). Deoxyribonucleotide and ribonucleotide monomers as used herein refer to monomer units containing a triphosphate group, an adenine (“A”), cytosine (“C”), guanine (“G”), thymine (“T”), or uracil (“U”) nitrogenous base, and respectively containing deoxyribose or ribose. Modified nucleotide bases are also included herein, wherein the nucleotide base is, for example, hypoxanthine, xanthine, 7-methylguanine, inosine, xanthine nucleoside, 7-methylguanosine, 5,6-dihydrouracil, 5-methylcytosine, pseudouridine, dihydrouridine, or 5-methylcytidine. In the context of this invention, when describing nucleotides, “N” represents any nucleotide, “Y” represents any pyrimidine, and “R” represents any purine. Nucleotide monomers are linked by internucleotide bonds, such as phosphodiester bonds or their phosphate analogs, and associated counterions (e.g., H+). + NH4 + Na + The nucleic acid molecules of this invention can be double-stranded or single-stranded, and most often double-stranded DNA. However, it should be understood that the invention is also applicable to perfectly or imperfectly paired single-stranded DNA-single-stranded DNA duplexes, or alternatively to perfectly or imperfectly paired single-stranded DNA-single-stranded RNA duplexes, or alternatively to perfectly or imperfectly paired single-stranded RNA-single-stranded RNA duplexes, as well as single-stranded DNA and single-stranded RNA. In particular, the invention is applicable to the secondary structures of single-stranded DNA or single-stranded RNA. When the nucleic acid molecule is single-stranded RNA (e.g., mRNA) or a single-stranded RNA-single-stranded RNA duplex (e.g., viral dsRNA), the RNA can be reverse transcribed before contact with a class 2 type V Cas protein-gRNA complex. The duplex can be composed of at least partially re-paired single nucleic acid strands obtained from samples from different sources. The nucleic acid molecule can be naturally occurring (e.g., of eukaryotic or prokaryotic origin) or synthetic. Nucleic acid molecules may include circular nucleic acid molecules, such as covalently closed circular DNA and / or circular RNA, including plasmids and / or circular chromosomes, or linear nucleic acid molecules. Nucleic acid molecules may specifically include genomic DNA (gDNA), cDNA, hnRNA, mRNA, rRNA, tRNA, microRNA, mtDNA, cpDNA, cfDNA (e.g., ctDNA or cffDNA), cfRNA, etc.
[0027] Nucleic acids can range in length from just a few monomeric units (e.g., oligonucleotides, which can range in length from, for example, about 15 to about 200 monomeric units) to thousands, tens of thousands, hundreds of thousands, or millions of monomeric units. Preferably, the nucleic acid molecule comprises one or more cfDNA molecules. In a first aspect, the length of the nucleic acid molecule is less than 300 bp, for example, including those between about 125 and 225 bp, preferably between 130 and 200 bp. In a second aspect, the length of the nucleic acid molecule is equal to or greater than 300 bp. In this application, unless otherwise stated, it should be understood that the nucleic acid molecule is expressed in a 5' to 3' direction from left to right.
[0028] As used herein, the term "nucleic acid molecular group" refers to more than one type of nucleic acid molecule. The group can comprise one or more distinct nucleic acid molecules of any length and sequence as defined above. Specifically, a nucleic acid molecular group can comprise more than 10 3 10 4 10 5 10 6 10 7 10 8 10 9 Or 10 10 Different nucleic acid molecules.
[0029] As used herein, “nucleic acid target region,” “target nucleic acid region,” or “target region” refers to a specific nucleic acid molecule present in a more complex sample or group of nucleic acid molecules, or a specific nucleic acid region present in a larger nucleic acid molecule, which will be specifically targeted for isolation or enrichment. As used herein, the term “region” refers to a continuous nucleotide polymer of any length. When the target region is present within a larger nucleic acid molecule, it is preferably side-mounted to a first site on its first side, the first site being at least partially complementary to a crRNA molecule or a guide segment of gRNA contained in a class 2 type V Cas protein-gRNA complex. In some cases, the nucleic acid target region is further side-mounted to a second site on its second side. Thus, the first site and the second site are located on either side of the target region. The first site (and the second site, when present) are located adjacent to, preferably immediately adjacent to, the target region. The first site and the target region (and the second site, when present) may further be side-mounted to non-target regions on one or both sides. As used herein, the term “adjacent” means the presence of a first nucleotide or nucleic acid region and a second nucleotide or nucleic acid region, wherein the two nucleotides and / or regions are present on the same continuous nucleotide polymer. Therefore, nucleotides are considered to be adjacent as long as at least partially complementary sites and target regions to the guide region of the crRNA molecule or gRNA contained in the type V Cas protein-gRNA complex exist on the same nucleotide polymer. As used herein, the term “closely adjacent” means that the first nucleotide or nucleic acid region is close to the second nucleotide or nucleic acid region, wherein the two nucleotides and / or regions are directly adjacent in the nucleotide polymer (i.e., there is no intermediate nucleotide present).
[0030] In the context of this invention, a "site" corresponds to a non-discontinuous nucleotide polymer of no more than 100 nucleotides in length, preferably no more than 60 nucleotides in length. The site is preferably a double-stranded nucleotide polymer. Preferably, the "site" is at least partially complementary to the guide region of the gRNA, preferably completely complementary to the guide region of the gRNA. Preferably, the site is about 12 to about 35 nucleotides in length, preferably 15 to 35 nucleotides in length. Preferably, the site comprises a sequence at least partially complementary to the guide region of the gRNA, more preferably a sequence completely complementary to the guide region of the gRNA, and a PAM. The PAM is preferably adjacent to the target region. The PAM is preferably located on a nucleic acid strand that does not hybridize with the type V Cas protein-gRNA complex (i.e., the "non-target" strand).
[0031] The second site is preferably at least partially complementary to the guide region of a crRNA molecule or gRNA contained in a type 2 Cas protein-gRNA complex (preferably a type 2 type V Cas protein-gRNA complex). Alternatively, the second site may contain or consist of restriction sites. In this case, the length of the second site is preferably about 4 to 8 nucleotides.
[0032] In some embodiments, two or more distinct nucleic acid target regions can be isolated. The “target nucleic acid region” of this invention can therefore comprise one or more distinct regions, preferably at least 2, 5, 10, 25, 50, 100, or more regions. The nucleic acid target region can be coding or non-coding, or a combination of both. The target region can be genomic or free-form. The target region can contain one or more repetitive regions, rearrangements, duplications, translocations, deletions, mismatches, SNPs, and / or modified bases, such as epigenetic modifications. In some cases, the nucleic acid target regions can be identical (e.g., corresponding to repetitive sequences). In other cases, the nucleic acid target regions can be different. Preferably, the target nucleic acid region will have a length of at least about 10, 20, 50, 100, 250, 500, 1,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, or 100,000 nucleotides. While a given gRNA may allow the separation of multiple nucleic acids containing the target region (e.g., due to nonspecific binding, or recognition of sites present more than once in a nucleic acid molecule), in the context of this invention, each gRNA preferably recognizes a single site within a group of nucleic acid molecules. In some cases, two or more class 2 type V Cas protein-gRNA complexes will bind at different sites adjacent to different target regions, thereby enabling the separation of two or more regions. Preferably, when the target regions are present on the same nucleic acid molecule, the nucleic acid target regions are separated from each other by at least 100, 200, 300, 500, 750, 1000, 2000, 5000 or 10000 nucleotides.
[0033] In some cases, two type 2 V Cas protein-gRNA complexes will bind to sites located flanking the target region, thereby separating a single target region located between the two sites. In other cases, when separating multiple target regions, the above two scenarios can be used in tandem (i.e., some target regions are adjacent to a single site, while other target regions are adjacent to two sites located flanking the target region). The number of type 2 V Cas protein-gRNA complexes binding to sites adjacent to a given target region can be selected based on the length of the nucleic acid molecule containing the target region, the desired structure of the separated nucleic acid molecule (e.g., the presence of single-stranded overhangs at one or two sites adjacent to the target region), and / or the downstream application of the target region.
[0034] According to a preferred embodiment, at least two target regions are separated, more preferably at least 5, at least 10, at least 25, at least 50, or at least 100 target regions are separated. Preferably, the nucleic acid molecule is contacted with at least two type V Cas protein-gRNA complexes, each complex containing a different gRNA. More preferably, the nucleic acid molecule is contacted with at least 5, at least 10, at least 25, at least 50, or at least 100 type V Cas protein-gRNA complexes, each type V Cas protein-gRNA complex capable of separating a different target region.
[0035] Cas protein
[0036] As used herein, the term "Cas protein" refers to an RNA-guided endonuclease that specifically recognizes and binds to sites within nucleic acid molecules, and in this case, specifically to sites adjacent to the target region. To recognize and bind to a specific site, the Cas protein complexes with a "guide RNA" or "gRNA" to form a "Cas protein-gRNA complex." The binding specificity of the Cas protein is determined by the gRNA, which contains a "guide region" whose sequence must be at least partially complementary to the sequence of the specific site in the nucleic acid molecule. The guide region in the Cas protein-gRNA complex hybridizes with said site, thereby forming a Cas protein-gRNA-nucleic acid complex. Successful binding of the Cas protein-gRNA complex to the site also requires the presence of a short, conserved sequence in the nucleic acid molecule located immediately adjacent to the hybridization region. This sequence is called the protospacer-associated motif, or "PAM." Therefore, the binding of the Cas protein-gRNA complex to a specific site within a nucleic acid molecule involves both hybridization of the guide region with the nucleic acid at that site and the interaction between the Cas protein itself and the PAM. After the Cas protein-gRNA complex binds to a site within the nucleic acid, the Cas protein typically cleaves the nucleic acid by breaking the phosphodiester bonds between two adjacent nucleotides in each strand of the double-stranded nucleic acid molecule. Specifically, one domain of the Cas protein cleaves the nucleic acid strand hybridized with the gRNA, while the second domain of the Cas protein cleaves the unhybridized nucleic acid strand. The cleavage of the two strands of the double-stranded molecule may be staggered, resulting in single-stranded overhangs or blunt ends.
[0037] To date, three main classes of Cas proteins (classes 1, 2, and 3) have been described. Within class 2 Cas proteins, at least five distinct types have been identified to date (i.e., types I, II, III, IV, and V). As a non-limiting example, Cas proteins may be selected from class 2 Cas proteins, particularly type V and type II Cas proteins, and more specifically from Cas9, Cas12a (also known as Cpf1), C2c1, C2c3, and C2c2 (Cas13a) proteins.
[0038] As a non-limiting example, class 2 Cas proteins can originate from one of the following species: *Streptococcus pneumoniae*, *Streptococcus pyogenes*, *Streptococcus thermophilus*, *Streptococcus canis*, *Staphylococcus aureus*, *Neisseria meningitidis*, *Treponema denticola*, *Francisella tularensis*, *Francisella novicida*, *Pasteurella multocida*, *Streptococcus mutans*, *Campylobacter jejuni*, *Campylobacter lari*, or *Mycoplasma gallisepticum*. The following bacteria are listed: gallisepticum, Nitratifractorsalsuginis, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria cinerea, Gluconacetobacter diazotrophicus, Azospirillum, Sphaerochaetaglobosa, Flavobacterium columnare, Fluviicola taffensis, Bacteroides coprophilus, Mycoplasma mobile, Lactobacillus farciminis, Streptococcus pasteurianus, Lactobacillus johnsonii, Staphylococcus pasteuri, Filifactoralocis, Veillonella. sp.), Suterella wadsworthensis, Leptotrichia sp.Corynebacterium diphtheriae, Acidaminococcus sp., or Lachnospiraceae sp., Prevotella albensis, Eubacterium eligens, Butyrivibrio fibrisolvens, Smithella sp., Flavobacterium sp., Porphyromonas crevioricanis, or Lachnospiraceae bacterium ND2006.
[0039] As a non-limiting example, class 2 Cas proteins could be J3F2B0, Q0P897, Q6NKI3, A0Q5Y3, Q927P4, A1IQ68, C9X1G5, Q9CLT2, J7RUA5, Q8DTE3, Q99ZW2, G3ECR1, Q73QW6, G1UFN3, Q7NAI2, E6WZS9, A7HP89, D4KTZ0, D0W2Z9 B5ZLK9, F0RSV0, A0A1L6XN42, F2IKJ5, S0FEG1, Q6KIQ7, A0A0H4LAU6, F5X275, F4AF10, U5ULJ7, D6GRK4, D6KPM9, U2SSY7, G4Q6A5, R9MHT9, A0A111NJ61, D3NT09, G4Q6A5, A0Q7Q2, or U2UMQ6. Accession number from UniProt (www.uniprot.org), last modified on January 10, 2017. As a non-limiting example, a gene encoding class 2 Cas proteins can be any gene containing a nucleotide sequence that produces the amino acid sequence of the corresponding Cas protein (e.g., one of the Cas proteins listed above). Those skilled in the art will readily understand that the nucleotide sequence of a gene can vary due to the degeneracy of the genetic code without altering the amino acid sequence. Class 2 Cas proteins can also be codon-optimized for expression in bacterial (e.g., Escherichia coli), insect, fungal, or mammalian cells.
[0040] Two classes of Cas proteins and their orthologs have also been identified in other bacterial species, and are specifically described in Example 1 of PCT application WO2015 / 071474, which is incorporated herein by reference. In some cases, the Cas protein may be a homolog or ortholog of one of the two classes of Cas proteins from one of the species listed above.
[0041] Variants and mutants of wild-type Cas proteins have been described. As a non-limiting example, Cas variants that retain endonuclease activity but have improved binding specificity have been described (e.g., class 2 Cas proteins eSpCas9, as described in Slaymaker et al., Science, 2015, 351(6268): 84-86).
[0042] Type 2 V Cas proteins
[0043] The Cas protein used in the context of the present invention’s method for isolating nucleic acid target regions is a class 2 V Cas protein. Indeed, as described above, the first step (step a) of the method for isolating the target region comprises contacting a group of nucleic acid molecules with a class 2 V Cas protein. Specifically, step a) comprises contacting the group of nucleic acid molecules with a class 2 V Cas protein-gRNA complex, wherein the gRNA contains a guide segment complementary to at least the first site adjacent to the target region, thereby forming a class 2 V Cas protein-gRNA nucleic acid complex. When catalytically active, class 2 V proteins complexed with appropriate gRNA typically produce staggered cleavages (e.g., short 5' overhangs of 4 to 6 nucleotides) located distal to the PAM sequence of the double-stranded nucleic acid molecule (e.g., at least 10 nucleotides from the PAM) (see, for example, Zetsche et al., Cell, 2015, 163(3): 759-771). It has been further observed that the class 2 V protein-gRNA complex retains its binding to the nucleic acid molecule after cleavage. Furthermore, the inventors have surprisingly demonstrated herein that when nucleic acid molecules bound to class 2 type V protein-gRNA complexes are contacted with enzymes having 3' to 5' single-stranded exonuclease activity, unexpectedly, 5' overhangs (e.g., at least 9 nucleotides in length, preferably at least 12 nucleotides in length) are generated. Without being limited by theory, the class 2 type V protein-gRNA complex can remain bound to the target nucleic acid strand (excluding the PAM sequence) after cleavage, while the reverse strand (i.e., the strand containing the PAM sequence) dissociates from the complex and can be used for digestion with 3' to 5' single-stranded exonucleases. The inventors have also surprisingly discovered that when the class 2 type V protein-gRNA complex is cleaved in the presence of 3' to 5' single-stranded exonucleases, the variability of the cleavage site is reduced (see, for example, Figure 1B , Figure 1C ).
[0044] The two types of V-type Cas proteins of the present invention possess catalytic activity (i.e., cleavage of both strands of a double-stranded molecule). Preferably, the two types of V-type Cas proteins of the present invention are selected from Cas12a and C2c1, more preferably Cas12a. The Cas12a protein is preferably a Cas12a protein from one of the suitable species listed above, and more preferably a Cas12a protein from one of the following species or strains: *F. novicida* U112 (accession number: AJI61006.1), *P. albensis* (accession number: WP_024988992.1), *Acidaminococcus sp.* BV3L6 (accession number: WP_021736722.1), *E. eligens* (accession number: WP_012739647.1), *B. fibrisolvens* (accession number: WP_027216152.1), *Smithella sp. SCADC* (accession number: KFO67988), and *Flavobacterium*. sp.)316 (accession number: WP_045971446.1), P. crevioricanis (accession number: WP_036890108.1), Bacteroidetes oral taxa 274 (accession number: WP_009217842.1), or Lachnospiraceae bacterium ND2006 (accession number: WP_051666128.1). In a preferred embodiment, the amino acid cocci BV3L6Cas12a is a variant containing the following amino acid substitutions: S542R / K607R or S542R / K548V / N552R (Gao et al., Nat Biotechnol. 2017, 35(8): 789-792). The Cas12a protein is even more preferably selected from the genus *AsCas12a*, the family *Dendrocalamus* ND2006 Cas12a (also known as "LbaCas12a"), and *Francis neoformans* U112 Cas12a (also known as "FnCas12a").
[0045] Guide RNA
[0046] As used herein, the terms "guide RNA" or "gRNA" generally refer to a crRNA molecule. This is especially true when the gRNA is a class 2 type V gRNA. However, in some cases, the term gRNA can refer to two guide RNA molecules, consisting of a crRNA molecule and a tracrRNA molecule, or the term gRNA can refer to a single guide RNA molecule or sgRNA, which comprises crRNA and tracrRNA sequence fragments, for example when the gRNA is a class 2 type II gRNA. In some cases, the crRNA may contain a tracr-mate region. The characteristics of tracr-mate regions and tracrRNAs (e.g., length, presence of secondary structures, etc.) are well known to those skilled in the art. Furthermore, those skilled in the art know when to include such regions or molecules in a gRNA molecule. In particular, such fragments or molecules do not need to be included in class 2 type V gRNAs.
[0047] gRNA molecules can be chemically modified, for example, by modifying the bases, sugars, or phosphates of one or more ribonucleotides. Optionally, the 5' and / or 3' ends of the gRNA molecule can be modified, for example, by covalently attaching to another molecule or chemical group.
[0048] The crRNA molecule or segment is preferably 20 to 75 nucleotides in length, more preferably 30 to 60 nucleotides, and even more preferably 40 to 45 nucleotides. The crRNA molecule or segment preferably includes a first region, referred to herein as a "guide region," whose sequence is at least partially complementary to a sequence present in a nucleic acid molecule, preferably a sequence located at a first site adjacent to the target region. Preferably, the guide region of the gRNA of the present invention has at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more preferably 100% sequence complementarity with the sequence present in the nucleic acid molecule. Preferably, when complementarity is less than 100%, the mismatch is located near the crRNA end furthest from the hybridized PAM. For example, when the type V Cas protein is Cas12a, the mismatch is preferably contained at the 3' end of the crRNA molecule or segment (e.g., within the last 7 nucleotides), because Cas12a recognizes PAM at the 5' end of the crRNA. The length of the guide segment is preferably at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides, more preferably 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides, and even more preferably 17, 18, 19, 20, 21, 22, 23 or 24 nucleotides. Alternatively, the length of the guide segment is preferably 10 to 30, more preferably 15 to 25, and even more preferably 17 to 24 nucleotides.
[0049] Preferably, when the type V Cas protein is Cas12a, the gRNA consists only of crRNA molecules. Therefore, the term "Cas12a-gRNA complex" can also be used interchangeably herein as "Cas12a-crRNA complex". When the gRNA is only a crRNA molecule, at least a guide segment must be present. An exemplary universal crRNA nucleotide sequence is shown in SEQ ID NO: 1, wherein the guide segment is represented by an extension of "N" nucleotides. Preferably, the crRNA molecule also contains secondary structures. As a non-limiting example, the "secondary structures" present in the gRNA can be stem-loops or hairpins, protrusions, tetraloops, and / or pseudoknots. The terms "hairpin" and "stem-loop" are used interchangeably herein in the context of gRNA and are defined as follows (see the "hairpin aptamer" section). According to a preferred embodiment, the gRNA contains at least one hairpin secondary structure.
[0050] Preferably, the crRNA molecule does not contain a tracr-mate region. Preferably, the guide region is located at the 3' end of the crRNA molecule. Preferably, the secondary structure is located at or near the 5' end of the crRNA molecule. As used herein, the term "located at or near the 5' end of a nucleic acid molecule" refers to the location of a segment or structure within the first half of the molecule from the 5' to the 3' end. Similarly, as used herein, the term "located at or near the 3' end of a nucleic acid molecule" refers to the location of a segment or structure within the first half of the molecule. Preferably, the crRNA is 40 to 50 nucleotides in length.
[0051] As used herein, the term "complementarity" refers to the ability of one nucleic acid sequence or molecule (e.g., gRNA) to interact with another nucleic acid sequence or molecule (e.g., a sequence within a nucleic acid molecule) through sequence-specific antiparallel nucleotide base pairing, resulting in the formation of a double helix or other higher-order structure. The primary types of interactions are nucleotide base-specific, such as A:T, A:U, and G:C via Watson-Crick and Hoogsteen-type hydrogen bonds. This is also known as "nucleic acid binding," "hybridization," or "annealing." The conditions for hybridization of nucleic acids with complementary regions of a target site are well known in the art (see, for example, Nucleic Acid Hybridization, A Practical Approach, Hames and Higgins, eds., IRL Press, Washington, DC (1985)). Hybridization conditions depend on the specific application and can be routinely determined by those skilled in the art.
[0052] In the context of this invention, complementary binding does not mean that two nucleic acid sequences or molecules (e.g., gRNA and target region) must be perfectly complementary to each other. Furthermore, the crRNA sequence segment or molecule need not be perfectly complementary to the sequence in the nucleic acid molecule. In fact, two types of Cas protein-gRNA complexes are known to specifically bind nucleic acid sequences having as few as 8 or 9 bases complementary to the gRNA. Preferably, there is no mismatch between the 10 bases of the gRNA closest to the PAM and the corresponding 10 bases of the complementary nucleic acid sequence closest to the PAM; more preferably, there is no mismatch between the 6 bases of the gRNA closest to the PAM and the corresponding 6 bases of the complementary nucleic acid sequence closest to the PAM; and even more preferably, there is no mismatch between the bases of the gRNA located 4, 5, and / or 6 bases from the PAM and the corresponding bases of the complementary nucleic acid sequence located 4, 5, and / or 6 bases from the PAM. In fact, if a mismatch exists at one or more of these base positions, binding will be unstable, and cleavage of the two types of Cas protein-gRNA complexes at the target site will be reduced or even eliminated. As described above, off-target hybridization can be reduced by increasing the length of the crRNA segment or by placing a mismatch at the end of the crRNA segment furthest from the PAM. Alternatively, gRNAs can be modified to have increased binding specificity by the presence of one or more modifying bases or chemical modifications, such as those described in Cromwell et al., Nat Commun. 2018 Apr 13; 9(1): 1448 or Orden Rueda et al., Nat Commun. 2017; 8: 1610, which are incorporated herein by reference. Furthermore, nucleic acids can hybridize on one or more regions such that intermediate regions do not participate in hybridization events (e.g., loops or hairpin structures). Those skilled in the art can readily design one or more gRNA molecules based on their general knowledge and the parameters detailed above, according to the nucleic acid sequence to be hybridized.
[0053] It has been previously shown that the ratio of nucleic acid to Cas protein and gRNA molecules (nucleic acid:Cas protein:gRNA) affects the separation efficiency of the target region. In this paper, the ratio of nucleic acid to Cas protein to gRNA molecules containing the target region can be significantly optimized based on the source and / or complexity of the nucleic acid target region and / or the nucleic acid molecular population. Without being theoretically limited, optimization can be specifically performed based on the complexity of the DNA, where more complex nucleic acid populations require a greater number of Cas proteins and gRNAs. As a non-limiting example, less complex nucleic acid populations may substantially contain repetitive sequences or PCR-amplified fragments, while more complex nucleic acid populations may contain genomic DNA. As a non-limiting example, when the nucleic acid molecules are present in a nucleic acid molecular population generated by PCR, a ratio of at least 1:10:20 can be used in the method of the present invention. In contrast, when the population contains or consists of E. coli genomic DNA, a ratio of at least 1:1600:3200 is preferred, and when the population contains or consists of human genomic DNA, a ratio of at least 1:100000:200000 is preferred. When using multiple gRNAs (e.g., where two gRNAs recognize a first site and a second site located adjacent to and flanking the target region, or when preparing multiple nucleic acid molecules containing different target regions in a multipath stream), a single optimized nucleic acid:Cas protein:gRNA ratio can be selected for all gRNAs. Alternatively, an optimized ratio can be selected individually for each gRNA.
[0054] According to a preferred embodiment, the ratio is at least 1:10:10, more preferably at least 1:10:20, and even more preferably at least 1:10:50. A ratio of at least 1:10:20 is particularly preferred when generating template DNA via PCR. Preferably, if necessary (e.g., if the template is different), a guide RNA is selected for efficiency using the PCR template, and then the ratio of nucleic acid:Cas protein:gRNA containing the target region is optimized on a suitable template. Preferably, the cleavage efficiency of the wild-type Cas protein-gRNA complex is at least 70%, more preferably at least 80%, and even more preferably at least 90%. Preferably, the protection efficiency of the Cas protein-gRNA complex for the target region is at least 70%, more preferably at least 80%, and even more preferably at least 90%. When nucleic acids are prepared from bacterial sources, such as Gram-negative bacteria like *Escherichia coli*, preferably, the ratio of nucleic acid target region:Cas protein:gRNA is at least 1:200:400, more preferably at least 1:400:800, even more preferably at least 1:800:1600, at least 1:1600:3200, or at least 1:3200:6400. According to an alternative preferred embodiment, the ratio of nucleic acid target region:Cas protein:gRNA is at least 1:10,000:20,000, more preferably at least 1:100,000:200,000. Based on the ratios provided above, those skilled in the art can easily adjust the ratio of nucleic acid target region:Cas protein:gRNA according to the source and / or complexity of the target region and / or nucleic acid molecule to be prepared. Although the ratio of Cas protein to gRNA may vary, gRNA is advantageously provided in an amount at least twice the excess of Cas protein to ensure successful loading of gRNA onto the Cas protein. Of course, higher ratios of Cas protein (e.g., 1:20:40, 1:50:100, etc. for PCR targets) and optionally gRNA (e.g., 1:10:30, 1:10:40, etc. for PCR targets) can be used. The above ratios are preferably used in step a) of the method of the present invention, particularly with respect to the ratio of nucleic acid containing the target region: type V Cas protein: gRNA. The above ratios can also be used in the presence of a second type II Cas protein.
[0055] Proximity Affiliation Motion (PAM) in the original spacer region
[0056] As used herein, the term "protospacer adjacent motif" or "PAM" refers to a short nucleotide sequence (e.g., 2 to 6 nucleotides) that is directly recognized by the Cas protein itself (e.g., type 2 V Cas proteins). The PAM sequence and its position will vary depending on the Cas protein and can be readily determined by those skilled in the art based on their general knowledge or using techniques such as those described in Karvelis et al., Genome Biology, 2015, 16:253. As an example, the Cas12a protein of *Francis neoformans* recognizes PAM 5'-TTTN-3' or 5'-YTN-3', while the Cas12a protein of *Aminococcus* recognizes PAM 5'-TTTN-3'. As another example, the Cas9 protein of *Streptococcus pyogenes* recognizes PAM 5'-NGG-3'. In contrast, the Cas9 protein of *Staphylococcus aureus* recognizes PAM 5'-NNGRRT-3', while engineered Cas9 proteins derived from *Francis neoformans* recognize PAM 5'-YG-3'. PAM motifs are typically located on the non-hybridized (or “non-target”) strand of a double-stranded nucleic acid molecule, immediately adjacent to the 5' or 3' end of the nucleic acid site that hybridizes with the gRNA. The desired PAM location depends on the Cas protein used (e.g., when using Cas9, the PAM is preferably located immediately adjacent to the 3' end of the gRNA, while when using Cas12a, the PAM is preferably located immediately adjacent to the 5' end of the gRNA). In some cases, the PAM motif may be contained within the gRNA molecule itself or added to the sample as a separate DNA oligonucleotide. For example, when isolating single-stranded RNA molecules using this method, it may be necessary to add PAM to the sample using one of these methods. The binding of class 2 Cas proteins to PAM is thought to slightly disrupt the stability of double-stranded nucleic acids, thereby allowing the gRNA to hybridize with the nucleic acid sequence.
[0057] When the Cas protein is type 2V, the PAM is preferably located on the non-hybridized strand of the target region immediately adjacent to the 5' end of the gRNA. In contrast, for type 2II Cas proteins, the PAM is preferably located on the non-hybridized strand of the target region immediately adjacent to the 3' end of the gRNA. However, in some cases, the PAM is preferably contained within the gRNA molecule itself or on a DNA oligonucleotide.
[0058] Enzymes with single-stranded 3' to 5' exonuclease activity
[0059] According to the method described herein, after contacting a group of nucleic acid molecules with at least one type 2 V Cas protein-gRNA complex (e.g., in step a), the method further includes a step of contacting the group of nucleic acid molecules with at least one enzyme having single-stranded 3' to 5' exonuclease activity (e.g., in step b). This step degrades the single-stranded region of a single-stranded nucleic acid molecule or a double-stranded nucleic acid molecule from its 3' end. Those skilled in the art will understand that when the target nucleic acid is a double-stranded molecule, this step can be performed simultaneously with step a), because the enzyme is specific to the single-stranded region or molecule. However, in the case that the target nucleic acid is single-stranded, this step must be performed after step a) to prevent undesirable degradation of the nucleic acid target.
[0060] As used herein, "enzyme having single-stranded 3' to 5' exonuclease activity" may refer to an exonuclease or an exonuclease or both. The enzyme having 3' to 5' exonuclease activity may have or not have one or more additional enzymatic activities (e.g., specific or non-specific endonuclease activity). However, the enzyme preferably does not contain double-stranded exonuclease activity. According to a preferred embodiment, the enzyme will also have single-stranded 5' to 3' exonuclease activity. As a non-limiting example, enzymes having single-stranded 3' to 5' exonuclease activity that can be used in this invention include exonuclease I (ExoI), S1 exonuclease, exonuclease T, exonuclease VII (ExoVII), etc. In some cases, combinations of two or more enzymes may be used, such as those selected from those listed above. As a specific example, ExoVII and ExoI may be used in combination. Enzymatic degradation can be partial (i.e., single-stranded nucleic acid regions or molecules may remain in the population even after contact with an enzyme having single-stranded 3' to 5' exonuclease activity) or complete. This can depend on incubation conditions, sample composition, the nucleic acid population itself (e.g., nucleic acid structure), or other variables known to the technician. Therefore, the term "degradation" includes at least partial degradation of single-stranded nucleic acid molecules or regions present in a population of nucleic acid molecules.
[0061] According to a preferred embodiment, the enzyme having exonuclease activity does not have endonuclease activity. This may be advantageous when the target region contains sites that can be recognized by site-specific endonucleases, or when the target region is readily degraded by non-specific endonucleases. According to a preferred embodiment, the at least one enzyme having single-stranded 3' to 5' exonuclease activity is selected from exonuclease I, S1 exonuclease, exonuclease T, exonuclease VII, or a combination of two or more thereof. Preferably, the enzyme having single-stranded 3' to 5' exonuclease activity is Exo I and / or Exo VII.
[0062] After incubating the nucleic acid cluster (containing two types of V Cas protein-gRNA-nucleic acid complexes at at least the first site) of step a) of the method provided herein with at least one enzyme having single-stranded 3' to 5' exonuclease activity, a 5' single-stranded overhang is generated at said at least the first site. The 5' single-stranded overhang may have a length of at least 9 nucleotides, preferably 9, 10, 11, 12 or more nucleotides, as further described herein.
[0063] Removal of type V Cas protein-gRNA complex
[0064] The class 2 type V Cas protein-gRNA complex binds stably and tightly to nucleic acid molecules at specific sites, forming a class 2 type V Cas protein-gRNA-nucleic acid molecule complex (also referred to herein as the class 2 type V Cas protein-gRNA-nucleic acid complex). Therefore, in order to provide an accessible 5' overhang for binding to the class 2 type V Cas protein-gRNA complex for downstream steps, it is necessary to remove the class 2 type V Cas protein-gRNA complex from the class 2 type V Cas protein-gRNA nucleic acid molecule complex. Therefore, the method provided herein further includes step c) of removing the class 2 type V Cas protein-gRNA complex from the group of steps b). As used herein, the term "removal" refers to the physical separation of the class 2 type V Cas protein-gRNA complex from the class 2 type V Cas protein-gRNA-nucleic acid molecule complex. Removal may be partial or complete. Preferably, removal is complete. In some cases, the class 2 type V Cas protein-gRNA complex may still be present in solution, although not bound to nucleic acid molecules, while in other cases, it may be eliminated from solution. In particular, the type 2 type V Cas protein-gRNA complex can be removed from solution by degrading type 2 type V proteins and / or gRNA. As a non-limiting example, the type 2 type V Cas protein can be degraded by contacting a group of nucleic acid molecules with at least one protease. Advantageously, contacting the group of nucleic acid molecules with at least one protease will also degrade any other proteins that may be present (i.e., other site-specific endonucleases, such as type 2 Cas proteins that may be present in other type 2 type V Cas protein-gRNA complexes, contaminating proteins remaining in the initial sample, etc.). As a non-limiting example, the protease may be selected from serine proteases, cysteine proteases, threonine proteases, aspartic proteases, glutamate proteases, metalloproteinases, and / or asparagine peptide lyases.
[0065] Alternatively, nucleic acid molecules can be chelated with divalent cations (especially Mg2+) 2+The type V Cas protein-gRNA complex is removed by contacting a compound (e.g., EDTA or EGTA). In fact, the inventors have surprisingly shown that the type V Cas protein-gRNA complex can be removed upon addition of such a compound without the need for enzymes or other chemicals that degrade the type V Cas proteins. Therefore, according to a preferred embodiment, the nucleic acid molecular group is removed by contacting a divalent cationic chelating agent (preferably chelating Mg...) 2+ A chelating agent (a cation) is used to remove type V Cas protein-gRNA complexes. Preferably, the chelating agent is EDTA or EGTA. The amount of EDTA or EGTA added is preferably at least 2 times, more preferably at least 3, 4, 5 times, and even more preferably at least 10 times the amount of the divalent cation to be chelated. A suitable amount of chelating agent can be readily determined by a person skilled in the art based on the composition of the solution containing the nucleic acid molecular group (e.g., based on the presence and quantity of cations) and further according to the embodiments provided herein. According to a specific example, EDTA is added at a concentration of at least 20 mM, more preferably at least 25 mM. In the case of using at least one protease and a divalent cation chelating agent in step c), the at least one protease and the divalent cation chelating agent can be added simultaneously, wherein the chelating agent does not inhibit the activity of the at least one protease.
[0066] Alternatively or additionally, gRNA can be degraded by contacting the nucleic acid molecular group with at least one RNase (e.g., RNase A, RNase H, or RNase I), thereby removing the type V Cas protein-gRNA complex. In another embodiment, since RNA is unstable at elevated temperatures, the sample can be heated (e.g., to at least 65°C), optionally in the presence of divalent metal ions and / or at an alkaline pH.
[0067] Oligonucleotide probes
[0068] According to the method provided herein, after removing the type V Cas protein-gRNA complex, a group of nucleic acid molecules is contacted with an oligonucleotide probe (e.g., as provided in step d). As used herein, the term "oligonucleotide probe" refers to a polynucleotide molecule containing a single-stranded region at least partially complementary to the 5' single-stranded overhang. The oligonucleotide probe may be a single-stranded polynucleotide or may contain both single-stranded and double-stranded regions. In some cases, the oligonucleotide probe may be a hairpin aptamer. In particular, the oligonucleotide probe may contain a DNA sequence equivalent to the guide segment of the gRNA molecule (i.e., uracil bases are replaced by thymine bases and ribose is replaced by deoxyribose). Therefore, the probe will contain a region complementary to the 5' overhang, thereby enabling the formation of a double-stranded region or duplex between the oligonucleotide probe and the 5' overhang through hybridization. As a non-limiting example, the oligonucleotide probe region at least partially complementary to the 5' single-stranded overhang contains fewer than 4, 3, or 2 mismatches with the 5' single-stranded overhang, more preferably fewer than 3 or 2 mismatches with the 5' single-stranded overhang, and even more preferably, the oligonucleotide probe region at least partially complementary to the 5' single-stranded overhang is completely complementary to the 5' single-stranded overhang (i.e., contains no mismatches). As a non-limiting example, the 5' single-stranded overhang contains 9 to 24 nucleotides. Therefore, the 5' single-stranded overhang preferably contains 9 to 24 nucleotides, preferably at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 nucleotides.
[0069] Considering that the 5' protruding end can vary (e.g.) Figure 1C As shown in the diagram, the sequence of the probe is preferably complementary to a region extending to and potentially including the PAM site. Alternatively, a group of nucleic acid molecules may be contacted with multiple oligonucleotide probes (e.g., two, three, four, or more oligonucleotide probes) containing sequences complementary to different possible or intended 5' overhangs, such that the probes will successfully hybridize to most (if not all) of the possible (e.g., at least 80%, 90%, 95%, 99%) 5' overhangs, thereby forming a double strand. Alternatively, the oligonucleotide probe may further contain a sequence complementary to the target region (i.e., extending beyond the PAM site) at its 5' end. The oligonucleotide probe may also, as needed, further contain any sequence and any length of additional nucleotides at its 3' end (i.e., the 3' region that does not bind to the 5' overhang), for example, for downstream applications. Thus, additional sequences may be included in the oligonucleotide probe, wherein said sequences are present on one or both sides of a sequence at least partially complementary to the 5' single-stranded overhang. Figure 8 Non-restrictive instructions for oligonucleotide probes are provided.
[0070] Although the oligonucleotide probe can be of any length, preferably, the length of the probe is equal to or less than 200 nucleotides, more preferably equal to or less than 100 nucleotides, and even more preferably equal to or less than 50 nucleotides.
[0071] After hybridization, the probe is preferably ligated to a nucleic acid using methods known in the art, preferably by an enzyme with ligase activity, such as Taq DNA ligase. This step is particularly preferred when the 5' overhang is 12 nucleotides or less in length, but it can also be used when the 5' overhang is longer than 12 nucleotides.
[0072] On the one hand, further, the nucleic acid molecule group obtained in step c) of the method described herein is contacted with an enzyme having ligase activity. Preferably, the nucleic acid molecule group is contacted with an enzyme having ligase activity during step d) or in a separate step after step d) but before step e). More preferably, the nucleic acid molecule group is contacted simultaneously with the oligonucleotide probe and the enzyme having ligase activity, thereby producing at least one continuous nucleic acid molecule in the form of a double strand (i.e., without nicks). The simultaneous use of the oligonucleotide probe and the enzyme having ligase activity advantageously reduces the time required to perform the method.
[0073] On the other hand, the nucleic acid molecule group obtained in step c) of the method described herein is contacted with at least one enzyme that catalyzes the cleavage of 5' nucleic acid flaps from a double-stranded nucleic acid molecule. The enzyme is therefore preferably phosphorylated at the 5' end of the oligonucleotide probe after cleavage. The 5' end can then be ligated to the nucleic acid molecule at the resulting cleavage site. If the enzyme does not phosphorylate the 5' end, the nucleic acid molecule group can be contacted with an enzyme that phosphorylates the 5' end, such as T4 polynucleotide kinase, prior to ligation.
[0074] In one particular implementation, the nucleic acid molecular group is advantageously contacted with the FEN1 enzyme, which, after cleavage hybridization, may be present in any 5' DNA lobe of the oligonucleotide probe (e.g., as shown in the image). Figure 5 and Figure 6(As shown). The step of contacting the nucleic acid molecule group with the FEN1 enzyme can be performed during or after step d) but before step e). After or simultaneously with contact with the FEN1 enzyme, the nucleic acid molecule group can be contacted with a ligase to close any remaining nicks after FEN1 cleavage. Therefore, in particular, the method provided herein can further include the steps of cleaving the 5' nucleic acid flap from the double-stranded nucleic acid molecule and / or a ligation step. According to a particular embodiment, the exact location of the 3' end of the overhang obtained after cleavage with the type 2 V Cas protein-gRNA complex can be determined, and the oligonucleotide probe is designed such that there is no extra oligonucleotide at its 5' end. In this case, contacting the nucleic acid molecule group with FEN1 is not required. Alternatively, the probe may contain modifications that increase the hybridization strength between the probe and the 5' overhang, such as base modifications, backbone modifications, or 5' or 3' oligonucleotide modifications (e.g., using LNA, acridine, etc.), and strand substitution can be performed so that the entire length of the oligonucleotide probe (including the sequence complementary to the target region) hybridizes. In these cases, ligation is not necessarily required.
[0075] Oligonucleotide probes may further comprise ligands capable of binding to a capture agent (e.g., a capture protein). As used herein, the term "ligand" refers to a molecule, peptide, substrate, etc., that has an affinity for the capture agent. The ligand may contain functional groups. When the ligand is a peptide, it may in particular contain natural, non-natural, and / or artificial amino acids. The ligand may form non-covalent or covalent bonds with the capture agent upon interaction. The ligand-capture agent interaction is further detailed below. In some cases, the ligand may be a single-stranded nucleotide overhang that is at least partially (preferably completely) complementary to an oligonucleotide anchored to a support. When multiple oligonucleotide probes are used to isolate multiple target regions, the probes may contain the same or different ligands (e.g., different ligands for each target region, or different ligands for each group of target regions). Preferably, when multiple target regions are isolated, all oligonucleotide probes will contain the same ligand. Therefore, the term "hybrid capture" as used herein refers to the hybridization of an oligonucleotide with its 5' overhang to form a double strand (e.g., in step d of the method provided herein), preferably, wherein the oligonucleotide is linked to the target nucleic acid to form a continuous double strand, and then the double strand is isolated from the nucleic acid molecule group (e.g., in step e of the method provided herein), preferably, by capturing the double strand directly or indirectly (e.g., by a ligand on the oligonucleotide probe or a ligand on another molecule that binds to the probe, such as a second oligonucleotide molecule, protein, etc.).
[0076] The oligonucleotide probe may further include cleavage sites, such as uracil bases, abase sites, or restriction sites. Preferably, the uracil bases or abase sites are contained within a single-stranded region of the oligonucleotide that is not complementary to the 5' overhang, and therefore are not contained in the duplex. Preferably, the restriction sites are contained within a double-stranded region of the oligonucleotide. In some cases, the restriction sites can be generated by forming a duplex. This is advantageous because the duplex can be separated from the ligand by cleavage at the cleavage site. Therefore, according to a preferred embodiment, the oligonucleotide probe preferably includes cleavage sites, and even more preferably includes uracil bases, abase sites, or restriction sites.
[0077] double chain
[0078] After contacting the nucleic acid cluster with the oligonucleotide probe as described above, the double strand is isolated from the nucleic acid cluster of step d), thereby also isolating the nucleic acid target region from the nucleic acid cluster (e.g., as in step e of the method provided herein). Preferably, the nucleic acid cluster of step d) (i.e., the double strand comprising the 5' overhang and the oligonucleotide probe) is contacted with a trapping agent. This is particularly preferred when the probe contains a ligand.
[0079] As used herein, the term "capture agent" refers to an oligonucleotide, protein, or other molecule that forms a non-covalent or covalent bond with its ligand upon interaction with that ligand. Non-covalent bonds typically include several weak interactions, such as hydrophobicity, van der Waals forces, and hydrogen bonds, which often occur simultaneously. In the first example, when the ligand is a substrate of an enzyme and the enzyme is a capture agent, the substrate can be non-covalently bonded to the enzyme. In the second example, a biotin ligand (or a biotin analog such as dethiobiotin) will bind to a streptavidin capture agent, forming a quasi-covalent bond. As a third example, a digoxigenin ligand will be non-covalently bonded to a capture agent that is an antibody against digoxigenin (anti-DIG). According to a specific example, the oligonucleotide probe is a hairpin aptamer containing a digoxigenin ligand, which can bind to an anti-DIG capture agent.
[0080] Alternatively, a “covalent bond” refers to a form of chemical bond characterized by the sharing of electron pairs between atoms or between atoms and other covalent bonds. As a first example, when the ligand contains an amino acid, the amino acid can be covalently bound to the scavenger. In another example, the ligand will contain a first functional group, and the scavenger will contain a second functional group, the first and second functional groups interacting to form a covalent bond, for example, through click chemistry. The covalent bond between the ligand and the scavenger can be, for example, an amide bond, an amine-thiol bond, a Cu(I)-catalyzed azide-alkyne cycloaddition, an alkyne-nitroketone cycloaddition, etc. As another example, the ligand may contain functional groups (--COOH, --NH2, -OH, etc.) capable of reacting with the carboxyl (--COOH) or amine (-NH2) terminus of a protein scavenger. The technique described in patent EP152886 uses enzymatic coupling to link DNA to a scavenger (such as cellulose). Patent EP146815 also describes various methods for linking DNA to a scavenger. Similarly, patent application WO92 / 16659 discloses a method for ligating DNA using a polymeric trapping agent. The oligonucleotide probe may further include a "spacer" molecule or region that connects the ligand, advantageously providing additional space for the ligand to bind its specific trapping agent. As a non-limiting example, the spacer may be a polynucleotide region, tetraethylene glycol (TEG), or any other spacer known to those skilled in the art. Alternatively, the "spacer" may include, or consist of, cleavage sites, such as TEV protease cleavage sites or uracil bases, abase sites, or the aforementioned restriction sites.
[0081] The trapping agent is preferably anchored to a solid support. In some embodiments, multiple trapping agents are anchored to the same solid support. As a non-limiting example, the trapping agent may more specifically be attached to the solid support via any suitable structure or mechanism (e.g., fusion, chemical linkage (e.g., direct or indirect), enzymatic linkage, fusion, tethering, linkage, etc. via adapters (e.g., peptides, nucleic acids, polymers, ester bonds, PEG adapters, carbon chains, etc.)). As a non-limiting example, possible supports include: glass and modified or functionalized glass, plastics (including acrylic resins, polystyrene and copolymers of styrene with other materials, polypropylene, polyethylene, polybutene, polyurethane, Teflon™, etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica or silica-based materials (including silicon and modified silicon), carbon, metals, inorganic glasses, plastics, fiber bundles, and various other polymers. The solid support may be selected from, for example, pores, tubes, glass slides, plates, resins, or beads. In some embodiments, the solid support (e.g., resin, beads, etc.) is magnetic. In some embodiments, the solid support (e.g., resin, beads, etc.) is paramagnetic. When the solid support is a bead, the bead size ranges from nanometers (e.g., 100 nm) to millimeters (e.g., 1 mm).
[0082] According to the first aspect, the capture agent is anchored to a solid support before the nucleic acid molecular group is contacted with the capture agent. In other cases, the nucleic acid molecular group is contacted with the capture agent before the capture agent is anchored to the solid support. In both cases, the nucleic acid target region is indirectly connected to the solid support.
[0083] After a bond is formed between the duplex and the trapping agent, the duplex can be separated from the nucleic acid molecule group, for example, by separating the duplex from the group. Preferably, the support to which the trapping agent is anchored (and thus the duplex and associated target region) are separated from the nucleic acid molecule group by mechanical separation, washing, or any method known to those skilled in the art. Preferably, the duplex is separated by washing according to methods known in the art, wherein non-target nucleic acid molecules (which are not bound to the solid support) are removed. Preferably, at least two washes are performed. As a non-limiting example, the buffer used for washing may or may not contain a detergent. The detergent may be an ionic or non-ionic detergent. As a non-limiting example, the detergent may be polysorbate-20 (commonly known as Tween 20), 4-(1,1,3,3-tetramethylbutyl)phenyl-polyethylene glycol (commonly known as Triton X100), or sodium dodecyl sulfate (SDS). The detergent may be included in the washing buffer at a concentration of about 0.05% to about 1.5%. The duplexes can be washed with a single wash buffer or multiple wash buffers. Each wash can use the same wash buffer or different wash buffers. For example, one wash can use a wash buffer containing detergent, while another wash can use a wash buffer without detergent.
[0084] In some cases, the isolated duplex containing the separated target region can be used directly for downstream applications. Alternatively, after separating the duplex, the target region is preferably released by any suitable mechanism (e.g., chemical or enzymatic cleavage). As a non-limiting example, the target region is released from the solid support (e.g., beads) by releasing the target region from the duplex. In particular, releasing the target region from the duplex can include site-specific cleavage within the duplex, such as cleavage at a restriction site within the duplex, for example, by contacting the duplex with a suitable restriction enzyme. Alternatively, the duplex (which binds the target region) itself is released from the solid support by any suitable mechanism. As a non-limiting example, cleavage (e.g., chemically or enzymatically) occurs at a cleavage site between the trap and the duplex (e.g., located within a single-stranded region or spacer region of an oligonucleotide probe) to release the duplex.
[0085] According to a preferred embodiment, the method provided herein further includes the step of releasing the bichain from the target region, wherein releasing the bichain from the target region preferably includes cutting within the bichain, more preferably as described above.
[0086] As an alternative, non-limiting example, in the case of non-covalent bonding between the capture agent and the ligand, an excess of the ligand may be added, wherein the ligand competes with the bound ligand, causing the bound ligand (and thus the double strand and target region) to dissociate from the capture agent. As a non-limiting example, the ligand and the excess ligand may be digoxigenin. In a particular example, the excess ligand may be a variant of the ligand, wherein the excess ligand has a higher affinity for the capture agent than the ligand contained in the oligonucleotide probe. This advantageously promotes the dissociation of the double strand, thereby promoting the dissociation of the target region. As a particular example, the ligand may be dethiobiotin, and the excess ligand is biotin. In other embodiments, the capture agent itself is released from the surface, thereby releasing the entire complex comprising the capture agent, the double strand, and the target region. Each of these alternatives results in the release of the target region from the solid support. Advantageously, enzymatic cleavage produces a 5' or 3' overhang, preferably a 3' overhang.
[0087] According to a preferred embodiment, step e) includes:
[0088] • Contact the double strand with a trapping agent, which is preferably a ligand that binds to an oligonucleotide probe.
[0089] Preferably, the trapping agent is anchored to a solid support.
[0090] According to a preferred embodiment, when the trapping agent is anchored to a solid support, the method of separating the target region further includes releasing the target region from the solid support, preferably by releasing the bichain from the target region.
[0091] In one specific implementation, when the target region is contacted with the second type 2 V Cas protein-gRNA complex in step a), the nucleic acid group may be contacted with the second oligonucleotide probe in step d), wherein the second probe may contain the same or different ligands as those contained in the first probe (e.g., both probes contain biotin ligands, or one probe contains biotin ligands while the second probe contains digoxigenin ligands).
[0092] When both probes contain the same ligand, the two duplexes (i.e., located on opposite sides of the target region) will bind to a given trapping agent. For example, when both duplexes contain biotin-binding oligonucleotide probes, both will bind to streptavidin-coated beads. This can advantageously improve enrichment because each target region binds simultaneously to the solid support via covalent or non-covalent bonds of the two ligands.
[0093] Therefore, according to a specific implementation scheme, step e) includes:
[0094] • Contact the first and second double strands with the trapping agent, which binds to the ligands of the first and second probes.
[0095] • Release the first and second duplexes from the target region.
[0096] According to the alternative preferred embodiment, step e) includes:
[0097] • Contact the first and second double-stranded structures with a trapping agent, wherein the trapping agent binds to ligands of the first and second probes, and wherein the trapping agent is anchored to a solid support.
[0098] • Release the nucleic acid target region from the solid support, preferably by releasing the first and second double strands from the target region.
[0099] In an alternative specific implementation, when the second oligonucleotide probe contains a different ligand, the first duplex is preferably separated before the second duplex is separated. In this case, step e) is repeated a second time. This is in Figure 5 The document specifically explains that each oligonucleotide probe contains a different ligand and is sequentially separated.
[0100] According to a preferred embodiment, step e) includes:
[0101] • The first double-stranded polymer is brought into contact with the first trapping agent, and the first trapping agent binds to the ligand of the first probe.
[0102] • Release the first double strand from the target region.
[0103] • The second double-stranded polymer is brought into contact with the second trapping agent, and the second trapping agent binds to the ligand of the second probe.
[0104] • Release the second double strand from the target region.
[0105] According to a specific implementation scheme, step e) includes:
[0106] • The first double-stranded polymer is brought into contact with the first trapping agent, and the first trapping agent binds to the ligand of the first probe.
[0107] • Release the first double strand from the target region.
[0108] • The hairpin aptamer is brought into contact with a second trapping agent, which binds to the ligands of the hairpin aptamer.
[0109] • Release the hair clip adapter from the target area.
[0110] Those skilled in the art will understand that, for example, when two or more trapping agents are used for successive separation according to any of the above embodiments, the separation of the first and second duplexes can be performed in any order.
[0111] Optional supplementary steps
[0112] Fragmentation
[0113] According to one specific embodiment of the invention, nucleic acid molecules can be fragmented before or after contact with the type 2 V Cas protein-gRNA complex in the above-described method, advantageously after contact with the type 2 V Cas protein-gRNA complex. As used herein, the term "fragmentation" refers to increasing the number of 5'- and 3'-free ends of a nucleic acid molecule by breaking it into at least two smaller molecules. Nucleic acid fragmentation is advantageous because the method according to the invention allows for easier separation of smaller nucleic acid molecules. In this step, it is advantageous to separate "non-target" nucleic acid regions (e.g., regions other than the target region and adjacent sites that bind to the type 2 Cas protein-gRNA complex) from the target, thereby improving the enrichment of the target region when performing the method of the invention.
[0114] Fragmentation can be performed by shearing, such as by sonication, water shearing, sonication, nebulization, or by enzymatic fragmentation, such as by using one or more site-specific endonucleases, for example, restriction enzymes. When fragmentation is performed by site-specific endonucleases, one or more site-specific endonucleases can be used, preferably 1, 2, 3, 4, 5, or more site-specific endonucleases. It should be understood that the ever-increasing number of available sequences in databases allows those skilled in the art to easily identify one or more restriction enzymes whose cleavage sites are located outside the nucleic acid containing the target region. Advantageously, when two or more enzymes are used simultaneously, the enzymes are compatible with each other (e.g., the same buffer requirements, inactivation conditions). Fragmentation can be partial (e.g., not all cleavage sites present in the group of nucleic acid molecules are cleaved by restriction enzymes) or complete. Thus, the term "fragmentation" includes at least partially fragmenting the nucleic acid outside the target region and adjacent sites that bind to the class 2 type V Cas protein-gRNA complex.
[0115] Therefore, in some embodiments, the method for preparing nucleic acids containing the target region further includes the following steps:
[0116] • Fragmenting a group of nucleic acid molecules, preferably by contacting the group with at least one site-specific endonuclease, more preferably with at least one restriction enzyme.
[0117] Particularly preferably, only the non-target region (i.e., the target region that binds to type V Cas proteins and the region other than the first site) is fragmented. More preferably, when the second site is also present adjacent to the target region, only the target region, the first site, and the region other than the second site are fragmented.
[0118] Therefore, according to a specific embodiment, the method provided herein further includes fragmenting the nucleic acid molecular group before or during step b), preferably by contacting the nucleic acid molecular group with at least one site-specific endonuclease, wherein the site-specific endonuclease:
[0119] • Do not cut within the target region or the first site, and
[0120] • When the molecule contains a second site, it is not cleaved within the second site.
[0121] Preferably, the site-specific endonuclease is a restriction enzyme.
[0122] Those skilled in the art will understand that the above-described fragmentation step can occur at any step, such as before or during step a), between steps a) and b), during step b), between steps b) and c), during step c), or during step e). When enzymatic fragmentation is performed and the site-specific endonuclease cleaves the type 2 V Cas protein-gRNA complex under the same conditions (e.g., buffer, temperature), the fragmentation step is preferably performed simultaneously with step a) which contacts the nucleic acid molecule cluster with the type 2 V Cas protein-gRNA complex.
[0123] According to a preferred embodiment, nucleic acid molecules are fragmented by contacting a group of nucleic acid molecules with at least one site-specific endonuclease, preferably at least 1, 2, 3, 4, 5 or more site-specific endonucleases. Preferably, the site-specific endonuclease is a restriction enzyme, more preferably type II, type III or artificial restriction enzymes, even more preferably type II restriction enzymes, and / or type V Cas protein-gRNA complexes, such as Cas12a-gRNA complexes. Type II restriction enzymes include the categories IIP, IIS, IIC, IIT, IIG, IIE, IIF, IIG, IIM and IIB, as described, for example, in Pingoud and Jeltsch, Nucleic Acids Res, 2001, 29(18): 3705–3727. Preferably, one or more enzymes from these categories are used for the nucleic acid molecule fragmentation in this invention. Those skilled in the art can select suitable enzymes. In cases where multiple restriction enzymes used are incompatible with each other, fragmentation may include multiple successive steps using different restriction enzymes and conditions (e.g., temperature, time, buffer). Preferably, at least one site-specific endonuclease generates non-palindromic overhangs. Preferably, at least 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the cleavage sites are cleaved by the site-specific endonuclease. In the case where the site-specific endonuclease is a type V Cas protein-gRNA complex, the protein-gRNA complex binds to and cleaves at a site outside the nucleic acid molecule containing the target region being prepared. Preferably, the site-specific endonuclease binds to sequences located 100 to 5000 bases from the nucleic acid containing the target region, more preferably 150 to 1000 bases from the nucleic acid region containing the target region, and even more preferably 250 to 750 bases from the nucleic acid region containing the target region. Preferably, the site-specific endonuclease targets specific sequences that are present multiple times within the nucleic acid molecule, such as the Alu element in the human genome, but not present in the nucleic acid molecule containing the target region.
[0124] According to one specific embodiment, after contacting a group of nucleic acid molecules with a type V Cas protein-gRNA complex, the group is simultaneously contacted with at least one enzyme having 3' to 5' single-stranded exonuclease activity and at least one site-specific endonuclease, fragmenting the nucleic acid molecules according to any of the above embodiments. This is particularly advantageous because it reduces the duration of the method. According to another preferred embodiment, the enzyme having 3' to 5' single-stranded exonuclease activity may also have site-specific endonuclease activity. According to this embodiment, the cleavage site of the enzyme having site-specific endonuclease activity is located outside the first site and the target region. When a second site is present, the cleavage site is also located outside the second site. Using a single enzyme having both exonuclease and site-specific endonuclease activities is particularly advantageous because it reduces the number of reagents required and the cost.
[0125] Site-specific endonucleases
[0126] In some cases, as described above, the method may further include contacting the nucleic acid molecule group with a site-specific endonuclease at any stage prior to step c) (i.e., before, during, or after step a) or step b). The site-specific endonuclease is preferably a Cas protein-gRNA complex, more preferably a class 2 Cas protein-gRNA complex, and even more preferably a class 2 type V Cas protein-gRNA complex. However, any other site-specific nuclease that stably binds nucleic acids at a specific site, such as transcription activator-like effector nucleases (TALENs) or zinc finger proteins, is also included within the scope of the invention. This step is particularly advantageous when modifying both ends of the molecule (e.g., generating two 5' protrusions, one on either side of the target region). In cases where it is desirable to obtain a molecule containing a single 5' protrusion on one side of the target region and a blunt end on the other side, the nucleic acid molecule group is advantageously contacted with a class 2 type II Cas protein-gRNA complex, preferably with a Cas9-gRNA complex.
[0127] The site-specific endonuclease will bind within the nucleic acid molecule.
[0128] Therefore, according to a preferred embodiment, the method of the present invention further includes contacting the nucleic acid molecular group with a site-specific endonuclease prior to step c), the site-specific endonuclease preferably being a TALEN, a zinc finger protein, or a type 2 Cas protein-gRNA complex, more preferably a second type 2 V Cas protein-gRNA complex, wherein the gRNA contains a guide segment complementary to the second site, wherein the second site is adjacent to the target region, and wherein the first site and the second site are located on either side of the target region. When the nucleic acid molecular group is contacted with multiple Cas protein-gRNA complexes (e.g., multiple type 2 V Cas protein-gRNA complexes), it is particularly preferred that the nucleic acid molecular group is contacted with all of the Cas protein-gRNA complexes simultaneously. This is advantageous because the duration of the method is reduced. Preferably, the site-specific endonuclease is a Cas12a-gRNA complex. Preferably, the PAM site is adjacent to the nucleic acid target region. Therefore, according to a specific embodiment of the method, step a) includes contacting the nucleic acid molecular group with first and second Cas12a-gRNA complexes, the first and second complexes binding to first and second sites, respectively, which are located near and on either side of the target region.
[0129] According to a preferred embodiment, the site-specific endonuclease binds to a site located at least 50 nucleotides away from the type 2 V Cas protein-gRNA-nucleic acid complex, preferably at least 100, 250, 500, 1,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 500,000, 750,000, or 1,000,000 nucleotides away. Therefore, according to the preferred embodiment, the length of the target region is at least 50, 100, 250, 500, 1,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 500,000, 750,000, or 1,000,000 nucleotides.
[0130] According to a preferred embodiment, when the site-specific endonuclease is a type 2 V site-specific endonuclease, the group is contacted with a second oligonucleotide probe containing at least a sequence complementary to the second 5' single-stranded overhang in the second site, thereby forming a second double strand. In this case, as discussed herein, the first and second probes preferably contain different ligands that bind different trapping agents.
[0131] Hair clip fit
[0132] In some cases, the method may further include the step of contacting a group of nucleic acid molecules with a hairpin aptamer. The hairpin aptamer binds to one or more free ends of one or more nucleic acid molecules present in the group. This step may be performed before or simultaneously with contact with the type 2 V Cas protein-gRNA complex. When fragmentation is performed, this step is preferably performed after fragmentation. As used herein, the term "hairpin" or "hairpin aptamer" refers to a molecule in which the bases themselves pair to form a structure having a double-stranded stem and loop, wherein the 5' end of one strand is physically connected to the 3' end of the other strand by an unpaired loop. The physical connection may be covalent or non-covalent. Preferably, the physical connection is a covalent bond. As used herein, the term "loop" refers to a chain of nucleotides of a nucleic acid strand that does not pair with nucleotides of the same or another strand of the nucleic acid by hydrogen bonds, and is therefore single-stranded. As used herein, "stem" refers to a pairing region within the strand. Preferably, the stem contains at least 3, 5, 10, or 20 base pairs, more preferably at least 5, 10, or 20 base pairs, and even more preferably at least 10 or 20 base pairs. When the hairpin binds to the free end of the double-stranded nucleic acid molecule, the 3' and 5' ends of the hairpin are attached to the 5' and 3' ends of the double-stranded nucleic acid molecule, respectively. The hairpin can bind to either a protruding end or a blunt end. Therefore, the hairpin aptamer does not need to bind to a specific sequence or site.
[0133] Preferably, the hairpin aptamer binds to one or both free ends of a nucleic acid molecule. The hairpin aptamer may specifically bind to one of the free ends of a nucleic acid molecule or non-specifically bind to all free ends of all nucleic acid molecules. As a non-limiting example, specific binding can be achieved by fragmenting the nucleic acid molecule with a non-palindromic restriction enzyme, thereby generating a different overhang at each new free end of the nucleic acid molecule. In other cases, the nucleic acid molecule may contain blunt ends, to which the hairpin aptamer may attach for non-specific binding. In another example, a group of nucleic acid molecules (fragmented or unfragmented) may undergo A-tailing, such that a hairpin aptamer containing a suitable "T" nucleotide overhang can then be hybridized and attached to the site. In a preferred embodiment, the first site bound by the type V Cas protein-gRNA complex and the free ends (and therefore the hairpin aptamer) are located on opposite sides of the target region. After removing the type 2 V Cas protein-gRNA complex, this structure can be obtained by contacting a nucleic acid mass with a hairpin aptamer that specifically binds to a free end adjacent to the nucleic acid target region, wherein the first side and the free end are located on opposite sides of the target region. Alternatively, the hairpin aptamer can bind to all free ends, wherein the nucleic acid mass is contacted with the hairpin aptamer before or simultaneously with contacting the type 2 V Cas protein-gRNA complex (see also...). Figure 3 ).
[0134] As used herein, the term "free end" refers to the end of a nucleic acid molecule, which may contain a phosphate group at the 5' end and / or a hydroxyl group at the 3' end. The free end may be blunt or include a single-stranded overhang. The single-stranded overhang may be a 3' or 5' overhang. The single-stranded overhang preferably has a length of less than 100, 50, 25, 10, 5, 4, 3, or 2 nucleotides, more preferably having a length of 1 nucleotide.
[0135] According to a preferred embodiment, the hairpin aptamer binds to the free end located at least 50 nucleotides away from the type 2 V Cas protein-gRNA nucleic acid complex, preferably at least 100, 250, 500, 1,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 500,000, 750,000, or 1,000,000 nucleotides away. Therefore, according to the preferred embodiment, the length of the target region is at least 50, 100, 250, 500, 1,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 500,000, 750,000, or 1,000,000 nucleotides.
[0136] Preferably, the hairpin aptamer is connected to a nucleic acid molecule. More preferably, step a) includes contacting the nucleic acid cluster with the hairpin aptamer and connecting the hairpin aptamer to the free end of the nucleic acid molecule. In a preferred embodiment, if desired (e.g., when the sample contains circular nucleic acid molecules), the method further includes a step of linearizing the nucleic acid cluster before contacting the cluster with the hairpin aptamer.
[0137] According to a specific implementation, when isolating multiple target regions, in addition to contacting the type 2 V Cas protein-gRNA complex, the nucleic acid molecular group can also be contacted with a site-specific endonuclease (preferably the type 2 Cas protein-gRNA complex as described herein) and a hairpin aptamer. The nucleic acid molecular group can be contacted simultaneously or sequentially in any order.
[0138] Incubation and / or storage
[0139] According to a specific embodiment of the present invention, nucleic acid molecules may be stored after step c), step d), or step e) of the method provided herein.
[0140] According to specific embodiments, nucleic acid molecules can be incubated before, during, or after any step of the method provided herein. Preferably, the nucleic acid molecule group can be incubated during or after contact with a type V Cas protein-gRNA complex (step a), during or after contact with at least one enzyme having 3' to 5' single-stranded exonuclease activity (step b), during removal of the type V Cas protein-gRNA complex, during step d) contacting the group in step c) with the oligonucleotide probe, and / or during step e) separating the duplex from the nucleic acid molecule group. Preferably, the incubation time for the nucleic acid molecule group ranges from 15 minutes to 2 hours. Those skilled in the art will particularly know that an appropriate incubation time / temperature is needed to maximize enzyme activity or to reduce the duration of the method while maximizing enzyme activity, and the incubation time / temperature can be further adjusted if necessary.
[0141] Downstream applications
[0142] The nucleic acid target regions isolated according to one of the methods described above are advantageously highly enriched. The isolated nucleic acids are particularly suitable for a wide range of applications. In fact, the nucleic acids isolated according to the invention can be further processed, reacted, or analyzed, whether in the same container or in different containers. For example, the nucleic acids isolated according to the invention can be used for detection, cloning, sequencing, amplification, hybridization, cDNA synthesis, diagnostics, and any other methods known to technicians that require nucleic acids. In some cases, the isolated nucleic acid target regions can be further isolated, enriched, or purified.
[0143] The method of this invention is particularly suitable for generating hairpin libraries after isolating one or more target regions, wherein each hairpin contains at least one nucleic acid target region. Therefore, this method is particularly convenient for detecting or determining the sequence of a target region of interest, such as a specific allele isolated from an entire population of nucleic acid molecules in a biological sample, or epigenetic modifications of the target region of interest.
[0144] According to a preferred aspect of the invention, the method of the invention may further include additional steps. As a non-limiting example, the isolated nucleic acid may be further purified using well-known purification methods (e.g., bead or column purification, such as purification with paramagnetic beads) to remove proteins, such as proteases, salts, EDTA, excess oligonucleotides, etc. As a non-limiting example, the nucleic acid molecule may hybridize and / or ligate to a target region or a site adjacent to the target region, and single-strand gaps in the nucleic acid molecule may be filled by synthesizing a complementary strand and / or performing strand substitution. One or more of these additional steps are particularly useful for generating hairpin libraries, but may also be necessary when preparing isolated nucleic acids for other downstream applications. In a specific example, when one or more double-stranded nucleic acid target regions or molecules containing said target regions are isolated according to the method of the invention, hairpin molecules, as previously defined herein, may then be ligated to one or both free ends of said target regions or molecules. Preferably, the hairpin is ligated to one free end of the isolated nucleic acid target region or molecule (see also...). Figure 6 Preferably, at least one free end of the isolated nucleic acid target region or molecule includes a 3' or 5' protrusion. Preferably, the hairpin includes a 3' or 5' protrusion that is at least partially complementary to at least one of the 5' or 3' protrusions of the isolated nucleic acid target region or molecule. Preferably, the hairpin is attached to a 3' protrusion at one end of the isolated nucleic acid target region or molecule. As an alternative example, it is advantageous to attach the hairpin to the protrusion in the presence of the FEN1 enzyme, which cleaves the 5' DNA flap. In fact, the inventors have found that, when preparing the fragment to be isolated using catalytically active Cas12a, attaching the hairpin to the protrusion in the presence of FEN1 promotes the cleavage of the protruding nucleotide at the 5' end of the oligonucleotide (see [link to product description]). Figure 7 In fact, the two types of V Cas proteins (such as Cas12a) are not always cleaved at the same location, such as... Figure 1C As shown. After the hairpin is attached, gap filling and connection reactions can be performed using methods known in the art.
[0145] Therefore, according to the first embodiment, the method of the present invention further includes the following steps:
[0146] • Hybridize and / or link one or more single-stranded or double-stranded nucleic acid molecules to isolated nucleic acid target regions.
[0147] Preferably, the single-stranded or double-stranded nucleic acid molecule hybridizes with the 5' or 3' overhang adjacent to the target region. After hybridization, ligation is preferably performed. In a particular embodiment, the method may include the following steps:
[0148] • Hybridize at least one single-stranded nucleic acid molecule to the 5' or 3' overhang adjacent to the target region, and
[0149] • Link the single-stranded nucleic acid molecule to the double-stranded region.
[0150] However, when a single-stranded nucleic acid molecule (e.g., an oligonucleotide) binds to a single-stranded region of a target that is directly adjacent to a double-stranded region, ligation can also occur directly without hybridization.
[0151] According to a preferred embodiment, the single-stranded nucleic acid molecule hybridizes with a single-stranded region of an oligonucleotide probe, which is retained within the nucleic acid molecule containing the target region. Preferably, the oligonucleotide probe contains a single-stranded sequence that can hybridize with all single-stranded nucleic acid molecules. The single-stranded sequence can be an artificial sequence (i.e., not naturally occurring). The presence of such a single-stranded sequence within the probe is particularly advantageous when isolating multiple nucleic acid target regions because it eliminates the need to design multiple sequence-specific single-stranded nucleic acid molecules, for example, for constructing hairpin molecules. In fact, all single-stranded nucleic acid molecules can bind the same sequence, which is present in all probes. Therefore, the sequence in the probe and the single-stranded nucleic acid molecule containing the complementary sequence are referred to as "universal." This is in... Figure 8 This is explained in detail in the text.
[0152] According to another implementation, the method includes the following steps:
[0153] • To hybridize at least one single-stranded nucleic acid molecule with the isolated target region,
[0154] • To extend single-stranded nucleic acid molecules into double-stranded regions, preferably by contacting the separated target region with a nucleic acid polymerase, and
[0155] • Link the extended single-stranded nucleic acid molecule to the double-stranded region.
[0156] According to a preferred embodiment, at least one single-stranded nucleic acid molecule hybridizes and polymerizes at a 3' overhang. Preferably, the single-stranded nucleic acid molecule hybridizes with a region located outside the first site and target region on the sequence provided by the oligonucleotide probe. Preferably, the single-stranded nucleic acid molecule hybridizes with a region located 2 or more bases away from the double strand, the double strand being formed by the oligonucleotide probe and a 5' overhang (see also, for example, Figure 6 (E) Figure 8 The methods of hybridization, extension, and joining are well known to technicians.
[0157] In some cases, any of the above implementation methods can be repeated, for example, by adding a second single-stranded nucleic acid molecule to the isolated target region (see, for example...). Figure 6The second single-stranded nucleic acid molecule can hybridize with the same or opposite strands and may or may not contain a label or ligand. The single-stranded nucleic acid molecule may be only partially complementary to the sequence of the isolated target region. The single-stranded nucleic acid molecule may preferably contain a spacer region, for example, a 12-carbon spacer region or any other spacer region described herein or known to those skilled in the art, which does not bind to the isolated target region (e.g., is not complementary to the sequence of the isolated target region). Preferably, the single-stranded nucleic acid molecule contains a 5' phosphate group for ligation.
[0158] Optionally, excess reagents, such as non-hybridized single-stranded nucleic acid molecules, are then removed. For example, non-hybridized single-stranded nucleic acid molecules are eliminated by contacting a sample containing the isolated target region with an enzyme having 3' to 5' exonuclease activity (more preferably ExoI).
[0159] Preferably, the hairpin structure is obtained according to any of the methods described herein, methods particularly suitable for downstream applications, such as those described in WO 2011 / 147931, WO 2011 / 147929, WO 2013 / 093005 and WO 2014 / 114687, all of which are incorporated herein by reference in their entirety. Alternatively, the hairpin structure produced herein may be particularly suitable for use as a hairpin precursor molecule (e.g., the HP2 molecule described in WO 2016 / 177808, the entire contents of which are incorporated herein by reference).
[0160] Preferably, one or more single-stranded nucleic acid molecules of any embodiment described herein have optimized hybridization specificity, as described by Zhang et al., Nat Chem, 2012, 4(3): 208-214, which is incorporated herein by reference in its entirety. Alternatively, the one or more single-stranded nucleic acid molecules of any embodiment described herein may be degenerate.
[0161] Preferably, one or more single-stranded nucleic acid molecules of any of the above embodiments comprise a label or ligand. As a non-limiting example, the label or ligand may be FITC, digoxigenin, biotin, or any other label known to the art. The label or ligand may be conjugated to a protein using techniques such as chemical coupling and chemical cross-linking agents. Advantageously, the target region can be detected and optionally quantified within a sample, for example by fluorescent labeling or other detectable labels or ligands known to the art. In some cases, the ligand may be used to further isolate or purify the target region. In a first aspect, according to methods known to the art, when an oligonucleotide is labeled with biotin, the target region can be isolated by a further pull-down reaction, for example on beads coated with streptavidin. In a second aspect, the target region may be linked to a support, such as beads or a chip, by the label or ligand. Preferably, the support is functionalized to facilitate the linking of the labeled target region, and the label or ligand reacts with functional groups present on the support (e.g., the support may be coated with streptavidin or COOH groups, which react with a suitable label or ligand).
[0162] According to a particular embodiment, at least one single-stranded nucleic acid molecule of any of the above embodiments contains a sequence complementary to an oligonucleotide bound to a solid support (e.g., a surface), when said sequence is not included in the oligonucleotide probe. Preferably, the oligonucleotide contains a modification at its 3' end to prevent elongation. Hybridization and attachment of the single-stranded nucleic acid molecule, with or without a tag, to the 3' protrusion advantageously produces a hairpin structure, particularly suitable for downstream applications, such as those described in WO2011 / 147931, WO 2011 / 147929, WO 2013 / 093005, and WO 2014 / 114687. Preferably, any embodiment described herein produces a hairpin having a "Y" shape.
[0163] This invention further allows those skilled in the art to enumerate the number of nucleic acid molecules carrying the sequence. According to a preferred embodiment, the method of this invention further includes detecting and quantifying nucleic acid molecules as described in WO2013 / 093005.
[0164] The isolated target regions of this invention are particularly suitable for downstream analysis using methods such as: single-molecule analysis methods, for example, those described in WO 2011 / 147931 and WO 2011 / 147929; nucleic acid detection and quantification methods, for example, those described in WO 2013 / 093005; and methods for detecting proteins bound to nucleic acids, as described in WO 2014 / 114687. Therefore, further embodiments and applications of the methods of this invention can be found in these applications, which are incorporated herein by reference in their entirety.
[0165] According to a preferred embodiment of the invention, the method includes enriching SNPs or genetic chimeras contained within isolated target regions. In some cases, the SNPs or genetic chimeras are located within the target region itself (and therefore not within the gRNA recognition site in the class 2 type V Cas protein-gRNA complex). However, in other cases, the SNPs or genetic chimeras may be located at the gRNA recognition site in the class 2 type V Cas protein-gRNA complex. In one specific embodiment, the gRNA contains nucleotide bases corresponding to the minor allele of the SNP. When multiple SNP alleles are present at a given locus, multiple gRNA molecules can be provided, corresponding to each allele, preferably each minor allele. When gRNA molecules corresponding to major and minor alleles are provided, the number of isolated target regions containing each allele can be quantified, for example, to determine whether a subject is homozygous or heterozygous at the SNP locus. Preferably, the bases corresponding to the SNP locus are located within the gRNA sequence at bases -1 to -10 relative to the PAM site, preferably -1 to -6, preferably any one of -4, -5, or -6. In fact, when one or more of these bases mismatch, the hybridization of the gRNA with the region containing the SNP locus is reduced. In this case, the protection of the nucleic acid region from exonuclease digestion is also reduced or eliminated. This localization is particularly advantageous because the presence or absence of the SNP can be determined by the reduced probability of error. In some cases, the target nucleic acid region is sequenced to determine alleles at the SNP locus. This sequencing is especially important when using gRNA containing degenerate bases at the SNP locus to identify alleles that may be present at adjacent SNP loci within the target region, or when the SNP is contained within the target region itself (i.e., the gRNA is complementary to a site adjacent to the target region). Indeed, as is well known to those skilled in the art, SNPs that are close to each other in the genome tend to be inherited together.
[0166] The extent to which target region cleavage is reduced or eliminated will vary depending on experimental conditions, the two classes of Cas proteins used, and / or the gRNA used. For example, Cas12a is known to have higher binding specificity than Cas9 (Strohkendl et al., Molecular Cell, 2018, 71: 1-9). Therefore, when using Cas9, the hybridization specificity, or protection against mismatched regions, will be greater than when using Cas12a. Depending on whether it is necessary to isolate mismatched regions, Cas12a proteins or their variants with optimized binding specificity can be used in particular.
[0167] According to a preferred embodiment of the invention, the method may further include sequencing the isolated target region. Many sequencing methods are available in the art. The method of the present invention is particularly suitable for generating hairpins for single-molecule sequencing methods, such as those described in WO 2011 / 147931 or WO 2011 / 147929. The isolated nucleic acid can be further used as a template for specific or nonspecific polymerase chain reactions, isothermal amplification such as loop-mediated isothermal amplification, strand displacement amplification, helicase-dependent amplification, nickase amplification reactions, reverse transcription, enzymatic digestion, nucleotide incorporation, oligonucleotide ligation, and / or strand invasion. Isolated nucleic acids can also be used as substrates for sequencing, such as Sanger dideoxy sequencing or chain termination, whole genome sequencing, hybridization-based sequencing, pyrosequencing, capillary electrophoresis, cyclic sequencing, single-base extension, solid-phase sequencing, high-throughput sequencing, massively parallel labeled sequencing, nanopore-based sequencing, transmission electron microscopy sequencing, optical sequencing, mass spectrometry, 454 sequencing, reversible terminator sequencing, "paired ends" or "paired" sequencing, exonuclease sequencing, ligation sequencing (e.g., SOLiD technology), short read sequencing, single-molecule sequencing, chemical degradation sequencing, sequencing-by-synthesis, massively parallel sequencing, real-time sequencing, semiconductor ion sequencing (e.g., Ion Torrent), dual-end dual-label multiplex sequencing (MS-PET), droplet microfluidic sequencing, partial sequencing, fragment mapping, and any combination of these methods.
[0168] According to a preferred embodiment, the method of the present invention further includes sequencing the target region by single-molecule sequencing, next-generation sequencing, partial sequencing, or fragment mapping, more preferably by single-molecule sequencing as described in WO 2011 / 147931 or WO2011 / 147929. According to a preferred embodiment of the invention, the method may further include detecting the binding of a protein to a specific nucleic acid sequence or site. A variety of methods for detecting protein binding can be used by those skilled in the art. The method of the present invention is particularly suitable for generating hairpins for use in single-molecule protein binding methods, such as those described in WO 2014 / 114687. The isolated target region can be further used as a substrate for detecting protein-nucleic acid binding, for example, as a substrate for detecting epigenetic modifications. The isolated target region can be used, for example, in bisulfite conversion, high-resolution melting analysis, immunoprecipitation (e.g., ChIP, enChIP), microarray hybridization, and other analyses of nucleic acid / protein interactions well known to those skilled in the art. As used herein, the term "epiggenetic modification" refers to modifications of the bases constituting the nucleic acid molecule that occur after the synthesis of the nucleic acid molecule. As a non-limiting example, base modifications may result from damage to said bases. Epigenetic modifications include, for example, especially, 3-methylcytosine (3mC), 4-methylcytosine (4mC), 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxycytosine (5caC) and 6-methyladenosine (m6A) in DNA, 5-hydroxymethyluracil (5hmU) and pseudouridine in RNA, and 3-methylcytosine (3mC) and N6-methyladenosine (m6A) in both DNA and RNA.
[0169] Similarly, this method can further include the detection of modified bases caused by nucleic acid damage, such as DNA damage. DNA damage occurs continuously due to chemical (i.e., intercalating agent), radiation, and other mutagenesis that may occur on isolated nucleic acids. DNA base modifications caused by these types of DNA damage are widespread and play an important role in influencing physiological states and disease phenotypes. Examples include 8-oxoguanine, 8-oxoadenine (oxidative damage; aging; Alzheimer's disease; Parkinson's disease), 1-methyladenine, 6-O-methylguanine (alkylation; glioma and colorectal cancer), benzo[a]pyrene glycol epoxide (BPDE), pyrimidine dimers (adduct formation; smoking, industrial chemical exposure, ultraviolet radiation exposure; lung cancer and skin cancer), and 5-hydroxycytosine, 5-hydroxyuracil, 5-hydroxymethyluracil, and thymine glycol (ionizing radiation damage; chronic inflammatory diseases, prostate, breast cancer, and colorectal cancer).
[0170] Preferably, the method of the present invention further includes detecting the binding of a protein to a specific nucleic acid sequence, as described in WO2014 / 114687.
[0171] Reagent test kit
[0172] Another object of the present invention is a kit for nucleic acid isolation that can be used according to any method or embodiment of the invention described herein. This kit will provide materials and methods for nucleic acid isolation and enrichment according to the invention, as described above. Therefore, the kit will include the materials necessary for isolating nucleic acids according to the methods described herein. According to any of the methods described herein, the kit contents may vary depending on the type V Cas proteins to be used (e.g., Cas12a, C2c1), the targeted nucleic acid region, the capture method, etc.
[0173] According to a particular embodiment, the kit of the present invention comprises:
[0174] a) Two types of V-type Cas proteins, preferably Cas12a with catalytic activity.
[0175] b) At least one gRNA, said gRNA being complementary to a site in a neighboring nucleic acid target region.
[0176] c) At least one 3' to 5' single-stranded enzyme with exonuclease activity, preferably exonuclease I.
[0177] d) Oligonucleotide probes,
[0178] e) Optionally, at least one protease, and
[0179] F) Optional, use notifications.
[0180] According to a further embodiment, the kit contains EDTA (preferably a solution of EDTA) instead of at least one protease, or the kit contains EDTA in addition to at least one protease.
[0181] In some cases, the kit may further contain two types of Cas proteins, preferably a second type 2 V Cas protein. Preferably, the kit contains at least two gRNAs, more preferably at least three, four, five, six, or ten gRNAs, each of which is complementary to a specific nucleic acid region.
[0182] According to a particular embodiment, the kit contains two gRNAs for each target region, wherein the gRNAs are complementary to first and second sites located on either side of the target region, as described above. Using two class 2 type V Cas protein-gRNA complexes is advantageous because this allows for the creation of 5' overhangs on either side of the target region.
[0183] In cases requiring downstream multiplexing, the kit may contain two or more gRNAs, each gRNA being at least partially complementary to a site adjacent to a different target region. According to a particular embodiment, the kit contains two types of V Cas proteins and two types of II Cas proteins, more preferably Cas9, and corresponding appropriate gRNAs, each gRNA being at least partially complementary to a first and second site located flanking the target region. Alternatively, the two gRNAs of the two types of V Cas proteins may recognize two distinct sites located flanking the target region described herein. In some cases, the kit contains one or more Cas proteins as described herein, which have been preloaded with gRNA to form one or more Cas protein-gRNA complexes, preferably one or more type II V Cas protein-gRNA complexes. According to a particular embodiment, when the kit contains multiple type II V Cas protein-gRNA complexes, the complexes are preferably mixed together in a single container. Preferably, the ratio of each Cas protein-gRNA complex contained in the kit has been predetermined for ease of use.
[0184] Preferably, the guide region of the gRNA is complementary to a region adjacent to a target region of interest in clinical diagnosis or genetic risk assessment. For example, the gRNA is complementary to a site adjacent to a non-coding target region downstream of the coding region of septin 9 (SEPT9) or epidermal growth factor receptor (EGFR). Indeed, the epigenetic status of these regions is known to be important for cancer outcomes. As another example, the gRNA is complementary to a site adjacent to a downstream region of the gene FMR1, which is involved in Fragile X syndrome. An increase in the copy number of the 5'-CGG-3' repeat sequence in this gene is a cause of the disease. The epigenetic status (e.g., methylation) of the upstream region of this CpG island is also known to be associated with the clinical severity of the disease. As yet another example, the gRNA is complementary to a site adjacent to the DMPK gene. Indeed, an increase in the number of the 5'CTG-3' repeat sequence in this gene is a characteristic of type 1 ankylosing dysplasia. As a further example, the gRNA is complementary to a region located within a cfDNA molecule, thereby enabling the isolation of adjacent target regions contained within the cfDNA molecule. In fact, the isolation of specific cfDNAs (e.g., cffDNA or ctDNA) is of particular importance in a variety of downstream applications, including prenatal testing (e.g., see Gahan, Int J Womens Health. 2013, 5: 177–186), and cancer diagnosis and / or monitoring (e.g., see Ghorbian and Ardekani, Avicenna J Med Biotech. 2012, 4(1): 3–13). One or more target regions contained within cfDNA can be advantageously isolated directly from biological samples (e.g., plasma, serum, or urine samples).
[0185] The kits described herein are preferably capable of isolating at least two distinct target regions. Indeed, the value of multiplexing in improving certain epigenetic cancer diagnostic tests has been demonstrated, where the sequence or structural features (e.g., methylation status) of two or more distinct target regions are analyzed in a single test. As a non-limiting example, the kits provided herein are capable of isolating target regions comprising, or composed of, the human GSTP1, APC, and / or RASSF1 genes or their appropriate regions methylated with DNA, according to any of the methods described herein. The isolated target regions can then be subjected to downstream analysis of methylation status, for example, according to the methods provided herein (e.g., those provided in WO2014 / 114687). Such kits are particularly advantageous in determining the risk of prostate cancer in subjects (Wojno et al., American health & drug benefits, 2014, 7(3): 129), and are superior to existing kits, especially those that use bisulfite treatment of sample DNA followed by PCR. In contrast to the method of the present invention, nucleic acids isolated using existing kits may be particularly prone to false positive and false negative signals, as well as sample loss due to harsh and inefficient chemical processing.
[0186] According to a specific implementation, the kit contains two gRNAs for each target region, said gRNAs being complementary to sites flanking the human genes GSTP1, APC, and RASSF1 as defined herein.
[0187] As another non-limiting example, the kit of the present invention is capable of isolating at least one of the following target regions located in the human genome: 65676359-65676418 on chromosome 17, 21958446-21958585 on chromosome 9, 336844-336903 on chromosome 6, 33319507-33319636 on chromosome 21, 166502151-166502220 on chromosome 6, 896902-897031 on chromosome 18, and 32747873-327480 on chromosome 5. The target regions are preferably all 15: 27949195-27949264 on chromosome 6, 27191603-27191672 on chromosome 7, 170170302-170170361 on chromosome 16, 30797737-30797876 on chromosome 15, 7936767-7936866 on chromosome 1, 170077565-170077634 on chromosome 1, 1727592-1727661 on chromosome 2, and 72919092-72919231 on chromosome 8. Isolation of these target regions is advantageous because downstream analysis of the DNA methylation status of these target regions can be used for the detection of bladder cancer. Existing kits use methylation-sensitive restriction enzymes followed by PCR to recognize methylated sequences, which can be limited by the presence of appropriate restriction sites in the target regions, complicating test design and limiting sensitivity. Therefore, an improved kit for isolating and detecting bladder cancer preferably comprises two gRNAs at least partially complementary to sequences flanking the target regions, and two additional gRNAs at least partially complementary to sequences located at the ends of said target regions, for isolating each of the 15 target regions according to the method described herein, preferably for isolating all 15 target regions.
[0188] As a non-limiting example, the target region may contain or not contain a specific sequence, the number of repeating sequences of the specific sequence, or one or more nucleotide base modifications. As a further non-limiting example, the target region used for isolation may have a specific length or a length different from said specific length. Preferably, the kit of the present invention also contains at least one restriction enzyme and / or RNase. Preferably, the kit also contains suitable type 2 type V Cas protein reaction buffer and suitable type 2 type II Cas protein reaction buffer, such as those detailed in the following examples.
[0189] The kit may further include additional components suitable for a given application. For example, the kit may also contain one or more hairpin aptamer or site-specific endonucleases, ligases and / or polymerases, oligonucleotides, dNTPs, suitable buffers, etc. According to one specific embodiment, the kit further includes a ligase and / or a 5' valve endonuclease, preferably FEN1.
[0190] Additional features and advantages of the invention are illustrated in the following figures and embodiments. Attached Figure Description
[0191] Figures 1A to 1C Adding ExoI to the molecular population bound to the Cas12a-gRNA complex produces longer 5' overhangs with less variability at the cleavage site. Figure 1A Fragments containing the target sequence guided by SEPT9.2 crRNA#1 (SEQ ID NO: 2) (SEQ ID NO: 3) were amplified by PCR using primers PS1462 and PS1464, with FITC fluorescence labeled at the 5' ends of the primers (SEQ ID NO: 4 and 5, respectively). The fragments were then incubated with the Cas12a-gRNA complex alone to determine the cleavage site on the non-target strand (left); or simultaneously incubated with the Cas12a-gRNA complex and ExoI to determine the number of cryptic bases at the 3' end via ExoI treatment (middle); or incubated with the Cas12a-gRNA complex and then filled in with T4 DNA polymerase to determine the cleavage site on the strand hybridized to the gRNA (right). Figure 1B )pass( Figure 1A The representative trajectories obtained from capillary electrophoresis of the three experimental setups described in [the text] allow us to determine the position of the 3' end of the non-hybridized strand and the cleavage position within the 5' hybridized strand at single-base resolution. Each reaction incorporated PCR fragments of known sizes (129, 273, and 503 bp) labeled with FITC (corresponding to SEQ ID NO: 6 to 8) as markers. The trajectories from the three different experiments are compared based on the 129 and 273 bp peaks. The arrow pointing to the left in the middle figure indicates that the fragment was shortened by adding ExoI (thus migrating closer to the lower-labeled 129 base). The arrow pointing to the right in the bottom figure indicates that the 5' overhang was filled in by T4 DNA polymerase (thus migrating closer to the upper marker, at 273 bp). Figure 1C ).based on( Figure 1AThe experimental setup provided in [the document] determined the cleavage site of SEPT9.2 crRNA#1 complexing with LbaCas12a(NEB). The location indicated in the figure is provided relative to the first base of the gRNA sequence (underlined), which begins with the first base adjacent to the PAM site.
[0192] Figure 2 : Schematic diagram of steps a) through c) of the method of the present invention using a single Cas12a-gRNA complex. (A) The nucleic acid molecular group may optionally be mechanically or enzymatically fragmented to produce random fragments. (B) After fragmentation, the nucleic acid molecular group is contacted with a Cas12a-gRNA complex containing a guide segment that is at least partially complementary to a site located in a neighboring target region. During incubation, any single-stranded region is targeted for digestion by ExoI treatment. (C) The Cas12a-gRNA complex is then removed (referred to as “purification” in the diagram). The resulting fragment contains a longer and more defined 5' overhang, which may subsequently be processed by steps d) and e) of the method of the present invention (not shown here).
[0193] Figure 3 : A schematic diagram of steps a) through c) of the method of the present invention, wherein the Cas12a-gRNA complex and hairpin aptamers are used as alternative strategies. After the first step of optional fragmentation and end repair (A), hairpins (or other) aptamers can be ligated to all free ends, for example using T / A clones (B). (C) After fragmentation, the nucleic acid molecule population is contacted with the Cas12a-gRNA complex containing a guide segment that is at least partially complementary to a site located in a neighboring target region. During incubation, any single-stranded region is targeted for digestion by ExoI treatment. Only sites containing sequences complementary to the gRNA are targeted for this treatment, and therefore only these sites will contain single-stranded 5' overhangs. (D) The Cas12a-gRNA complex is then removed. The resulting nucleic acid molecule population containing 5' single-stranded overhangs can then be subjected to steps d) and e) (not shown here) of the method of the present invention to isolate the nucleic acid target region.
[0194] Figure 4: Schematic diagram of steps a) to c) of the method of the present invention, wherein two Cas12a-gRNA complexes are used as an alternative strategy. (A) Illustration of gRNA recognition sites of the Cas12a-gRNA complexes, wherein the gRNAs are designed such that the PAM sequence is located adjacent to the target region, as indicated by the inward-pointing arrows. (B) Two Cas12a-gRNA complexes target the first and second sites on the flanking sides of the target region, respectively, on the first and second sides of the target region. The two Cas12a-gRNA complexes may be used simultaneously with ExoI treatment to generate 5' overhangs on both sides of the target region to be separated. Fragmentation is not described here because this step is optional. (C) The Cas12a-gRNA complexes are then removed. The resulting population of nucleic acid molecules containing 5' overhangs can then be subjected to steps d) and e) (not shown here) of the method of the present invention to separate the nucleic acid target region, which can then be used in downstream applications, such as, but not limited to, cloning, library preparation, or hairpin generation.
[0195] Figure 5 The target region can be separated using two oligonucleotide probes located on either side of it. (A) to (C) according to Figure 4 The procedure is as illustrated in the diagram. (D) After treatment with Cas12a and ExoI, the nucleic acid molecule clusters are contacted with synthetic oligonucleotide probes containing various ligands (e.g., but not limited to biotin and digitoxin), which are complementary at their 5' ends to sites at least partially complementary to Cas12a gRNA. The 5' lobes of the oligonucleotide probes are removed using the FEN1 endonuclease, and the resulting nicks are closed with a ligase. This allows for covalent ligation of the oligonucleotide probes at each end of the molecule. (E) The target region can then be isolated as described herein. (F) The isolation steps can be repeated using a second ligand. This advantageously increases the specificity of the method. (G) The resulting molecules can be used for downstream applications, such as hairpin or library generation.
[0196] Figure 6: Schematic diagram of a method for producing molecules (hairpin structures) suitable for SIMDEQ instruments from *E. coli* genomic DNA. (A) Illustration of gRNA recognition sites via the Cas12a-gRNA complex, wherein the gRNA is designed such that the PAM sequence is located adjacent to the target region, as indicated by the inward-pointing arrows. (B) Two distinct Cas12a-crRNA complexes (SEQ ID NO: 11 to 14) are flanked by the regions of interest (targets #1 and 2, SEQ ID NO: 9 and 10), which bind to the first and second sites, respectively. Simultaneously, 3' to 5' single-stranded exonucleases are added. This generates 5' overhangs on both sides of the target region. (C) The Cas12a-gRNA complex is then removed (referred to as "purification" in the diagram). Preferably, salts and any other proteins are also removed. (D) Oligonucleotide probes PS1466 and PS1468 (SEQ ID NO: 15 and 16) contain the necessary sequences (hence the designation as surface oligonucleotides) that allow hybridization with the support, and these oligonucleotide probes have a sequence at their 5' end corresponding to the Cas12a crRNA sequence (thus they are complementary to the 5' overhang of the adjacent target region). The reaction containing these oligonucleotides is supplemented with FEN1 and Taq DNA ligase to remove the 5' flap region of the oligonucleotides and close the nicks created by 5' flap digestion. (E) A second oligonucleotide (containing biotin (indicated by a square) at its 5' end) with a sequence complementary to the single-stranded region of the oligonucleotide probe is added to the reaction tube. The second oligonucleotide is PS1467 and PS1469 (SEQ ID NO: 17 and 18 for targets #1 and #2, respectively). The gap between the second oligonucleotide and the duplex is filled and ligated by DNA polymerase. (F) The resulting fragment is digested with BsaI, and a synthetic hairpin is attached to the other side of the target region. In some cases, step F can be performed as the first step of this method (i.e., before generating 5' overhangs with the Cas12a-gRNA complex and ExoI). The desired molecules are then isolated from the nucleic acid population using streptavidin beads and loaded onto the SIMDEQ platform. A typical fingerprint trajectory of the hairpin generated by this method is shown below. Figure 10 A and Figure 10 As shown in B.
[0197] Figure 7 Alternative schematic diagrams of methods for producing molecules suitable for SIMDEQ instruments (hairpin structures) from E. coli genomic DNA. (A) to (C) use with Figure 6The same method is shown in (D). In (D), a hairpin is attached to a second site in addition to contacting the nucleic acid molecule cluster with the oligonucleotide probe. In both cases, any 5' lobes can be removed by FEN1, and the remaining nick is closed by ligation. This is advantageous because two sites present on either side of the target region can be processed simultaneously, thus reducing the duration of the method.
[0198] Figure 8 Exemplary structures of oligonucleotide probes used in this invention. The gRNA-targeted sequence is shown in bold and capital letters. The 5' end of the oligonucleotide probe contains a sequence at least partially complementary to the gRNA-targeted sequence, corresponding to the 5' overhang. This sequence can be extended to the PAM sequence to address variability in the Cas12a-gRNA cleavage site on the non-target strand. Any valve structure can be removed via FEN1 processing. The single-stranded region of the probe contains a "universal" sequence, a single-stranded oligonucleotide (e.g., required to generate a hairpin). Figure 6 (E) Biotin oligonucleotides can bind to this sequence. Optionally, the oligonucleotide probe may further include a specific sequence for anchoring the probe (and thus the target region) to a solid support (e.g., a flow cell surface). In some cases, a "universal" sequence may be included in the loop region of the hairpin molecule.
[0199] Figure 9 The 5' overhang generated by treatment with Cas12a-gRNA and ExoI facilitated hybridization of oligonucleotide probes. A 287 bp PCR fragment (SEQ ID NO: 19) containing a site recognized by human NDGR4.1 crRNA#1 (SEQ ID No: 20) was FITC-labeled at the 5' end of the non-target strand containing the 5'-TTTV-3' PAM sequence and analyzed by capillary electrophoresis after incubation under the following conditions: The fragment was incubated with (A) LbaCas12a crRNA#1 in the presence of FEN1 and Taq DNA ligase, (B) LbaCas12a crRNA#1 and ExoI, (C) LbaCas12a crRNA#1 and ExoVII, (D) LbaCas12acrRNA#1 and ExoVII in the presence of FEN1 and Taq DNA ligase, or (E) LbaCas12a crRNA#1 and ExoI in the presence of FEN1 and Taq DNA ligase. In all cases, oligonucleotide probes containing biotin ligands are also available, which can be processed by FEN1 and ligated to the target site via Taq DNA ligase. The ligation product will increase the size of the fragment (marked with an asterisk) (SEQ ID NO: 21), bringing the peak closer to the 273 bp marker. Also incorporated... Figure 1BThe same FITC markers were used for trajectory alignment. In particular, trajectories were aligned based on 129 and 273 bp fragments.
[0200] Figure 10 This is a schematic diagram of a qPCR quantification assay performed to determine the efficiency of Cas12a-gRNA cleavage and ligation of oligonucleotide probes using FEN1 and ligase. As shown, three different primer sets (A / B, A / C, and A / D, SEQ ID NO: 94-97) were designed near the Cas12a-crRNA#2 cleavage site (SEQ ID NO: 29) close to FMR1. A TaqMan probe (SEQ ID NO: 93) was designed for use in any of the three qPCR reactions, allowing for quantification of cleavage and ligation efficiency of the oligonucleotide probe with genomic DNA. Each reaction was normalized (also referred to as “total”) using the products obtained with primers A and B, which amplified fragments regardless of whether Cas12a-crRNA#2 cleaved the nucleic acid molecule. To quantify the Cas12a cleavage efficiency, a second primer C was designed located outside the target region to be enriched. Therefore, when the Cas12a-crRNA#2 complex cleaves its target, this fragment can no longer serve as a template for amplification, resulting in a reduction in amplifiable starting material. To estimate the cleavage rate, we calculated the ratio of the remaining fragments after incubation with the Cas12a-crRNA#2 complex and divided that number by the calculated "total". This provides the amount of uncleaved fragments and can be converted to the amount of cleaved fragments by taking a fractional return (i.e., 100 - amount of uncleaved fragments = number of cleaved fragments). Finally, to quantify the ligation efficiency of the oligonucleotide probe with the FMR1 fragment, primers A and D were used, which amplify fragments specific to the ligated oligonucleotides. The portion obtained using primers A / D is also referred to as the "FEN1 ligation".
[0201] Figure 11 The cleavage efficiency of the Cas12a-gRNA complex varied with reaction time and concentration. To determine the cleavage efficiency of the Cas12-crRNA#2 complex as a function of reaction time, 2 μg of human genomic DNA was incubated with 600 fmol of FMR1-specific Cas12a-crRNA#2 gRNA (SEQ ID NO: 29) in NEB2.1 buffer supplemented with 10 mM DTT at 37°C for 10, 20, 30, and 60 minutes (left panel). qPCR was performed using two sets of primers (A / B and A / C, SEQ ID NO: 94-95 and 94-96, respectively) with a Taqman probe (SEQ ID NO: 93). A control reaction was performed in the absence of the Cas12a-crRNA#2 complex and used as an uncleaved reference. The cleavage efficiency was determined by the "total amount" of the material (primer sets A / B; see also...). Figure 10 The relative quantification of the cleaved fragments (primer set A / C) was normalized using the legend and then normalized against an uncleaved reference. 78% cleavage of the target was achieved when the reaction time was increased from 10 min to 30 min. The cleavage efficiency (68%) of the Cas12a-crRNA#2 complex was also evaluated by increasing the amount of the complex. 2 μg of human genomic DNA was incubated with 300, 600, 1200, and 1500 fmol of FMR1-specific Cas12a-crRNA#2 guide RNA (SEQ ID NO: 29) in NEB2.1 buffer supplemented with 10 mM DTT at 37°C for 10 min (right figure). Doubling the amount of the Cas12a-crRNA#2 complex from 600 to 1200 fmol increased the amount of cleaved target to 84%.
[0202] Figures 12A to 12B Adding exonuclease I does not affect the cleavage of Cas12a-gRNA, but it improves the binding of the probe to FEN1. Figure 12A Two μg of human genomic DNA was incubated with 600 fmol of FMR1-specific Cas12a-crRNA#2 gRNA (SEQ ID NO: 29) in NEB2.1 buffer supplemented with 10 mM DTT at 37°C for 30 min, with or without exonuclease I to increase the number of crypted bases at the 3' end. qPCR was performed using two sets of primers (A / B and A / C, SEQ ID NO: 94 / 95 and 94 / 96, respectively) with a Taqman probe (SEQ ID NO: 93). A control reaction performed in the absence of the Cas12a-crRNA#2 complex served as an uncut reference. The relative quantification of the cut fragments (primer set A / C) was normalized by the "total amount" of the material (primer set A / B), and then normalized by the uncut reference. The presence of exonuclease I did not affect the cleavage of Cas12a-crRNA#2; in the study, DNA was cleaved at 78% and 76% in the presence and absence of exonuclease I, respectively. Figure 12B The cleavage reaction was performed using Cas12a-crRNA#2, with or without exonuclease I. Figure 12AAs described, the reaction was incubated in ThermoPol buffer supplemented with 1 mM NAD, 40 units of Taq DNA ligase, 32 units of FEN1, and 10 pmol of a target-specific oligonucleotide probe (SEQ ID NO: 98), along with a 20-base sequence complementary to the crRNA recognition site and containing a known sequence for qPCR. The reaction was incubated at 37°C for 30 min. A control reaction performed in the absence of the Cas12a-crRNA#2 complex was used as a reference for unligated fragments. The relative quantification of probe-ligated fragments (primer A / D, SEQ ID NO: 94 / 97) was normalized by the “total amount” of materials (primer A / B, SEQ ID NO: 94 / 95), and then normalized by a “standard reference reaction” (a 30-minute cleavage reaction with Cas12a-crRNA#2 at 37°C in the presence of exonuclease I and FEN1). The efficiency of FEN1 and probe ligation was increased by two-fold in the presence of exonuclease I.
[0203] Figure 13The effect of temperature on FEN1 processing and oligonucleotide probe ligation efficiency. 2 μg of human genomic DNA was incubated with 600 fmol of FMR1-specific Cas12a-crRNA#2 gRNA (SEQ ID NO: 29) in NEB2.1 buffer supplemented with 10 mM DTT at 37°C for 30 min, with or without exonuclease I, to determine the effect of the length of the increased 5' single-stranded overhang on the ligation efficiency of the target-specific oligonucleotide probe. With or without exonuclease I, the Cas12a-crRNA#2 cleavage reaction was incubated with 10 pmol of oligonucleotide probe oligonucleotide (SEQ ID NO: 98) in ThermoPol buffer supplemented with 1 mM NAD, 40 units of Taq DNA ligase, and 32 units of FEN1 (or without FEN1) for qPCR. The oligonucleotide probe had a 20-base sequence complementary to the site recognized by crRNA#2 and contained a known sequence. Three incubation temperatures were tested to observe probe ligation efficiency at 37°C, 45°C, and 55°C for 30 minutes. qPCR reactions were performed using two sets of primers (A / B and A / D, SEQ ID NO: 94 / 95 and 94 / 97, respectively) and a Taqman probe (SEQ ID NO: 93). Relative quantification of the probe ligation fragments (primer set A / D) was normalized by the “total amount” of the material (primer set A / B), and then normalized by our “standard reference reaction” (Cas12a cleavage at 37°C for 30 minutes in the presence of exonuclease I and FEN1). For all tested temperatures, the presence of exonuclease I increased ligation efficiency by 2-fold. Furthermore, compared to 37°C, ligation efficiency and / or FEN1 activity were increased by 1.8-fold when the reaction was incubated at 45°C.
[0204] Figure 14A schematic diagram of a method for generating molecules (hairpin structures) suitable for SIMDEQ instruments from human genomic DNA. (A) shows oligonucleotide probes (SEQ ID NO: 100 and 99) specific for crRNA#1 and crRNA#2. The first probe contains the necessary sequence (hence the name surface oligonucleotide) that allows hybridization with the support and has a sequence corresponding to the Cas12a crRNA#1 sequence at its 5' end (shown on the left side of the image), while the second probe contains biotin (indicated by a square) and a base-degrading site modification (indicated by a triangle) at its 5' end and has a sequence corresponding to the Cas12a crRNA#2 sequence at its 5' end, thus forming a double strand. The reaction containing these oligonucleotides is supplemented with FEN1 and Taq DNA ligase to remove the 5' flap region of each oligonucleotide and close the nicks produced by 5' flap digestion. (B) The DNA molecules were hybridized to a second set of oligonucleotides (SEQ ID NO: 101 and 97) complementary to the probes linked by FEN1 and ligase in (A). The DNA molecules were then filled in and ligated by DNA polymerase to produce continuous dsDNA molecules with intact target regions and a Y-shaped structure at one end. (C) The duplexes were then isolated from the nucleic acid molecule group using streptavidin beads. (D) The duplexes were eluted from the beads using endonuclease IV, which catalyzes the cleavage of the DNA phosphodiester backbone at the debasement site with a 3'-hydroxyl terminus. A 4 bp overhang was generated at the 5' end of the previously added oligonucleotides to generate the dsDNA molecules in (B). (E) A synthetic hairpin (SEQ ID NO: 102) was attached to this 5' overhang. The duplexes were isolated using streptavidin beads coated with specific adapters (SEQ ID NO: 103 and 104) and loaded onto a SIMDEQ platform.
[0205] Figures 15A to 15B .exist Figure 6 The method provided allows for the successful production of hairpin molecules containing target regions. Using... Figure 6 The method shown isolates target #1 from E. coli genomic DNA. Figure 15A Target #2 Figure 15B Following this, a tetrabasic oligonucleotide (CAAG) fingerprinting assay was performed. All expected peaks were correctly identified in both targets using CAAG tetrabasic oligonucleotides.
[0206] Figures 16A to 16B Exemplary results of Illumina sequencing of 15 different human targets performed on libraries prepared from human genomic DNA using the method of the present invention. Figure 16A Target regions with corresponding genomic coordinates were selected for enrichment. These regions were chosen either because of the presence of epigenetic biomarkers or because loci with extended repeats are known to be involved in human diseases. Figure 16B A screenshot of read alignment obtained from Illumina sequencing after enriching the target region using the method of this invention. Specifically, using... Figure 5 The method shown generates the library, except that it uses only one oligonucleotide probe containing a ligand (biotin in this case) (therefore step F is omitted). The black box represents the enriched target region corresponding to the SNCACpG islands (chromosome 4: from 89,836,538 to 89,837,940 bp). Reads for each mapping are represented by gray horizontal lines. As shown in dark gray, this region shows high coverage (up to 614x coverage) compared to the surrounding areas (showing coverage less than 1x). Detailed Implementation
[0207] Example
[0208] The following embodiments are included to illustrate preferred embodiments of the invention. All subject matter set forth or illustrated in the following embodiments and drawings should be interpreted as illustrative rather than restrictive. The following embodiments include any substitutions, equivalents, and modifications that can be determined by those skilled in the art.
[0209] Example 1: Methods for selecting gRNA
[0210] For all the strategies described below, one or more guide RNAs are designed using available online tools. The RNA guides can then be synthesized in vitro using a viral transcription system (e.g., T7, SP6, or T3 RNA polymerase), or chemically generated using an automated synthesizer as a single crRNA guide, containing a sequence complementary to the target region. The efficiency of each gRNA is evaluated in vitro using wild-type Cas nucleases on standardized / control samples (e.g., PCR fragments) to ensure efficient cleavage of each Cas protein-gRNA complex (e.g., cleavage of at least 80% of the initial PCR fragment).
[0211] In this embodiment, an automated synthesizer can be used to chemically generate Cas12a guide RNA based on a common universal sequence (SEQ ID NO: 1).
[0212] gRNA was incubated in the appropriate buffer at 95°C for 5 minutes, and then incubated at 80°C, 50°C, 37°C and room temperature for 10 minutes using a progressive gradient. Each step was used for annealing and / or secondary structure formation.
[0213] Example 2: Reaction protocol for isolating nucleic acid molecules containing target regions
[0214] 1. Guide RNA (e.g., crRNA) is loaded onto type V Cas proteins by incubating for 10 minutes at room temperature (e.g., 25°C) in an appropriate reaction buffer (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 10 mM DTT, 100 μg / ml BSA, pH 7.9) to form a protein-RNA complex.
[0215] 2. Add the loading complex prepared in step 1 to a sample containing nucleic acid molecules in NEB buffer 2.1 supplemented with an additional 10 mM DTT. Add 1 μl of Exo I (100 units) to the reaction tube and incubate at 37°C for at least 1 hour. This will allow the type V Cas protein-gRNA complex to bind to and cleave the nucleic acid at a site adjacent to the target region, while Exo I clarifies the 3' end of the non-target strand.
[0216] 3. The reaction is stopped by adding a "stop buffer" (a mixture of 1.2 units of proteinase K and 20 mM EDTA), thereby removing the class V Cas protein-gRNA complex. In some cases, RNase A can be added to digest the gRNA. RNase A and proteinase K treatment can be performed continuously at 37°C for 15 minutes. In this case, the addition of EDTA is optional.
[0217] 4. The sample can then be purified using any known technique, such as bead or column purification, for example, purification with paramagnetic beads.
[0218] 5. Then, the eluted DNA is mixed with oligonucleotide probes (in the case of only one target) or multiple probes simultaneously (in the case of multiple targets) in ThermoPol buffer (20 mM Tris-HCl, 10 mM (NH4)2SO4, 10 mM KCl, 2 mM MgSO4, 0.1%). Incubate at 37°C for 30 minutes in X-100, 1 mM NAD+, pH 8.8 (NEB). If the second site is also processed by the second Cas12a protein, a probe containing a hairpin loop (or multiple probes, depending on the number of targets) can be added to the reaction mixture simultaneously. The concentration of each probe added is between 20 and 50 nM, depending on the starting material. The reaction can also be supplemented with 1 μl FEN1 (e.g., 32 units) and 1 μl Taq DNA ligase (e.g., 40 units).
[0219] 6. The sample can then be purified using any known technique, such as bead or column purification, for example, purification with paramagnetic beads.
[0220] 7. If the oligonucleotide probe does not contain a ligand, a universal oligonucleotide probe containing a 5' biotin ligand can be added to the DNA eluted from step 6 at a concentration of 100 nM to generate a Y-shaped structure. The gap between the oligonucleotide and the 5' end of the double-stranded region (and the hairpin loop, if used) is then filled using Bst full-length DNA polymerase (0.2 units, containing 0.2 mM dNTPs), and blocked by reacting with 1 mM NAD and 40 units of Taq DNA ligase in ThermoPol buffer (NEB) at 50°C for 30 minutes.
[0221] 8. The sample can then be purified using any known technique, such as bead or column purification, for example, purification with paramagnetic beads.
[0222] 9. At room temperature (25°C), incubate the sample with streptavidin-coated paramagnetic beads for 30 minutes in a binding buffer recommended by the magnetic bead manufacturer (e.g., but not limited to 0.5M NaCl, 20mM Tris-HCl (pH 7.5), 1mM EDTA).
[0223] 10. Then wash the beads with a recommended washing buffer (e.g., but not limited to 0.5M NaCl, 20mM Tris-HCl (pH 7.5), 1mM EDTA, with or without 0.5% Tween-20) for downstream applications such as sequencing or detection of epigenetic modifications.
[0224] Example 3: Determining the protrusion generated by LbaCas12a in conjunction with ExoI.
[0225] To identify the 5' overhangs generated by simultaneous treatment of nucleic acids with Cas12a and ExoI, primers PS1462 and PS1464 (SEQ ID NO: 4 and 5) were used with OneTaq DNA polymerase (NEB) to amplify a 221-base-pair (bp) PCR fragment (SEQ ID NO: 3) containing the site recognized by the SEPT9.2crRNA#1 sequence. Since PS1464 contains 5' FITC, the strand containing the PAM sequence 5'TTTV 3' was FITC-labeled. Using the same PS1464 reverse primer (and therefore also FITC-labeled), but with three different forward primers PS1461, PS1463, and PS1131 (oligonucleotides of SEQ ID NO: 22, 23, and 24, PCR fragments of SEQ ID NO: 6 to 8), using OneTaq (NEB), three additional PCR fragments of 128, 273, and 503 bp were generated. These three PCR fragments were used as markers to analyze the length of the resulting reaction fragments.
[0226] To determine the cleavage sites of LbaCas12a (NEB) on the target and non-target strands with and without ExoI treatment, 150 ng of a FITC-labeled 221 bp PCR fragment was incubated with an LbaCas12a:crRNA complex (DNA:Cas12a:crRNA ratio of 1:10:20), or simultaneously with the LbaCas12a complex and 25 units of ExoI. A third reaction was prepared in which the PCR fragment was incubated only with LbaCas12a:crRNA, and the 5' overhang was filled with T4 DNA polymerase (NEB) to determine the cleavage site on the 5' end of the target strand of the crRNA hybridization (see [link to reaction]). Figures 1A-1C The three resulting reactions were run on a capillary electrophoresis system (Abi3730) to resolve FITC-labeled fragments present in the reactions at single-base resolution. Three FITC-labeled PCR fragments of known size were incorporated into a spiked-in reaction as markers to determine the size of unknown fragments.
[0227] like Figure 1B As shown, ExoI treatment surprisingly allows for longer 5' overhangs. This contrasts sharply with 5' overhang lengths described in the art and those obtained here using only LbaCas12a (e.g., lengths as short as 6 bases, ranging from -13 to -19 relative to PAM). In fact, ExoI treatment replaces the cleavage on the non-target strand containing the PAM sequence with at least two additional bases, compared to the absence of ExoI at the cleavage site closest to PAM (where cleavage sites on the non-target strand can be further replaced by up to 5 bases). The 5' overhangs obtained with Cas12a-gRNA and ExoI treatment also surprisingly exhibit better definition. Figure 1B and Figure 1C As shown, less cleavage variability was observed in the non-target strand compared to the cleavage sites observed without ExoI treatment. The 5' overhangs generated by simultaneous treatment with LbaCas12a and ExoI were advantageously 11 or 12 nucleotides in length.
[0228] Example 4: The 5' overhangs generated by treatment with Cas12a-gRNA and ExoI can promote the hybridization of oligonucleotide probes.
[0229] A 287 bp PCR fragment (SEQ ID NO: 19) containing a site recognized by human NDGR4.1crRNA#1 (SEQ ID No: 20) was FITC-labeled at the 5' end of the non-target strand containing the 5'-TTTV-3' PAM sequence. The PCR fragment was first incubated with the LbaCas12a-crRNA#1 complex to generate overhangs. The reaction was stopped and purified. An oligonucleotide probe containing biotin at its 3' end and a sequence complementary to the 5' overhang at its 5' end was added to the eluted DNA and ligated to the PCR fragment by FEN1 treatment and Taq DNA ligase. Independent reactions were performed by incubation with the LbaCas12a-crRNA#1 complex in the presence or absence of exonuclease VII (ExoVII, with 3' to 5' and 5' to 3' single-stranded exonuclease activities) or ExoI (only 3' to 5' single-stranded exonuclease activity). Larger fragments resembling the 273 bp marker appeared in capillary electrophoresis, revealing successful ligation of the biotinylated oligonucleotide probe. Different trajectories were compared using the same FITC markers (129, 273, and 503 bp PCR fragments) as in Example 3. Five different experiments were compared based on the 129 and 273 bp fragments.
[0230] The percentage of successful ligation is calculated by dividing the fluorescence intensity of the corresponding ligation peak by the total fluorescence of all fragments with 5' overhangs. For example... Figure 9 As shown in (A), after incubating the PCR fragment with the LbaCas12a-crRNA#1 complex, oligonucleotide probe, FEN1, and Taq DNA ligase, approximately 24% of the resulting fragments corresponded to ligation products. Conversely, no ligation products were observed in the absence of FEN1 and Taq DNA ligase. Figure 9 As shown in (B), the ExoI treatment allows all the peaks present in A (corresponding to different ends generated by Cas12a activity) to split into two main peaks, corresponding to positions -9 and -10 from the PAM sequence. Although the ExoVII treatment also produces peaks corresponding to positions -9 and -10 from the PAM sequence (…), Figure 9 (C)), but the 3' end may not be well hidden because, in addition to the -9 and -10 peaks, there are additional peaks corresponding to Cas12a cleavage activity.
[0231] and Figure 9 The results obtained in (A) are the opposite; when the fragments were further incubated with exonucleases with 3' to 5' activity (either ExoI or ExoVII) in the presence of FEN1 and Taq DNA ligase, a higher percentage (i.e., approximately 50%) of the fragments now corresponded to the PCR products already ligated to the probe (see [link to PCR product]). Figure 9 (D) and Figure 9 (E) Therefore, when a group of nucleic acid molecules comes into contact with an enzyme having at least 3' to 5' exonuclease activity, the efficiency of isolating target nucleic acid regions using the method of the present invention is advantageously improved. In fact, since the oligonucleotide probe forms stable double strands with the 5' overhangs of a large number of fragments, the number of double strands (and the number of target regions) isolated by the method of the present invention will also be increased.
[0232] Example 5: Quantification of cleavage efficiency by qPCR.
[0233] We developed a quantitative PCR assay (qPCR, e.g.) Figure 10 (As shown), qPCR was used to determine the cleavage efficiency of Cas12a in the human genome. In short, qPCR using oligonucleotides A and B (SEQ ID NO: 94 and 95) allowed us to measure the amount of material present in the tube, while qPCR using oligonucleotides A and C (SEQ ID NO: 96) only produced a product when the Cas12a-crRNA complex did not cleave the target. To improve specificity, we included a TaqMan probe (SEQ ID NO: 93) within the amplicon, allowing the same probe to be used for both qPCR reactions. The percentage of molecules cleaved by the complex was determined by calculating the Ct ratio (ΔΔCt) of qPCR A / C versus A / B.
[0234] We incubated Cas12a-crRNA complexes containing the CGG repeat region targeting the human FMR1 promoter (SEQ ID NO: 57) for different durations at varying complex-to-human genomic DNA ratios to estimate the fractions of cleaved and uncleaved fragments. To prepare the complexes, AsCas12a protein was incubated with 2 molar amounts of FMR1-specific crRNA#2gRNA (SEQ ID NO: 29) in NEB2.1 buffer supplemented with 10 mM DTT (50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl2, 100 μg / ml BSA, pH 7.9, at 25 °C) at room temperature for 10 minutes.
[0235] In the first experiment, we estimated the time required for the Cas12a-crRNA complex to cleave its target in the human genome. 2 μg of human genomic DNA purified from the human embryonic kidney cell line (HEK293) was incubated with 600 fmol of the Cas12a-crRNA complex (a ratio of 640,000 Cas12a-crRNA complex per human genome) in NEB2.1 buffer supplemented with 10 mM DTT at 37°C for 10, 20, 30, and 60 minutes. As a control, we incubated the same amount of genomic DNA at 37°C under the same buffer conditions (but without the Cas12a complex). We performed qPCR reactions on all samples (triples) using primer pairs A and B, or A and C, and the quantitative methods described above. We observed that only 78% of the target was cleaved by this Cas12a-crRNA complex after 30 minutes. Figure 11 Increasing the incubation time to 60 minutes did not result in an increase in cutting efficiency.
[0236] Next, based on the obtained maximum cleavage efficiency, the optimal ratio of the Cas12a-crRNA complex per DNA molecule was determined for a given incubation time (10 minutes). 2 μg of human genomic DNA was incubated with 300, 600, 1200, or 1800 fmol of Cas12a complex (corresponding to ratios of 320,000, 640,000, 1,280,000, and 1,600,000 complexes per human genome, respectively, in NEB2.1 buffer supplemented with 10 mM DTT at 37°C for 10 minutes. Figure 11 As shown, the optimal ratio of the Cas12a-crRNA complex to the human genome was 1,280,000, as 84% of the cleavage target was obtained under these conditions. Further increasing the amount of complex (1800 pmol) did not improve cleavage efficiency. Since the efficiency did not increase above 84% (similar to an incubation time of 30 minutes at 600 pmol), we determined that the preferred reaction can be incubated for at least 30 minutes at a ratio of 640,000 Cas12a-crRNA complex per human genome.
[0237] Example 6: Quantifying the ligation efficiency of oligonucleotide probes at the overhangs generated by Cas12a with or without ExoI using qPCR.
[0238] To confirm the results obtained using capillary electrophoresis in Example 4, we performed qPCR assays to quantify the efficiency of FEN1 and oligonucleotide probe ligation to the overhangs generated by Cas12a-crRNA alone or by Cas12a-crRNA and exonuclease ExoI. To quantify this efficiency, two sets of primers and TaqMan probes (e.g., ...) were used. Figure 10 (As shown). Oligonucleotides A and B (SEQ ID NO: 94 and 95) were used as internal controls to quantify the amount of DNA present in the reaction tubes (to normalize all reactions), and oligonucleotides A and D (specific to the ligated probe; SEQ ID NO: 94 and 97) were used to determine ligation efficiency, as qPCR products were only amplified using oligonucleotides A and D when ligation was present. The relative quantification of the probe ligation fragment amount (primer set A / B, SEQ ID NO: 94 / 95) was normalized to the total amount of material (primer set A / D, SEQ ID NO: 94 / 97). The relative quantification was then normalized to 100% using a "standard reference reaction" (in the presence of ExoI and FEN1, at 37°C, with Cas12a cutting performed for 30 minutes). As described in these examples, this ratio was used for comparisons between different experimental conditions.
[0239] Initially, 2 μg of human genomic DNA purified from the human embryonic kidney cell line (HEK293) was incubated with 600 fmol of Cas12a complex in NEB2.1 buffer supplemented with 10 mM DTT at 37°C for 30 min, with or without 100 units of exonuclease I. As a control, the same amount of DNA was incubated at 37°C under the same buffer conditions (but without the Cas12a complex). The reaction was stopped using a stop buffer (1.2 units of proteinase K and 20 mM EDTA), and purification was then performed using paramagnetic beads according to the manufacturer's recommendations (KAPA beads, Roche).
[0240] The target-specific oligonucleotide (SEQ ID NO: 98) and 20 bases complementary to the site recognized by crRNA (therefore complementary to the 5' overhang generated by Cas12a and exonuclease I, see...) Figure 8 The reagents were added to reaction tubes containing ThermoPol buffer with Thermostable Flap endonuclease 1 (FEN1, NEB) and Taq DNA ligase (NEB) and incubated at 37°C for 30 minutes. The oligonucleotide probe also contained a universal sequence at its 3' end that was complementary to qPCR oligonucleotide D. All these reactions were quantified using primer pairs A / B and A / D in two separate tubes. Each qPCR reaction was performed in triplicate.
[0241] When exonuclease I was added to the reaction tube along with Cas12a-crRNA, no difference in cleavage efficiency was observed, because almost 80% of the fragments were cleaved under both conditions. Figure 12AHowever, in the absence of exonuclease I, a 50% decrease in the ligation efficiency of FEN1 and Taq DNA ligases for oligonucleotide probes was observed. Figure 12B The control reaction (without Cas12a) was used as a control for relative quantification. These results are consistent with those observed using capillary electrophoresis (from 24% efficiency without ExoI to 54% efficiency with ExoI treatment during Cas12a lysis). Figure 9 ).
[0242] Example 7: Effect of temperature on the ligation of target-specific oligonucleotide probes using FEN1 and Taq DNA ligase
[0243] To determine the effect of incubation temperature on the activities of FEN1 and the ligase, we tested three temperatures (37°C, 45°C, and 55°C) within their activity range and determined the efficiency of probe ligation when the overhangs generated by Cas12a cleavage were amplified by exonuclease I. To determine this effect, two reaction conditions were prepared by incubating 2 μg of human genomic DNA (HEK293) with 600 fmol of FMR1-specific Cas12a / crRNA#2 guide RNA (SEQ ID NO: 29) in NEB2.1 buffer supplemented with 10 mM DTT at 37°C for 30 min, with or without 100 units of exonuclease I. The reaction was stopped using a stop buffer (1.2 units of proteinase K and 20 mM EDTA), and purification was then performed using paramagnetic beads according to the manufacturer's recommendations (KAPA beads, Roche). 10 pmol of target-specific oligonucleotide (SEQ ID NO: 98) and 20 bases complementary to the site recognized by crRNA were added to a reaction tube containing 1 mM NAD, 40 units of Taq DNA ligase (NEB), and 32 units of Thermostable Flap endonuclease 1 (FEN1, NEB), and incubated at 37°C, 45°C, or 55°C for 30 minutes. A separate reaction without FEN1 was performed at each temperature to determine nonspecific amplification. Quantitative PCR was performed using two sets of primers (A / B and A / D, SEQ ID NO: 94-95 and 94-97, respectively) with a Taqman probe (SEQ ID NO: 93), and each qPCR reaction was performed in triplicate. The relative quantification of probe ligation fragments (primer sets A / D) was standardized by the total amount of material (primer set A / B), and then the quantification was standardized by a "standard reference reaction" (Cas12a cleavage in the presence of exonuclease I and FEN1, reaction at 37°C for 30 minutes), as described above (Example 6).
[0244] like Figure 13As shown, temperature did not affect the ligation improvement brought about by exonuclease I (exonuclease I doubled the ligation efficiency). However, when the reaction was incubated at 45°C, an overall increase in the efficiency of ligation and / or FEN1 activity was observed (probe ligation was 1.8 times greater than at 37°C).
[0245] Example 8: Constructing a hairpin from a separated target region
[0246] Using 1.6 μg of E. coli genomic DNA Figure 6 The method described herein and according to Example 2. Specifically, genomic DNA was incubated with 2 pmol of LbaCas12a-crRNA complex and regions of interest (targets #1 and 2, SEQ ID NO: 9 and 10), each region of interest flanked by two distinct Cas12a-crRNA complexes (SEQ ID NO: 11-14) that bind to the first and second sites, respectively. Using two Cas12a-gRNA complexes for each target region advantageously increases specificity and limits the number of steps in the method (e.g., fragmentation is not required and Cas12a-gRNA complexes can be added simultaneously). Furthermore, using two distinct Cas12a-gRNA complexes advantageously allows for the generation of two distinct, specific 5' single-stranded overhangs that can be used in subsequent steps to produce molecules suitable for the desired downstream application. The LbaCas12a-crRNA reaction was supplemented with 100 units of ExoI, which produced the 5' overhangs required for the method to function effectively. The reaction was stopped using a termination buffer (1.2 units of proteinase K and 20 mM EDTA), and then purified using paramagnetic beads according to the manufacturer's recommendations (KAPA beads, Roche).
[0247] Add a target-specific oligonucleotide probe to a reaction tube containing Thermostable Flap endonuclease 1 (FEN1) and Taq DNA ligase. The target-specific oligonucleotide probe has a 20-base pair at its 5' end that is complementary to the site recognizing crRNA (and therefore complementary to the 5' overhang generated by Cas12a and exonuclease I). (See also...) Figure 6 Oligonucleotide probes also contain a region of known sequence at their 3' end, which remains single-stranded, thus forming a 3' overhang (see also...). Figure 8 Advantageously, all oligonucleotide probes can contain the same 3' sequence and are therefore considered "universal," such as... Figure 8As shown. In this case, oligonucleotides PS1465 or PS1467 (SEQ ID NO: 17 and 18) with a 5' biotin ligand hybridized to a known sequence on an oligonucleotide probe, and gaps were filled and closed using Bst full-length DNA polymerase and Taq ligase. This allowed for the generation of Y-shaped aptamers, which enabled the molecules to be used for downstream analysis on the SIMDEQ platform. The resulting nucleic acid fragments were purified using paramagnetic beads using the manufacturer's recommendations (KAPA beads, Roche).
[0248] Both targets contain the non-palindromic restriction enzyme BsaI site and have the same four-base overhang at a site on the opposite side of the target region from the Y aptamer. The hairpin aptamer PS421 (SEQ ID NO: 69) hybridizes and ligates to them. The BsaI reaction is carried out at 37°C. The mixture was incubated for 30 minutes in buffer (50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, 100 μg / ml BSA, pH 7.9), and then purified with paramagnetic beads using the manufacturer's recommended method (KAPA beads, Roche). Ligation was then performed using T4 DNA ligase in T4 DNA ligase buffer (50 mM Tris-HCl, 10 mM MgCl2, 1 mM ATP, 10 mM DTT, pH 7.5) with 10 pmol hairpin aptamers at room temperature for 30 minutes. The ligated product was then purified with paramagnetic beads using the manufacturer's recommended method (KAPA beads, Roche).
[0249] The prepared target region was then analyzed on our SIMDEQ platform to evaluate the specificity of the method, and the expected peak was detected using CAAG tetranucleotides. The trajectory of the binding site is shown... Figure 15A and Figure 15B The presence of these peaks indicates that all expected peaks were correctly identified. The absence of other hairpin molecules in the flow cell indicates 100% specificity. This contrasts sharply with existing hybridization capture methods, which exhibit 15-25% off-target capture.
[0250] Example 9: Constructing a hairpin using human genomic DNA
[0251] We used the above-provided... Figure 14The protocol shown uses Cas12a and exonuclease I, followed by FEN1 and ligation steps, to construct hairpin molecules suitable for analysis on a SIMDEQ instrument. Briefly, 5 μg of human genomic DNA was incubated at 37°C for 30 min with 600 fmol of each Cas12a-crRNA complex (SED ID NO: 28 and 29) flanking the FMR1 target region (SEQ ID NO: 65), and with NEB2.1 buffer supplemented with 10 mM DTT containing 100 units of exonuclease I. The reaction was stopped using a stop buffer (1.2 units of proteinase K, 20 mM EDTA), and purification was performed using paramagnetic beads according to the manufacturer's recommendations (KAPA beads, Roche).
[0252] The resulting DNA was then incubated with 10 pmol of each of two oligonucleotide probes. The first probe was complementary to the 5' overhang generated by exonuclease I at the cleavage site of Cas12a-crRNA#1 (SEQ ID NO 100), and the second probe was complementary to the 5' overhang of Cas12a-crRNA#2 (SEQ ID NO 99), which contains a biotin ligand at its 3' end and has a base-degrading site (THF or tetrahydrofuran) within its sequence. The 5' flap was cleaved with 30 units of FEN1 to create a nick. The nick was blocked using 40 units of Taq DNA ligase in ThermoPol reaction buffer supplemented with 1 mM NAD+. After incubation at 45°C for 30 minutes, the two specific oligonucleotides (SEQ ID NO: 101 and 97) hybridized to the known sequences on the oligonucleotide probes, and the nicks were filled and blocked using Bst full-length DNA polymerase supplemented with 200 nM dNTPs. This allows for the generation of Y-shaped aptamers (which will enable the molecules to be used for downstream analysis on the SIMDEQ platform) and dsDNA duplexes around the abase site (3 bp after the abase site is dsDNA). The resulting nucleic acid fragments are captured and enriched by streptavidin paramagnetic beads.
[0253] Specifically, the target molecule was eluted from the beads using endonuclease IV at 37°C for 30 minutes in NEB3 buffer (100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, pH 7.9, 25°C). This elution catalyzes the cleavage of the DNA phosphodiester backbone at the abase site, leaving a 1 nucleotide gap and a 3'-hydroxyl terminus, generating a 4 bp 5' overhang. Then, using T4 DNA ligase in T4 DNA ligase buffer (50 mM Tris-HCl, 10 mM MgCl2, 1 mM ATP, 10 mM DTT, pH 7.5), 10 pmol of hairpin aptamer was reacted with the overhang for 30 minutes at room temperature, and the overhang was used to ligate the hairpin aptamer.
[0254] Example 10: Enrichment of target regions from human genomic DNA
[0255] Human genomic DNA was purified from the human embryonic kidney cell line (HEK293). Cas12a guide RNAs (SEQ ID NOs: 2, 20, and 25 to 52) were designed to target the first and second sites flanking the target region (SEQ ID NOs: 53 to 68). We selected 15 different human targets known to be cancer-associated epigenetic markers or composed of STRs (short tandem repeats) known to cause disease in humans (see [link to relevant documentation]). Figure 16A The first step of this method is as follows: 10 μg of genomic DNA is incubated with 390 fmol of each Cas12a-crRNA complex and 800 units of exonuclease I at 37°C for 2 hours to generate 5' overhangs. The reaction is stopped by adding stop buffer, and purification is performed using KAPA beads (Roche) according to the manufacturer's instructions.
[0256] Oligonucleotide probes containing a biotin ligand at their 3' end (SEQ ID NO: 21 and 70 to 83) were synthesized such that their 5' end was complementary to the generated 5' overhang. Due to the variability of the non-target strand cleavage position (see... Figure 1C The probe was designed to contain a PAM sequence and an additional 5 bases following the PAM sequence at its 5' end. This 5' lobe sequence, complementary to the site recognized by crRNA, corresponds to a typical substrate of the FEN1 enzyme. These oligonucleotides also contain restriction sites recognized by three restriction enzymes: DdeI (C^TNAG), HinflI (G^ANTC), and AluI (AG^CT).
[0257] DNA eluted by treatment with the LbaCas12a-crRNA complex and ExoI, supplemented with 6 pmol probe and 30 units of FEN1, was cleaved at 5' lobes to produce a nick. The nick was then cleaved using 40 units of Taq DNA ligase in a reaction buffer (20 mM Tris-HCl, 10 mM (NH4)2SO4, 10 mM KCl, 2 mM MgSO4, 0.1%... The nick was blocked in X-100, 1 mM NAD+, pH 8.8. This resulted in a double strand forming between the probe and the 5' overhang. After incubation at 37°C for 30 minutes, 30 units of RecJF (a 5' to 3' ssDNA exonuclease) were added to digest the unligated oligonucleotides. The ligated product was then purified with paramagnetic beads using the manufacturer's recommended method (KAPA beads, Roche), and the resulting DNA preparation was captured for 1 hour using streptavidin-coated paramagnetic beads (Ocean Nanotech) in the same reaction buffer as the FEN1 reaction. The beads bound to the double strand (and therefore the DNA target region) were split into three reactions, each treated with a different restriction enzyme (DdeI, AluI, or HinfI), which cleaved the DNA target region from the beads and made it a template for Illumina library preparation. We used the NEB Ultra II kit for low starting material and then prepared the library according to the manufacturer's protocol. Sequencing was performed on a NextSeq 500 Illumina sequencer using paired-end sequencing (150 base pairs per side). Read alignment was performed on the human reference genome using the Bowtie algorithm, and coverage was calculated using Samtools. Reads were visualized using IGV software. Representative coverage of one of the target regions (SNCA) is shown below. Figure 16B As shown in the screenshot extracted from the IGV software, the maximum coverage obtained in this target region is very high (614×) compared to the coverage of the rest of the genomic DNA (less than 1× on average).
Claims
1. A method for isolating a nucleic acid target region from a population of nucleic acid molecules, the method comprising the following steps: a) Contact the nucleic acid molecule group with a type V Cas protein-gRNA complex, wherein the gRNA contains a guide segment complementary to a first site adjacent to the target region, thereby forming a type V Cas protein-gRNA nucleic acid complex. b) Contact the nucleic acid molecule group containing the type 2 V Cas protein-gRNA nucleic acid complex with at least one enzyme having single-stranded 3' to 5' exonuclease activity, thereby forming a 5' single-stranded overhang at the first site. c) Remove the type V Cas protein-gRNA complex from the group in step b). d) Contact the group from step c) with an oligonucleotide probe, the probe containing a sequence at least partially complementary to the overhang, thereby forming a double strand between the probe and the overhang. e) Isolate the double strand from the nucleic acid molecule population of step d), thereby isolating the nucleic acid target region; The type V Cas protein mentioned above is Cas12a.
2. The method of claim 1, further comprising contacting the nucleic acid molecular group with a site-specific endonuclease prior to step c), wherein the gRNA contains a guide segment complementary to a second site, wherein the second site is adjacent to the target region, and wherein the first site and the second site are located on opposite sides of the target region.
3. The method according to claim 2, wherein the site-specific endonuclease is TALEN, zinc finger protein, or a type 2 Cas protein-gRNA complex.
4. The method according to claim 2, wherein the site-specific endonuclease is a type 2 V Cas protein-gRNA complex.
5. The method according to claim 1, wherein the at least one enzyme having single-stranded 3' to 5' exonuclease activity is selected from exonuclease I, S1 exonuclease, exonuclease T and exonuclease VII.
6. The method according to claim 5, wherein the at least one enzyme having single-stranded 3' to 5' exonuclease activity is exonuclease I.
7. The method of claim 1, further comprising fragmenting the nucleic acid molecular population before or during step b).
8. The method of claim 7, wherein before or during step b), the nucleic acid molecular group is fragmented by contacting it with at least one site-specific endonuclease, wherein the site-specific endonuclease: • Do not cut within the target region or the first site, and • When the molecule contains a second site, it is not cleaved within the second site.
9. The method according to claim 8, wherein the site-specific endonuclease is a restriction enzyme.
10. The method of claim 1, wherein step c) comprises contacting the nucleic acid molecular group with EDTA and / or at least one protease.
11. The method of claim 10, wherein the at least one protease is a serine protease.
12. The method of claim 10, wherein the at least one protease is proteinase K.
13. The method of claim 1, further comprising releasing the bichain from the target region.
14. The method of claim 13, wherein releasing the bichain from the target region comprises cutting within the bichain.
15. The method of claim 2, wherein the site-specific endonuclease is a second type 2 type V site-specific endonuclease, which contacts the group with a second oligonucleotide probe to form a second double strand, the second oligonucleotide probe comprising a sequence at least partially complementary to the second 5' single-stranded overhang in the second site.
16. The method of claim 15, wherein the first probe comprises a ligand, the second probe comprises a different ligand, the ligand of the first probe binds to a first trapping agent, the ligand of the second probe binds to a second trapping agent, and the first trapping agent and the second trapping agent are different.
17. The method of claim 15, wherein the length of the 5' single-stranded overhang at the first site and / or the second site is at least 9 nucleotides.
18. The method of claim 17, wherein the length of the 5' single-stranded overhang at the first site and / or the second site is 9, 10, 11, 12 or more nucleotides.
19. The method of claim 16, wherein step e) comprises: • The first double-stranded polymer is brought into contact with the first trapping agent, and the first trapping agent binds to the ligand of the first probe. • Release the first double strand from the target region. • The second double-stranded polymer is brought into contact with the second trapping agent, and the second trapping agent binds to the ligand of the second probe. • Release the second double strand from the target region.
20. The method of claim 1, wherein the target region comprises repeating regions, rearrangements, replications, translocations, deletions, or modified bases.
21. The method of claim 20, wherein the modified base is an epigenetic modification, a mismatch, or an SNP.
22. The method of claim 1, wherein at least two target regions are separated.
23. The method of claim 22, wherein at least five target regions are separated.
24. The method of claim 22, wherein at least 50 target regions are separated.
25. The method of claim 22, wherein at least 100 target regions are separated.
Citation Information
Patent Citations
Immobilized nucleic acid-containing probes
EP0152886A2
Element and method for nucleic acid amplification and detection using adhered probes
WO1992016659A1
Method of DNA sequencing by polymerisation
WO2011147929A1
Method of DNA sequencing by hybridisation
WO2011147931A1
Method of DNA detection and quantification by single-molecule hybridization and manipulation
WO2013093005A1