Targeted cleavage of nucleic acids by argonaute proteins
The use of thermostable Argonaute proteins with SSB to generate single-stranded DNA guides addresses the inefficiencies of existing sequencing methods, enabling efficient and cost-effective sequencing of plant genomes by selectively depleting repetitive sequences and capturing comprehensive genetic variation.
Patent Information
- Application Number
- PCT/EP2025/075879
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-12
- Filing Date
- 2025-09-11
- Publication Date
- 2026-03-19
AI Technical Summary
Current methods for sequencing plant genomes are costly and inefficient due to high complexity from repetitive sequences, with existing technologies like PCR, probe-based enrichment, and CRISPR-Cas having limitations such as bias, size constraints, and high synthesis costs, making it difficult to achieve comprehensive genetic variation analysis.
A method using thermostable prokaryotic Argonaute proteins (e.g., TtAgo) generates single-stranded DNA guides from double-stranded DNA pre-guides with the aid of Single-Strand Binding Protein (SSB) to inhibit annealing, enabling efficient and specific cleavage of target sequences without off-target activity.
This approach allows for robust and cost-effective sequencing of non-repetitive genomic regions, enhancing genetic variation capture and reducing sequencing costs by selectively depleting repetitive sequences, thus improving genome-wide genetic variation analysis.
Smart Images

Figure IMGF000026_0001 
Figure IMGF000027_0001 
Figure IMGF000028_0001
Abstract
Description
[0001] P36819PCOO / MJO
[0002] Targeted cleavage of nucleic acids by Argonaute proteins
[0003] Technical field
[0004] The present invention relates to the fields of molecular biology and genetics. In particular, the invention relates to the generation of ssDNA guides from a dsDNA pre-guide using an Argonaute protein. Said guides may be useful in applications, like selectively reducing repetitive sequences in sequencing libraries using Argonaute proteins. The invention further relates to a method for cleaving a nucleic acid molecule, a method for reducing the number of nucleic acid molecules comprising a target sequence and a method for selectively enriching and / or sequencing nucleic acid molecules in a sample.
[0005] Background of the invention
[0006] Unraveling the genetic diversity of plant genomes via (whole-genome) sequencingbased methods presents significant challenges due to their high complexity resulting from large variations in genome size, large fractions of repetitive sequences and differences in ploidy levels. For instance, the genome of onion (Allium cepa) is 15 Gb, compared to the 3.1 Gb human genome. Hence, achieving comparable sequencing depth for the onion genome incurs whole-genome sequencing costs which are approximately five times higher than those for the human genome. Similarly, a large fraction of the DNA of most plant genomes comprises of highly repetitive regions and / or are polyploid (such as tetrapioid, hexapioid or octoploid), and the genomes frequently undergo (structural) variations. In view of these complexities, resolving them at the haplotype level, i.e. assembling sequences of alleles located on the same chromosome and identifying (structural) genetic variants is particularly difficult and costly. Cost reduction is therefore essential for widespread use of plant genome sequencing to support breeding processes.
[0007] Sequencing costs can be reduced by sequencing only a portion of the genome instead of whole genome sequencing. Preferably, such “complexity reduction” methods aim to enrich for the non-repetitive fraction of the genome as this is generally believed to contain the most relevant genetic variation for breeding purposes, compared to the repetitive genome sequences which are generally considered as non-informative. For example, the tomato (Solanum lycopersicum) genome contains 65.66% repetitive sequences. Depleting these repeats from a sequencing library can, therefore, theoretically more than halve sequencing costs without much reduction in the amount of information about genetic variation in sequence data obtained. Another genome complexity- and cost-reduction strategy is targeted sequencing. In targeted sequencing methods regions of interest in the genome are selectively enriched and subsequently sequenced.
[0008] Methods for targeted sequencing include PCR amplification of regions of interest, capturing selected regions with long DNA probes that anneal to the DNA of interest, and using CRISPR-Cas to excise the selected region from the genomic DNA. Each method has its advantages and limitations.
[0009] For instance, PCR-based enrichment effectively enriches specific loci and is reproducible and cost-efficient. However, during amplification by the PCR bias can occur, where one allele is preferentially amplified over others, leading to incomplete sequencing data in case an allele is underrepresented in the library and no longer sequenced. This may lead to an incorrect assessment of the genetic variation at a locus in the genome. Similarly, PCR amplification bias may also lead to different enrichment levels between specific loci during multiplexed amplification, which in turn leads to differences in the representation of these loci in sequence datasets and may also lead to inaccurate calls of genetic diversity (i.e. incorrect genotypes). Additionally, PCR amplification is limited to regions smaller than 20 kb, as DNA polymerases generally struggle with larger regions.
[0010] Probe-based enrichment is not driven by selective amplification and employs multiple biotinylated DNA probes that hybridize to the target loci, which are then captured with streptavidin-coated magnetic beads. While effective, the reproducibility of the method is affected by specific hybridization conditions, the enrichment levels may vary between targeted loci, and the size of the purified DNA fragments is relatively small (around 3 kb). Another disadvantage is that upfront investments are required to design and validate (highly) multiplexed probe enrichment panels for every species, which is a limitation that also applies to multiplexed PCR methods. Targeted sequencing methods based on PCR and probe enrichment are therefore not well suited to capture most of the genetic diversity of a plant genome outside repetitive regions, but rather to enrich for just a small fraction, typically 1 to 10% of the whole genome.
[0011] CRISPR-Cas allows for the highly (sequence-)specific digestion of particular DNA sequences using RNA guide molecules. Like PCR, CRISPR-Cas methods can be used to enrich for specific target fragments and an advantage is that these targets are not limited in size to a few kilobases. However, due to their chemical instability and the complexity of synthesis, generating the required RNA guide molecules with specific sequences is labor- intensive and costly. Furthermore, CRISPR-mediated cleavage is constrained to specific sequences due to design limitations because the presence of a specific recognition motif is required, such as PAM (Protospacer Adjacent Motif) sequence for Cas9. Finally, in the art highly multiplexed application of CRISPR-Cas to target thousands of fragments is less advanced compared to PCR and probe-based enrichment methods. Besides using CRISPR-Cas for selective enrichment of target fragments, CRISPR- Cas can also be used to selectively deplete (as opposed to enrich) fragments from complex fragments mixtures, because its mode of action is based on digestion (or nicking) target DNA fragments in a sequence-specific manner. In view of this, reduction of sequencing costs of (whole) plant genomes may be achieved by selective depletion of unwanted sequences (i.e. repetitive fragments) instead of enriching desired fragments in a targeted fashion using PCR and probe-based methods. A significant advantage of such selective depletion compared to target enrichment methods is that it has the potential to retain all informative sequences in the genome, whereas PCR- or probe-based enrichment methods are limited to multiple (small) regions targeted by primers or probes, which collectively typically represent only a small fraction of the genome. For sequencing information to support decisions in breeding programs, retaining as much informative genome sequence as possible is preferred to increase accuracy and maximize genetic gain. However, in practice, this remains challenging and often not feasible, particularly for crops with large and complex genomes.
[0012] Besides CRISPR-Cas, DNA-guided endonuclease Argonaute proteins (pAgo) have also been used to selectively enrich certain nucleic acids. Argonaute proteins, such as Thermus thermophilus Argonaute (TtAgo) are a family of endonucleases involved in gene silencing and defense mechanisms in a wide range of organisms. Prokaryotic Argonaute proteins, like TtAgo, bind to single-stranded DNA (ssDNA) guides and cleave complementary DNA targets, making them useful tools for genome editing and sequencing applications.
[0013] WO 2023 / 148235 discloses methods for selectively fragmenting and enriching certain nucleic acids of known or unknown sequences which have low abundance in a sample of nucleic acids. This method involves generating a library of nucleic acid guides using DNasel, which randomly fragments DNA. After denaturation, these random fragments are then used as guides for a Pyrococcus furiosus Argonaute (PfAgo). A 1 :2 ratio of PfAgo:guide is used to achieve PfAgo saturation and thereby suppress its unguided cleavage activity (i.e. by chopping). If no guides are made for the low-abundance sequences, these will not be specifically cleaved, and the low-abundance sequences are effectively enriched in the sample. This enriched sample is subsequently used in further steps for sequence detection and sequencing.
[0014] A limitation associated with prokaryotic Argonautes (compared to CRISPR-Cas) is the lack of understanding regarding the design of effective ssDNA guides. Unlike RNA-guided systems like CRISPR-Cas, where guide design principles are well-established and guide efficacy can be reliably predicted, the rules governing the effectiveness of ssDNA guides for Argonaute proteins remain unclear. This uncertainty makes it difficult to predict with confidence which ssDNA guides will efficiently guide Argonaute to cleave target DNA sequences and limits the use of Argonautes for targeted DNA cleavage purposes. However, an advantage associated with prokaryotic Argonaute proteins, like TtAgo, is that they are known to autonomously generate single-stranded DNA (ssDNA) guides from double-stranded DNA molecules as part of its guide loading process. These ssDNA guides do not suffer from the chemical instability and high synthesis costs of CRISPR-Cas RNA guides. This autonomous process is referred to as chopping. Swarts et al. 2017 (Molecular Cell, volume 65, issue 6, p985-998.e6, 2017) has described the mechanism by which guide-free Argonaute from Thermus thermophilus (apo-TtAgo) can degrade dsDNA, thereby generating small dsDNA fragments that subsequently are loaded onto TtAgo. Besides the fact that this process is rather inefficient, in the context of target enrichment methods the chopping activity of Argonaute proteins (pAgo) is considered undesirable as it leads to off-target cleavage. In selective enrichment applications where highly specific cleavage of DNA at selected target sites is required, the chopping activity is therefore suppressed by providing an excess of guide molecules relative to the number of Argonaute proteins present. For selective depletion methods, however, this limitation is of lesser concern and outweighed by the advantages of the higher stability and lower synthesis costs of ssDNA guides compared to synthesis of RNA guides required for CRISPR-Cas.
[0015] Hence, there is a need to reduce whole genome sequencing costs, in particular for plant genomes, and to provide a more efficient method for sequence library preparation, particularly a robust, broadly applicable method that enables selective sequencing the nucleic acid molecules in a sample derived from the vast majority of the non-repetitive fraction of the genome to capture genome-wide genetic variation.
[0016] Summary
[0017] The present disclosure provides a method for generating guides for prokaryotic Argonaute proteins. In particular, the present disclosure improves the efficiency of providing functional guides for pAgo, which makes the use of prokaryotic Argonaute proteins, such as TtAgo, more practical and reliable for various applications.
[0018] Thereto, in a first aspect, the present disclosure provides a method for providing a single-stranded deoxyribonucleic nucleic acid (ssDNA) guide for a prokaryotic Argonaute protein (pAgo), the method comprising the step of combining:
[0019] - an unloaded prokaryotic Argonaute protein (apo-pAgo);
[0020] - a double-stranded deoxyribonucleic nucleic acid (dsDNA) pre-guide; and
[0021] - a means for inhibiting the annealing of single-stranded DNA (ssDNA) to a complementary ssDNA; to provide the ssDNA guide.
[0022] It is known that apo-pAgo may produce ssDNA guides from dsDNA, but this activity is inefficient when performed without a means for inhibiting the annealing of single-stranded DNA (ssDNA) to a complementary ssDNA, in particular in comparison to the cleaving activity of the pAgo when it is loaded with a ssDNA guide. Before the present disclosure, it was therefore necessary to first generate a ssDNA guide using other methods, like other nucleases, enzymatic methods or DNA synthesis, and then load the ssDNA guide onto the Argonaute protein.
[0023] In the prior art, it is known that Argonaute proteins have off-target cleavage activity, which is known as “star” activity. This “star” activity is in part due to prokaryotic Argonautes possessing an intrinsic nuclease activity that enables them to cleave DNA in the absence of a ssDNA guide. This intrinsic nuclease activity is referred to as chopping. Chopping and DNA- guided cleavage have different molecular mechanisms.
[0024] US 11,466,264 discloses that adding a Single-Strand Binding protein (SSB) reduces the off-target DNA-guided cleavage of DNA by prokaryotic DNA guided Argonaute proteins.
[0025] The present inventors unexpectedly found that the presence of a means for stabilizing single-strand DNA (ssDNA), in particular single-strand DNA binding protein (SSB), significantly improves the ability and efficiency of apo-pAgo to generate ssDNA guides from a dsDNA pre-guide, thereby making the generation of ssDNA guides by Argonaute proteins a feasible approach for various applications.
[0026] The ability to generate ssDNA guides from dsDNA pre-guides using apo-pAgo has several advantages. There is no need to design ssDNA guides. Multiple different ssDNA guides may be generated in parallel for the same target DNA sequence improving the efficiency of the targeting of the target DNA sequence. The ssDNA guides may be generated with pAgo from any source of dsDNA, including, for example, but not limited to, (enriched) genomic dsDNA, plasmid DNA, plastid DNA and dsDNA generated using DNA synthesis methods or polymerase chain reaction (PCR).
[0027] In a second aspect, the disclosure further provides a composition comprising an apo- pAgo, means for inhibiting the annealing of single-stranded DNA (ssDNA) to a complementary ssDNA, preferably a Single-Strand Binding protein (SSB), and a dsDNA preguide. These components of the composition may be included in a kit. The kit may further comprise an aqueous buffer.
[0028] In a third aspect, the disclosure provides a method for nicking or cleaving a nucleic acid at a target sequence, the method comprising the steps of: a) generating an ssDNA guide with a method according to the present disclosure; b) providing a prokaryotic Argonaute protein (apo-pAgo); c) providing the nucleic acid comprising the target sequence, wherein the target sequence is complementary to at least part of the ssDNA guide; preferably, the nucleic acid is DNA; d) combining the generated ssDNA guide, the apo-pAgo, and the nucleic acid comprising the target sequence to nick or cleave the nucleic acid comprising the target sequence.
[0029] In a fourth aspect, the disclosure provides a method for reducing the number of nucleic acid molecules comprising a target sequence in a sample, the method comprising the steps of: a) generating an ssDNA guide with a method according to the present disclosure; b) providing a prokaryotic Argonaute protein (pAgo); c) providing nucleic acid molecules, wherein at least one nucleic acid molecule is a nucleic acid molecule comprising a target sequence that is complementary to at least part of the ssDNA guide; preferably the nucleic acid molecules are DNA molecules; d) combining the ssDNA guide, pAgo and the nucleic acid molecules to nick or cleave the nucleic acid molecule comprising the target sequence.
[0030] In a fifth aspect, the present disclosure further provides a method for selectively amplifying and / or sequencing nucleic acids in a sample, the method comprising the steps of: a) generating an ssDNA guide with a method according to the present disclosure; b) providing a prokaryotic Argonaute protein (pAgo); c) providing nucleic acids, wherein at least one nucleic acid is a target nucleic acid comprising a target sequence that is complementary to at least part of the ssDNA guide; d) combining the ssDNA guide, pAgo and the nucleic acids to nick or cleave the at least one target nucleic acid; e) amplifying and / or sequencing the unnicked and uncleaved nucleic acids.
[0031] In one embodiment, the nucleic acid is DNA, preferably dsDNA, and cleaving of a strand of the dsDNA molecules results in nicked DNA.
[0032] In one embodiment, the nucleic acids may be dsDNA molecules and the method further comprises cleaving the opposing strand of a dsDNA molecule at the complement of the target sequence, thereby creating a double-stranded break or two nicks in the dsDNA molecule.
[0033] In one embodiment, the ssDNA guide and prokaryotic Argonaute protein (pAgo) are provided as an Argonaute / ssDNA guide complex (holo-pAgo), i.e. , the pAgo is loaded with the ssDNA guide. Pre-loading of the ssDNA guide prior to step d) increases the cleaving efficiency and accuracy of pAgo.
[0034] In preferred embodiments of any method according to the present disclosure, the pAgo and SSB are thermostable proteins, i.e. the pAgo and SSB function at temperatures above 50 °C.
[0035] Detailed description
[0036] Definitions
[0037] In the context of the present disclosure, the term 'a' shall "a" refer to both singular and plural instances. The term “a” may thus be interchangeably used with terms such as “one or more” or “at least one”.
[0038] Any of the proteins described herein (e.g., the Argonaute protein, SSB etc.) is preferably thermostable. The term “thermostable” refers to a protein that retains at least 95% of its activity after 10 minutes at a temperature of 50°C, preferably 65 °C. Extreme thermostable refers to a protein that retains at least 95% of its activity after 10 minutes at a temperature of 95°C. The (retained) activity is preferably tested at optimal conditions for the protein.
[0039] Biologically functional means that the protein (e.g., the Argonaute protein, SSB etc.) is capable of performing its (catalytic) activity when placed in the appropriate physiological or experimental conditions necessary for its activity.
[0040] Binding refers to an interaction between two molecules, where one molecule (e.g., an Argonaute protein or a SSB) attaches to another molecule (e.g., a ssDNA guide or ssDNA) through one or more or various types of chemical bonds or interactions. This binding can involve hydrogen bonds, ionic bonds, van der Waals forces, and hydrophobic interactions. Binding implies a selective and reversible interaction. Examples of equivalent terms for binding are attaching, interacting, associating, and engaging.
[0041] The terms “nucleic acid”, “nucleic acid molecule”, or “oligonucleotide” are used interchangeably and refer to any polymer or oligomer of pyrimidine and purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively. A nucleic acid or oligonucleotide may be DNA or RNA. DNA and RNA can be synthesized naturally (e.g., by DNA replication or transcription of DNA) or chemically (e.g. solid-phase synthesis). DNA and RNA can be single-stranded (i.e., ssDNA and ssRNA) or double-stranded (i.e., dsDNA and dsRNA).
[0042] The term “RNA”, “RNA molecule” or “ribonucleic acid molecule” refers to a polymer of ribonucleotides. The term “DNA”, “DNA molecule” or “deoxyribonucleic acid molecule” refers to a polymer of deoxyribonucleotides.
[0043] The term “complementary” refers to the ability of nucleotides or analogues thereof to form Watson-Crick base pairs. Complementary (single-stranded) nucleotide sequences will form Watson-Crick base pairs (under suitable conditions) and non-complementary (singlestranded) nucleotide sequences will not.
[0044] The terms “target sequence”, “target nucleotide sequence”, and “target nucleic acid”, all refer to a target nucleic acid to be targeted. These terms may be interchangeable with the terms “sequence of interest”, “nucleotide sequence of interest”, and “nucleic acid of interest”. A target sequence can be present in a dsDNA or ssDNA; a target nucleic acid may also be present in an RNA molecule.
[0045] A k-mer is a sequence of “k” nucleotides (bases) in a DNA or RNA molecule. For example, a 2-mer is NN (k=2), where N represents adenine (A), thymine (T), guanine (G), or cytosine (C) in DNA, and uracil (II) replaces thymine in RNA.
[0046] The term “unloaded” in the context of prokaryotic DNA-guided Argonaute proteins, refers to the state of the Argonaute protein when it does not have a DNA guide bound to it, also referred to as apo-pAgo.
[0047] Repetitive DNA sequences are segments of DNA that are repeated at least 2 times within a sample or within a genome, preferably at least 3 times, more preferably at least 5 times. Repetitive DNA may be satellite DNA or a transposable element. Satellite DNA is an array of tandemly repeating, non-coding DNA and may be found in centromeres or telomeres. A transposable element is a DNA sequence capable of changing its position within the genome of an organism. The transposable element may be a retrotransposon belonging to Class I (Retrotransposon) that moves within the genome of an organism via an RNA intermediate, such as for example long interspersed nuclear elements (LINEs) and short interspersed nuclear elements (SINEs). The transposable element may be a transposon belonging to class II (DNA Transposons) that moves within the genome of an organism directly through DNA, such as for example an endogenous retrovirus.
[0048] Method for providing a single-stranded deoxyribonucleic nucleic acid (ssDNA) guide
[0049] According to the present disclosure, a combination of a (thermostable) unloaded prokaryotic Argonaute protein, a dsDNA pre-guide, a means for inhibiting the annealing of single-stranded DNA (ssDNA) to a complementary nucleic acid strand, and a thermostable SSB is used to provide a single-stranded DNA guide (ssDNA).
[0050] In one embodiment, the present disclosure relates to a method for providing a singlestranded deoxyribonucleic nucleic acid (ssDNA) guide for a prokaryotic Argonaute protein (pAgo), the method comprising the step of combining an unloaded prokaryotic Argonaute protein, a single-strand binding protein (SSB) and a double-stranded deoxyribonucleic nucleic acid (dsDNA) pre-guide, to generate the ssDNA guide.
[0051] An advantage of the methods according to the present disclosure is that the generation of ssDNA guides by apo-pAgo is improved compared to known methods in the prior art. It is known that apo-pAgo may generate ssDNA guides, but this activity is inefficient compared to the cleaving activity of the pAgo when it is loaded with a ssDNA guide. Before the present disclosure, it was therefore considered at least more practical, if not necessary, to use other methods, such as other nucleases, enzymatic methods or DNA synthesis, to generate ssDNA guides first and then load these guides onto the Argonaute protein. The present inventors surprisingly found that the presence of SSB improved the ssDNA guide generating activity of apo-pAgo, thereby making guide formation by apo-pAgo a feasible approach for various applications.
[0052] The ability to generate ssDNA guides using apo-pAgo has several advantages. There is no need to design ssDNA guides. Multiple different ssDNA guides may be generated in parallel for the same target DNA sequence further increasing the chance that the dsDNA comprising the target sequence is cleaved. Generating ssDNA by chopping using pAgo is not limited to any particular source of dsDNA from which ssDNA is generated, i.e., (enriched) genomic dsDNA, plasmid DNA, plastid DNA as well as synthetic dsDNA generated using DNA synthesis methods or polymerase chain reaction (PCR) may be used.
[0053] In one embodiment, the method for providing a single-stranded deoxyribonucleic nucleic acid (ssDNA) guide for a prokaryotic Argonaute protein (pAgo), comprises the steps of: a) providing an unloaded prokaryotic Argonaute (apo-pAgo); b) providing means for inhibiting the annealing of single-stranded DNA (ssDNA), preferably a single-strand binding protein (SSB); c) providing a double-stranded deoxyribonucleic nucleic acid (dsDNA) pre-guide; d) combining the apo-pAgo, the means for inhibiting the annealing of single-stranded DNA (ssDNA), preferably the SSB and the dsDNA pre-guide so as to allow apo-pAgo to provide the single-stranded deoxyribonucleic nucleic acid (ssDNA) guide; and e) preferably, loading the apo-pAgo with the generated ssDNA guide to obtain a loaded pAgo (holo-pAgo).
[0054] In one embodiment, the method for providing a ssDNA guide according to the present disclosure comprises the steps of (a) providing a thermostable, unloaded, DNA- guided prokaryotic Argonaute (apo-pAgo); (b) providing a thermostable SSB; (c) providing at least one dsDNA pre-guide; (d) combining the apo-pAgo, the SSB and the dsDNA pre-guide to allow pAgo to generate at least one ssDNA guide. In one embodiment, at least 10% of the Argonaute proteins are unloaded. Preferably at least 20%, 30%, 40%, 50%, 60%, 70%, 80, 90% or 95% of the Argonaute proteins are unloaded.
[0055] In one embodiment, at least 10% of the Argonaute proteins have chopping activity. Preferably at least 20%, 30%, 40%, 50%, 60%, 70%, 80, 90% or 95% of the Argonaute proteins have chopping activity.
[0056] In a preferred embodiment, no ssDNA guides are combined with the apo-pAgo, the means for inhibiting the annealing of single-stranded DNA (ssDNA), preferably the SSB protein, and the dsDNA pre-guide prior to allowing apo-pAgo to provide the ssDNA guides in step d).
[0057] In some embodiments, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80, preferably 90% or 95% of the Argonaute proteins are unloaded and do not have DNA-guided cleavage activity prior to allowing pAgo to provide the ssDNA guide in step d).
[0058] In one embodiment, the generation of ssDNA guide is performed by incubating the combination of apo-pAgo, SSB and dsDNA pre-guide at a temperature of between 40-99 °C; preferably 50-95 °C; more preferably 65 to 85 °C in an aqueous medium.
[0059] In one embodiment, the generation of ssDNA guide is accomplished by incubating the combination of apo-pAgo, SSB and dsDNA for 1-48 hours; preferably 2-24 hours; more preferably 5 to 20 hours.
[0060] In one embodiment, the generation of ssDNA guide is accomplished by combining an apo-pAgo, a dsDNA pre-guide and a means for inhibiting ssDNA from annealing (i.e. , SSB) in a buffered aqueous medium. Preferably, the buffered aqueous medium comprises a buffering agent at a concentration of 1 mM to about 200 mM (e.g., Tris-HCI), salt (e.g., KC1 or NaCI) and a divalent cation (e.g., Mg2+or Mn2+). In these embodiments, the composition can have a pH of about 7 to about 8.8 at 25°C, such as 7.4 to 7.5. In some embodiments, the composition may further contain a reducing agent (e.g., dithiothreitol), a detergent, glycerol, sugar or dNTPs, as needed. An example of a suitable aqueous buffer is a buffer comprising 5-40 mM Tris-HCI, 5-20 mM (NH4)2SO4; 5-20 mM KCI, 0.5-5 mM MgSO4, and / or 0.005-2 w / v% polyethylene glycol p-(1,1,3,3-tetramethylbutyl)-phenyl ether (Triton®-X-100).
[0061] In one embodiment, the generation of ssDNA guide is accomplished by incubating the combination of a thermostable apo-pAgo (preferably TtAgo), thermostable SSB and dsDNA pre-guide:
[0062] - at a temperature of between 40-99 °C; preferably 50-95 °C; more preferably 65 to
[0063] 85 °C;
[0064] - for 1-48 hours; preferably 2-36 hours; more preferably 8 to 24 hours; - in a buffered aqueous medium; preferably a buffered aqueous medium comprising a buffering agent at a concentration of 1 mM to about 200 mM (e.g., Tris-HCI), salt (e.g., KC1 or NaCI) and a divalent cation (e.g., Mg2+or Mn2+).
[0065] In one embodiment, the amount of SSB is at least 0.1 pM; preferably at least 0.5 pM; more preferably at least 1 pM. In addition or alternatively, the amount of SSB is at most 10 pM; preferably at most 5 pM. The addition of such amounts of SSB improved the efficiency of ssDNA guide generation by a prokaryotic Argonaute, in particular TtAgo. The optimal concentration of SSB is between 1.0 pM and 2 pM.
[0066] In one embodiment, apo-pAgo is present in a molar ratio of dsDNA pre-guide:apo- pAgo of at least 1 (dsDNA pre-guide): 1 (apo-pAgo), preferably 1 (dsDNA pre-guide):2 (apo- pAgo). In other words, per 1 dsDNA pre-guide molecule there is at least 1 apo-pAgo protein present. Preferably there are at least 2 apo-pAgo proteins present per 1 dsDNA pre-guide molecule. The processing of dsDNA is (almost) complete at these ratios. The number of apo- pAgo proteins is thus preferably equal or greater than the number of dsDNA pre-guide molecules in the sample. However, ratios outside these ranges may also be used as SBB enhances the production of ssDNA guides, irrespective of whether all dsDNA pre-guides are processed or a fraction thereof.
[0067] In one embodiment, the apo-pAgo is present in a molar ratio of dsDNA pre- guides:apo-pAgo of 5 or less, preferably 2 or less, more preferably 1 or less, most preferably 0.5 or less. In other words, 1 apo-pAgo protein per 5 dsDNA pre-guide molecules or less dsDNA pre-guide molecules, preferably 1 apo-pAgo protein per 2 dsDNA pre-guide molecules or less, more preferably 1 apo-pAgo protein per 1 dsDNA pre-guide molecule or less, most preferably 2 apo-pAgo proteins per 1 dsDNA pre-guide molecule or less, e.g. 4 apo-pAgo proteins per 1 dsDNA pre-guide molecule.
[0068] In preferred embodiments, the apo-pAgo is present in a molar ratio of dsDNA pre- guide:apo-pAgo of between 0.01 (dsDNA pre-guide):1 (apo-pAgo) and 5 (dsDNA pre-guide):1 (apo-pAgo), preferably 0.05 (dsDNA pre-guide):1 (apo-pAgo) and 2 (dsDNA pre-guide):1 (apo-pAgo), more preferably 0.1 (dsDNA pre-guide): 1 (apo-pAgo) and 1 (dsDNA pre-guide): 1 (apo-pAgo) and most preferably, 0.125 (dsDNA pre-guide) :1 (apo-pAgo) and 1 (dsDNA pre- guide):2 (apo-pAgo). Preferred amounts of dsDNA pre-guide in a sample with a 20 pl reaction volume may be 0.01-10 pmol, preferably 0.05-5 pmol, more preferably 0.1-2.5 pmol, most preferably 0.5-1.5 pmol.
[0069] The total amount of dsDNA pre-guide in a 20 pl reaction volume is preferably 10-500 ng, more preferably 20-250 ng, most preferably 40-125 ng.
[0070] In some embodiments, the method of providing ssDNA guides may comprise the step of removing the Argonaute protein and SSB, preferably using proteinase K, after generating the ssDNA guides by apo-pAgo, to obtain a mixture comprising ssDNA guides and dsDNA pre-guides. Such ssDNA guides may be used in further downstream methods, e.g., for cleavage of a target nucleic acid, by combining them with apo-pAgo.
[0071] In some embodiments, the dsDNA pre-guides are removed, preferably using a DNase, after the generation of the ssDNA guides to obtain a mixture of ssDNA guides bound to Argonaute proteins.
[0072] An ssDNA guide and an Argonaute protein (Argonaute) can form a complex, wherein the ssDNA guide provides targeting specificity to the complex by comprising a DNA sequence that can hybridize to DNA comprising a target sequence. The formation of the ssDNA guide:pAgo complex is referred to as loading. Once loaded with the ssDNA guide, pAgo can target and cleave DNA comprising a target sequence complementary to the ssDNA.
[0073] In one embodiment, the ssDNA guide is generated by a first Argonaute protein and subsequently loaded onto a second Argonaute protein to obtain a loaded pAgo (holo-pAgo).
[0074] The method for providing ssDNA guides according to the present may be employed in a variety of methods, particularly in vitro (i.e. , cell free) methods, including, but not limited to digestion, depletion, or enrichment methods; or as an initiator of amplification; for purposes that include recognizing a nucleic acid for editing; detecting nucleic acid sequence via binding to Argonaute; and sequencing of target nucleic acids.
[0075] Argonaute protein (pAgo)
[0076] Argonaute (pAgo) proteins are a class of nucleic acid-guided endonucleases found in eukaryotes, bacteria and archaea. The prokaryotic Argonaute used in a method according to the present disclosure is preferably a DNA-guided endonuclease. A DNA-guided prokaryotic Argonaute as used in a method according to the present disclosure is a protein capable of cleaving a nucleic acid at a target sequence using a single-stranded DNA (ssDNA) guide to identify and cleave the nucleic acid at the target sequence.
[0077] Argonautes can exist in two forms, namely an apo-form and an holo-form. In the apo-form, the Argonaute protein is referred to as unloaded, meaning the Argonaute protein is not bound to a guide nucleic acid. An Argonaute protein may be loaded with a ssDNA guide to form a complex with a ssDNA guide. Such complexes of pAgo / ssDNA guide complexes are also referred to as holo-pAgo. Once loaded with a guide, pAgo can target and cleave DNA comprising a target sequence that is at least in part complementary to the ssDNA guide.
[0078] In one embodiment, the Argonaute (pAgo) protein is a prokaryotic Argonaute protein. A prokaryotic Argonaute is obtained, obtainable or derived from bacteria and archaea.
[0079] In one embodiment, the Argonaute protein (pAgo) is a thermostable, prokaryotic Argnonaute protein. The thermostable, prokaryotic Argnonaute protein may be obtained, obtainable or derived from a thermophilic prokaryote. An Argonaute protein obtained or derived from a mesophilic prokaryote may be stable and active at temperatures in the range of 25-50°C. A thermophilic Argonaute may be stable and active at temperatures of at least 50°C. Examples of thermophilic Argonautes are TtAgo from Thermus thermophilus, which may be stable and active between 50-75°C; PfAgo from Pyrococcus furiosus, which may be stable and active between 90-99.9°C; and MjAgo from Methanocaldococcus jannaschii, which may be stable and active between 85-95°C.
[0080] In one embodiment, the Argonaute (pAgo) is obtained or derived from a prokaryote selected from the group consisting of Microsystis aeruginosa, Methanocaldococcus jannaschii, Clostridium bartlettii, Exiguobacterium, Aromatoleum aromaticum, Synechococcus, Synechococcus elongatus, Thermus thermophilus, Pyrococcus furiosus, Aquifex aeolicus, Anoxybacillus flavithermus, and Thermosynechococcus elongatus.
[0081] In one embodiment, the Argonaute (pAgo) is obtained or derived from a prokaryote selected from the group consisting of Thermus thermophilus, Methanocaldococcus jannaschii, Pyrococcus furiosus, Aquifex aeolicus, Anoxybacillus flavithermus, and Thermosynechococcus elongatus. More preferably, the Argonaute is selected from the group consisting of Thermus thermophilus, Methanocaldococcus jannaschii or Pyrococcus furiosusi. Thermophilic Arognautes have the advantage that they may be stable and active at temperatures of at least 50°C, preferably 65 °C to 85 °C, which is beneficial for a wide range of applications.
[0082] In one embodiment, the Argonaute is obtained or derived from a thermophilic prokaryote. Such an Argonaute may be selected from the group consisting of Thermus thermophilus, Methanocaldococcus jannaschii, Pyrococcus furiosus, Aquifex aeolicus, Anoxybacillus flavithermus, and Thermosynechococcus elongatus. Preferably, the Argonaute is selected from the group consisting of Thermus thermophilus, Methanocaldococcus jannaschii or Pyrococcus furiosus.
[0083] In one embodiment, the Argonaute protein is selected from the group consisting of Thermus thermophilus (Genbank: WVY30436. 1; Genbank Release 261.0; June 14 2024), Methanocaldococcus jannaschii (Genbank: WP_010870838; Genbank Release 261.0; June 142024) or Pyrococcus furiosus (Genbank: 1U04_A; Genbank Release 261.0; June 142024.
[0084] In a preferred embodiment, the Argonaute protein comprises an amino acid sequence with at least 90% sequence identity with an Argonaute obtained, obtainable or (derived) from Thermus thermophilus (Genbank: WVY30436.1; Genbank Release 261.0; June 14 2024). Preferably 95%, more preferably 98%, most preferably 99%. The Argonaute protein preferably is biologically functional. Biologically functional in the context of Argonaute proteins means that the Argonaute protein is capable of cleaving a nucleic acid when placed in the appropriate physiological or experimental conditions necessary for its activity.
[0085] In one embodiment, the Argonaute is obtained, obtainable or is (derived) from Thermus thermophilus. Means for inhibiting the annealing of single-stranded DNA (ssDNA)
[0086] A means for inhibiting the annealing of single-stranded DNA (ssDNA) to a complementary nucleic acid strand is defined as a means that slows down the annealing process of two (partially) complementary single DNA strands compared to when the means is absent. Preferably, it relates to the annealing of a ssDNA to a (partially) complementary ssDNA.
[0087] Suitable means for inhibiting the annealing of ssDNA may be a DNA-binding protein, like single-strand binding protein (SSB), recombinase A (RecA) or DNA helicase, small molecule stabilizers, peptide nucleic acids (PNAs), locked nucleic acids (LNAs) and DNA mimic proteins.
[0088] In one embodiment, the means for inhibiting the annealing of ssDNA is a DNA- binding protein, preferably a thermostable DNA-binding protein.
[0089] In one embodiment, the means for inhibiting the annealing of ssDNA is a DNA- binding protein selected from the group consisting of single-strand binding protein (SSB), replication protein A (RPA), recombinase A (RecA), recombinase mediator protein (RMP) and DNA helicase. Examples of single-strand binding protein (SSB) are E. coli SSB, T7 gp2.5 SSB, and Phage phi29 SSB. Examples of Recombination Mediator Proteins (RMPs) are T4 Gene 32 Protein, T7 gene 2.5 product, (Phage lambda) RedB and (Rac prophage) RecT. Examples of DNA helicases are UV-specific DNA helicase UvrD and E. coli helicase II UvrD.
[0090] In one embodiment, the means for inhibiting the annealing of ssDNA is selected from the group consisting of single-strand binding protein (SSB), Tth RecA, E. coli RecA, T4 Gene 32 Protein, E. coli helicase II UvrD, E. coli SSB, T7 gp2.5 SSB, phage phi29 SSB, T7 gene 2.5 product, (phage lambda) RedB and (Rac prophage) RecT. UV-specific DNA helicase UvrD.
[0091] In one embodiment, the means for inhibiting the annealing of ssDNA is a singlestrand binding protein (SSB). A single-strand binding protein (SSB) is a protein that binds to single-stranded DNA (ssDNA). There is a strong preference for SSB as means to inhibit annealing of ssDNA in a method according to the present disclosure. SSB binds specifically to ssDNA, prevents secondary structure formation in ssDNA and facilitates the action of DNA processing enzymes. SSBs may be obtained or derived from a variety of organisms, including but not limited to bacteria, archaea, and eukaryotes, as long as the SSB inhibits the annealing of ssDNA.
[0092] In a preferred embodiment, the means for inhibiting the annealing of ssDNA is a thermostable single-strand binding protein (SSB), preferably extreme thermostable SSB (ET- SSB). A thermostable single-strand binding protein (SSB) is stable and biologically functional at a temperature of 50, preferably 65 °C, which is beneficial for a wide range of applications, in particular in combination with a thermostable Argonaute protein. ET-SSB is a singlestranded DNA-binding protein isolated from hyperthermophilic microorganisms and is known in the art for its stability at high temperatures, such as 95°C for up to 60 minutes. A thermostable Argonaute is an Argonaute protein that remains catalytically active for at least 10 minutes, preferably 30 minutes, at a temperature of at least 50° C, preferably 65° C.
[0093] In one embodiment, the SSB protein is obtained or derived from a thermophilic organism. The advantage of thermophilic SSB proteins is that they maintain their capacity to bind single-stranded DNA up to higher temperatures compared to mesophilic SSB proteins. Thermophilic SSB proteins can be active and stable up to temperatures of up to 100 °C. Preferably, the thermophilic SSB is obtained or is (derived) from an organism selected from the group consisting of Thermus thermophilus SSB (TtSSB), Pyrococcus furiosus SSB (PfSSB), Methanocaldococcus jannaschii SSB (MjSSB), Saccharolobus solfataricus SSB and Deinococcus radiodurans (DrSSB).
[0094] In a preferred embodiment, the SSB protein is obtained, obtainable or is (derived) from Thermus thermophilus SSB (TtSSB), in particular when used in a method according to the present disclosure in combination with an Argonaute protein obtained, obtainable or derived from Thermus thermophilus dsDNA pre-guide
[0095] In a method for providing a ssDNA guide according to the present disclosure, dsDNA pre-guide is dsDNA obtained or derived from biological samples (including human and animal samples), e.g., blood, urine, saliva, buccal swab, bodily fluid, bodily secretion, tissue, cultured cells, hair, hair follicle or fecal matter; plant samples, e.g., leaves, seeds, roots, cultured plant cells, or pollen; microbial samples, e.g., bacterial cultures; viral samples; environmental samples, e.g., soil samples and water samples; ancient DNA, i.e. , DNA extracted from ancient remains such as bones, teeth, or archaeological samples; and synthetic DNA, e.g., DNA synthesized using solid-phase synthesis. A person skilled in the art knows how to obtain double-stranded DNA (dsDNA) from these different types of samples.
[0096] The dsDNA (pre-guide) may be dsDNA obtained from a first sample and a further second sample. The dsDNA from the first and the second sample may be provided in the form of a mixture. For example, the dsDNA may be obtained from a first leaf sample and a second leaf sample. Alternatively, the mixture of dsDNA may comprise dsDNA obtained from a biological sample, such as a human blood sample, and a further second sample, such as a bacterial sample.
[0097] In one embodiment, the dsDNA pre-guide is dsDNA obtained or derived from a plant, mammal, fungus, or microorganism. Specimens or biopsy samples arising from diagnostic, therapeutic or surgical procedures may provide suitable sample material from which dsDNA can be obtained. Any kind of cell culture may provide dsDNA, whether entirely or in part. The cells may be of prokaryotic or eukaryotic origin. Examples of prokaryotic cell cultures are bacteria (including cyanobacteria) and archaea. Eukaryotic cell cultures may be any of protist, plant, fungi, algae, or animal, e.g. insect, bird, fish mammalian or human. More complex biological samples may be used, such as those taken from the environment, e.g. water samples, ice samples, soil samples, rock samples. Also within the scope of the disclosure are samples wherein there is viral or other nucleic acid containing material, which may be at a low level undetectable by current methods.
[0098] In one embodiment, the dsDNA pre-guide may be genomic DNA, plastid DNA, mitochondrial DNA, chloroplast DNA, plasmid DNA, viral DNA, extrachromosomal DNA, complementary DNA (cDNA), and / or synthetic DNA (i.e. , artificial DNA). The number of dsDNA pre-guides may be at least 1 , 5, 10, 10A2,10A3, 10A4 and / or at most 10A5, 10A6.
[0099] In one embodiment, the dsDNA pre-guide is dsDNA obtained or derived from a (polyploid) plant. Such plant can be any type of plant, e.g. a monocot or dicot. Non-limiting examples include Cucurbitaceae, Solanaceae and Gramineae, maize / corn (Zea spp.), wheat (Triticum spp.), barley (e.g. Hordeum vulgare), oat (e.g. Avena sativa), sorghum (Sorghum bicolor), rye (Secale cereale), soybean (Glycine spp., e.g., G. max), cotton (Gossypium spp., e.g., G. hirsutum, G. barbadense), Brassica spp. (e.g., B. napus, B. juncea, B. oleracea, B. rapa), sunflower (Helianthus annus), safflower, yam, cassava, alfalfa (Medicago sativa), rice (Oryza spp., e.g. O. sativa indica or japonica), forage grasses, pearl millet (Pennisetum spp. e.g. P. glaucum), tree species (Pinus, poplar, fir, plantain, etc), tea, coffea, oil palm, coconut, vegetable species, such as pea, zucchini, beans (e.g. Phaseolus spp.), cucumber, artichoke, asparagus, broccoli, garlic, leek, lettuce, onion, radish, lettuce, turnip, Brussels sprouts, carrot, cauliflower, chicory, celery, spinach, endive, fennel, beet, fleshy fruit bearing plants (grapes, peaches, plums, strawberry, mango, apple, plum, cherry, apricot, banana, blackberry, blueberry, citrus, kiwi, figs, lemon, lime, nectarines, raspberry, watermelon, orange, grapefruit, etc.), ornamental species (e.g. Rose, Petunia, Chrysanthemum, Lily, Gerbera species), herbs (mint, parsley, basil, thyme, etc.), woody trees (e.g. species of Populus, Salix, Quercus, Eucalyptus), fibre species e.g. flax (Linum usitatissimum) and hemp (Cannabis sativa), or model organisms, such as Arabidopsis thaliana.
[0100] In preferred embodiments, the dsDNA pre-guide is dsDNA obtained or derived from a (polyploid) plant selected from the group consisting of pearl millet, parsley, cotton, corn, lettuce, hop, ginger, tea, pepper, sunflower, lentil, barley, chrysanthemum, rye, wheat, oat, sugarcane, fir, leek, onion, triticale, and lily. The wheat may be durum wheat, winter wheat, or bread wheat. These crops may be of particular interest for repeat depletion due to their large genome sizes. In one embodiment, the dsDNA pre-guide is provided by converting ssDNA into dsDNA, e.g., by using DNA polymerase or allowing complementary ssDNA to anneal with each other. The ssDNA may for example, but not limited to, be obtained by enzymatic or chemical synthesis of ssDNA (single-stranded artificial or synthetic DNA), denaturing of dsDNA, or converting RNA into ssDNA, e.g., using reverse transcriptase. The dsDNA may also be provided or obtained using a polymerase chain reaction (PCR).
[0101] In one embodiment, the dsDNA pre-guide is dsDNA obtained from a synthetic oligonucleotide pool, which has been converted into double-stranded DNA, preferably using DNA polymerase I.
[0102] In one embodiment, the dsDNA pre-guide is provided by a method comprising the steps of: a) providing an in silico list of k-mers that represent sequences in a (genomic) sequence; preferably, the k-mers are 50-500 nucleotides in length, more preferably 80-300 nucleotides in length. b) selecting k-mers that are present at least 2 times in the in silico list, to obtain a selection of k-mers; preferably the k-mers are present at least 3 times, more preferably at least 5 times. c) providing (a plurality of) dsDNA (each) comprising at least one k-mer sequence of the selection of k-mers, which is / are the dsDNA pre-guide(s).
[0103] The provided dsDNA may be synthetic DNA, e.g., produced by PCR amplification or DNA synthesis.
[0104] In one embodiment, the dsDNA is selectively enriched for dsDNA comprising a sequence of interest (for example, a repetitive sequence) prior to generating ssDNA guides using an Argonaute protein. Selective enrichment refers to the process of selectively increasing the relative amount of at least one nucleic acid comprising a sequence of interest within a mixture of (genomic) dsDNA compared to the mixture before enrichment. Such enrichment may be performed using known methods in the art, such as, but not limited to, hybridization-based enrichment, PCR-based methods to amplify specific regions using primers, denaturing double-stranded DNA into single strands and then allowing them to reanneal, and circularization techniques. The advantage of such enrichment is that dsDNA enriched for a sequence of interest, e.g., a repetitive sequence, may be used to generate ssDNA guides with a method according to the present disclosure. This results in the generation of ssDNA guides comprising the sequence of interest, e.g., a repetitive sequence. The advantage of this method is that ssDNA guides may be generated directed at a sequence of interest without the need to be aware of - and comply with - specific design rules such as those applicable for the design of, for example, CRISPR-Cas RNA guides. ssDNA guide
[0105] As used herein, the term “guide DNA” refers to a single stranded oligonucleotide composed of at least 50% deoxyribonucleotides (e.g., at least 60%, at least 70%, at least 80%, at least 90%, or 100% deoxyribonucleotides). The guide DNA is capable of directing an Argonaute polypeptide:guide DNA complex to a target nucleic acid. More specifically, the guide DNA is believed to bind an Argonaute protein and to hybridize to a target nucleic acid.
[0106] In one embodiment, the DNA guide length suitable for Argonaute cleavage of dsDNA in the presence of a single strand binding protein may comprise at least 12 nucleotides, for example having a size range of 12-60 nucleotides, 14-50 nucleotides, 15-40 nucleotides, 16- 35 nucleotides, 15-24 nucleotides, or 16-21 nucleotides.
[0107] In another embodiment, the guide DNA may be greater than 21 nucleotides or at least 24 nucleotides in length. In many embodiments, the guide DNA can be 16-21 nucleotides in length (i.e., 16, 17, 18, 19, 20 or 21 nucleotides). The number of dsDNA guides may be at least 1, 5, 10, 100, 1000, 10A4 and / or at most 10A5, 10A6.
[0108] As the ssDNA guide is generated using apo-pAgo, the length of the ssDNA guide may vary between ssDNA guides depending on the application and the specific pAgo used. Generally, the ssDNA guides comprise at least 15 nucleotides in length. Preferably, the ssDNA guides are 13-25 nucleotides long.
[0109] In some embodiments, the guide DNA may comprise a modified nucleotide or nucleotide analogue. Modified nucleotides refer to any nucleotide that has been chemically altered from its natural form, either by modifying the base, sugar, or phosphate backbone. Modified nucleotides include artificial nucleotides specifically engineered for particular applications, such as diagnostics, therapeutics, or research tools. Nucleotide analogs may cover simple base substitutions to complex chemical alterations that confer new properties or functions to nucleic acids. In particular, the DNA guide may be 5’-phosphorylated. In some embodiments, the nucleotide sugar modification comprises a 2' sugar modification and maybe selected from the group consisting of a 2'-0 — CH3, a 2'-F, and a 2’-MOE modification. In other embodiments, the nucleotide substitution comprises one selected from the group consisting of locked nucleic acid (LNA), an unlocked nucleic acid (UNA), deoxyuridine, pseudouridine, 5- methylcytosine, 2-aminopurine, 2,6-diaminopurine, deoxyinosine, 5-hy- droxybutynl-2'- deoxyuridine, 8-aza-7-deazaguanosine, and 5-nitroindole. In further embodiments, the guide molecule comprises a sugar modification and a nucleotide substitution. Guide DNAs should have a 5' phosphate.
[0110] In one embodiment, the DNA guide may have a 5' phosphate.
[0111] A DNA guide may be generated from a synthetic or from a natural source such as genomic DNA, cDNA, extrachromosomal DNA, microbial DNA or viral DNA. The DNA guide preferably is single-stranded when used with an Argonaute protein although it may be derived from dsDNA.
[0112] Method for nicking or cleaving a nucleic acid
[0113] In one embodiment, the disclosure provides a method for nicking or cleaving a target nucleic acid, the method comprising the steps of: a) providing an ssDNA guide according to the present disclosure; b) providing a prokaryotic Argonaute protein (apo-pAgo); c) providing a target nucleic acid comprising a target sequence, wherein (one strand of) the target sequence is complementary to at least part of the ssDNA guide; preferably the ssDNA guide is fully complementary to the nucleic acid comprising the target sequence; d) combining the generated ssDNA guide, the apo-pAgo, and the nucleic acid comprising the target sequence to nick or cleave the nucleic acid comprising the target sequence.
[0114] In preferred embodiments, the apo-pAgo is loaded with the ssDNA guide to obtain holo-pAgo before step d).
[0115] In the context of the present disclosure, the target nucleic acid may be DNA or RNA. Preferably, the target nucleic acid is DNA. The target nucleic acid may be ssDNA or dsDNA.
[0116] The target nucleic acid, which may be DNA or RNA, may be provided within a mixture containing further nucleic acids, such as a library of DNA molecules. An example of such a library is a sequencing library. Short-read sequencing libraries typically contain DNA fragments between 300-600 base pairs. The library may be provided by fragmenting (genomic) DNA to the desired size-range prior to library construction.
[0117] In one embodiment, the target nucleic acid is DNA comprising one or more sequences that are at least partially complementary to the ssDNA guide. The target nucleic acid can be part of a gene, a 5' end of a gene, a 3' end of a gene, a regulatory element (e.g. promoter, enhancer), a pseudogene, non-coding DNA, repetitive DNA, a microsatellite, an intron, an exon, chromosomal DNA, mitochondrial DNA, sense DNA, plasmid DNA, antisense DNA, nucleoid DNA, chloroplast DNA, or RNA among other nucleic acid entities.
[0118] In one embodiment, the target nucleic acid may be obtained or is (obtainable) from biological samples (including human and animal samples), e.g., blood, urine, saliva, buccal swab, bodily fluid, bodily secretion, tissue, cultured cells, hair, hair follicle or fecal matter; plant samples, e.g., leaves, seeds, roots, cultured plant cells, or pollen; microbial samples, e.g., bacterial cultures; viral samples; environmental samples, e.g., soil samples and water samples; ancient DNA, i.e. , DNA extracted from ancient remains such as bones, teeth, or archaeological samples; and synthetic DNA or RNA. A person skilled in the art knows how to obtain DNA and / or RNA from these different types of samples. In one embodiment, the target nucleic acid may be obtained or is (obtainable) from a first sample and a further second sample. For example, RNA may be obtained from a first leaf sample and a second leaf sample. Alternatively, a mixture of nucleic acids may comprise dsDNA obtained from a first biological sample, such as a human blood sample, and RNA from a further second sample, such as a bacterial sample. Alternatively, a mixture of nucleic acids may comprise dsDNA originating from different parts of the cell, such as dsDNA derived from the nuclear genome and from the chloroplasts of (cultured) plant cells.
[0119] In one embodiment, the target nucleic acid is obtained or derived from a plant, mammal, fungus, or microorganism. Specimens or biopsy samples arising from diagnostic, therapeutic or surgical procedures may provide suitable sample material from which nucleic acid can be obtained. Any kind of cell culture may provide nucleic acid, whether entirely or in part. The cells may be of prokaryotic or eukaryotic origin. Examples of prokaryotic cell cultures are bacteria (including cyanobacteria) and archaea. Eukaryotic cell cultures may be any of protist, plant, fungi, algae, or animal, e.g. insect, bird, fish mammalian or human. More complex biological samples may be used, such as those taken from the environment, e.g. water samples, ice samples, soil samples, rock samples. Also within the scope of the disclosure are samples wherein there is viral or other nucleic acid containing material, which may be at a low level undetectable by current methods.
[0120] In one embodiment, the (target) nucleic acid may be genomic DNA, plastid DNA, mitochondrial DNA, chloroplast DNA, plasmid DNA, viral DNA, extrachromosomal DNA, complementary DNA (cDNA), and / or synthetic DNA (i.e. , artificial DNA). The nucleic acid may be fragmented.
[0121] In one embodiment, the (target) nucleic acid is obtained or derived from a (polyploid) plant. Such plant can be any type of plant, e.g. a monocot or dicot. Non-limiting examples include Cucurbitaceae, Solanaceae and Gramineae, maize / corn (Zea spp.), wheat (Triticum spp.), barley (e.g. Hordeum vulgare), oat (e.g. Avena sativa), sorghum (Sorghum bicolor), rye (Secale cereale), soybean (Glycine spp., e.g., G. max), cotton (Gossypium spp., e.g., G. hirsutum, G. barbadense), Brassica spp. (e.g., B. napus, B. juncea, B. oleracea, B. rapa), sunflower (Helianthus annus), safflower, yam, cassava, alfalfa (Medicago sativa), rice (Oryza spp., e.g. O. sativa indica or japonica), forage grasses, pearl millet (Pennisetum spp. E.g. P. glaucum), tree species (Pinus, poplar, fir, plantain, etc), tea, coffea, oil palm, coconut, vegetable species, such as pea, zucchini, beans (e.g. Phaseolus spp.), cucumber, artichoke, asparagus, broccoli, garlic, leek, lettuce, onion, radish, lettuce, turnip, Brussels sprouts, carrot, cauliflower, chicory, celery, spinach, endive, fennel, beet, fleshy fruit bearing plants (grapes, peaches, plums, strawberry, mango, apple, plum, cherry, apricot, banana, blackberry, blueberry, citrus, kiwi, figs, lemon, lime, nectarines, raspberry, watermelon, orange, grapefruit, etc.), ornamental species (e.g. Rose, Petunia, Chrysanthemum, Lily, Gerbera species), herbs (mint, parsley, basil, thyme, etc.), woody trees (e.g. species of Populus, Salix, Quercus, Eucalyptus), fibre species e.g. flax (Linum usitatissimum) and hemp (Cannabis sativa), or model organisms, such as Arabidopsis thaliana.
[0122] In preferred embodiments, the (target) nucleic acid (DNA) is obtained or derived from a (polyploid) plant selected from the group consisting of pearl millet, parsley, cotton, corn, lettuce, hop, ginger, tea, pepper, sunflower, lentil, barley, chrysanthemum, rye, wheat, oat, sugarcane, fir, leek, onion, triticale, and lily. The wheat may be durum wheat, winter wheat, or bread wheat. These crops may be of particular interest for repeat depletion due to their large and complex genomes.
[0123] In preferred embodiments, the (target) nucleic acid is DNA. The DNA may be obtained, obtainable or may be (derivable) from plant.
[0124] In one embodiment, the (target) nucleic acid, preferably DNA, comprising a target nucleic acid may comprise one or more sequences that are at least partially complementary to the ssDNA guide. The target nucleic acid can be part of a gene, a 5' end of a gene, a 3' end of a gene, a regulatory element (e.g. promoter, enhancer), a pseudogene, non-coding DNA, repetitive DNA, a microsatellite, an intron, an exon, chromosomal DNA, mitochondrial DNA, sense DNA, antisense DNA, nucleoid DNA, chloroplast DNA, or RNA among other nucleic acid entities. The target nucleic acid can be part or all of a plasmid DNA. The plasmid DNA or a portion thereof may be negatively supercoiled. The target nucleic acid can be in vitro or in vivo.
[0125] In a preferred embodiment of step d), SSB is also combined with the generated ssDNA guide, the apo-pAgo, and the nucleic acid comprising the target sequence to nick or cleave the nucleic acid comprising the target sequence. This enhances cleavage by the ssDNA guide:pAgo complex, in particular if the target nucleic acid is dsDNA.
[0126] In one embodiment the target nucleic acid is DNA, preferably dsDNA. In some embodiments, the nucleic acid may be DNA obtained or derived from a plant, preferably a polyploid plant.
[0127] Method of reducing the amount of a target nucleic acid in a composition
[0128] In a further aspect, the disclosure relates to a method of reducing the amount of a target nucleic acid in a composition, the method comprising the steps of: a) generating the ssDNA guide with a method according to the disclosure; b) providing a prokaryotic Argonaute protein (pAgo); c) providing nucleic acid molecules, wherein at least one nucleic acid molecule is a target nucleic acid molecule comprising a target sequence that is complementary to at least part of the ssDNA guide; d) combining the ssDNA guide, apo-pAgo and the nucleic acid molecules to nick or cleave the at least one target nucleic acid molecule.
[0129] In a further aspect, the present disclosure relates to a method for selectively amplifying and / or sequencing (low-copy or non-repetitive) sequences in a composition comprising dsDNA fragments, the method comprising the steps of: a) generating an ssDNA guide with a method according to the present disclosure; b) providing a prokaryotic Argonaute protein (pAgo); c) providing a composition comprising nucleic acids, wherein at least one nucleic acid is a target nucleic acid comprising a target sequence that is complementary to at least part of the ssDNA guide; d) combining the pAgo, ssDNA guide and the nucleic acids to nick or cleave the at least one target nucleic acid; e) amplifying and / or sequencing the unnicked and uncleaved nucleic acids.
[0130] In some embodiments, the method according to the present disclosure comprises the steps of: a) providing DNA fragments, preferably obtained by fragmenting genomic DNA; b) providing pAgo; c) providing a ssDNA guide corresponding to repetitive sequences in the genomic DNA, preferably provided by a method according to the present disclosure; d) providing a means for inhibiting the annealing of single-stranded DNA (ssDNA) to a complementary ssDNA, preferably SSB; e) ligating adapters to both ends of the DNA fragments to obtain adapter-ligated DNA fragments, wherein the adapter comprises an amplification primer-binding site and / or a sequencing primer-binding recognition site; f) combining the adapter-ligated DNA fragments with apo-pAgo, the ssDNA guide, and single stranded binding protein (SSB) so as to allow cleavage of repetitive sequences in the sample of DNA fragments; g) amplifying and / or sequencing the unnicked and uncleaved DNA fragments by using primers that bind the adapters.
[0131] Composition
[0132] In a further aspect, the present disclosure provides a composition comprising an unloaded prokaryotic DNA-guided Argonaute (apo-pAgo) as disclosed herein; means for inhibiting the annealing of single-stranded DNA (ssDNA) to a complementary ssDNA as disclosed herein, preferably a Single-Strand Binding protein (SSB) as disclosed herein; and a dsDNA pre-guide as disclosed herein. These components of the composition may be included in a kit. The kit may further comprise an aqueous buffer as disclosed herein.
[0133] The kit may also include a reaction buffer. The components in the kit may be in the same or different tubes.
[0134] In preferred embodiments, the composition comprises a thermostable, unloaded prokaryotic DNA-guided Argonaute (apo-pAgo) as disclosed herein, a thermostable SingleStrand Binding protein (SSB) as disclosed herein, preferably an Extreme Thermostable Single-Stranded DNA Binding Protein (ET-SBB) and a dsDNA pre-guide as disclosed herein.
[0135] In one embodiment, the composition comprises all the components necessary to perform a method for providing a ssDNA guide as described herein.
[0136] As used herein, the term “composition” refers to a combination of reagents that may contain other reagents, e.g., glycerol, salt, dNTPs, etc., in addition to those listed. A composition may be in any form, e.g., aqueous or lyophilized, and may be at any state (e.g., frozen or in liquid form).
[0137] In some embodiments, the composition may be an aqueous solution that comprises a non-naturally occurring buffering agent at a concentration of 1 mM to about 200 mM, salt (e.g., KC1 or NaCI), a divalent cation (e.g., Mg2+or Mn2+). In these embodiments, the composition can have a pH of about 7 to about 8.8, such as 7.4 to 7.5.
[0138] In some embodiments, the composition may further contain a reducing agent (e.g., dithiothreitol), a detergent, glycerol, sugar or dNTPs, as needed.
[0139] The composition described above may be employed in a variety of methods, particularly in vitro (i.e. , cell free) methods, including, but not limited to digestion, depletion, or enrichment methods; or as an initiator of amplification; for purposes that include recognizing a nucleic acid for editing; detecting nucleic acid sequence via binding to Argonaute; and sequencing of target nucleic acids.
[0140] In a further aspect, the present disclosure provides use of the composition as disclosed above for the generation of ssDNA guides.
[0141] Other applications
[0142] The methods according to the present disclosure may be used for repeat depletion or depletion of another selected target sequence.
[0143] In a method according to the present disclosure, the dsDNA pre-guide may be (derived from) plastid DNA (chloroplast and / or mitochondrial DNA) in addition or instead of nuclear DNA. In such methods, the generated ssDNA guides may be used for depleting plastid DNA and / or mitochondrial DNA in a sample, such as a whole genome DNA sequencing library. In some embodiments, the sequencing is skim sequencing, also known as low pass sequencing.
[0144] In some embodiments, a nucleic acid library, preferably a DNA library, is depleted for a particular target sequence with a method according to the present disclosure at least once, preferably twice.
[0145] In one embodiment, the sequencing is skim sequencing and the nucleic acid library, preferably a DNA library, is depleted for a particular target sequence with a method according to the present disclosure at least once, preferably twice.
[0146] In one embodiment, a method according to the present disclosure is used to remove a microbial sequence in a DNA library, preferably the DNA is from ancient archaeological specimens. Ancient DNA fragments are short molecules with significant chemical damage acquired over time. Sequencing of ancient DNA helps addressing questions in anthropology, evolutionary biology, and environmental and archaeological sciences, improving the understanding of major prehistoric and historic events. Detection of ancient DNA is challenging due to genetic material obtained from ancient sources is largely comprised of microbial DNA. If a small fraction of the ancient DNA library is sequenced, TtAgo pre-guides against the detected microbial contaminating sequences may be designed according to the present disclosure. The guides can partially remove the contaminating DNA, thereby enriching for the small fraction of ancient DNA of interest.
[0147] In one embodiment, the target nucleic acid may be RNA or cDNA, e.g. to remove ribosomal DNA. Preferably, the RNA is converted to cDNA. Subsequently, ribosomal RNA may be removed from cDNA libraries, in particular after its conversion to DNA.
[0148] Brief description of the figures
[0149] Figure 1 : Two examples of a dsDNA pre-guide library being chopped by TtAgo in two independent experiments. With SSB (30 pmol in a 20 pl reaction volume), the chopping is much more effective than without SSB.
[0150] Figure 2: Different concentrations of dsDNA pre-guide library being chopped by TtAgo. The conversion of pre-guides into guides was assessed using 0.25 to 4 pmol of pre-guides with a constant 2 pmol of TtAgo. Across all conditions, the inclusion of SSB significantly increased the amount of ssDNA guides produced, while reducing the amount of pre-guide dsDNA. Only when 1 pmol of pre-guides or less (a 1 :2 ratio of dsDNA pre-guide to TtAgo) was combined with SSB, all or almost all of the dsDNA pre-guides appeared processed.
[0151] Figure 3: Overview of protocol to deplete repeats from an Illumina whole genome sequencing library. TtAgo and dsDNA are mixed (step 1) and TtAgo generates ssDNA guides from the dsDNA through a process called chopping (step 2). A standard DNA library preparation is performed. After ligation of stubby adapters (step 3), TtAgo loaded with ssDNA guides is incubated with the sequencing library. TtAgo will cleave repetitive DNA in the library, leading to repeat depletion (step 4). Subsequently, the library is purified and PCR amplified (step 5). The fragments that are cut will not be amplified. The repeat library is sequenced (step 6). Figure 4: Repeat depletion in papaya at an example locus. Normalized coverage is visualized for the four experimental conditions (top to bottom, left to right). Vertical grey bars represent expected depletion targets. Several large clusters of depletion targets are predicted. At these loci, the normalized coverage is strongly reduced for the experiment with TtAgo loaded with pre-guides.
[0152] The present disclosure will be further detailed in the following examples. Unless stated otherwise all experiments were carried out according to standard protocols.
[0153] EXAMPLES
[0154] Example 1 : Production of pre-guides
[0155] Two libraries targeting repetitive DNA in papaya and cucumber were generated as a proof of principle.
[0156] Synthetic DNA pre-guides were designed in the following manner. An in silico list of k-mers that represent the whole genome sequences of the nuclear genome of cucumber and papaya was generated. The k-mers were 80-300 nucleotides in length.
[0157] A selection was made of k-mers present at least 3 times in the in silico list of the cucumber and papaya genomes to obtain a selection of kmers per species. A total of 254 selected k-mers were synthesized as ssDNA and then converted into dsDNA to produce an oligonucleotide pool (oPool) library of dsDNA pre-guides using methods known in the art, e.g., DNA Polymerase I and Large (Klenow) Fragment.
[0158] Example 2: Generation of ssDNA guides
[0159] The generated dsDNA pre-guides were combined with water, a reaction buffer, apo- TtAgo and ET-SBB to form a reaction mixture and so as to allow the generation of ssDNA guides by apo-pAgo. The reaction buffer (1x) comprises 20 mM Tris-HCI, 10 mM (NH4)2SO4, 10 mM KCI, 2 mM MgSO4, and 0.1% polyethylene glycol p-(1,1,3,3-tetramethylbutyl)-phenyl ether (Triton®-X-100). The pH is 8.8 at 25°C.
[0160] The reaction mixture was incubated for 20 hours at 75 °C.
[0161] Table 1.
[0162] After incubation of the reaction mixtures, the mixtures were subjected to gel electrophoresis to visualize the DNA. The observed reduction in the intensity of the dsDNA band indicates that TtAgo has effectively cleaved the DNA pre-guide molecules. The results thus demonstrate that adding SBB significantly increased the chopping activity of the Argonaute protein (see Figure 1).
[0163] Example 3: Generation of ssDNA guides using varying reaction conditions
[0164] The concentrations of SSB, pre-guides and TtAgo, the GC-content of the pre-guides and the number of pre-guides per library to observe the effect on the chopping reaction.
[0165] SSB concentration: The generation of ssDNA guides was repeated as described in Example 2, with modifications to the SSB concentration, which was 10, 20, 30, or 40 pmol. Subsequently, the mixtures were subjected to electrophoresis to visualize the DNA. It was observed that increasing the SSB concentration above 10 pmol did not significantly impact the cleavage efficiency of the chopping process (data not shown).
[0166] Amount of DNA pre-guides: The generation of ssDNA guides was repeated as described in Example 2, with modifications to the dsDNA pre-guide concentration (0.25, 0.5, 1 , 2, and 4 pmol). Subsequently, the mixtures were subjected to electrophoresis to visualize the DNA (see Figure 2).
[0167] Pre-guides were almost completely processed at TtAgo: pre-guide ratios of 2:2, 1:2, 0.5:2 and 0.25:2. Increasing the pre-guide concentration to a ratio of 4:2 for DNA pre- guides:TtAgo: resulted in partially unprocessed dsDNA pre-guides. Omitting SSB from the incubations led to a significantly higher proportion of unprocessed dsDNA pre-guides, showing that SBB enhanced production of ssDNA guides by TtAgo, despite the presence of unprocessed dsDNA pre-guides (Figure 2). Unprocessed dsDNA pre-guides may be removed in a subsequent step, e.g., by treatment with a DNase.
[0168] Number of DNA pre-guides: The number of pre-guides per library was adjusted (10, 17, 77). The total amount of pre-guide DNA remained 1 pmol and the GC content was similar for each reaction mixture. After incubation of the reaction mixtures, proteinase K digestion was performed to remove proteins from the mixtures. Subsequently, the mixtures were subjected to electrophoresis to visualize the DNA.
[0169] A similar intensity for all bands was observed meaning the number of different DNA pre-guides in a reaction did not significantly impact the efficiency of the chopping process (data not shown).
[0170] All the experiments showed that the addition of SSB significantly improves chopping activity.
[0171] Example 4: Repeat depletion in DNA sequencing library
[0172] A DNA sequencing library was generated by fragmenting isolated plant DNA into approximately 550 bp fragments and ligating (stubby) adapters to the generated fragments.
[0173] TtAgo proteins were loaded with 77 different ssDNA guides targeting repetitive sequences in the papaya and cucumber genomes. These libraries were expected to target 10.2% and 15% of the genome for papaya and cucumber, respectively, based on the number of instances these repeats occur in the respective (nuclear) genomes.
[0174] The loaded TtAgo and ssDNA guides were mixed with the components listed in Table 2.
[0175] Table 2.
[0176] The reaction buffer (1x) comprises 20 mM Tris-HCI, 10 mM (NH^SOt, 10 mM KCI, 2 mM MgSC and 0.1% polyethylene glycol p-(1,1,3,3-tetramethylbutyl)-phenyl ether (Triton®-X-100). The pH is 8.8 at 25°C.
[0177] The reaction was incubated for 12h at 80 °C.
[0178] After an overnight (ON) incubation at 80°C, the reaction was cleaned up to remove the buffer and proteins and the standard library preparation protocol (IDT xGen NGS library preparation kit) was continued with a PCR amplification using primers that target the stubby adapters. Note that the DNA fragments cleaved by TtAgo will not be amplified by PCR as they lack adapter sequences on both flanks which are required for PCR amplification (see Figure 3, step 5).
[0179] Example 5: Sequencing and analysis of library depleted for repetitive sequences
[0180] The library depleted for repetitive sequences was subsequently sequenced using a standard protocol (Illumina).
[0181] The sequencing reads were processed using a standard procedure. Reads were first trimmed to remove low-quality base calls and adapter sequences. Afterwards, reads were aligned to the reference genome. Only primary alignments were returned. This means that the best alignment was returned for every read, even when additional alignments were possible. If multiple alignments were equally good, the read was randomly assigned to one alignment. The reads that align to repetitive sequences were thus assigned to one of the repeats, which ensured that every read was only counted once in our downstream processing.
[0182] A coverage profile described for every genomic locus how many reads align at a given locus and were' calculated using all aligned reads. To normalize for the library size (the number of sequenced reads), coverages were normalized to RPKM values (Reads Per Kilobase per Million mapped reads). This standard approach ensured that coverages between samples could be directly compared.
[0183] To accurately quantify the amount of repeat depletion, the normalized coverage at expected genomic targets was compared between a repeat-depleted library and a WGS library. This analysis could be performed for all genomic targets combined, reflecting the effective repeat depletion of an experiment. Alternatively, the analysis could be performed for the genomic targets of a given pre-guide. The latter can be used to determine the effectiveness of pre-guides in depleting target sequences.
[0184] A strong depletion in normalized coverage at the expected depletion targets can be observed for the condition with TtAgo and pre-guides. To better quantify the extent of repeat depletion, the sequencing coverage was further normalized relative to the WGS sample (Figure 4). In the latter visualization, local “spikes” of depletion can be seen besides the depletion at the large blocks of repetitive DNA. Similar patterns are observed in the cucumber experiment (data not shown).
[0185] The amount of depletion at target and non-target loci were quantified. For the papaya and cucumber repeat depletion libraries, the maximum amount of depletion that could be achieved was 10.2% and 15.0%, corresponding to the combined size of the depletion targets with 150 bps flanks. Approximately 50% of the signal is depleted at all targets for both the papaya and cucumber experiments. Combined, these pre-guide libraries resulted in a 4- 8% reduction in sequencing costs compared to sequencing whole genome libraries without depletion and equal depth of coverage. To conclude, repeat depletion was achieved in multiple species, resulting in an increased sequencing coverage of genomic loci of interest when a similar number of reads is sequenced.
Claims
CLAIMS1. Method for providing a single-stranded deoxyribonucleic nucleic acid (ssDNA) guide for a prokaryotic Argonaute protein (pAgo), the method comprising the steps of: a) providing an unloaded prokaryotic Argonaute (apo-pAgo); b) providing a single-strand binding protein (SSB); c) providing a double-stranded deoxyribonucleic nucleic acid (dsDNA) pre-guide; d) combining the apo-pAgo, the SSB and the dsDNA pre-guide so as to allow apo- pAgo to provide the single-stranded deoxyribonucleic nucleic acid (ssDNA) guide.
2. Method according to claim 1 , wherein the apo-pAgo is from a thermophilic prokaryote; preferably the apo-pAgo is from a thermophilic prokaryote selected from the group consisting of Thermus thermophilus, Methanocaldococcus jannaschii, and Pyrococcus furios us more preferably Thermus thermophilus.
3. Method according to claim 1 or 2, wherein the SSB is from a eukaryote, archaea or virus; preferably the SSB is from a thermophilic organism; from an organism selected from the group consisting of Thermus thermophilus SSB (TtSSB), Pyrococcus furiosus SSB (PfSSB), Methanocaldococcus jannaschii SSB (MjSSB), Saccharolobus solfataricus and Deinococcus radiodurans (DrSSB), preferably extreme thermostable SSB (ET-SSB).
4. Method according to any one of the preceding claims, wherein the method further comprises the step of: e) loading the apo-pAgo with the provided ssDNA guide to obtain a loaded pAgo (holo-pAgo).
5. Method according to any one of the preceding claims, wherein in step d) apo- pAgo is in a molar ratio of dsDNA pre-guide:apo-pAgo of 5 or less, preferably 2 or less, more preferably 0.5 or less.
6. Method according to any one of the preceding claims, wherein in step d) apo- pAgo is in a molar ratio of dsDNA pre-guide:apo-pAgo of between 0.05 (dsDNA pre-guide):1 (apo-pAgo) and 2 (dsDNA pre-guide):1 (apo-pAgo), preferably 0.125 (dsDNA pre-guide):1 (apo-pAgo) and 1 (dsDNA pre-guide):2 (apo-pAgo).
7. Method according to any one of the preceding claims, wherein in step d) apo- pAgo is in a molar ratio of dsDNA pre-guide:apo-pAgo of between 1 (dsDNA pre-guide): 1 (apo-pAgo) and 1 (dsDNA pre-guide):2 (apo-pAgo).
8. Method according to any one of the preceding claims, wherein SSB is present in an amount of at least 5 pmol; preferably at least 10 pmol.
9. Method according to any one of the preceding claims, wherein step d) is performed:- at a temperature of between 40-99 °C; preferably 50-95 °C; more preferably 65 to 85 °C; and / or- for 1-48 hours; preferably 2-24 hours; more preferably 5 to 16 hours.
10. Method according to any one of the preceding claims, wherein the dsDNA preguide is provided by a method comprising the steps of: a) providing an in silico list of kmers that represent sequences in a (genomic) sequence; preferably, the kmers are 50-500 nucleotides in length; b) selecting kmers that are present at least 2 times in the in silico list, to obtain a selection of kmers; c) providing dsDNA comprising a kmer sequence of the selection of kmers.
11. Method for providing an ssDNA guide comprising the step of combining dsDNA with an unloaded prokaryotic Ago protein (apo-pAgo) and an SSB to provide the ssDNA guide.
12. Method for nicking or cleaving a target nucleic acid, the method comprising the steps of: a) providing an ssDNA guide with a method according to any one of the preceding claims; b) providing a prokaryotic Argonaute protein (pAgo); c) providing a target nucleic acid comprising a target sequence, wherein the target sequence is complementary to at least part of the ssDNA guide; d) combining the ssDNA guide, the apo-pAgo, and the target nucleic acid to nick or cleave the target nucleic acid.
13. Method for selectively amplifying and / or sequencing nucleic acids in a sample, the method comprising the steps of:a) providing an ssDNA guide with a method according to any one of the preceding claims; b) providing a prokaryotic Argonaute protein (pAgo); c) providing nucleic acids, wherein at least one nucleic acid is a target nucleic acid comprising a target sequence that is complementary to at least part of the ssDNA guide; d) combining the pAgo, ssDNA guide and the nucleic acids to nick or cleave the at least one target nucleic acid; e) amplifying and / or sequencing the unnicked and uncleaved nucleic acids.
14. Method according to claim 12 or 13, wherein the ssDNA guide is fully complementary to the target sequence in the target nucleic acid.
15. Method according to any one of claims 12 to 14, wherein before and / or during step d) the apo-pAgo is loaded with the ssDNA guide to obtain holo-pAgo.
16. Method according to any one of claims 12 to 15, wherein the target nucleic acid is DNA.
17. Method according to any one of claims 12 to 16, wherein the target nucleic acid is obtained from a (polyploid) plant, preferably from a plant selected from the group consisting of pearl millet, parsley, cotton, corn, lettuce, hop, ginger, tea, pepper, sunflower, lentil, barley, chrysanthemum, rye, wheat, oat, sugarcane, fir, leek, onion, triticale, and lily.
Citation Information
Patent Citations
In vitro cleavage of DNA using argonaute
US11466264B2
Methods of enriching nucleic acids
WO2023148235A1