Mutant cpf1 endonucleases
By developing variant Cpf1-terminal onucleases that can introduce single-stranded or double-stranded breaks, the limitations of existing CRISPR-terminal onucleases are solved in the application, and the effect of efficient detection and quantification of nucleic acid sequences and diagnosis of infectious diseases is achieved at extremely low concentrations.
Patent Information
- Application Number
- JP2024226441
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-08-27
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-14
AI Technical Summary
Existing CRISPR-sided procureases have limitations in applications and have failed to fully realize their potential, especially in optimized use in new and known applications.
The variant Cpf1-terminal onucleases was developed, which can introduce single-stranded or double-stranded breaks into single-stranded or double-stranded nucleic acid target sequences and bind to single-stranded or double-stranded DNA targets without cleavage, and has the ability to nonspecifically cleave single-stranded DNA.
These variants Cpf1-terminal nucleases can effectively detect and quantify nucleic acid sequences at very low concentrations, are suitable for the diagnosis of infectious diseases, and show efficient cleavage capabilities in double-stranded or single-stranded target nucleic acids.
Smart Images

Figure 2025075025000001_ABST
Abstract
Description
[Technical field]
[0001] Technical Field The present invention relates to mutant Cpf1 (also known as Cas12a) endonucleases with altered activity compared to wild-type Cpf1, and their use for inducing single-strand or double-strand breaks in nucleic acid target sequences, either single-strand or double-strand, and for inducing single-strand DNA breaks after binding (but not cleaving) to DNA target in a single-strand or double-strand non-specific manner.Methods for detecting and quantifying nucleic acid sequences are also disclosed.Methods for diagnosing infectious diseases are also disclosed. [Background technology]
[0002] background Cpf1 is a DNA endonuclease belonging to the CRISPR-Cas (clustered regularly interspaced short palindromic repeats) class 2 VB type adaptive immune system used by bacteria and archaea to protect cells against the invasion of potentially dangerous DNA molecules (Makarova et al., 2011; Makarova and Koonin, 2015; Vestergaard et al., 2014). Cpf1 forms a ribonucleoprotein complex together with a relatively small (40-45 base) crRNA (CRISPR-RNA) molecule that can recognize, unwind, and cut DNA targets complementary to the crRNA with high specificity. The DNA target sequence of Cpf1 is stored in CRISPR-arrays, regions of DNA containing repeated sequences flanked by unique sequences (spacers) that are a record of potentially dangerous DNA that the cell has previously encountered (Zetsche et al., 2015). CRISPR arrays are transcribed as a single pre-crRNA, which is then processed by Cpf1 to generate mature crRNAs (Fonfara et al., 2016). These RNAs therefore contain a conserved portion, identical for all crRNAs, that folds into a pseudoknot structure recognized by Cpf1, and a variable region that is complementary to the target DNA. To maintain the integrity of the CRISPR-array, cleavage of the DNA target is coupled to the recognition of a short PAM sequence (protospacer adjacent motif) 3 or 4 nucleotides long upstream of the target site. The PAM sequence recognized by Cpf1 is present in the invading DNA molecule but not in the CRISPR-array, thus protecting this DNA region from cleavage (Mojica et al., 2005).
[0003] Several CRISPR endonucleases, such as Cas9, Cpf1 and Cas13, have been used to modify the genomes of several organisms, as well as as diagnostic molecules to identify infections and also as potential tools for gene therapy (Chen et al., 2018; Gootenberg, 2018).
[0004] However, CRISPR endonucleases have so far been used in the same way that they act in nature, and there is a need to discover the full potential of these enzymes and optimize them for use in known as well as new applications. Summary of the Invention
[0005] Abstract The present disclosure relates to mutant Cpf1 endonucleases that can introduce single-stranded or double-stranded breaks in nucleic acid target sequences that are either single-stranded or double-stranded. Furthermore, the mutant Cpf1 endonucleases of the present disclosure can bind (without cutting) either single-stranded or double-stranded DNA targets and can cleave single-stranded DNA in a non-specific manner. Furthermore, mutant Cpf1 endonucleases and the complexes they form with crRNA are disclosed that can introduce one or more single-stranded or double-stranded breaks in nucleic acid sequences that are different from the nucleic acid sequence recognized and hybridized by crRNA.
[0006] The novel mutant Cpf1 endonucleases disclosed herein offer several advantages over wild-type Cpf1 endonuclease and can be advantageously used in the detection and quantification of target nucleic acid sequences in a specific or non-specific manner at very low concentrations, even at concentrations below the picomolar range.
[0007] The novel mutant Cpf1 endonucleases disclosed herein offer several advantages over wild-type Cpf1 endonucleases and can be advantageously used to cleave both double-stranded and single-stranded target nucleic acid sequences. Furthermore, in double-stranded targets, cleavage can be performed on both strands or only on one strand.
[0008] Finally, the novel mutant Cpf1 endonucleases disclosed herein can also be used for the diagnosis of infectious diseases by detection of genetic material derived from disease-causing infectious pathogens.
[0009] One aspect of the present disclosure is a method for producing a method for manufacturing a semiconductor device comprising the steps of: i) a polypeptide sequence having at least 95% sequence identity to the sequences corresponding to residues 1 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO:2; wherein the polypeptide sequence is a. at least two amino acid mutations in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, where each mutation is independently an amino acid substitution or deletion; and / or b. at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, where at least one of the residues at positions 917 and 1006 is glutamic acid (E) or aspartic acid (D); Further includes; and / or ii) a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2; wherein the polypeptide sequence comprises at least two amino acid substitutions at two positions independently selected from positions corresponding to residues 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2; The present invention relates to a mutant Cpf1 endonuclease or its orthologue, comprising:
[0010] Another aspect of the present disclosure pertains to polynucleotides encoding a mutant Cpf1 endonuclease or an ortholog thereof as disclosed herein.
[0011] Another aspect of the present disclosure relates to a recombinant vector comprising a polynucleotide or nucleic acid sequence encoding a mutant Cpf1 endonuclease or an orthologue thereof as disclosed herein, said polynucleotide or nucleic acid sequence being optionally operably linked to a promoter.
[0012] A further aspect of the present disclosure relates to a cell expressing a mutant Cpf1 or an orthologue thereof, or a polynucleotide, nucleic acid, or recombinant vector as disclosed herein.
[0013] Yet a further aspect of the present disclosure relates to a system for expression of crRNA-Cpf1 complex, comprising: a. a polynucleotide or recombinant vector comprising a polynucleotide or nucleic acid encoding a mutant Cpf1 endonuclease or an ortholog thereof as disclosed herein; and b. a polynucleotide or recombinant vector comprising a polynucleotide encoding a guide RNA (crRNA) operably linked to a promoter; Includes.
[0014] Another aspect of the present disclosure relates to the use of a crRNA-Cpf1 complex for introducing a single-stranded break in a first target nucleic acid, wherein: a. contacting a Cpf1 endonuclease or its orthologue with a guide RNA (crRNA), thereby obtaining a crRNA-Cpf1 complex capable of recognizing a second target nucleic acid, wherein the second target nucleic acid comprises a protospacer adjacent motif (PAM); and b. contacting the crRNA-Cpf1 complex with a first target nucleic acid; This creates a single-stranded break in the first target sequence.
[0015] Another aspect of the disclosure relates to a method for introducing a single-stranded break in a first target nucleic acid, the method comprising the steps of: a. designing a guide-RNA (crRNA) capable of recognizing a second target nucleic acid containing a protospacer adjacent motif (PAM); b. contacting the crRNA of step a. with Cpf1 endonuclease or an orthologue thereof, thereby obtaining a crRNA-Cpf1 complex capable of binding to said second target nucleic acid; and c. contacting crRNA and Cpf1 with the first target nucleic acid; thereby introducing one or more single-stranded breaks in the first target nucleic acid.
[0016] A further aspect of the present disclosure relates to an in vitro method of introducing a site-specific double-stranded break in a second target nucleic acid in a mammalian cell, the method comprising introducing a crRNA-Cpf1 complex into the mammalian cell, wherein Cpf1 is a mutant Cpf1 endonuclease or an orthologue as disclosed herein, and wherein the crRNA is specific for the second target nucleic acid.
[0017] A further aspect of the present disclosure relates to a method for detection of a second target nucleic acid in a sample, the method comprising: a. providing a crRNA-Cpf1 complex, where Cpf1 is a Cpf1 endonuclease as disclosed herein or an orthologue thereof, and where the crRNA is specific for a second target nucleic acid; b. providing labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one dye and at least one quencher; c. contacting the crRNA-Cpf1 complex and the ssDNA with a sample, wherein the sample contains at least one second target nucleic acid; and d. detecting ssDNA breaks by detecting a fluorescent signal from the fluorophore; thereby detecting the presence of a second target nucleic acid in the sample.
[0018] Another aspect of the present disclosure relates to an in vitro method for diagnosing an infectious disease in a subject, the method comprising: a. providing a crRNA-Cpf1 complex, where Cpf1 is a Cpf1 endonuclease as disclosed herein or an orthologue thereof, and where the crRNA is specific for a second target nucleic acid; b. providing labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one dye and at least one quencher; c. providing a sample from a subject, wherein said sample contains or is suspected of containing a second target nucleic acid; and d. determining the level and / or concentration of a second target nucleic acid by the methods disclosed herein; wherein the second target nucleic acid is a nucleic acid or a fragment thereof of the genome of a disease-causing infectious pathogen; thereby diagnosing an infectious disease in a subject. [Brief description of the drawings]
[0019] [Figure 1] Representation of the crRNA and double-stranded DNA targets.
[0020] [Diagram 2] Cleavage products of double-stranded DNA target obtained by using wild-type Cpf1 and mutant 1 (SEQ ID NO:3).
[0021] [Diagram 3]Left panel: cleavage products of double-stranded DNA target obtained by using wild-type Cpf1 and mutant 1 (SEQ ID NO: 4, 918-1013 mutant) (filled circles). Right panel: cleavage products of non-specific ssDNA obtained by using wild-type Cpf1 and mutant 2 (SEQ ID NO: 4, 918-1013 mutant) (filled circles).
[0022] [Figure 4A] (A) Cleavage activity assays with target DNA substrates for REC-linker deletions and substitutions (SEQ ID NO: 34 and 36), lid substitutions (SEQ ID NO: 38) and Q1025G, E1028G and Q1025G-E1028G mutants (SEQ ID NO: 30 and 32). [Figure 4B] (B) Cleavage activity assay monitoring the cut of the t-strand complementary to the crRNA for REC-linker deletions and substitutions, lid substitutions and Q1025G, E1028G and Q1025G-E1028G mutants. [Figure 4C] (C) Cleavage activity assays after activation by ssDNA complementary to crRNA using non-specific ssDNA as substrate for REC-linker deletions and substitutions, lid substitutions and Q1025G, E1028G and Q1025G-E1028G mutants. [Figure 4D] (D) Cleavage activity assay monitoring the cut of dsDNA substrates (t-strand, nt-strand and nonspecific ssDNA for wild-type Cpf1, (D917)Cpf1, (E1006-D917A)Cpf1 and (Lid-substituted-D917A)Cpf1).
[0023] [Diagram 5] Protein expression and purification profile mutant Cpf1 with deletions or substitutions in the REC-linker, lid and fingers regions.
[0024] [Figure 6]A. Quantification of cleavage activity: NTS (non-target strand), TS (target strand) and ssDNA (non-specific single strand) of Cas12a wild type (cas12aWT); B. Cpf1 Q1025G (SEQ ID NO: 30); C. Cpf1 Q1025G-E1028G (SEQ ID NO: 32); D. Nickase Cpf1 K1013G-R1014G (SEQ ID NO: 3). Each bar represents the average of at least three independent replicates with standard deviations shown. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0025] Detailed Description The invention is as defined in the claims. The present disclosure relates to mutant Cpf1 endonucleases or orthologues thereof and their use to introduce single-stranded or double-stranded breaks in a nucleic acid target sequence, either single-stranded or double-stranded, in a specific or non-specific manner. Throughout the present disclosure, a "mutant Cpf1 endonuclease" may be a naturally occurring mutant, e.g., a mutant encoded by a Cpf1 gene carrying one or more single nucleotide polymorphisms (SNPs), or a non-naturally occurring mutant, e.g., a mutant obtained by direct or random mutagenesis of the Cpf1 gene.
[0026] The inventors have identified mutants of Cpf1 with altered activity. Some of the mutants disclosed herein can function as nickases, thus recognizing specific target DNA sequences and introducing single-strand breaks therein. The inventors have also identified mutants of Cpf1 that recognize target DNA sequences and introduce single-strand breaks in non-target DNA sequences after binding to the target nucleic acid. Some mutants can cleave DNA in single-stranded sequences, others in double-stranded sequences. Cleavage can be specific or non-specific. Cleavage can be activated by double-stranded target DNA, single-stranded target DNA, or both.
[0027] The present inventors also provide a method for detecting and quantifying a target DNA sequence using a mutant Cpf1 endonuclease.The present inventors also provide a method for diagnosing an infectious disease by detecting and quantifying a target DNA sequence using a mutant Cpf1 endonuclease.
[0028] definition Codon The term "codon" as used herein refers to a triplet of contiguous nucleotides that codes for a specific amino acid. CRISPR-Cas system The term refers to members of the CRISPR-Cas family. CRISPR-Cas (clustered regularly interspaced short palindromic repeats and CRISPR-associated proteins), a prokaryotic adaptive immune system, can bind and cleave target DNA sequences through RNA-guided recognition. According to their molecular structure, different members of CRISPR-Cas systems have been classified into two classes: class 1 encompasses several effector proteins, whereas class 2 systems use a single element (Makarova et al., 2015). Cpf1 ("CRISPR from Prevotella and Francisella") has been described as a new member of the class 2·V type of CRISPR-Cas endonucleases present in many bacterial genomes (Zetsche et al., 2015).
[0029] Endonuclease The term refers to an enzyme that can cleave phosphodiester bonds in a polynucleotide chain. Some endonucleases are specific, i.e., they recognize a given nucleotide sequence that directs the site of cleavage, and some are non-specific. The present disclosure is directed to both specific and non-specific endonucleases. One example of an endonuclease is a nicking endonuclease. Nicking endonuclease, as used herein, refers to an enzyme that cuts one strand of double-stranded DNA to generate a "nicked" DNA molecule. Nicking endonucleases, as used herein, also refer to endonucleases that cut one strand of single-stranded DNA.
[0030] Fragment: The term is used to indicate a non-full length portion of a nucleic acid or polypeptide. Thus, a fragment is itself also a nucleic acid or polypeptide, respectively. DNA fragments are specified throughout this disclosure starting from the 5' end.
[0031] Gene editing: The term refers to the use of genetic engineering techniques to insert, delete or replace one or more nucleotides in a nucleotide sequence.
[0032] Guide RNA: This term is used interchangeably with "crRNA" herein and refers to an RNA molecule required for recognition of a target nucleic acid sequence by CRISPR-Cas proteins, in particular Cpf1.
[0033] Homologs Homologs or functional homologs may be any polypeptide that exhibits at least some sequence identity with a reference polypeptide and retains at least one aspect of the original function. As used herein, a functional homolog of Cpf1 is a polypeptide that shares at least some sequence identity with Cpf1 or a fragment thereof, and has the ability to function as an endonuclease similar to Cpf1, i.e., it can specifically bind to crRNA and can specifically recognize, bind to and cleave a target nucleic acid.
[0034] Protospacer adjacent motif (PAM) This term refers to the DNA sequence immediately downstream of the DNA sequence targeted by a CRISPR-Cas system, such as Cpf1. The crRNA of the crRNA-Cpf1 complex can only recognize and hybridize to target DNA sequences that contain a PAM.
[0035] Recognition The term "recognition" as understood herein refers to the ability of a molecule to identify a nucleotide sequence. For example, an enzyme or a DNA binding domain can recognize and bind to a nucleic acid sequence as a potential substrate. Preferably, the recognition is specific.
[0036] Sequence identity As used herein, the term refers to two polynucleotide sequences that are identical (i.e., nucleotide by nucleotide) over a window of comparison.The term "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a window of comparison, determining the number of positions in both sequences where the same nucleic acid base (e.g., A, T, C, G, U or I) exists, to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the window of comparison (i.e., the size of the window), and multiplying the result by 100 to obtain the percentage of sequence identity.
[0037] When applied to polypeptides, peptides or proteins, the degree of identity of amino acid sequences is a function of the number of identical amino acids at positions shared by the amino acid sequences. The degree of homology or similarity of amino acid sequences is a function of the number of identical, i.e. structurally related, amino acids at positions shared by the amino acid sequences.
[0038] Global percentages of sequence identity are determined by the algorithms GAP, BESTFIT or FASTA in the Wisconsin Genetics Software Package Release 7.0, using default gap weights.
[0039] The term "interactive label" or "set of interactive labels" as used herein refers to at least one fluorophore and at least one quencher that can interact when they are adjacently positioned.When the interactive labels are adjacently positioned, the quencher can quench the signal of the fluorophore.The interaction can be mediated by fluorescence resonance energy transfer (FRET).
[0040] The term "disposed adjacently" as used herein refers to the physical distance between two objects.When a fluorophore and a quencher are disposed adjacently, the quencher can partially or completely quench the signal of the fluorophore.Quenching by FRET can typically occur over a distance of up to about 100 Å.Disposed adjacently as used herein can refer to a distance of less than 100 Å and / or about 100 Å.
[0041] The term "fluorescent label" or "fluorophore" as used herein refers to a fluorescent chemical compound that can re-emit light upon photoexcitation. Fluorophores absorb light energy at specific wavelengths and re-emit light at longer wavelengths. The absorbed wavelength, energy transfer rate, and time to emission depend on both the structure of the fluorophore and its chemical environment (as the molecule interacts with surrounding molecules in its excited state). The wavelengths of maximum absorption (≈excitation) and emission (e.g., absorption / emission=485 nm / 517 nm) are typical terms used to refer to a given fluorophore, but the entire spectrum can be important to consider.
[0042] The term "quench" or "quenching" as used herein refers to any process that reduces the fluorescence intensity of a given substance, such as a fluorophore. Quenching can be mediated by fluorescence resonance energy transfer (FRET). FRET is based on classical dipole-dipole interactions between the transfer dipoles of a donor (e.g., a fluorophore) and an acceptor (e.g., a quencher), and depends on the donor-acceptor distance. FRET can typically occur over distances up to 100 Å. FRET also depends on the donor-acceptor spectral overlap and the relative orientation of the transfer dipole moments of the donor and acceptor. Quenching of a fluorophore can also occur as a result of the formation of a non-fluorescent complex between the fluorophore and another fluorophore or a non-fluorescent molecule. This mechanism is known as "contact quenching," "static quenching," or "ground state complex formation." The term "quencher," as used herein, refers to a chemical compound that is capable of quenching a given substance, such as a fluorophore.
[0043] Target strand and non-target strand The target strand refers to the strand of nucleic acid that interacts with the crRNA to form the crRNA-DNA hybrid. The non-target strand is complementary to the target strand and is transferred to the RuvC / NuC pocket of the Cpf1 endonuclease.
[0044] The term "ortholog" as used herein refers to genes (and the proteins encoded by said genes) that are inferred to be derived from the same ancestral sequence separated by a speciation event: when a species diverges into two separate species, the copies of a single gene in the two resulting species are said to be orthologous. Orthologs, or orthologous genes, are genes in different species that are derived by vertical descent from a single gene in an immediate common ancestor. Orthologs of Cpf1 can be identified and characterized based on sequence similarity to current systems, as the wild-type II system is described. For example, orthologs of Cpf1 include U1 12 from F. novicida, Prevotella albensis, Acidaminococcus sp. BV3L6, CAG:72 from Eubacterium eligens, Butyrivibrio fibrisolvens, SCADC from Smithella sp., Flavobacterium sp. 316, Porphyromonas crevioricanis, oral taxon 274 from Bacteroidetes, and ND2006 from Lachnospiraceae.
[0045] Mutant Cpf1 The present inventors have identified three domains in Cpf1 that are likely responsible for different activities of the enzyme. These three domains are: ·REC Domain ·Lid domain Finger Domain It is.
[0046] In FnCpf1 (SEQ ID NO:2), these three domains are defined as follows: ·REC domain: residues 324 to 336; ·Lid domain: residues 1006–1018; · Finger domain: residues 298–309.
[0047] Substitution or deletion of amino acids in any of these three domains results in altered enzyme activity, as will be described in detail herein below.In addition, five residues that are considered to be important for enzyme activity have been identified.That is, mutation or deletion of any of these three residues also alters enzyme activity.These residues are at positions 918, 1013, 1014, 1025 and 1028, of which 1013 and 1014 are part of the Lid region.The present disclosure therefore relates to modified Cpf1 proteins with altered activity.
[0048] One aspect of the present disclosure therefore comprises: i) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity, to a sequence corresponding to residues 1 to 297, 310 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO: 2, wherein said polypeptide sequence further comprises at least one amino acid substitution or deletion in the Finger domain (residues 298 to 309; SEQ ID NO: 12), the REC domain (residues 324 to 336; SEQ ID NO: 16) or the Lid domain (residues 1006 to 1018; SEQ ID NO: 20) compared to SEQ ID NO: 2; and / or ii) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2; The present invention relates to a mutant Cpf1 endonuclease or its orthologue, comprising:
[0049] One aspect of the present disclosure therefore comprises: i) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, for example at least 99% sequence identity, to a sequence corresponding to residues 1 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO: 2; wherein the polypeptide sequence is a. at least two amino acid mutations in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, where each mutation is independently an amino acid substitution or deletion; and / or b. at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, where at least one of the residues at positions 917 and 1006 is glutamic acid (E) or aspartic acid (D); Further includes; and / or ii) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2; wherein the polypeptide sequence comprises at least two amino acid substitutions at two positions independently selected from positions corresponding to residue positions 918, 1013, 1014, 1025, and 1028 of SEQ ID NO:2; The present invention relates to a mutant Cpf1 endonuclease or its orthologue, comprising:
[0050] One aspect of the present disclosure relates to a mutant Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, or 1014 of SEQ ID NO:2.
[0051] One aspect of the present disclosure relates to a mutant Cpf1 endonuclease or orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at two positions independently selected from positions corresponding to residues 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2. In one embodiment, the Cpf1 endonuclease is derived from Francisella novicida.
[0052] In some embodiments, the mutant Cpf1 comprises a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity, such as 100% identity, to a sequence corresponding to residues 1-323 and 337-1329 of SEQ ID NO:2, wherein said polypeptide sequence further comprises at least one amino acid substitution or deletion in the Fingers domain (residues 298-309; SEQ ID NO:12), the REC domain (residues 324-336; SEQ ID NO:16) or the Lid Finger domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2. The mutant Cpf1 may further comprise at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0053] In some embodiments, the mutant Cpf1 comprises a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity, such as 100% identity, to a sequence corresponding to residues 1-323 and 337-1329 of SEQ ID NO:2, wherein said polypeptide sequence further comprises at least two mutations (wherein each mutation is independently an amino acid substitution or deletion) in the REC domain (residues 324-336; SEQ ID NO:16), or at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20), compared to SEQ ID NO:2. A mutant Cpf1 comprising a mutation in the REC and / or Lid domain may further comprise a substitution of at least one amino acid at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0054] In one embodiment, the mutant Cpf1 has at least one amino acid substitution or deletion in the fingers domain. The at least one amino acid substitution or deletion may be a substitution or deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 consecutive or non-consecutive amino acids in the fingers domain defined as residues 298-309 of SEQ ID NO:2. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to a sequence corresponding to residues 1-297 and 310-1329 of SEQ ID NO:2. The mutant Cpf1 may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0055] In another embodiment, the mutant Cpf1 has at least one amino acid substitution or deletion in the REC domain. The at least one amino acid substitution or deletion may be a substitution or deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 consecutive or non-consecutive amino acids in the REC domain defined as residues 324-336 of SEQ ID NO:2. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to a sequence corresponding to residues 1-323 and 337-1329 of SEQ ID NO:2. The mutant Cpf1 may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0056] In another embodiment, the mutant Cpf1 has at least two amino acid mutations in the REC domain, where each mutation is independently an amino acid substitution or deletion. The at least two amino acid mutations may be substitutions and / or deletions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 consecutive or non-consecutive amino acids in the REC domain defined as residues 324-336 of SEQ ID NO: 2. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to a sequence corresponding to residues 1-323 and 337-1329 of SEQ ID NO: 2. A mutant Cpf1 comprising at least two mutations in the REC domain may further comprise at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0057] In another embodiment, the mutant Cpf1 has at least one amino acid substitution or deletion in the lid domain. The at least one amino acid substitution or deletion may be a substitution or deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 consecutive or non-consecutive amino acids in the Lid domain defined as residues 1006-1013 of SEQ ID NO:2. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to a sequence corresponding to residues 1-1005 and 1019-1329 of SEQ ID NO:2. The mutant Cpf1 may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0058] In another embodiment, the mutant Cpf1 has at least two amino acid substitutions in the Lid domain. The at least two amino acid substitutions may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 consecutive or non-consecutive amino acid substitutions in the Lid domain defined as residues 1006-1013 of SEQ ID NO:2. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to a sequence corresponding to residues 1-1005 and 1019-1329 of SEQ ID NO:2. The mutant Cpf1 comprising at least two substitutions in the Lid domain may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0059] In another embodiment, Cpf1 has at least one amino acid substitution or deletion in the fingers domain and in the REC domain, where the at least one amino acid substitution or deletion is as defined above. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence corresponding to residues 1-297, 310-323 and 337-1329 of SEQ ID NO:2. The mutant Cpf1 may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0060] In another embodiment, Cpf1 has at least two amino acid mutations in the Lid domain and in the REC domain, where each mutation is independently an amino acid substitution or deletion as defined above. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequence corresponding to residues 1-297, 310-323 and 337-1329 of SEQ ID NO:2. The mutant Cpf1 comprising at least two amino acid mutations in the Lid domain and in the REC domain may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0061] In another embodiment, Cpf1 has at least one amino acid substitution or deletion in the fingers domain and in the lid domain, where the at least one amino acid substitution or deletion is as defined above. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequences corresponding to residues 1-297, 310-1005 and 1019-1329 of SEQ ID NO:2. The mutant Cpf1 may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0062] In another embodiment, Cpf1 has at least one amino acid substitution or deletion in the REC domain and in the lid domain, where the at least one amino acid substitution or deletion is as defined above. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequences corresponding to residues 1-323, 336-1005 and 1019-1329 of SEQ ID NO:2. The mutant Cpf1 may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0063] In another embodiment, Cpf1 has at least one amino acid substitution or deletion in the fingers domain, in the REC domain, and in the lid domain, where the at least one amino acid substitution or deletion is as defined above. Thus, the mutant Cpf1 may comprise a sequence having at least 80%, such as at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to the sequences corresponding to residues 1-297, 310-323, 337-1005 and 1019-1329 of SEQ ID NO:2. The mutant Cpf1 may further comprise at least one amino acid substitution at positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0064] In a particular embodiment, the Cpf1 mutant is as described in SEQ ID NO: 42, i.e., the finger domain is deleted. In another embodiment, the Cpf1 mutant is as described in SEQ ID NO: 34, i.e., the REC domain is deleted. In another embodiment, the Cpf1 mutant is as described in SEQ ID NO: 36, i.e., the REC domain is substituted. In another embodiment, the Cpf1 mutant is as described in SEQ ID NO: 40, i.e., the Lid domain is deleted. In another embodiment, the Cpf1 mutant is as described in SEQ ID NO: 38, i.e., the Lid domain is substituted.
[0065] It will be understood that the substitution or deletion of at least one amino acid as defined above may refer to the deletion of some amino acids in a domain, while other amino acids may be substituted.
[0066] All of the above variants may further comprise at least one amino acid substitution and / or deletion at one or more of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2, as further detailed herein below. In one embodiment, the variant further comprises at least one amino acid substitution at one of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2, with the other positions remaining unaltered. In one embodiment, the variant thus further comprises a single substitution at one of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2. For example, the variant further comprises a single substitution at position 918, or at position 1013, or at position 1014, or at position 1025, or at position 1028.
[0067] In another embodiment, the variant further comprises an amino acid substitution and / or deletion at two of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO: 2. For example, the variant further comprises a substitution at positions 918 and 1013; or at positions 918 and 1014; or at positions 918 and 1025; or at positions 918 and 1028; or at positions 1013 and 1014; or at positions 1013 and 1025; or at positions 1013 and 1028; or at positions 1014 and 1025; or at positions 1014 and 1028; or at positions 1025 and 1028.
[0068] In another embodiment, the variant further comprises amino acid substitutions and / or deletions at three of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO: 2. For example, the variant further comprises substitutions at positions 918, 1013 and 1014; or at positions 918, 1013 and 1025; or at positions 918, 1013 and 1028; or at positions 918, 1014 and 1025; or at positions 918, 1014 and 1028; or at positions 918, 1025 and 1028; or at positions 1013, 1014 and 1025; or at positions 1013, 1014 and 1028; or at positions 1014, 1025 and 1028.
[0069] In another embodiment, the variant further comprises amino acid substitutions and / or deletions at four of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO: 2. For example, the variant further comprises substitutions at positions 918, 1013, 1014 and 1025; or at positions 918, 1013, 1025 and 1028; or at positions 1013, 1014, 1025 and 1028; or at positions 918, 1014, 1025 and 1028. In another embodiment, the variant further comprises amino acid substitutions and / or deletions at five of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2.
[0070] The at least one amino acid substitution, whether in the Finger, REC or Lid domain, or in one or more of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2, in some embodiments may be a substitution of an amino acid having a charged side chain with an amino acid having an uncharged or non-polar side chain. In some embodiments, the amino acid is substituted with an amino acid selected from the group consisting of glycine, alanine, valine, leucine, isoleucine, serine or threonine. In some embodiments, the amino acid is substituted with glycine.
[0071] The at least two amino acid substitutions, whether in the Finger, REC or Lid domains or in one or more of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO: 2, in some embodiments may be substitutions of amino acids having a charged side chain with amino acids having an uncharged or non-polar side chain. In some embodiments, the amino acids are substituted with amino acids selected from the group consisting of glycine, alanine, valine, leucine, isoleucine, serine or threonine. In some embodiments, the amino acid is substituted with glycine.
[0072] For example, the finger domain of SEQ ID NO:2, KGINEYINLY SQ (SEQ ID NO:20), may be mutated to GGGAGAAGGA SG (SEQ ID NO:22). The REC domain of SEQ ID NO:2, LFKQILSDTE SKS (SEQ ID NO:12), may be mutated to GGGGAGASAGG SGS (SEQ ID NO:14). The Lid domain of SEQ ID NO:2, EDLNFGFKRG RFK (SEQ ID NO:16), may be mutated to GGGAGGAAGG GAG (SEQ ID NO:18).
[0073] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or orthologue thereof comprises a polypeptide sequence orthologue thereof, including a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, or 1014 of SEQ ID NO:2 of an amino acid residue having a charged side chain for an amino acid residue having an uncharged side chain.
[0074] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or ortholog thereof comprises a polypeptide sequence ortholog thereof, comprising a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions of two amino acids having charged side chains to amino acid residues having two uncharged side chains at two positions independently selected from positions corresponding to residues 918, 1013, or 1014 of SEQ ID NO:2.
[0075] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or orthologue thereof comprises a polypeptide sequence orthologue thereof, including a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, or 1014 of SEQ ID NO:2 of an amino acid residue having a charged side chain for an amino acid residue having a non-polar side chain.
[0076] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or orthologue thereof comprises a polypeptide sequence orthologue thereof, comprising a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, or 1014 of SEQ ID NO:2 of an amino acid having a charged side chain with an amino acid residue selected from the group consisting of glycine, alanine, valine, leucine, isoleucine, serine, or threonine.
[0077] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or ortholog thereof comprises a polypeptide sequence ortholog thereof, including a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution of an amino acid having a charged side chain at position 918, 1013, or 1014 of SEQ ID NO:2 with glycine.
[0078] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or ortholog thereof comprises a polypeptide sequence ortholog thereof, including a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution selected from R918G, K1013G, R1014G, Q1025G or E1028G.
[0079] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or ortholog thereof comprises a polypeptide sequence ortholog thereof, including a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions selected from R918G, K1013G, R1014G, Q1025G or E1028G.
[0080] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or ortholog thereof comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises an amino acid substitution at residue K1013.
[0081] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or orthologue thereof comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises amino acid substitutions at residues R918 and K1013.
[0082] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or its orthologue comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises amino acid substitutions at residues K1013 and R1014.
[0083] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or orthologue thereof comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises the amino acid substitution K1013G.
[0084] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or its orthologue comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises the amino acid substitutions R918G and K1013G.
[0085] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or its orthologue comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises the amino acid substitutions K1013G and R1014G.
[0086] In some embodiments of the present disclosure, the mutant Cpf1 endonuclease or its orthologue comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said Cpf1 polypeptide comprises amino acid substitutions at residues K1013 and R1014 as defined in the above embodiments, and is a nickase capable of introducing a single-stranded break in a target nucleic acid such as one in a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical.
[0087] In some embodiments of the present disclosure, the mutant Cpf1 endonuclease or its orthologue comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2, such as at least 96% sequence identity to SEQ ID NO:2, such as at least 97% sequence identity to SEQ ID NO:2, such as at least 98% sequence identity to SEQ ID NO:2, such as at least 99% sequence identity to SEQ ID NO:2, such as about 100% sequence identity to SEQ ID NO:2, wherein said Cpf1 polypeptide comprises amino acid substitutions at residues R918 and K1013 as defined in the above embodiments, and is a non-specific endonuclease that can introduce single-strand breaks at a non-specific site of a nucleic acid, such as a non-specific site of a first target nucleic acid. In some embodiments of the present disclosure, the mutant Cpf1 endonuclease or its orthologue is a non-specific endonuclease that recognizes a target DNA sequence but only introduces single-strand breaks at a non-specific site of a nucleic acid sequence. In some embodiments of the present disclosure, the mutated, non-specific Cpf1 endonuclease remains active for a longer period of time and introduces single-stranded breaks in several non-specific single-stranded nucleic acid sequences before reverting to an inactive state.
[0088] In some embodiments, the mutant Cpf1 endonuclease or its orthologue comprises a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:3 or SEQ ID NO:4, such as at least 96% sequence identity to SEQ ID NO:3 or SEQ ID NO:4, such as at least 97% sequence identity to SEQ ID NO:3 or SEQ ID NO:4, such as at least 98% sequence identity to SEQ ID NO:3 or SEQ ID NO:4, such as at least 99% sequence identity to SEQ ID NO:3 or SEQ ID NO:4, such as 100% sequence identity to SEQ ID NO:3 or SEQ ID NO:4.
[0089] In some embodiments, the mutant Cpf1 endonuclease or ortholog thereof comprises SEQ ID NO:3 or SEQ ID NO:4. In some embodiments, the mutant Cpf1 endonuclease or its ortholog consists of SEQ ID NO:3 or SEQ ID NO:4.
[0090] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or orthologue thereof comprises a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises an amino acid substitution at residue Q1025.
[0091] In another embodiment of the present disclosure, the mutant Cpf1 endonuclease or orthologue thereof comprises a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises an amino acid substitution at residue E1028.
[0092] In one embodiment of the present disclosure, the mutant Cpf1 endonuclease or its orthologue comprises a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises amino acid substitutions at residues Q1025 and E1028.
[0093] The substitution at position 1025 or 1028 may be as described herein. In some embodiments, position 1025 is substituted with glycine. In some embodiments, position 1028 is substituted with glycine. In some embodiments, positions 1025 and 1028 are both substituted with glycine. Cpf1 contains an RNAse active site that is used to process its own pre-crRNA to assemble an active ribonucleoprotein to achieve interference.
[0094] Modification of the Finger, REC and Lid domains of Cpf1 results in altered properties of Cpf1. The mutant protein may therefore have altered properties compared to the wild-type protein, in particular the mutant Cpf1 may be able to: ·Cleaving double-stranded target nucleic acids; Cleave double-stranded non-target nucleic acids; ·Cleaving single-stranded target nucleic acids; Cleave single-stranded non-target nucleic acids; Cleave single-stranded DNA in a non-specific manner, optionally enhanced by double-stranded target DNA; · cleaves single-stranded DNA in a non-specific manner, optionally enhanced by single-stranded target DNA; Cleavage of one strand of a double-stranded nucleic acid target; cleaves only the target double-stranded nucleic acid; and / or -Cuts only the target single-stranded nucleic acid.
[0095] The mutant Cpf1 may have one or more of the above activities. In some embodiments, the mutant Cpf1 exhibits improved activity compared to wild-type Cpf1. The mutants disclosed herein may therefore have different uses.For example, mutants that cleave target and / or non-target nucleic acids in a more efficient manner than the wild-type Cpf1 from which they are derived, whether the target and / or non-target nucleic acids are single-stranded or double-stranded, may generally be more useful for efficient gene editing.Mutants that preferentially cleave double-stranded target and / or non-target nucleic acids may show higher specificity than the wild-type Cpf1 from which they are derived.Mutants that can cleave single-stranded DNA nonspecifically but cannot cleave target or non-target strands may retain activity for longer than wild-type protein.
[0096] The mutants disclosed herein may therefore have different uses.For example, a mutant that cleaves targets, whether single-stranded or double-stranded, in a more efficient manner than the wild-type Cpf1 from which it is derived, but does not cleave non-target single-stranded nucleic acid, may generally be more useful for efficient gene editing.A mutant that preferentially cleaves double-stranded target and / or non-target nucleic acid may show higher specificity than the wild-type Cpf1 from which it is derived.A mutant that can non-specifically cleave single-stranded DNA but cannot cleave target nucleic acid may retain activity for longer than wild-type protein and may be more useful in target detection.
[0097] In one embodiment, the Cpf1 endonuclease is a nicking endonuclease. Polynucleotides and nucleic acid sequences encoding the mutant Cpf1 disclosed herein are also provided. Those skilled in the art know how to design such nucleic acid sequences that encode desired Cpf1 mutants.
[0098] In some embodiments, the present disclosure provides the following: i) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, for example at least 99% sequence identity, to a sequence corresponding to residues 1 to 297, 310 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO: 2; wherein the polypeptide sequence further comprises at least one amino acid substitution or deletion in the Finger domain (residues 298-309; SEQ ID NO: 12), the REC domain (residues 324-336; SEQ ID NO: 16) or the Lid domain (residues 1006-1018; SEQ ID NO: 20) compared to SEQ ID NO: 2; and / or ii) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2; wherein the polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2; The present invention provides a nucleic acid encoding a mutant Cpf1 endonuclease, or an ortholog thereof, comprising:
[0099] In some embodiments, the present disclosure provides the following: i) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, for example at least 99% sequence identity, to a sequence corresponding to residues 1 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO: 2; wherein the polypeptide sequence is: a. at least two amino acid mutations in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, where each mutation is independently an amino acid substitution or deletion; and / or b. At least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, where at least one of the residues at positions 917 and 1006 is glutamic acid (E) or aspartic acid (D). Further includes; and / or ii) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2; wherein the polypeptide sequence comprises at least two amino acid substitutions at two positions independently selected from positions corresponding to residues 918, 1013, 1014, 1025, and 1028 of SEQ ID NO:2; The present invention provides a nucleic acid encoding a mutant Cpf1 endonuclease, or an ortholog thereof, comprising:
[0100] The nucleic acid sequence encoding FnCpf1 is set forth in SEQ ID NO:1. In some embodiments, the nucleic acid sequence encodes a Cpf1 mutant lacking the finger domain. In certain embodiments, the nucleic acid sequence is set forth in SEQ ID NO: 41 and encodes the mutant Cpf1 of SEQ ID NO: 42.
[0101] In other embodiments, the nucleic acid sequence encodes a Cpf1 mutant in which the REC domain is deleted. In certain embodiments, the nucleic acid sequence is set forth in SEQ ID NO: 33 and encodes the mutant Cpf1 of SEQ ID NO: 34.
[0102] In other embodiments, the nucleic acid sequence encodes a Cpf1 mutant in which an amino acid residue in the REC domain is substituted. In a particular embodiment, the nucleic acid sequence is set forth in SEQ ID NO: 35 and encodes the mutant Cpf1 of SEQ ID NO: 36.
[0103] In other embodiments, the nucleic acid sequence encodes a Cpf1 mutant in which the lid domain is deleted. In a particular embodiment, the nucleic acid is set forth in SEQ ID NO:39 and encodes the mutant Cpf1 of SEQ ID NO:40.
[0104] In other embodiments, the nucleic acid sequence encodes a Cpf1 mutant in which an amino acid residue in the lid domain has been substituted. In a particular embodiment, the nucleic acid sequence is set forth in SEQ ID NO:37 and encodes the mutant Cpf1 of SEQ ID NO:38.
[0105] In some embodiments, Cpf1 endonuclease or its ortholog is encoded by a nucleic acid sequence comprising or consisting of a sequence that is at least 70% identical, such as at least 75% identical, such as at least 80% identical, such as at least 85% identical, such as at least 86% identical, such as at least 87% identical, such as at least 88% identical, such as at least 89% identical, such as at least 90% identical, such as at least 91% identical, such as at least 92% identical, such as at least 93% identical, such as at least 94% identical, such as at least 95% identical, such as at least 96% identical, such as at least 97% identical, such as at least 98% identical, such as at least 99% identical to the sequence that codes for parent Cpf1, i.e. parent sequence.In one embodiment, parent sequence is the sequence that codes for FnCpf1 as described in SEQ ID NO:1.
[0106] One aspect of the present disclosure relates to a vector comprising a polynucleotide or nucleic acid sequence encoding a mutant Cpf1 endonuclease or its orthologue as defined above in operable combination with a promoter, wherein the Cpf1 endonuclease or its orthologue has at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to a polypeptide of SEQ ID NO:2, and said polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, or 1014 of SEQ ID NO:2.
[0107] One aspect of the present disclosure relates to a vector comprising a polynucleotide or nucleic acid sequence encoding a mutant Cpf1 endonuclease or its orthologue as defined above in operable combination with a promoter, wherein the Cpf1 endonuclease or its orthologue has at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to a polypeptide of SEQ ID NO:2, and said polypeptide sequence comprises at least two amino acid substitutions at positions 918, 1013, or 1014 of SEQ ID NO:2.
[0108] One aspect of the present disclosure relates to a vector comprising a polynucleotide or nucleic acid sequence encoding a mutant Cpf1 endonuclease or its orthologue as defined above in operable combination with a promoter, wherein the Cpf1 endonuclease or its orthologue has at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to a polypeptide of SEQ ID NO:2, and said polypeptide sequence comprises at least two amino acid substitutions at positions 1025 and 1028 of SEQ ID NO:2.
[0109] In one embodiment, the vector further comprises a nucleic acid sequence encoding a guide RNA (crRNA) in operable combination with a promoter, where the crRNA is linked to the encoded Cpf1 endonuclease and a segment having sufficient base pairs to hybridize to a target nucleic acid. The crRNA is further described below in the section "Guide RNA (crRNA)."
[0110] One aspect of the present disclosure relates to a system for expression of crRNA-Cpf1 complex, the system comprising: a. a polynucleotide as disclosed herein, or a recombinant vector as disclosed herein, comprising a polynucleotide encoding a mutant Cpf1 endonuclease or an ortholog thereof; b. A polynucleotide or recombinant vector comprising a polynucleotide encoding a guide RNA (crRNA) operably linked to a promoter.
[0111] One aspect of the present disclosure relates to a cell expressing a polynucleotide or recombinant vector as described herein above. Cpf1 can be expressed from a cell, particularly a suitable expression host as known to those skilled in the art. In one embodiment, Cpf1 is expressed from Escherichia coli. This can be done as known in the art, for example, by introducing a vector containing a nucleic acid sequence encoding a desired Cpf1 effector protein or homolog as described herein above into an E. coli cell. The protein can be purified as known in the art. In one embodiment, the polynucleotide is codon-optimized for expression in a host cell.
[0112] Without being bound by theory, it is believed that the Finger, REC and Lid domains are conserved. Thus, the Cpf1 variant may be a variant of Cpf1 from Francisella novicida, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium COE1, Succicrasticum ruminis, Candidatus Methanoplasma termitum, uncultured Clostridium sp. or Butyrivibrio fibrisolvens. In one embodiment, the Cpf1 may be FnCpf1 from Francisella novicida, the domains being as defined herein above.
[0113] In another embodiment, Cpf1 is derived from Acidaminococcus sp. BV3L6 as set forth in SEQ ID NO:24. The Fingers domain is defined by residues 275-285 of SEQ ID NO:24. The REC domain is defined by residues 305-317 of SEQ ID NO:24. The Lid domain is defined by residues 993-1005 of SEQ ID NO:24. Position 918 of SEQ ID NO:2 corresponds to position 909 of SEQ ID NO:24. Position 1013 of SEQ ID NO:2 corresponds to position 1000 of SEQ ID NO:24. Position 1014 of SEQ ID NO:2 corresponds to position 1001 of SEQ ID NO:24. Position 1025 of SEQ ID NO:2 corresponds to position 1013 of SEQ ID NO:24. Position 1028 of SEQ ID NO:2 corresponds to position 1016 of SEQ ID NO:24.
[0114] In another embodiment, Cpf1 is derived from a bacterial COE1 of the Lachnospiraceae family as set forth in SEQ ID NO:25. The Fingers domain is defined by residues 260-270 of SEQ ID NO:25. The REC domain is defined by residues 283-295 of SEQ ID NO:25. The Lid domain is defined by residues 936-948 of SEQ ID NO:25. Position 918 of SEQ ID NO:2 corresponds to position 843 of SEQ ID NO:25. Position 1013 of SEQ ID NO:2 corresponds to position 943 of SEQ ID NO:25. Position 1014 of SEQ ID NO:2 corresponds to position 944 of SEQ ID NO:25. Position 1025 of SEQ ID NO:2 corresponds to position 955 of SEQ ID NO:25. Position 1028 of SEQ ID NO:2 corresponds to position 958 of SEQ ID NO:25.
[0115] In another embodiment, Cpf1 is derived from Succiniclasticum ruminis as set forth in SEQ ID NO:26. The Fingers domain is defined by residues 279-289 of SEQ ID NO:26. The REC domain is defined by residues 309-321 of SEQ ID NO:26. The Lid domain is defined by residues 997-1009 of SEQ ID NO:26. Position 918 of SEQ ID NO:2 corresponds to position 913 of SEQ ID NO:26. Position 1013 of SEQ ID NO:2 corresponds to position 1004 of SEQ ID NO:26. Position 1014 of SEQ ID NO:2 corresponds to position 1005 of SEQ ID NO:26. Position 1025 of SEQ ID NO:2 corresponds to position 1017 of SEQ ID NO:26. Position 1028 of SEQ ID NO:2 corresponds to position 1020 of SEQ ID NO:26.
[0116] In another embodiment, Cpf1 is derived from Candidatus Methanoplasma termitum as set forth in SEQ ID NO:43. The Fingers domain is defined by residues 256-266 of SEQ ID NO:43. The REC domain is defined by residues 281-293 of SEQ ID NO:43. The Lid domain is defined by residues 944-956 of SEQ ID NO:43. Position 918 of SEQ ID NO:2 corresponds to position 860 of SEQ ID NO:43. Position 1013 of SEQ ID NO:2 corresponds to position 951 of SEQ ID NO:43. Position 1014 of SEQ ID NO:2 corresponds to position 952 of SEQ ID NO:43. Position 1025 of SEQ ID NO:2 corresponds to position 963 of SEQ ID NO:43. Position 1028 of SEQ ID NO:2 corresponds to position 966 of SEQ ID NO:43.
[0117] In another embodiment, Cpf1 is derived from an uncultured Clostridium species as set forth in SEQ ID NO:44. The Fingers domain is defined by residues 266-270 of SEQ ID NO:44. The REC domain is defined by residues 285-297 of SEQ ID NO:44. The Lid domain is defined by residues 959-971 of SEQ ID NO:44. Position 918 of SEQ ID NO:2 corresponds to position 8776 of SEQ ID NO:44. Position 1013 of SEQ ID NO:2 corresponds to position 966 of SEQ ID NO:44. Position 1014 of SEQ ID NO:2 corresponds to position 967 of SEQ ID NO:44. Position 1025 of SEQ ID NO:2 corresponds to position 979 of SEQ ID NO:44. Position 1028 of SEQ ID NO:2 corresponds to position 982 of SEQ ID NO:44.
[0118] In another embodiment, Cpf1 is derived from Butyrivibrio fibrisolvens as set forth in SEQ ID NO:27. The Fingers domain is defined by residues 234-242 of SEQ ID NO:27. The REC domain is defined by residues 260-264 of SEQ ID NO:27. The Lid domain is defined by residues 925-937 of SEQ ID NO:27. Position 918 of SEQ ID NO:2 corresponds to position 835 of SEQ ID NO:27. Position 1013 of SEQ ID NO:2 corresponds to position 932 of SEQ ID NO:27. Position 1014 of SEQ ID NO:2 corresponds to position 933 of SEQ ID NO:27. Position 1025 of SEQ ID NO:2 corresponds to position 944 of SEQ ID NO:27. Position 1028 of SEQ ID NO:2 corresponds to position 947 of SEQ ID NO:27.
[0119] Guide RNA (crRNA) To function as an endonuclease, the crRNA-Cpf1 complex requires not only the Cpf1 effector protein but also a guide RNA (crRNA), which is responsible for recognizing the target nucleic acid to be cleaved.
[0120] The crRNA comprises or consists of a constant region and a variable region. The constant region consists of 19-20 nucleotides and is constant for all complexes derived from a given organism. For optimal activity of the crRNA-Cpf1 complex, it may be important to design the crRNA based on a constant region specific for the organism from which Cpf1 or its homologue is derived. The constant region may be specific for Francisella novicida and may have the sequence set forth in SEQ ID NO:5.
[0121] The variable region consists of 12 to 24 nucleotides, e.g., 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24 nucleic acids. The variable region is the region of the crRNA that is believed to be responsible for target recognition. Thus, modifying the sequence of the variable region can be used to enable the crRNA-Cpf1 complex to specifically cleave various target nucleic acids. In contrast to the constant region, the variable region is not organism-specific.
[0122] Thus, in one embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 12 nucleotides, and the crRNA has a total length of 31 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 12 nucleotides, and the crRNA has a total length of 32 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 13 nucleotides, and the crRNA has a total length of 32 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 13 nucleotides, and the crRNA has a total length of 33 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 14 nucleotides, and the crRNA has a total length of 33 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 14 nucleotides, and the crRNA has a total length of 34 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 15 nucleotides, and the crRNA has a total length of 34 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 15 nucleotides, and the crRNA has a total length of 35 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 16 nucleotides, and the crRNA has a total length of 35 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 16 nucleotides, and the crRNA has a total length of 36 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 17 nucleotides, and the crRNA has a total length of 36 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 17 nucleotides, and the crRNA has a total length of 37 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 18 nucleotides, and the crRNA has a total length of 37 nucleotides.In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 18 nucleotides, and the crRNA has a total length of 38 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 19 nucleotides, and the crRNA has a total length of 38 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 19 nucleotides, and the crRNA has a total length of 39 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 20 nucleotides, and the crRNA has a total length of 39 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 20 nucleotides, and the crRNA has a total length of 40 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 21 nucleotides, and the crRNA has a total length of 40 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 21 nucleotides, and the crRNA has a total length of 41 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 22 nucleotides, and the crRNA has a total length of 41 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 22 nucleotides, and the crRNA has a total length of 42 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 23 nucleotides, and the crRNA has a total length of 42 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 23 nucleotides, and the crRNA has a total length of 43 nucleotides. In another embodiment, the crRNA consists of a constant region of 19 nucleotides and a variable region of 24 nucleotides, and the crRNA has a total length of 43 nucleotides. In another embodiment, the crRNA consists of a constant region of 20 nucleotides and a variable region of 24 nucleotides, and the crRNA has a total length of 44 nucleotides.
[0123] A person skilled in the art would have no difficulty in designing a variable region capable of binding to a desired target nucleic acid. The variable region has a sequence that is the reverse complement of the target nucleic acid.
[0124] Thus, the crRNA consists of a constant region of 19 or 20 nucleotides and a variable region of 12-24 nucleotides, such that the crRNA is 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44 nucleotides long, with the proviso that the 5'-terminal nucleotide consists of a PAM, as described in sections "First target nucleic acid" and "Second target nucleic acid."
[0125] Once the guide RNA sequence is designed, the guide RNA can be synthesized by known methods. For example, DNA oligonucleotides corresponding to the reverse complement of the target site can be ordered from commercially available oligonucleotides. These oligonucleotides can include a 24-base long T7 priming sequence. These DNA duplexes can then be used as templates in a transcription reaction carried out by T7 RNA polymerase. For example, the reaction can consist of incubation at 37°C for at least 1 hour. The reaction can be stopped using 2x stop solution, for example, 50 mM EDTA, 20 mM Tris-HCl (pH 8.0) and 8 M urea. RNA can be purified by methods known in the art, such as LiCl precipitation.
[0126] crRNA-Cpf1 complex When associated with the appropriate crRNA, Cpf1 forms a crRNA-Cpf1 complex. The crRNA-Cpf1 complex formed by some of the Cpf1 mutants of the present disclosure can: - introducing a single-stranded break in a first target nucleic acid; and - specifically recognizing a second target nucleic acid. This is in contrast to the crRNA-Cpf1 complex formed by wild-type Cpf1 endonuclease, which is able to recognize a specific nucleic acid and cleave both strands of said specific nucleic acid in a staggered manner.
[0127] The crRNA may be modified to allow the crRNA-Cpf1 complex to recognize various target nucleic acids. In particular, the variable region of the crRNA can be designed to recognize various targets. The sequence of the variable region of the crRNA must be complementary to the target nucleic acid.
[0128] It will be understood that Cpf1 or its homologues from a given organism should preferably be used in combination with a crRNA having a constant region as naturally found in the same organism. Thus, FnCpf1 or its homologues are preferably assembled with a crRNA having a constant region as set forth in SEQ ID NO:5.
[0129] In order for the crRNA-Cpf1 complex to function efficiently as an endonuclease or restriction enzyme, the ratio of Cpf1 to crRNA is preferably adjusted. Preferably, the molar ratio of Cpf1 to crRNA is 0.5:3.0 to 1.0:1.0, such as 0.7:2.5, such as 0.8:2.0, such as 0.9:1.75, such as 0.95:1.5, such as 1.0:1.4, such as 1.0:1.3, such as 1.1:1.2.
[0130] In some embodiments, a Cpf1 mutant capable of forming a crRNA-Cpf1 complex that introduces a single-strand break in a first target nucleic acid and has the ability to specifically recognize and optionally cleave a second target nucleic acid as described herein is selected from the group consisting of: i) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, for example at least 99% sequence identity, to a sequence corresponding to residues 1 to 297, 310 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO: 2; wherein the polypeptide sequence further comprises at least one amino acid substitution or deletion in the Finger domain (residues 298-309; SEQ ID NO: 12), the REC domain (residues 324-336; SEQ ID NO: 16) or the Lid domain (residues 1006-1018; SEQ ID NO: 20) compared to SEQ ID NO: 2; and / or ii) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2; wherein the polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2; and Cpf1 mutants, including
[0131] In some embodiments, the cleavage of the second target nucleic acid is a single-stranded cleavage. In other embodiments, the cleavage of the second target nucleic acid is a double-stranded cleavage. Preferably, the cleavage of the second target nucleic acid is a single-stranded cleavage. Cleavage can be further enhanced by the presence of activator DNA, which can be single-stranded or double-stranded.
[0132] As can be seen in the example, mutant Cpf1 with Q1025E substitution can perform such non-specific DNA cleavage of single-stranded DNA. Cleavage is enhanced by the presence of single-stranded or double-stranded target DNA. Similarly, mutant Cpf1 with REC domain deletion as described in SEQ ID NO: 34 and mutant Cpf1 with REC domain replaced in its entirety (SEQ ID NO: 36) can also perform such non-specific DNA cleavage of single-stranded DNA. Cleavage is enhanced by the presence of single-stranded or double-stranded target DNA. Cpf1 mutant with Lid domain replaced in its entirety can perform such non-specific DNA cleavage of single-stranded DNA, albeit to a lower extent than the above mutant; however, in this case, cleavage is thought to be enhanced by the presence of only single-stranded target DNA. K1013G, R1014G mutants can also perform such non-specific cleavage.
[0133] Thus, provided herein are Cpf1 mutants that, when in complex with crRNA, cleave nucleic acid sequences more efficiently than the wild-type Cpf1 endonuclease from which they are derived. In some embodiments, the mutant Cpf1 endonuclease cuts both the target and non-target strands more efficiently than the wild-type Cpf1 endonuclease. In other embodiments, the mutant Cpf1 endonuclease alternatively or additionally cuts only one of the target and non-target strands more efficiently than the wild-type Cpf1 endonuclease. In other embodiments, the mutant Cpf1 endonuclease can alternatively or additionally introduce single-strand breaks in a non-specific manner, as detailed herein below.
[0134] For example, a Cpf1 mutant containing a mutation or substitution at position 1025 of SEQ ID NO:2 can cleave a nucleic acid sequence more efficiently than the wild-type Cpf1 endonuclease of SEQ ID NO:2.
[0135] A Cpf1 mutant containing a mutation or substitution at positions 1025 and 1028 of SEQ ID NO:2 is capable of cleaving a nucleic acid sequence more efficiently than the wild-type Cpf1 endonuclease of SEQ ID NO:2.
[0136] A Cpf1 mutant comprising at least one mutation or substitution in the fingers domain defined by residues 298-309 of SEQ ID NO:2 is capable of cleaving a nucleic acid sequence more efficiently than the wild-type Cpf1 endonuclease of SEQ ID NO:2.
[0137] A Cpf1 mutant comprising at least one mutation or substitution in the REC domain defined by residues 324-336 of SEQ ID NO:2 is capable of cleaving a nucleic acid sequence more efficiently than the wild-type Cpf1 endonuclease of SEQ ID NO:2.
[0138] Cpf1 mutants containing mutations or substitutions in the Lid domain defined by residues 1006-1018 of SEQ ID NO:2 are capable of cleaving nucleic acid sequences more efficiently than the wild-type Cpf1 endonuclease of SEQ ID NO:2.
[0139] First target nucleic acid The mutant Cpf1 endonuclease of the present disclosure, after forming the crRNA-Cpf1 complex, can introduce a single-strand break in the first target nucleic acid, which is a nucleic acid sequence that cannot be specifically recognized by the crRNA-Cpf1 complex of the present disclosure, but can nevertheless be cleaved.
[0140] In some embodiments, the first target nucleic acid is identical to the second target nucleic acid, which is recognized and cleaved by the crRNA-Cpf1 complex. Thus, the first target nucleic acid may be as defined in the section "Second target nucleic acid" below.
[0141] In some embodiments, the first target nucleic acid is different from the second target nucleic acid and is cleaved by the crRNA-Cpf1 complex only after hybridization of the crRNA-Cpf1 complex to the second target nucleic acid. Thus, in some embodiments, the first target nucleic acid does not contain a protospacer adjacent motif (PAM).
[0142] Second target nucleic acid The recognition and binding of the crRNA-Cpf1 complex to the second target nucleic acid depends on the binding of the crRNA to the second target nucleic acid. This depends on the presence of a PAM (protospacer adjacent motif) in the target nucleic acid. In some embodiments, Cpf1 is FnCpf1 and the PAM sequence is 5'-YRN-3' or 5'-RYN-3', where Y is a pyrimidine such as A or G, R is a purine such as T, U or C, and N is any nucleotide. Preferably, R is T or C.
[0143] In one embodiment, the PAM consists of the sequence 5'-TTN-3'. In another embodiment, the PAM consists of the sequence 5'-CCN-3'. In another embodiment, the PAM consists of the sequence 5'-TAN-3'. In another embodiment, the PAM consists of the sequence 5'-TCN-3'. In another embodiment, the PAM consists of the sequence 5'-TGN-3'. In another embodiment, the PAM consists of the sequence 5'-CTN-3'. In another embodiment, the PAM consists of the sequence 5'-CAN-3'. In another embodiment, the PAM consists of the sequence 5'-CGN-3'. In another embodiment, the PAM consists of the sequence 5'-ATN-3'. In another embodiment, the PAM consists of the sequence 5'-AAN-3'. In another embodiment, the PAM consists of the sequence 5'-ACN-3'. In another embodiment, the PAM consists of the sequence 5'-AGN-3'. In another embodiment, the PAM consists of the sequence 5'-TTN-3'. In another embodiment, the PAM consists of the sequence 5'-TAN-3'. In another embodiment, the PAM consists of the sequence 5'-TCN-3'. In another embodiment, the PAM consists of the sequence 5'-TGN-3'.
[0144] The crRNA preferably hybridizes to the PAM itself. The choice of PAM sequence may be determined by the organism from which the Cpf1 to be used to cleave the target nucleic acid is derived.
[0145] The second target nucleic acid may comprise or consist of a recognition sequence comprising a sequence of at least 15 consecutive nucleotides, such as at least 16 consecutive nucleotides, such as at least 17 consecutive nucleotides, such as at least 18 consecutive nucleotides, such as at least 19 consecutive nucleotides, such as at least 20 consecutive nucleotides, such as at least 21 consecutive nucleotides, such as at least 22 consecutive nucleotides, such as at least 23 consecutive nucleotides, such as at least 24 consecutive nucleotides, such as at least 25 consecutive nucleotides, such as at least 26 consecutive nucleotides, such as at least 27 consecutive nucleotides, provided that the 5'-terminal three nucleic acids consist of a PAM sequence.
[0146] The second target nucleic acid sequence is DNA or RNA. The second target nucleic acid sequence may be DNA selected from the group consisting of genomic DNA, chromatin, nucleosomes, plasmid DNA, methylated DNA, synthetic DNA, and DNA fragments, such as PCR products.
[0147] In some embodiments, the second target nucleic acid is RNA. In some embodiments, the second target nucleic acid is preferably DNA. The second target nucleic acid may be any nucleic acid for which it may be desirable to specifically cleave. The second target nucleic acid may be purified by methods known in the art before contacting it with the crRNA-Cpf1 complex. In some embodiments, the second target nucleic acid may be one that is recognized and hybridized in vivo.
[0148] Use of the crRNA-Cpf1 endonuclease complex to introduce single-strand breaks One aspect of the present disclosure relates to the use of a crRNA-Cpf1 complex for introducing a single-stranded break in a first target nucleic acid, comprising: a. contacting a Cpf1 endonuclease or its orthologue with a guide RNA (crRNA), thereby obtaining a crRNA-Cpf1 complex capable of recognizing a second target nucleic acid containing a protospacer adjacent motif (PAM); and b. contacting the crRNA-Cpf1 complex with a first target nucleic acid; causes a single-stranded break in the first target sequence.
[0149] The crRNA-Cpf1 complex: i) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, for example at least 99% sequence identity, to a sequence corresponding to residues 1 to 297, 310 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO: 2; wherein the polypeptide sequence further comprises at least one amino acid substitution or deletion in the Finger domain (residues 298-309; SEQ ID NO: 12), the REC domain (residues 324-336; SEQ ID NO: 16) or the Lid domain (residues 1006-1018; SEQ ID NO: 20) compared to SEQ ID NO: 2; and / or ii) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2; wherein the polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2; The present invention relates to a mutant Cpf1 endonuclease or its ortholog, comprising:
[0150] The crRNA-Cpf1 complex: i) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, for example at least 99% sequence identity, to a sequence corresponding to residues 1 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO: 2; wherein the polypeptide sequence is: a. at least one amino acid mutation in the REC domain (residues 324-336; SEQ ID NO: 16) compared to SEQ ID NO: 2, where each mutation is independently an amino acid substitution or deletion; and / or b. at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, where at least one of the residues at positions 917 and 1006 is glutamic acid (E) or aspartic acid (D); Further includes; and / or ii) a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2; The present invention relates to a mutant Cpf1 endonuclease or its ortholog, comprising:
[0151] In another embodiment, the crRNA-Cpf1 complex comprises a mutant Cpf1 endonuclease or its orthologue comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least one amino acid substitution at position 918, 1013 or 1014 of SEQ ID NO:2, wherein said mutant Cpf1 is described in detail herein above in the section "Mutant Cpf1".
[0152] In one embodiment, the crRNA hybridizes to a second target nucleic acid and Cpf1 introduces a single-stranded break at a specific recognition nucleotide sequence of the first target nucleic acid.
[0153] In one embodiment, the crRNA hybridizes to the second target nucleic acid, and Cpf1 introduces a single-strand break in a specific recognition nucleotide sequence of the first target nucleic acid, where the first and second target nucleic acids are identical or overlapping. For example, the first and second target nucleic acids may both contain the same recognition sequence and the same PAM.
[0154] In one embodiment, the crRNA hybridizes to a second target nucleic acid and Cpf1 introduces a single-stranded break in a specific recognition nucleotide sequence of the first target nucleic acid, wherein the first and second target nucleic acids are the same nucleic acid. The characteristics of crRNA are described herein above in the section "Guide RNA (crRNA)."
[0155] Thus, the first nucleic acid and the second nucleic acid may be the same or different. In one embodiment, the first nucleic acid and the second nucleic acid are the same. For example, the first nucleic acid and the second nucleic acid are the same target nucleic acid.
[0156] In one embodiment, the crRNA hybridizes to a second target nucleic acid and Cpf1 introduces single-stranded breaks at one or more non-specific nucleotide sequences of the first target nucleic acid.
[0157] In one embodiment, the first and second nucleic acids are different, and the first target nucleic acid does not contain a PAM sequence and is therefore not recognized by the crRNA-Cpf1 complex, whereas the second target nucleic acid contains a PAM sequence and is recognized by the crRNA-Cpf1 complex.
[0158] The mutant Cpf1 contained in the crRNA-Cpf1 complex useful for introducing single-strand breaks in nucleic acid sequences may be any of the mutants described herein. However, some mutants may be particularly advantageous. For example, mutant Cpf1 with Q1025G substitution, mutant Cpf1 deleted for the REC domain as described in SEQ ID NO: 34, mutant Cpf1 with the entire REC domain replaced (SEQ ID NO: 36), mutant Cpf1 mutant with the entire Lid domain replaced are all capable of such non-specific DNA cleavage of single-stranded DNA and are particularly useful for cleaving a second target nucleic acid that is different from the first target nucleic acid. Mutant Cpf1 with mutations at positions 1013 and 1014 or at positions 918 and 1013 may also be of note.
[0159] In one aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises an amino acid substitution at position 1025 of SEQ ID NO:2.
[0160] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises a deletion of one or more amino acid residues of the REC domain defined by residues 298-309 of SEQ ID NO:2.
[0161] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises a deletion of one or more amino acid residues of the Lid domain defined by residues 324 to 336 of SEQ ID NO:2.
[0162] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises mutations of two or more amino acid residues in the REC domain defined by residues 324-336 of SEQ ID NO:2, wherein each mutation is independently an amino acid substitution or deletion.
[0163] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises a deletion of one or more amino acid residues of the Lid domain defined by residues 1006 to 1018 of SEQ ID NO:2.
[0164] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises a substitution of two or more amino acid residues in the Lid domain defined by residues 1006-1018 of SEQ ID NO:2.
[0165] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises an amino acid substitution at one or more of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2.
[0166] In one aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions 1013 and 1014 of SEQ ID NO:2.
[0167] The substitution may in some embodiments be a substitution of an amino acid having a charged side chain with an amino acid having an uncharged or non-polar side chain. In some embodiments, the amino acid is substituted with an amino acid selected from the group consisting of glycine, alanine, valine, leucine, isoleucine, serine, or threonine. In some embodiments, the amino acid is substituted with glycine.
[0168] In some embodiments, the deletion may be one, more than one, or all amino acid residues of the REC domain, the Lid domain, or the finger domain. In some embodiments, some amino acids are substituted, while others are deleted. In some embodiments, the Cpf1 mutant comprises amino acid deletions and / or substitutions in more than one domain, for example, in the REC domain and the Lid domain, in the REC domain and the finger domain, in the Lid domain and the finger domain, or in the REC domain, the Lid domain, and the finger domain. The Cpf1 mutant may further comprise an amino acid substitution or deletion at one or more of positions 918, 1013, 1014, 1025, or 1028.
[0169] For example, the crRNA-Cpf1 complex is used to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, where the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and where the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions 1013 and 1014 of SEQ ID NO:2.
[0170] In one aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids having charged side chains with amino acids having uncharged side chains at positions 1013 and 1014 of SEQ ID NO:2.
[0171] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids having charged side chains with amino acids having uncharged side chains at positions 1013 and 1014 of SEQ ID NO:2.
[0172] In one aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids having charged side chains with amino acids having non-polar side chains at positions 1013 and 1014 of SEQ ID NO:2.
[0173] In one aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids having charged side chains with amino acids having non-polar side chains at positions 1013 and 1014 of SEQ ID NO:2.
[0174] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains relative to the strand at positions 1013 and 1014 of SEQ ID NO:2 with glycine, alanine, valine, leucine, isoleucine, serine or threonine.
[0175] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains relative to the strand at positions 1013 and 1014 of SEQ ID NO:2 with glycine, alanine, valine, leucine, isoleucine, serine or threonine.
[0176] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains to glycine at positions 1013 and 1014 of SEQ ID NO:2 relative to the strand.
[0177] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains to glycine at positions 1013 and 1014 of SEQ ID NO:2 relative to the strand.
[0178] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains to glycine at positions 1013 and 1014 of SEQ ID NO:2 relative to the strand.
[0179] In one aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises the substitutions K1013G and R1014G of SEQ ID NO:2.
[0180] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of a polypeptide having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity, such as 100% sequence identity to SEQ ID NO:3.
[0181] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of a polypeptide having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity, such as 100% sequence identity to SEQ ID NO:3.
[0182] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:3.
[0183] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:3.
[0184] Some of the mutant Cpf1 endonucleases described herein are capable of introducing a single-stranded break in a target nucleic acid, referred to herein as the first target nucleic acid, which is different from the nucleic acid sequence recognized and hybridized by the crRNA, referred to herein as the second target nucleic acid.
[0185] In some embodiments, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-strand breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, where the first target nucleic acid and the second target nucleic acid are different, and where the Cpf1 endonuclease or its orthologue has a mutation at position 1025 of SEQ ID NO: 2. In a particular embodiment, the mutation is a Q1025G mutation.
[0186] In other embodiments, the disclosure relates to the use of a crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are different, and wherein the Cpf1 endonuclease or its orthologue has a substitution or deletion of at least one amino acid in the REC domain, corresponding to residues 324-336 of SEQ ID NO:2. For example, a mutant Cpf1 has a complete deletion of the REC domain and has the sequence set forth in SEQ ID NO:34. In a particular embodiment, the entire REC domain is replaced, and the mutant Cpf1 has, for example, the sequence set forth in SEQ ID NO:36.
[0187] In other aspects, the disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are different, and wherein the Cpf1 endonuclease or its orthologue has a substitution or deletion of at least one amino acid in the Lid domain corresponding to residues 324-336 of SEQ ID NO: 2. For example, the entire Lid domain is replaced and the mutant Cpf1 has the sequence set forth in SEQ ID NO: 42.
[0188] In other aspects, the disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are different, and wherein the Cpf1 endonuclease or its orthologue has a substitution or deletion of at least one amino acid in the Lid domain corresponding to residues 1006-1018 of SEQ ID NO: 2. For example, the entire Lid domain is replaced and the mutant Cpf1 has the sequence set forth in SEQ ID NO: 38.
[0189] In some embodiments, the crRNA-Cpf1 complex comprises an amino acid substitution or deletion in more than one of the REC, Lid or finger domains, for example, in the REC domain and in the Lid domain: in the REC domain and in the finger domain; in the Lid domain and in the finger domain; or in the REC domain, finger domain and Lid domain. The mutant Cpf1 may further comprise an amino acid substitution or deletion at one or more of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0190] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and a second target nucleic acid are different, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids having charged side chains with amino acids having uncharged side chains at positions 918 and 1013 of SEQ ID NO:2.
[0191] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids having charged side chains with amino acids having uncharged side chains at positions 918 and 1013 of SEQ ID NO:2.
[0192] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids having charged side chains with amino acids having non-polar side chains at positions 918 and 1013 of SEQ ID NO:2.
[0193] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids having charged side chains with amino acids having non-polar side chains at positions 918 and 1013 of SEQ ID NO:2.
[0194] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single strand breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains relative to the strand at positions 918 and 1013 of SEQ ID NO:2 with glycine, alanine, valine, leucine, isoleucine, serine or threonine.
[0195] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single strand breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains relative to the strand at positions 918 and 1013 of SEQ ID NO:2 with glycine, alanine, valine, leucine, isoleucine, serine or threonine.
[0196] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains relative to the strand at positions 918 and 1013 of SEQ ID NO:2 with glycine.
[0197] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single strand breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains relative to the strand at positions 918 and 1013 of SEQ ID NO:2 with glycine.
[0198] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions of amino acids with charged side chains to glycine at positions 918 and 1013 of SEQ ID NO:2 relative to the strand.
[0199] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises substitutions R918G and K1013G of SEQ ID NO:2.
[0200] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing single strand breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of a polypeptide having at least 80% sequence identity to SEQ ID NO:4, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity, such as about 100% sequence identity.
[0201] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing single strand breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of a polypeptide having at least 80% sequence identity to SEQ ID NO:4, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity, such as about 100% sequence identity.
[0202] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:4.
[0203] In one embodiment, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce single-stranded breaks in one or more non-specific nucleotide sequences of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are the same nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:4.
[0204] The mutant may further comprise additional mutations or deletions as described herein above in the section "Mutant Cpf1." The crRNA-Cpf1 complex can be used to introduce a single-strand break in a first target nucleic acid in an ex-vivo manner. The crRNA-Cpf1 complex can also be used to introduce a single-strand break in a first target nucleic acid in an in-vivo manner.
[0205] In some embodiments, the efficiency of the introduction of single-strand breaks as described above is enhanced by the presence of activator DNA.The activator DNA can be single-stranded or double-stranded.In some embodiments, the activator DNA is the target DNA that corresponds to the second target nucleic acid.
[0206] Use of the crRNA-Cpf1 endonuclease complex for genome editing Some of the mutant Cpf1 endonucleases of the present disclosure can only introduce single-strand breaks in the target sequence hybridized by the crRNA of the crRNA-Cpf1 complex. Thus, in some embodiments, when the crRNA of the crRNA-Cpf1 complex recognizes and hybridizes to a second target sequence, the DNAse activity of Cpf1 of said complex is activated, which will specifically introduce single-strand breaks at a specific site of the second target nucleic acid sequence. Moreover, non-specific nucleic acid sequences will not be cleaved by mutant Cpf1. This constitutes a great advantage over wild-type Cpf1, which instead has non-specific ssDNA activity. Indeed, single-strand non-specific DNA cleavage is a dangerous activity for cells, since it targets DNA that opens during normal cellular activities such as replication and transcription. The main consequence of this non-specific ssDNA activity is the generation of random breaks, which can lead to undesired mutagenesis and even cell death.
[0207] The mutant Cpf1 endonucleases of the present disclosure can be advantageously used for genome editing in a safer manner compared to wild-type Cpf1. Thus, one aspect of the present disclosure relates to an in vitro method of introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell, the method comprising introducing into the mammalian cell a crRNA-Cpf1 complex, wherein Cpf1 is a mutant Cpf1 endonuclease or an orthologue thereof as described herein, and wherein the crRNA is specific for the second target nucleic acid.
[0208] In some embodiments, the mammalian cell is not a germ cell or an embryonic cell. In some embodiments, the mutant Cpf1 endonuclease disclosed herein and used in the in vitro method of introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof, comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 1025 and 1028 of SEQ ID NO:2.
[0209] In some embodiments, the mutant Cpf1 endonuclease disclosed herein and used in the in vitro method of introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof, comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 1025 and 1028 of SEQ ID NO:2 of amino acids having charged side chains to amino acids having uncharged side chains.
[0210] In some embodiments, the mutant Cpf1 endonuclease disclosed herein and used in the in vitro method of introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof, comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 1025 and 1028 of SEQ ID NO:2 of amino acids having charged side chains with amino acids having non-polar side chains.
[0211] In some embodiments, the mutant Cpf1 endonuclease disclosed herein and used in the in vitro method of introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof, comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions of amino acids having charged side chains at positions 1025 and 1028 of SEQ ID NO:2 with glycine, alanine, valine, leucine, isoleucine, serine or threonine.
[0212] In some embodiments, the mutant Cpf1 endonuclease disclosed herein and used in the in vitro method of introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof, comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions of amino acids with charged side chains to glycine at positions 1025 and 1028 of SEQ ID NO:2.
[0213] For example, a mutant Cpf1 endonuclease for use in an in vitro method for introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises the amino acid substitutions Q1025G and E1028G.
[0214] For example, a mutant Cpf1 endonuclease for use in an in vitro method for introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, for example at least 99% sequence identity to SEQ ID NO: 32.
[0215] The mutant may further comprise additional mutations or deletions as described herein above in the section "Mutant Cpf1." For example, a mutant Cpf1 endonuclease used in an in vitro method for introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof comprising at least two amino acid substitutions in the Lid domain as defined herein and one or more amino acid substitutions at any one of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0216] For example, a mutant Cpf1 endonuclease used in an in vitro method for introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof comprising at least two amino acid mutations in the REC domain as defined herein (wherein each mutation is independently an amino acid substitution or deletion) and one or more amino acid substitutions at any one of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0217] For example, a mutant Cpf1 endonuclease used in an in vitro method for introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof comprising at least two amino acid substitutions in the REC domain as defined herein and one or more amino acid substitutions at any one of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0218] For example, a mutant Cpf1 endonuclease used in an in vitro method for introducing a site-specific double-stranded break in a second target DNA sequence in a mammalian cell may be a Cpf1 endonuclease or an orthologue thereof comprising at least two amino acid deletions in the REC domain as defined herein and one or more amino acid substitutions at any one of positions 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0219] Use of the crRNA-Cpf1 endonuclease complex for detection and / or quantification of target DNA sequences Some of the mutant Cpf1 endonucleases of the present disclosure can only introduce single-strand breaks in one target sequence that is not hybridized by the crRNA of the crRNA-Cpf1 complex. Thus, in some embodiments, when the crRNA of the crRNA-Cpf1 complex recognizes and hybridizes to a second target sequence, the DNAse activity of Cpf1 of the complex is activated, which will non-specifically introduce single-strand breaks at random sites in the nucleic acid sequence, e.g., at random sites in the first target sequence. Furthermore, the second target nucleic acid will not be cleaved by Cpf1, which will therefore remain active for a longer period of time and cleave more than one first target sequence. If the first target sequence is labeled in such a way that a signal is emitted upon cleavage of the first target sequence, the described method will therefore allow detection of the second target sequence.
[0220] These mutant Cpf1 endonucleases, when in the crRNA-Cpf1 complex, can be used to detect and quantify the second target sequence with the aid of the provided labeled first target sequence.
[0221] Accordingly, one aspect of the present disclosure relates to a method for detection of a second target nucleic acid in a sample, the method comprising: a. providing a crRNA-Cpf1 complex, where Cpf1 is a Cpf1 endonuclease as described herein or an orthologue thereof, and where the crRNA is specific for a second target nucleic acid; b. providing labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one dye and at least one quencher; c. contacting the crRNA-Cpf1 complex and the ssDNA with a sample, wherein the sample contains at least one second target nucleic acid; and d. detecting ssDNA breaks by detecting a signal from the dye; thereby detecting the presence of a second target nucleic acid in the sample.
[0222] In step c., the crRNA-Cpf1 complex and the ssDNA are contacted with at least one second target nucleic acid, and recognition and binding of the crRNA to the second target nucleic acid results in activation of the crRNA-Cpf1 complex, which can then introduce a single-strand break, e.g., a cleavage, into the ssDNA. Thus, step c. may involve activation of the crRNA-Cpf1 complex.
[0223] The method may further comprise determining the level and / or concentration of a second target nucleic acid, wherein the level and / or concentration of the second target nucleic acid correlates to the cleaved ssDNA.
[0224] As explained above, in some embodiments, the mutant Cpf1 endonuclease disclosed herein will not cleave the second target nucleic acid, and therefore will remain active for a sufficient period of time to cleave two or more first target nucleic acids. The first target nucleic acid can be ssDNA or a fragment thereof in the method described herein. After hybridization of ccRNA-Cpf1 complex to the second target nucleic acid, the more first target nucleic acid molecules are cleaved by ccRNA-Cpf1 complex, the higher the signal, and therefore the more sensitive the method. This is the advantage of the mutant Cpf1 disclosed herein over other Cpf1 endonucleases.
[0225] Thus, the methods disclosed herein have high sensitivity and may allow detection of a second target nucleic acid at concentrations in the nanomolar range or below, e.g., at concentrations in the picomolar range or below, e.g., at concentrations in the femtomolar range or below, for example, the methods disclosed herein allow detection of a second target nucleic acid at concentrations in the attomolar range.
[0226] In some embodiments, the mutant Cpf1 endonucleases disclosed herein cleave a second target nucleic acid and thus will remain active only until the cleaved second target nucleic acid is released.
[0227] The ssDNA may be labeled at least at one base at any position along the strand.For example, the ssDNA is labeled at one base at any position along the strand, for example at least two bases at any position along the strand, for example at least three bases at any position along the strand, for example at least four bases at any position along the strand.
[0228] The ssDNA may be labeled with at least one set of interactive labels comprising at least one dye and at least one quencher. In one embodiment, at least one dye is a fluorophore. Thus, cleavage of the ssDNA in step d. of the method comprises detecting a fluorescent signal resulting from cleavage of the ssDNA.
[0229] In one embodiment, the at least one fluorophore is selected from the group including Black Hole Quencher (BHQ)1, BHQ2 and BHQ3, Cosmic Quencher (e.g., from Biosearch Technologies, Novato, USA), Excellent Bioneer Quencher (EBQ) (e.g., from Bioneer, Daejeon, Korea) or a combination thereof.
[0230] In one embodiment, the at least one quencher is selected from the group comprising black hole quenchers (BHQ)1, BHQ2 and BHQ3 (Biosearch Technologies, Novato, USA).
[0231] Fluorophores that may be useful in the present invention may include any fluorescent molecule known in the art. Examples of fluorophores are: Cy2™ Cfflfi), YO-PRn™-1 (509), YDYO™-1 (509), Calrein (517), FITC (518), FluorX™ (519), Alexa™ (520), Rhodamine 110 (520), Oregon Green™ 500 (522), Oregon Green™ 488 (524), RiboGreen™ (525), Rhodamine Green™ (527), Rhodamine 123 (529), Magnesium Green™ (531), Calcium Green™ (533), TO-PRO™-l (533), TOTOl (533), JOE (548), 30 BODIPY530 / 550 (550), Dil (565), BODIPY TMR (568), BODIPY558 / 568 (568), BODIPY564 / 570 (570), Cy3™ (570), Alexa™ 546 (570), TRITC (572), Magnesium Orange™ (575), Phycoerythrin R&B (575), Rhodamine phalloidin (575), Calcium Orange™ (576), Pyronin Y (580), Rhodamine B (580), TAMRA (582), Rhodamine Red™ (590), Cy3.5™ (596), ROX (608), Calcium Crimson™ (615), Alexa™ 594 35(615), Texas Red(615), Nile Red(628), YO-PRO™-3(631), YOYO™-3(631), RP3649PC00 phycocyanin(642), C-phycocyanin(648), TO-PRO™-3(660), TOT03(660), DiD DilC(5)(665), Cy5™(670), Thiadicarbocyanine(671), Cy5.5(694), HEX(556), TET(536), Biosearch Blue(447), CAL Fluor Gold 540(544), CAL Fluor Orange 560(559), CAL Fluor Red 590(591), CAL Fluor Red 610(610), CAL Fluor Red 635(637), FAM(520), 6-Carboxyfluorescein (6-FAM), Fluorescein(520), Fluorescein-C3(520), Pulsar 650(566), Quasar 570(667), Quasar 670(705), and Quasar 705(610). The numbers in parentheses are the maximum emission wavelengths in nanometers.
[0232] It should be noted that non-fluorescent, black quencher molecules capable of quenching fluorescence over a broad range or at specific wavelengths can be used in the present invention. Suitable fluorophore / quencher pairs are known in the art.
[0233] The method can be used to detect the presence and level of any nucleic acid, and thus the sample can be any sample that contains nucleic acid and has been appropriately treated, for example to remove proteases. The sample can contain DNA and / or RNA. The sample can be a sample suspected of containing a second target nucleic acid. The sample can be any prokaryotic or eukaryotic cell culture culture extract, mammalian body fluid, such as human.
[0234] The second target nucleic acid may be a nucleic acid fragment of a viral genome, a microbial genome, a gene such as an oncogene, or the genome of a pathogen.
[0235] The second target nucleic acid may also be a mutated nucleic acid sequence, such as a single nucleotide polymorphism (SNP). The mutant Cpf1 endonuclease used in the method for detection of a second target nucleic acid in a sample may be any of the mutants described herein above, particularly in the section "Mutant Cpf1".
[0236] The mutant Cpf1 endonuclease used in the method for detection of a second target nucleic acid in a sample may be a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2.
[0237] For example, a mutant Cpf1 endonuclease for use in a method for detection of a second target nucleic acid in a sample is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2 of amino acids having charged side chains to amino acids having uncharged side chains.
[0238] For example, a mutant Cpf1 endonuclease for use in a method for detection of a second target nucleic acid in a sample is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2 of amino acids having a charged side chain to amino acids having a non-polar side chain.
[0239] For example, a mutant Cpf1 endonuclease used in a method for detection of a second target nucleic acid in a sample is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises substitutions of at least two amino acids having a charged side chain at positions 918 and 1013 of SEQ ID NO:2 with glycine, alanine, valine, leucine, isoleucine, serine or threonine.
[0240] For example, a mutant Cpf1 endonuclease used in a method for detection of a second target nucleic acid in a sample is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions of amino acids with charged side chains to glycine at positions 918 and 1013 of SEQ ID NO:2.
[0241] For example, a mutant Cpf1 endonuclease used in a method for detection of a second target nucleic acid in a sample is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises the amino acid substitutions R918G and K1013G.
[0242] For example, the mutant Cpf1 endonuclease used in the method for detection of a second target nucleic acid in a sample is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity to SEQ ID NO:4, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, such as at least 99% sequence identity, for example about 100% sequence identity.
[0243] For example, a mutant Cpf1 endonuclease used in a method for the detection of a second target nucleic acid in a sample is a Cpf1 endonuclease or an orthologue thereof, which comprises or consists of SEQ ID NO:4.
[0244] The mutant Cpf1 endonuclease used in the method for detection of a second target nucleic acid in a sample may also contain other amino acid modifications.
[0245] Use of crRNA-Cpf1 endonuclease complex for diagnosis of disease The present disclosure also relates to methods for the diagnosis of any disease associated with increased / decreased gene expression and / or with the presence of foreign genetic material.
[0246] One aspect of the present disclosure relates to an in vitro method for diagnosing a disease in a subject, the method comprising: a. providing a crRNA-Cpf1 complex, where Cpf1 is a mutant Cpf1 endonuclease as defined herein or an orthologue thereof, and where the crRNA is specific for a second target nucleic acid; b. providing labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one fluorophore and at least one quencher; c. providing a sample from a subject, wherein said sample contains or is suspected of containing a second target nucleic acid; and d. determining the level and / or concentration of the second target nucleic acid according to the methods of detection described herein; wherein the second target nucleic acid is a nucleic acid fragment correlated with a disease; thereby diagnosing the disease in the subject.
[0247] The method for diagnosis of a disease in a subject may further comprise a step of treating said disease. For example, the method may further comprise treating said disease by administering a therapeutically effective agent.
[0248] The mutant Cpf1 endonuclease used in the method for detection of a second target nucleic acid in a sample may be any of the mutants described herein above, particularly in the section "Mutant Cpf1".
[0249] Use of crRNA-Cpf1 endonuclease complex for the diagnosis of infectious diseases One aspect of the present disclosure relates to an in vitro method for diagnosing an infectious disease in a subject, the method comprising: a. providing a crRNA-Cpf1 complex, where Cpf1 is a mutant Cpf1 endonuclease as defined herein or an orthologue thereof, and where the crRNA is specific for a second target nucleic acid; b. providing a labeled ssDNA, wherein the ssDNA is labeled with at least one set of interactive labels, wherein said interactive labels generate a signal upon cleavage of the ssDNA; c. providing a sample from a subject, wherein said sample contains or is suspected of containing a second target nucleic acid; and d. determining the level and / or concentration of the second target nucleic acid according to the methods of detection described herein; wherein the second target nucleic acid is a nucleic acid fragment of the genome of a disease-causing infectious pathogen; thereby diagnosing an infectious disease in a subject.
[0250] One aspect of the present disclosure relates to an in vitro method for diagnosing an infectious disease in a subject, the method comprising: a. providing a crRNA-Cpf1 complex, where Cpf1 is a mutant Cpf1 endonuclease as defined herein or an orthologue thereof, and where the crRNA is specific for a second target nucleic acid; b. providing labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one fluorophore and at least one quencher; c. providing a sample from a subject, wherein said sample contains or is suspected of containing a second target nucleic acid; and d. determining the level and / or concentration of the second target nucleic acid according to the methods of detection described herein; wherein the second target nucleic acid is a nucleic acid fragment of the genome of a disease-causing infectious pathogen; thereby diagnosing an infectious disease in a subject. The interactive labels may include, for example, luminescent labels.
[0251] The method for diagnosis of an infectious disease in a subject may further comprise a step of treating said infectious disease. For example, the method may further comprise treating said infectious disease by administering a therapeutically effective agent.
[0252] The method for diagnosis of an infectious disease in a subject may further comprise the step of comparing the level and / or concentration of said second target nucleic acid to a cut-off value, wherein the cut-off value is determined from a concentration range of the second target nucleic acid in a healthy subject, e.g., a subject not exhibiting an infectious disease; Here, a level and / or concentration higher than the cut-off value indicates the presence of an infectious disease.
[0253] An infectious disease is any disease caused by an infectious agent such as a virus, viroid, prion, bacteria, nematodes, parasitic roundworms, pinworms, arthropods, fungi, ringworm and macroparasites.
[0254] Thus, the second target nucleic acid may be the genome or a fragment thereof of an infectious agent selected from the group consisting of a virus, a viroid, a prion, a bacterium, a nematode, a parasitic roundworm, a pinworm, an arthropod, a fungus, a ringworm, and a macroparasite.
[0255] The methods disclosed herein can be used to diagnose infectious diseases in humans. Thus, the sample containing the second target nucleic acid may be a sample taken from a human body, for example, the sample may be a human body fluid selected from the group consisting of blood, whole blood, plasma, serum, urine, saliva, tears, cerebrospinal fluid and semen.
[0256] The mutant Cpf1 endonuclease used in the method for detection of a second target nucleic acid in a sample may be any of the mutants described herein above, particularly in the section "Mutant Cpf1".
[0257] In one aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises an amino acid substitution at position 1025 of SEQ ID NO:2.
[0258] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises a deletion of one or more amino acid residues of the REC domain defined by residues 298-309 of SEQ ID NO:2.
[0259] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex to introduce a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises a deletion of one or more amino acid residues of the Lid domain defined by residues 324 to 336 of SEQ ID NO:2.
[0260] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises a deletion of one or more amino acid residues of the Lid domain defined by residues 1006 to 1018 of SEQ ID NO:2.
[0261] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises a substitution of two or more amino acid residues in the Lid domain defined by residues 1006-1018 of SEQ ID NO:2.
[0262] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing a single-stranded break at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises an amino acid substitution at one or more of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2.
[0263] In another aspect, the present disclosure relates to the use of crRNA-Cpf1 complex for introducing single-stranded breaks at a specific recognition nucleotide sequence of a first target nucleic acid, wherein the first target nucleic acid and the second target nucleic acid are identical, and wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at two or more of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2.
[0264] The mutant Cpf1 endonuclease used in the method for diagnosis of an infectious disease is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2.
[0265] For example, a mutant Cpf1 endonuclease for use in a method for the diagnosis of an infectious disease is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2 of amino acids having charged side chains to amino acids having uncharged side chains.
[0266] For example, a mutant Cpf1 endonuclease for use in a method for diagnosis of an infectious disease is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises at least two amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2 of amino acids having a charged side chain to amino acids having a non-polar side chain.
[0267] For example, a mutant Cpf1 endonuclease for use in a method for the diagnosis of an infectious disease is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises substitutions of at least two amino acids having a charged side chain at positions 918 and 1013 of SEQ ID NO:2 with glycine, alanine, valine, leucine, isoleucine, serine or threonine.
[0268] For example, a mutant Cpf1 endonuclease for use in a method for diagnosis of an infectious disease in a sample is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises substitutions of at least two amino acids with glycine for amino acids having a charged side chain at positions 918 and 1013 of SEQ ID NO:2.
[0269] For example, a mutant Cpf1 endonuclease for use in a method for diagnosis of an infectious disease is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, such as at least 90% sequence identity, such as at least 95% sequence identity, such as at least 96% sequence identity, such as at least 97% sequence identity, such as at least 98% sequence identity, such as at least 99% sequence identity to SEQ ID NO:2, wherein said polypeptide sequence comprises the amino acid substitutions R918G and K1013G.
[0270] For example, a mutant Cpf1 endonuclease for use in a method for diagnosis of an infectious disease is a Cpf1 endonuclease or an orthologue thereof comprising a polypeptide sequence having at least 80% sequence identity, such as at least 85% sequence identity, for example at least 90% sequence identity, such as at least 95% sequence identity, for example at least 96% sequence identity, such as at least 97% sequence identity, for example at least 98% sequence identity, such as at least 99% sequence identity, for example about 100% sequence identity to SEQ ID NO:4.
[0271] For example, a mutant Cpf1 endonuclease used in a method for the diagnosis of an infectious disease may be a Cpf1 endonuclease or an orthologue thereof, which comprises or consists of SEQ ID NO:4.
[0272] The mutant Cpf1 endonuclease used in the method for the diagnosis of an infectious disease may also contain other amino acid modifications.
[0273] item 1. The following: i) a polypeptide sequence having at least 95% sequence identity to sequences corresponding to residues 1 to 297, 310 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO:2; wherein the polypeptide sequence further comprises at least one amino acid substitution or deletion in the Finger domain (residues 298-309; SEQ ID NO: 12), the REC domain (residues 324-336; SEQ ID NO: 16) or the Lid domain (residues 1006-1018; SEQ ID NO: 20) compared to SEQ ID NO: 2; and / or ii) a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2; wherein the polypeptide sequence comprises at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2; A mutant Cpf1 endonuclease or its ortholog, including
[0274] 2. The following: i) a polypeptide sequence having at least 95% sequence identity to the sequences corresponding to residues 1 to 323, 337 to 1005, and 1019 to 1329 of SEQ ID NO:2; wherein the polypeptide sequence is a. at least two amino acid mutations in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, where each mutation is independently an amino acid substitution or deletion; and / or b. at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, where at least one of the residues at positions 917 and 1006 is glutamic acid (E) or aspartic acid (D); Further includes; and / or ii) a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2; wherein the polypeptide sequence comprises at least two amino acid substitutions at two positions independently selected from positions corresponding to residues 918, 1013, 1014, 1025, and 1028 of SEQ ID NO:2; A mutant Cpf1 endonuclease or its ortholog, including
[0275] 3. The mutant Cpf1 endonuclease or its orthologue according to item 1, wherein the Cpf1 endonuclease is derived from Francisella novicida.
[0276] 4. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding paragraphs, wherein the mutant Cpf1 comprises a polypeptide sequence having at least 95% sequence identity to a sequence corresponding to residues 1-323, 337-1005, and 1019-1329 of SEQ ID NO:2, wherein the polypeptide sequence is a. at least two amino acid substitutions or deletions in the REC domain (residues 324-336; SEQ ID NO: 16) compared to SEQ ID NO: 2; and / or b. at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, except that at least one of the residues at positions 917 and 1006 of SEQ ID NO:2 is a glutamic acid (E) or aspartic acid (D); Further comprising: wherein said polypeptide sequence further comprises at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:2.
[0277] 5. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein one of the at least two amino acid substitutions is a substitution of an amino acid having a charged side chain for an amino acid having an uncharged side chain.
[0278] 6. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein at least one amino acid substitution is a substitution of an amino acid having a charged side chain for an amino acid having an uncharged side chain.
[0279] 7. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein at least one amino acid substitution is a substitution of an amino acid having a charged side chain for an amino acid residue having a non-polar side chain.
[0280] 8. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein one of the at least two amino acid substitutions is a substitution of an amino acid residue having a charged side chain for an amino acid residue having a non-polar side chain.
[0281] 9. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein one of the at least two amino acid substitutions is a substitution of an amino acid having a charged side chain with glycine, alanine, valine, leucine, isoleucine, serine, or threonine.
[0282] 10. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein at least one amino acid substitution is a substitution of an amino acid having a charged side chain with glycine, alanine, valine, leucine, isoleucine, serine, or threonine.
[0283] 11. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein one of the at least two amino acid substitutions is a substitution of an amino acid having a charged side chain with a glycine.
[0284] 12. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein one of the at least two amino acid substitutions is a substitution of an amino acid having a charged side chain with a glycine.
[0285] 13. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein at least one amino acid substitution or deletion is in the REC domain.
[0286] 14. The mutant Cpf1 endonuclease or orthologue thereof according to any one of the preceding items, wherein the substitution or deletion of at least one amino acid is a substitution or deletion of at least one residue in the REC domain, such as a substitution or deletion of at least 2 residues, such as a substitution or deletion of at least 3 residues, such as a substitution or deletion of at least 4 residues, such as a substitution or deletion of at least 5 residues, such as a substitution or deletion of at least 6 residues, such as a substitution or deletion of at least 7 residues, such as a substitution or deletion of at least 8 residues, such as a substitution or deletion of at least 9 residues, such as a substitution or deletion of at least 10 residues, such as a substitution or deletion of at least 11 residues, such as a substitution or deletion of at least 12 residues, such as a substitution or deletion of at least 13 residues, of the REC domain.
[0287] 15. The mutant Cpf1 endonuclease or orthologue thereof according to any one of the preceding items, wherein the substitution or deletion of at least two amino acids is a substitution or deletion of at least two residues in the REC domain, such as a substitution or deletion of at least 3 residues, such as a substitution or deletion of at least 4 residues, such as a substitution or deletion of at least 5 residues, such as a substitution or deletion of at least 6 residues, such as a substitution or deletion of at least 7 residues, such as a substitution or deletion of at least 8 residues, such as a substitution or deletion of at least 9 residues, such as a substitution or deletion of at least 10 residues, such as a substitution or deletion of at least 11 residues, such as a substitution or deletion of at least 12 residues, such as a substitution or deletion of at least 13 residues, of the REC domain.
[0288] 16. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein at least one substitution or deletion is a substitution or deletion of all residues in the REC domain.
[0289] 17. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein the at least two substitutions or deletions are substitutions or deletions of all residues in the REC domain.
[0290] 18. The mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein at least one amino acid substitution or deletion is in the Lid domain.
[0291] 19. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein at least two amino acid substitutions are in the Lid domain.
[0292] 20. The mutant Cpf1 endonuclease or orthologue thereof according to any one of the preceding items, wherein the substitution or deletion of at least one amino acid is a substitution or deletion of at least one residue in the Lid domain, such as a substitution or deletion of at least 2 residues, such as a substitution or deletion of at least 3 residues, such as a substitution or deletion of at least 4 residues, such as a substitution or deletion of at least 5 residues, such as a substitution or deletion of at least 6 residues, such as a substitution or deletion of at least 7 residues, such as a substitution or deletion of at least 8 residues, such as a substitution or deletion of at least 9 residues, such as a substitution or deletion of at least 10 residues, such as a substitution or deletion of at least 11 residues, such as a substitution or deletion of at least 12 residues, such as a substitution or deletion of at least 13 residues, of the Lid domain.
[0293] 21. The mutant Cpf1 endonuclease or orthologue thereof according to any one of the preceding items, wherein the substitution of at least two amino acids is a substitution of at least two residues in the Lid domain, such as a substitution or deletion of at least 3 residues, such as a substitution or deletion of at least 4 residues, such as a substitution or deletion of at least 5 residues, such as a substitution or deletion of at least 6 residues, such as a substitution or deletion of at least 7 residues, such as a substitution or deletion of at least 8 residues, such as a substitution or deletion of at least 9 residues, such as a substitution or deletion of at least 10 residues, such as a substitution or deletion of at least 11 residues, such as a substitution or deletion of at least 12 residues, such as a substitution or deletion of at least 13 residues, in the Lid domain.
[0294] 22. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein at least one substitution or deletion is a substitution or deletion of all residues in the Lid domain.
[0295] 23. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein at least two of the substitutions are substitutions of all residues in the Lid domain.
[0296] 24. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein at least one amino acid substitution or deletion is in the fingers domain.
[0297] 25. The mutant Cpf1 endonuclease or orthologue thereof according to any one of the preceding items, wherein the substitution or deletion of at least one amino acid is a substitution or deletion of at least one residue in the fingers domain, such as a substitution or deletion of at least 2 residues, such as a substitution or deletion of at least 3 residues, such as a substitution or deletion of at least 4 residues, such as a substitution or deletion of at least 5 residues, such as a substitution or deletion of at least 6 residues, such as a substitution or deletion of at least 7 residues, such as a substitution or deletion of at least 8 residues, such as a substitution or deletion of at least 9 residues, such as a substitution or deletion of at least 10 residues, such as a substitution or deletion of at least 11 residues, such as a substitution or deletion of at least 12 residues, in the fingers domain.
[0298] 26. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein at least one substitution or deletion is a substitution or deletion of all residues in the fingers domain.
[0299] 27. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein at least one amino acid substitution is R918G, K1013G, R1014G, Q1025G or E1028G.
[0300] 28. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the at least two amino acid substitutions are independently selected from R918G, K1013G, R1014G, Q1025G and E1028G.
[0301] 29. A mutant Cpf1 endonuclease or an ortholog thereof according to any one of the preceding items, wherein the Cpf1 endonuclease is a nicking endonuclease.
[0302] 30. The mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein the amino acid substitution is at position K1013.
[0303] 31. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, comprising a substitution at a position corresponding to K1013 of SEQ ID NO:2.
[0304] 32. The mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein said amino acid substitutions are at positions R918 and K1013.
[0305] 33. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, comprising substitutions at positions corresponding to R918 and K1013 of SEQ ID NO:2.
[0306] 34. The mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein said amino acid substitutions are at positions K1013 and R1014.
[0307] 35. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, comprising substitutions at positions corresponding to 1013 and R1014 of SEQ ID NO:2.
[0308] 36. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the amino acid substitution is K1013G.
[0309] 37. A mutant Cpf1 endonuclease according to any one of the preceding items, comprising a substitution corresponding to K1013G in SEQ ID NO:2, or an ortholog thereof.
[0310] 38. A mutant Cpf1 endonuclease or an ortholog thereof according to any one of the preceding items, wherein the Cpf1 endonuclease comprises the amino acid substitutions K1013G and R1014G.
[0311] 39. A mutant Cpf1 endonuclease or an ortholog thereof according to any one of the preceding items, comprising substitutions corresponding to K1013G and R1014G in SEQ ID NO:2.
[0312] 40. The mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein the amino acid substitution is at position Q1025.
[0313] 41. A mutant Cpf1 endonuclease according to any one of the preceding items, comprising a substitution at a position corresponding to Q1025 of SEQ ID NO:2, or an ortholog thereof.
[0314] 42. The mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein the amino acid substitution is at position E1028.
[0315] 43. A mutant Cpf1 endonuclease according to any one of the preceding items, or an ortholog thereof, comprising a substitution at a position corresponding to E1028 of SEQ ID NO:2.
[0316] 44. The mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein said amino acid substitutions are at positions Q1025 and E1028.
[0317] 45. A mutant Cpf1 endonuclease according to any one of the preceding items, or an ortholog thereof, comprising substitutions at positions corresponding to Q1025 and E1028 of SEQ ID NO:2.
[0318] 46. The mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein the amino acid substitution is Q1025G.
[0319] 47. A mutant Cpf1 endonuclease according to any one of the preceding items, comprising a substitution corresponding to Q1025G in SEQ ID NO:2, or an ortholog thereof.
[0320] 48. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the amino acid substitution is E1028G.
[0321] 49. A mutant Cpf1 endonuclease according to any one of the preceding items, comprising a substitution corresponding to E1028G in SEQ ID NO:2, or an ortholog thereof.
[0322] 50. A mutant Cpf1 endonuclease or an orthologue thereof according to any one of the preceding items, wherein the Cpf1 endonuclease comprises the amino acid substitutions Q1025G and E1028G.
[0323] 51. A mutant Cpf1 endonuclease according to any one of the preceding items, or an ortholog thereof, comprising substitutions corresponding to Q1025G and E1028G in SEQ ID NO:2.
[0324] 52. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the Cpf1 endonuclease comprises at least two amino acid substitutions or deletions in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, as well as at least one amino acid substitution at any one of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2.
[0325] 53. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the Cpf1 endonuclease comprises at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, as well as at least one amino acid substitution at any one of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2.
[0326] 54. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the Cpf1 endonuclease comprises at least two amino acid substitutions or deletions in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, as well as at least two amino acid substitutions at any one of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2.
[0327] 55. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the Cpf1 endonuclease comprises at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, as well as at least two amino acid substitutions at any one of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:2.
[0328] 56. A mutant Cpf1 endonuclease or an orthologue thereof according to any one of the preceding items, wherein the Cpf1 endonuclease comprises or consists of a polypeptide having at least 95% sequence identity to SEQ ID NO:3 or SEQ ID NO:4.
[0329] 57. A mutant Cpf1 endonuclease or an orthologue thereof according to any one of the preceding items, wherein the Cpf1 endonuclease comprises or consists of the polypeptide of SEQ ID NO:3 or SEQ ID NO:4.
[0330] 58. A mutant Cpf1 endonuclease or an orthologue thereof according to any one of the preceding items, wherein the REC domain of the mutant Cpf1 endonuclease comprises or consists of SEQ ID NO:14.
[0331] 59. A mutant Cpf1 endonuclease or an ortholog thereof according to any one of the preceding items, wherein the Lid domain of the mutant Cpf1 endonuclease comprises or consists of SEQ ID NO:18.
[0332] 60. A mutant Cpf1 endonuclease or an orthologue thereof according to any one of the preceding items, wherein the fingers domain of the mutant Cpf1 endonuclease comprises or consists of SEQ ID NO:22.
[0333] 61. A mutant Cpf1 endonuclease or ortholog thereof according to any one of the preceding items, wherein the mutant Cpf1 has one or more altered activities compared to wild-type Cpf1, said activities being selected from the group consisting of: double-stranded cleavage of target or non-target nucleic acids; single-stranded cleavage of stranded target or non-target nucleic acids; single-stranded cleavage of nucleic acids in a non-specific manner, optionally enhanced by single-stranded or double-stranded target DNA.
[0334] 62. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the residue at position 917 of SEQ ID NO:2 is glutamic acid (E) or aspartic acid (D).
[0335] 63. The mutant Cpf1 endonuclease or ortholog thereof of any one of the preceding items, wherein the residue at position 1006 of SEQ ID NO:2 is glutamic acid (E) or aspartic acid (D).
[0336] 64. A polynucleotide encoding a mutant Cpf1 endonuclease or an ortholog thereof according to any one of the preceding items.
[0337] 65. The polynucleotide according to item 64, wherein the polynucleotide is codon-optimized for expression in a host cell.
[0338] 66. A recombinant vector comprising a nucleic acid sequence encoding a polynucleotide according to any one of items 1 to 64 or a mutant Cpf1 endonuclease or an ortholog thereof according to any one of items 1 to 64.
[0339] 67. The recombinant vector according to item 66, wherein the polynucleotide or nucleic acid sequence is operably linked to a promoter.
[0340] 68. The recombinant vector according to any one of items 66 to 67, further comprising a nucleic acid sequence encoding a guide RNA (crRNA) operably linked to a promoter, wherein the crRNA binds to the encoded Cpf1 endonuclease and a fragment having sufficient base pairs to hybridize to a target nucleic acid.
[0341] 69. The recombinant vector according to any one of items 66 to 68, wherein said crRNA consists of a constant region of 19 or 20 nucleotides and a variable region of 12 to 24 nucleotides, such that the crRNA is at least 40 nucleotides long, such as 41 nucleotides long, such as 42 nucleotides long, such as 43 nucleotides long, such as 44 nucleotides long.
[0342] 70. The recombinant vector according to any one of items 66 to 69, wherein the constant region of the crRNA is as set out in SEQ ID NO: 5.
[0343] 71. A cell capable of expressing a mutant Cpf1 or an orthologue thereof according to any one of items 1 to 63, a polynucleotide according to item 64 or 65, or a recombinant vector according to any one of items 66 to 70.
[0344] 72. A system for expression of the crRNA-Cpf1 complex, comprising: a. the polynucleotide according to item 64 or 65, or the recombinant vector according to any one of items 66 to 70, comprising a polynucleotide encoding a mutant Cpf1 endonuclease or an ortholog thereof; b. a polynucleotide or recombinant vector comprising a polynucleotide encoding a guide RNA (crRNA), optionally operably linked to a promoter; The system comprising:
[0345] 73.c. Cells for expressing the polynucleotides or recombinant vectors described in a. and b. above. 73. The system according to item 72, further comprising:
[0346] 74. The cell according to item 71 or the system according to items 72 or 73, wherein the cell is a prokaryotic or eukaryotic cell.
[0347] 75. Use of a crRNA-Cpf1 complex for introducing a single-strand break in a first target nucleic acid, comprising: a. contacting a Cpf1 endonuclease or its orthologue with a guide RNA (crRNA), thereby obtaining a crRNA-Cpf1 complex capable of recognizing a second target nucleic acid, the second target nucleic acid comprising a protospacer adjacent motif (PAM); and b. contacting the crRNA-Cpf1 complex with a first target nucleic acid; This results in a single strand break in the first target sequence. The above uses.
[0348] 76. Use according to item 75, wherein the Cpf1 endonuclease or an orthologue thereof is encoded by a polynucleotide or vector according to any one of items 1 to 63 or according to any one of items 64 to 70.
[0349] 77. Use according to any one of items 75 to 76, wherein the crRNA hybridizes to a second target nucleic acid.
[0350] 78. The use according to any one of items 75 to 77, wherein hybridization of the crRNA to the second target nucleic acid activates the crRNA-Cpf1 complex.
[0351] 79. The use according to any one of items 75 to 78, wherein a single-strand break is made in a specific recognition nucleotide sequence of the first target nucleic acid.
[0352] 80. The use according to any one of items 75 to 79, wherein single-strand breaks are made in one or more non-specific nucleotide sequences of the first target nucleic acid.
[0353] 81. The use according to any one of items 75 to 80, wherein the crRNA is as defined in any one of the preceding items.
[0354] 82. Use according to any one of items 75 to 81, wherein the molar ratio of Cpf1 to crRNA is 0.5:3.0 to 1.0:1.0, such as 0.7:2.5, for example 0.8:2.0, for example 0.9:1.75, for example 0.95:1.5, such as 1.0:1.4, for example 1.0:1.3, for example 1.1:1.2.
[0355] 83. Use according to any one of items 75 to 82, wherein the molar ratio of Cpf1:crRNA:second target nucleic acid is 0.5:3.0:5.0, such as 0.7:2.5:4.0, for example 0.8:2.0:3.5, for example 0.9:1.75:3.0, such as 0.95:1.5:2.0, for example 1.0:1.4:1.9, for example 1.0:1.3:1.7.
[0356] 84. Use according to any one of items 75 to 83, wherein the second target nucleic acid comprises or consists of a recognition sequence comprising a sequence of at least 15 consecutive nucleotides, such as at least 16 consecutive nucleotides, such as at least 17 consecutive nucleotides, such as at least 18 consecutive nucleotides, such as at least 19 consecutive nucleotides, such as at least 20 consecutive nucleotides, such as at least 21 consecutive nucleotides, such as at least 22 consecutive nucleotides, such as at least 23 consecutive nucleotides, such as at least 24 consecutive nucleotides, such as at least 25 consecutive nucleotides, such as at least 26 consecutive nucleotides, such as at least 27 consecutive nucleotides, with the proviso that the 5'-terminal three nucleic acids consist of a PAM sequence.
[0357] 85. Use according to any of items 75 to 84, wherein the PAM comprises or consists of the sequence 5'-YRN-3' or 5'-RYN-3'.
[0358] 86. Use according to any of items 75 to 85, wherein the PAM comprises or consists of the sequence 5'-TTN-3'.
[0359] 87. The use according to any one of items 75 to 86, wherein the first target nucleic acid and the second target nucleic acid are DNA or RNA.
[0360] 88. The use according to any one of items 75 to 87, wherein the second target nucleic acid is double-stranded DNA.
[0361] 89. The use according to any one of items 75 to 88, wherein the second target nucleic acid is DNA selected from the group consisting of genomic DNA, chromatin, nucleosomes, plasmid DNA, methylated DNA, synthetic DNA, and DNA fragments.
[0362] 90. The use according to any one of items 75 to 89, wherein the second target nucleic acid is RNA.
[0363] 91. The use according to any one of items 75 to 90, wherein the first target nucleic acid and the second target nucleic acid are identical.
[0364] 92. The use according to any one of items 75 to 91, wherein the first target nucleic acid and the second target nucleic acid are the same molecule, and wherein the Cpf1 endonuclease or an orthologue thereof comprises amino acid substitutions at positions 1013 and 1014 of SEQ ID NO:2.
[0365] 93. The use according to any one of items 75 to 92, wherein the first target nucleic acid and the second target nucleic acid are the same molecule, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO: 3.
[0366] 94. The use according to any one of items 75 to 93, wherein a single-strand break is made at a specific recognition nucleotide sequence of the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises amino acid substitutions at positions 1013 and 1014 of SEQ ID NO:2.
[0367] 95. The use according to any one of items 75 to 94, wherein Cpf1 specifically or non-specifically performs single-strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO:3.
[0368] 96. The use according to any one of items 75 to 95, wherein the first target nucleic acid and the second target nucleic acid are different and the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2.
[0369] 97. The use according to any one of items 75 to 96, wherein the first target nucleic acid and the second target nucleic acid are different and the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO: 4.
[0370] 98. The use according to any one of items 75 to 97, wherein Cpf1 specifically or non-specifically causes single-strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2.
[0371] 99. Use according to any one of items 75 to 98, wherein Cpf1 specifically or non-specifically performs single-strand breaks in a first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO: 4.
[0372] 100. The use according to any one of items 75 to 99, wherein the first target nucleic acid and the second target nucleic acid are different and wherein the Cpf1 endonuclease or its orthologue comprises two or more amino acid substitutions or deletions at any one of positions 324 to 336 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at all of the residues corresponding to positions 324 to 336 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its orthologue comprises amino acid deletions at all of the residues corresponding to positions 324 to 336 of SEQ ID NO:2.
[0373] 101. The use according to any one of items 75 to 100, wherein the first target nucleic acid and the second target nucleic acid are different and the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO: 34 or SEQ ID NO: 36.
[0374] 102. The use according to any one of items 75 to 101, wherein Cpf1 specifically or non-specifically performs single-strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises two or more amino acid substitutions or deletions at any one of positions 324 to 336 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at all of the residues corresponding to positions 324 to 336 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its orthologue comprises amino acid deletions at all of the residues corresponding to positions 324 to 336 of SEQ ID NO:2.
[0375] 103. Use according to any one of items 75 to 102, wherein Cpf1 specifically or non-specifically performs single-strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO: 34 or SEQ ID NO: 36.
[0376] 104. The use according to any one of items 75 to 103, wherein the first target nucleic acid and the second target nucleic acid are different, and wherein the Cpf1 endonuclease or its orthologue comprises two or more amino acid substitutions at any one of positions 1006 to 1018 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at all of the residues corresponding to positions 1006 to 1018 of SEQ ID NO:2.
[0377] 105. The use according to any one of items 75 to 104, wherein the first target nucleic acid and the second target nucleic acid are different, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO: 38.
[0378] 106. The use according to any one of items 75 to 105, wherein Cpf1 specifically or non-specifically performs single-strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises two or more amino acid substitutions at any one of positions 1006 to 1018 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at all of the residues corresponding to positions 1006 to 1018 of SEQ ID NO:2.
[0379] 107. Use according to any one of items 75 to 106, wherein Cpf1 specifically or non-specifically causes single-strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO: 38.
[0380] 108. The use according to any one of items 75 to 107, wherein the first target nucleic acid and the second target nucleic acid are identical or different, and wherein the Cpf1 endonuclease or an orthologue thereof comprises amino acid substitutions at positions corresponding to 1025 and 1028 of SEQ ID NO:2.
[0381] 109. The use according to any one of items 75 to 108, wherein the first target nucleic acid and the second target nucleic acid are identical or different, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO: 32.
[0382] 110. Use according to any one of items 75 to 109, wherein Cpf1 specifically performs single-strand breaks in a first target nucleic acid, wherein the target nucleic acid is in double-stranded form, and wherein the Cpf1 endonuclease or an orthologue thereof comprises two or more amino acid substitutions at positions corresponding to 1025 and 1028 of SEQ ID NO:2.
[0383] 111. Use according to any one of items 75 to 110, wherein Cpf1 specifically performs a single-strand break in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO: 32.
[0384] 112. The use according to any one of items 75 to 111, wherein the method is carried out ex-vivo.
[0385] 113. A method for introducing a single-strand break in a first target nucleic acid, comprising the steps of: a. designing a guide-RNA (crRNA) capable of recognizing a second target nucleic acid containing a protospacer adjacent motif (PAM); b. contacting the crRNA of step a. with Cpf1 endonuclease or an orthologue thereof, thereby obtaining a crRNA-Cpf1 complex capable of binding to said second target nucleic acid; and c. contacting crRNA and Cpf1 with the first target nucleic acid; thereby introducing one or more single-stranded breaks in the first target nucleic acid. The method comprising:
[0386] 114. The method according to item 113, wherein the Cpf1 endonuclease or an orthologue thereof is encoded by a polynucleotide or vector according to any one of items 1 to 63 or any one of items 64 to 70.
[0387] 115. The method according to item 113 or 114, wherein steps b. and c. are performed simultaneously or one after the other.
[0388] 116. The method according to any one of items 113 to 115, which is carried out in a cell in vitro.
[0389] 117. The method according to any one of items 113 to 116, wherein a single-strand break is made in a specific recognition nucleotide sequence of the first target nucleic acid.
[0390] 118. The method according to any one of items 113 to 117, wherein single-strand breaks are made specifically or non-specifically in the first target nucleic acid.
[0391] 119. The method according to any one of items 113 to 118, wherein the first and second target nucleic acids are as defined in any one of the preceding items.
[0392] 120. An in vitro method for introducing a site-specific double-strand break in a second target nucleic acid in a mammalian cell, comprising introducing a crRNA-Cpf1 complex into the mammalian cell, wherein Cpf1 is a mutant Cpf1 endonuclease or an orthologue according to any one of items 1 to 63, and wherein the crRNA is specific for the second target nucleic acid.
[0393] 121. A method for detection of a second target nucleic acid in a sample, comprising: a. providing a crRNA-Cpf1 complex, wherein Cpf1 is the Cpf1 endonuclease of any one of items 1 to 63 or an orthologue thereof, and wherein the crRNA is specific for a second target nucleic acid; b. providing labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one dye and at least one quencher; c. contacting the crRNA-Cpf1 complex and the ssDNA with a sample, wherein the sample contains at least one second target nucleic acid; and d. detecting ssDNA breaks by detecting a fluorescent signal from the fluorophore; Thereby detecting the presence of the second target nucleic acid in the sample. The method comprising:
[0394] 122. The method according to item 121, wherein step c. comprises activation of the crRNA-Cpf1 complex, e.g., activation by single-stranded or double-stranded target DNA.
[0395] 123. e. Determining the level and / or concentration of a second target nucleic acid; wherein the level and / or concentration of the second target nucleic acid correlates with the cleaved ssDNA. 123. The method of claim 121 or 122, further comprising:
[0396] 124. The method according to any one of items 121 to 123, wherein the second target nucleic acid can be detected at a concentration in the sub-nanomolar range, such as in the sub-picomolar range, such as in the sub-femtomolar range, such as in the sub-attomolar range.
[0397] 125. The method according to any one of items 121 to 124, wherein the ssDNA is labeled at least at one base at any position along the strand.
[0398] 126. The method according to any one of items 121 to 125, wherein at least one dye is a fluorophore.
[0399] 127. The method according to any one of items 121 to 126, wherein step d. comprises detecting a fluorescent signal resulting from cleavage of the ssDNA.
[0400] 128. The method according to any one of items 121 to 127, wherein the sample comprises DNA and / or RNA.
[0401] 129. The method of any one of items 121 to 128, wherein the sample is suspected of containing a second target nucleic acid.
[0402] 130. The method according to any one of items 121 to 129, wherein the second target nucleic acid is a nucleic acid fragment of a viral genome, a microbial genome, a gene, or a genome of a pathogen.
[0403] 131. An in vitro method for diagnosing an infectious disease in a subject, comprising: a. providing a crRNA-Cpf1 complex, wherein Cpf1 is the Cpf1 endonuclease of any one of items 1 to 63 or an orthologue thereof, and wherein the crRNA is specific for a second target nucleic acid; b. providing labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one dye and at least one quencher; c. providing a sample from a subject, wherein said sample contains or is suspected of containing a second target nucleic acid; and d. Determining the level and / or concentration of a second target nucleic acid as defined in any one of the preceding items; wherein the second target nucleic acid is a nucleic acid or a fragment thereof of the genome of a disease-causing infectious pathogen; Thereby, diagnosing an infectious disease in a subject. The method comprising:
[0404] 132. The method of item 131, further comprising treating the infectious disease.
[0405] 133. The method according to item 132, further comprising treating the infectious disease by administration of a therapeutically effective compound.
[0406] 134. The method according to any one of items 131 to 133, further comprising a step of comparing the level and / or concentration of the second target nucleic acid with a cut-off value; wherein the cut-off value is determined from a concentration range of the second target nucleic acid in a healthy subject, e.g., a subject not exhibiting an infectious disease; wherein a level and / or concentration greater than the cutoff value indicates the presence of an infectious disease; The method.
[0407] 135. The method according to any one of items 131 to 134, wherein the infectious disease is caused by an infectious pathogen, and wherein the infectious pathogen comprises a virus, a viroid, a prion, a bacterium, a nematode, a parasitic roundworm, a pinworm, an arthropod, a fungus, a ringworm, and a macroparasite.
[0408] 136. The method according to any one of items 131 to 135, wherein the subject is a human.
[0409] 137. The method according to any one of items 131 to 136, wherein the sample body fluid is selected from the group consisting of blood, whole blood, plasma, serum, urine, saliva, tears, cerebrospinal fluid and semen.
[0410] 138. The method according to any one of items 131 to 137, wherein the Cpf1 endonuclease or an ortholog thereof comprises amino acid substitutions at positions 918 and 1013 of SEQ ID NO:2.
[0411] 139. The method according to any one of items 131 to 138, wherein the Cpf1 endonuclease or an ortholog thereof comprises or consists of SEQ ID NO: 4.
[0412] example Example 1. Cleavage assay A double-stranded DNA target was prepared by annealing the following: For the labeled DNA, target strand (t strand) 5'-atgcagtggccttattaaatgacttctcTAA CG-3'-6FAM (SEQ ID NO: 6) and Non-target strand (nt strand) 6FAM-5'-cgTTAgagaagtcatttaataaggccactgcat-3' (SEQ ID NO: 7), and For unlabeled DNA targets, t-strand 5'-atgcagtggccttattaaatgacttctcTAACG-3' (SEQ ID NO: 8) and nt-target strand 5'-cgTTAgagaagtcatttaataaggccactgcat-3' (SEQ ID NO: 9).
[0413] The cleavage assay was performed as follows: Four picomoles of FnCpf1 protein (wild type protein was used as a control) and 5.2 picomoles of crRNA (rArArUrUrUrCrUrArCrUrGrUrUrGrUrArGrArUrCrArCrCrGrArGrUrGrCrArGrCrArCrUrCrGrGrCrGrU) (SEQ ID NO: 10) were mixed in 20 mM bicine-HCl (pH 8), 150 mM KCl, 0.5 mM TCPE (pH 8) and 5 mM MgCl for 30 min at 25° C. Then, 6 picomoles of double-stranded labeled DNA target were added to the mixture, which was incubated for another 30 min at 25° C. The reaction was stopped by adding an equal volume of stop solution (8 M urea and 100 mM EDTA at pH 8) followed by incubation at 95° C. for 5 min. Samples were loaded onto 15% Novex TBE-urea gels (Invitrogen) and gels were visualized using an Odyssey FC Imaging System (Li-Cor).
[0414] For non-specific ssDNA, cleavage assays were performed as follows: 4 pmoles of FnCpf1 protein (wild type protein was used as control) and 5.2 pmoles of crRNA were mixed in 20 mM bicine-HCl (pH 8), 150 mM KCl, 0.5 mM TCPE (pH 8) and 5 mM MgCl for 30 min at 25°C. 6 pmoles of double-stranded unlabeled DNA target was then added to the mixture, which was incubated for another 30 min at 25°C. Labeled non-specific ssDNA ( / 5' 6-FAM / AGCATGCCCAAATTGCTTACATATGTGTTACGACGGT) (SEQ ID NO: 23) was then added and the mixture was incubated for 30 min at 25°C. The reaction was stopped by adding an equal volume of stop solution (8 M urea and 100 mM EDTA at pH 8) followed by incubation at 95°C for 5 min. Samples were loaded onto 15% Novex TBE-urea gels (Invitrogen) and run according to the manufacturer's instructions. Gels were visualized using an Odyssey FC Imaging System (Li-Cor).
[0415] The endonuclease activity was calculated from the intensity of the DNA band (I). The intensity was quantified using ImageStudio. uc The intensity of the NTS (non-target strand), TS (target strand) and SSDNA (non-specific single strand) cleavage was calculated using the following formula: (100*I n ) / I uc , where I n is the intensity of the band of interest. The average of at least three independent experiments is used in the plots and the standard deviation is represented in the error bars.
[0416] result: The DNA duplex contains a PAM sequence (5'-TTA-3') and a target sequence in the t strand that is complementary to a portion of the crRNA (Figure 1). The FnCpf1 wild-type protein can cleave both the t strand and the nt strand (Figures 2-4 and 6). Mutant 1 (SEQ ID NO: 3; K1013-R1014Cpf1) can cleave only the t strand (Figures 2, 4 and 6D). This mutant is useful as a nickase.
[0417] FnCpf1 wild-type protein can cut non-specific ssDNA after activation with double-stranded DNA target (Figures 3, 4). Mutant 2 (SEQ ID NO: 4; R918G-K1013G) cannot cut specific DNA target, neither as ssDNA nor as ddDNA, nor as nt strand, but can cut non-specific ssDNA (Figures 3, 4).
[0418] Cpf1 Q1025G mutant (SEQ ID NO: 30) can cut both target and non-target strands more efficiently than wild type when they are in double-stranded DNA form. The mutant shows increased ability to cut only target strand as single-stranded DNA (Figures 4A and B and 6B). This mutant can cut non-specific ssDNA after activation by double-stranded and single-stranded DNA targets. This mutant is useful for genome editing.
[0419] The Cpf1 Q1025G E1028G mutant (SEQ ID NO: 32) can cut both the target and non-target strands when they are in double-stranded DNA form. The mutant shows a reduced ability to cut only the target strand as single-stranded DNA (Figures 4A, B and 6C). It is unlikely that this mutant can cut non-specific single-stranded DNA after activation with double-stranded DNA target. This mutant does not have non-specific ssDNA activity that may be dangerous to cells, so it can be used for genome editing in a safer way. The presence of non-specific nucleases that induce breaks in ssDNA can cut DNA during duplication, replication and transcription, thus generating random breaks in the genome of the cell, thus inducing undesirable mutations, genome damage and cytotoxicity.
[0420] The Cpf1 mutant carrying the deletion of the entire REC domain (SEQ ID NO: 34) is unable to cut either the target or non-target strand. The mutant shows a reduced ability to cut only the target strand as single-stranded DNA (Figure 4A, B). This mutant is able to cut non-specific ssDNA after activation by a double-stranded DNA target. Since it does not cut the activator, this protein will retain non-specific ssDNA activity longer compared to the wild type.
[0421] The Cpf1 mutant (SEQ ID NO: 36) carrying a substituted REC domain is unable to cut the target or non-target strands when they are in double-stranded DNA form. The mutant shows a reduced ability to cut only the target strand as single-stranded DNA. The mutant can cut non-specific ssDNA after activation by double-stranded and single-stranded DNA targets. Since it does not cut the activator, the protein will retain non-specific ssDNA activity longer compared to the wild type.
[0422] The Cpf1 mutant carrying the Lid domain substitution (SEQ ID NO: 38) is unable to cut the target or non-target strands when they are in double-stranded DNA form. The mutant is unable to cut only the target strand as single-stranded DNA (Fig. 4A, B). The mutant is able to cut non-specific ssDNA, albeit with low activity. The entire Lid region is thought to be important for the activity of the protein, and mutations in this region would induce changes in the activity of the protein. Interestingly, the activity of this mutant carrying the Lid domain substitution is preserved despite the substitution of E1006. The inventors surprisingly found that when the entire Lid domain including E1006 is replaced, the activity is preserved if the residue corresponding to position 917 of wild-type FnCpf1 is not replaced or is only replaced from aspartic acid (D) to glutamic acid (E).
[0423] A Cpf1 mutant carrying a deletion of the entire Lid domain (SEQ ID NO: 40) is not soluble (FIG. 5). A Cpf1 mutant carrying a deletion of the entire fingers domain (SEQ ID NO: 42) is not soluble (Figure 5).
[0424] References [Table 1] JPEG2025075025000003.jpg62163
[0425] Overview of Arrays SEQ ID NO:1 FnCpf1 DNA sequence SEQ ID NO:2 FnCpf1 amino acid sequence SEQ ID NO:3 Double mutant-nickase, Cpf1 K1013G-R1014G SEQ ID NO: 4 Double mutant-non-specific endonuclease, Cpf1 R918G-K1013G SEQ ID NO:5 crRNA constant region SEQ ID NO:6 t-strand SEQ ID NO: 7 nt strand SEQ ID NO:8 t-strand SEQ ID NO: 9 nt strand SEQ ID NO:10 crRNA SEQ ID NO: 11 REC domain of FnCpf1, DNA sequence SEQ ID NO: 12. REC domain of FnCpf1, amino acid sequence SEQ ID NO: 13 Mutant REC domain of FnCpf1, DNA sequence SEQ ID NO: 14. Mutant REC domain of FnCpf1, amino acid sequence SEQ ID NO: 15 Lid domain of FnCpf1, DNA sequence SEQ ID NO: 16. Lid domain of FnCpf1, amino acid sequence SEQ ID NO: 17 Mutant Lid domain of FnCpf1, DNA sequence SEQ ID NO: 18. Mutant Lid domain of FnCpf1, amino acid sequence SEQ ID NO: 19 Finger domain of FnCpf1, DNA sequence SEQ ID NO: 20 Finger domain of FnCpf1, amino acid sequence SEQ ID NO: 21 Mutant finger domain of FnCpf1, DNA sequence SEQ ID NO: 22 Mutant finger domain of FnCpf1, amino acid sequence SEQ ID NO: 23 Labeled non-specific ssDNA SEQ ID NO: 24 Cpf1 of Acidaminococcus sp. BV3L6 SEQ ID NO: 25 Cpf1 of bacterium COE1 of the Lachnospiraceae family SEQ ID NO:26 Cpf1 of Succicrasticum ruminis SEQ ID NO: 27 Cpf1 from Butyrivibrio fibrisolvens SEQ ID NO:28 Non-specific ssDNA SEQ ID NO: 29 FnCpf1 Q1025G mutant SEQ ID NO: 30 FnCpf1 Q1025G mutant SEQ ID NO: 31 FnCpf1 Q1025G E1028G mutant SEQ ID NO: 32 FnCpf1 Q1025G E1028G mutant SEQ ID NO: 33 Rec deletion mutant FnCpf1 SEQ ID NO: 34 Rec deletion mutant FnCpf1 SEQ ID NO: 35 Rec-substituted FnCpf1 mutant SEQ ID NO: 36 Rec-substituted FnCpf1 mutant SEQ ID NO: 37 LID-substituted FnCpf1 mutant SEQ ID NO: 38 LID-substituted FnCpf1 mutant SEQ ID NO: 39 LID deleted FnCpf1 mutant SEQ ID NO: 40 LID deleted FnCpf1 mutant SEQ ID NO: 41 Finger-deleted FnCpf1 mutant SEQ ID NO: 42 Finger-deleted FnCpf1 mutant SEQ ID NO: 43 Cpf1 of Candidatus Methanoplasma termitum SEQ ID NO: 44 Uncultured Clostridium sp. Cpf1
Claims
1. below: i) a polypeptide sequence having at least 95% sequence identity to sequences corresponding to residues 1-323, 337-1005, and 1019-1329 of SEQ ID NO:2; wherein the polypeptide sequence is a. at least two amino acid mutations in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, where each mutation is independently an amino acid substitution or deletion; and / or b. at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, where at least one of the residues at positions 917 and 1006 is a glutamic acid (E) or an aspartic acid (D); Further comprising: and / or ii) a polypeptide sequence having at least 95% sequence identity to SEQ ID NO:2; wherein said polypeptide sequence comprises at least two amino acid substitutions at two positions independently selected from positions corresponding to residues 918, 1013, 1014, 1025, and 1028 of SEQ ID NO:2; A mutant Cpf1 endonuclease or its ortholog, comprising:
2. The Cpf1 endonuclease or its orthologue according to claim 1 , wherein the Cpf1 endonuclease is derived from Francisella novicida.
3. 3. The mutant Cpf1 endonuclease or orthologue thereof of claim 1 or 2, wherein the mutant Cpf1 comprises a polypeptide sequence having at least 95% sequence identity to a sequence corresponding to residues 1-323, 337-1005, and 1019-1329 of SEQ ID NO:2, wherein the polypeptide sequence is a. at least two amino acid substitutions or deletions in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2; and / or b. at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, except that at least one of the residues at positions 917 and 1006 of SEQ ID NO:2 is a glutamic acid (E) or aspartic acid (D); Further comprising: wherein the polypeptide sequence further comprises at least one amino acid substitution at position 918, 1013, 1014, 1025 or 1028 of SEQ ID NO:
2. The mutant Cpf1 endonuclease or its ortholog.
4. 4. The Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 3, wherein one of the at least two amino acid substitutions is a substitution of an amino acid having a charged side chain for an amino acid having an uncharged side chain.
5. 5. The mutant Cpf1 endonuclease or its orthologue according to any one of claims 1 to 4, wherein one of the at least two amino acid substitutions is a substitution of an amino acid residue having a charged side chain for an amino acid residue having a non-polar side chain.
6. 6. The mutant Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 5, wherein one of the at least two amino acid substitutions is a substitution of an amino acid having a charged side chain with glycine, alanine, valine, leucine, isoleucine, serine or threonine.
7. 7. The mutant Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 6, wherein one of the at least two amino acid substitutions is a substitution of an amino acid having a charged side chain with glycine.
8. The mutant Cpf1 endonuclease or its orthologue according to any one of claims 1 to 7, wherein at least two amino acid substitutions or deletions are in the REC domain.
9. 9. The mutant Cpf1 endonuclease or orthologue thereof according to any one of claims 1 to 8, wherein the substitution or deletion of at least 2 amino acids is a substitution or deletion of at least 2 residues in the REC domain, such as a substitution or deletion of at least 3 residues, such as a substitution or deletion of at least 4 residues, such as a substitution or deletion of at least 5 residues, such as a substitution or deletion of at least 6 residues, such as a substitution or deletion of at least 7 residues, such as a substitution or deletion of at least 8 residues, such as a substitution or deletion of at least 9 residues, such as a substitution or deletion of at least 10 residues, such as a substitution or deletion of at least 11 residues, such as a substitution or deletion of at least 12 residues, such as a substitution or deletion of at least 13 residues, in the REC domain.
10. The mutant Cpf1 endonuclease or its orthologue according to any one of claims 1 to 9, wherein at least two of the substitutions or deletions are substitutions or deletions of all residues in the REC domain.
11. The mutant Cpf1 endonuclease or its orthologue according to any one of claims 1 to 10, wherein at least two amino acid substitutions are in the Lid domain.
12. 12. The mutant Cpf1 endonuclease or orthologue thereof according to any one of claims 1 to 11, wherein the substitution of at least two amino acids is a substitution of at least two residues in the Lid domain, such as a substitution or deletion of at least three residues, such as a substitution or deletion of at least four residues, such as a substitution or deletion of at least five residues, such as a substitution or deletion of at least six residues, such as a substitution or deletion of at least seven residues, such as a substitution or deletion of at least eight residues, such as a substitution or deletion of at least nine residues, such as a substitution or deletion of at least ten residues, such as a substitution or deletion of at least 11 residues, such as a substitution or deletion of at least 12 residues, such as a substitution or deletion of at least 13 residues, in the Lid domain.
13. The mutant Cpf1 endonuclease or its orthologue according to any one of claims 1 to 12, wherein at least two of the substitutions are substitutions of all residues in the Lid domain.
14. 14. The Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 13, wherein the at least two amino acid substitutions are independently selected from R918G, K1013G, R1014G, Q1025G and E1028G.
15. The mutant Cpf1 endonuclease or its orthologue according to any one of claims 1 to 14, wherein the Cpf1 endonuclease is a nicking endonuclease.
16. A mutant Cpf1 endonuclease or an orthologue thereof according to any one of claims 1 to 15, comprising a substitution at the position corresponding to K1013 of SEQ ID NO:
2.
17. 17. A mutant Cpf1 endonuclease according to any one of claims 1 to 16, or an orthologue thereof, comprising substitutions at positions corresponding to R918 and K1013 of SEQ ID NO:
2.
18. 18. A mutant Cpf1 endonuclease or orthologue thereof according to any one of claims 1 to 17, comprising substitutions at positions corresponding to 1013 and R1014 of SEQ ID NO:
2.
19. 19. A mutant Cpf1 endonuclease according to any one of claims 1 to 18, comprising a substitution corresponding to K1013G in SEQ ID NO:2, or an orthologue thereof.
20. 20. A mutant Cpf1 endonuclease according to any one of claims 1 to 19, or an orthologue thereof, comprising substitutions corresponding to K1013G and R1014G in SEQ ID NO:
2.
21. 21. A mutant Cpf1 endonuclease according to any one of claims 1 to 20, comprising a substitution at the position corresponding to Q1025 of SEQ ID NO:2, or an orthologue thereof.
22. 22. A mutant Cpf1 endonuclease or orthologue thereof according to any one of claims 1 to 21, comprising a substitution at the position corresponding to E1028 of SEQ ID NO:
2.
23. 23. A mutant Cpf1 endonuclease or orthologue thereof according to any one of claims 1 to 22, comprising substitutions at positions corresponding to Q1025 and E1028 of SEQ ID NO:
2.
24. 24. A mutant Cpf1 endonuclease according to any one of claims 1 to 23, comprising a substitution corresponding to Q1025G in SEQ ID NO:2, or an orthologue thereof.
25. 25. A mutant Cpf1 endonuclease according to any one of claims 1 to 24, comprising a substitution corresponding to E1028G of SEQ ID NO:2, or an orthologue thereof.
26. 26. A mutant Cpf1 endonuclease according to any one of claims 1 to 25, or an orthologue thereof, comprising substitutions corresponding to Q1025G and E1028G in SEQ ID NO:
2.
27. 27. A mutant Cpf1 endonuclease or orthologue thereof according to any one of claims 1 to 26, wherein the Cpf1 endonuclease comprises at least two amino acid substitutions or deletions in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, and at least one amino acid substitution at any one of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:
2.
28. 28. The mutant Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 27, wherein the Cpf1 endonuclease comprises at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, and at least one amino acid substitution at any one of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:
2.
29. 29. A mutant Cpf1 endonuclease or orthologue thereof according to any one of claims 1 to 28, wherein the Cpf1 endonuclease comprises at least two amino acid substitutions or deletions in the REC domain (residues 324-336; SEQ ID NO:16) compared to SEQ ID NO:2, as well as at least two amino acid substitutions at any one of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:
2.
30. 30. The mutant Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 29, wherein the Cpf1 endonuclease comprises at least two amino acid substitutions in the Lid domain (residues 1006-1018; SEQ ID NO:20) compared to SEQ ID NO:2, as well as at least two amino acid substitutions at any one of positions 918, 1013, 1014, 1025 and 1028 of SEQ ID NO:
2.
31. 31. The Cpf1 endonuclease or its orthologue according to any one of claims 1 to 30, comprising the mutations R918G K1013G or Q1025G E1028G or K1013G R1014G.
32. 32. The mutant Cpf1 endonuclease of any one of claims 1 to 31, wherein the Cpf1 endonuclease comprises or consists of a polypeptide having at least 95% sequence identity to SEQ ID NO:3 or SEQ ID NO:4, or an orthologue thereof.
33. 33. The mutant Cpf1 endonuclease of any one of claims 1 to 32, wherein the Cpf1 endonuclease comprises or consists of the polypeptide of SEQ ID NO:3 or SEQ ID NO:4, or an orthologue thereof.
34. 34. The mutant Cpf1 endonuclease or an orthologue thereof according to any one of claims 1 to 33, wherein the REC domain of the mutant Cpf1 endonuclease comprises or consists of SEQ ID NO:
14.
35. 35. The mutant Cpf1 endonuclease or its orthologue according to any one of claims 1 to 34, wherein the Lid domain of the mutant Cpf1 endonuclease comprises or consists of SEQ ID NO:
18.
36. 36. The mutant Cpf1 endonuclease or its orthologue according to any one of claims 1 to 35, wherein the fingers domain of the mutant Cpf1 endonuclease comprises or consists of SEQ ID NO:
22.
37. 37. The mutant Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 36, wherein the mutant Cpf1 has one or more altered activities compared to wild-type Cpf1, said activities being selected from the group consisting of: double-stranded cleavage of target or non-target nucleic acids; single-stranded cleavage of stranded target or non-target nucleic acids; single-stranded cleavage of nucleic acids in a non-specific manner, optionally enhanced by single-stranded or double-stranded target DNA.
38. 38. The mutant Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 37, wherein the residue at position 917 of SEQ ID NO:2 is glutamic acid (E) or aspartic acid (D).
39. 39. The mutant Cpf1 endonuclease or orthologue thereof of any one of claims 1 to 38, wherein the residue at position 1006 of SEQ ID NO:2 is glutamic acid (E) or aspartic acid (D).
40. A polynucleotide encoding a mutant Cpf1 endonuclease or an orthologue thereof according to any one of claims 1 to 39.
41. 41. The polynucleotide of claim 40, which is codon optimized for expression in a host cell.
42. A recombinant vector comprising a nucleic acid sequence encoding the polynucleotide of claim 40 or 41, or the mutant Cpf1 endonuclease or its orthologue of any one of claims 1 to 39.
43. 43. The recombinant vector of claim 42, wherein the polynucleotide or nucleic acid sequence is operably linked to a promoter.
44. 44. The recombinant vector of any one of claims 42-43, further comprising a nucleic acid sequence encoding a guide RNA (crRNA) operably linked to a promoter, wherein the crRNA is linked to an encoded Cpf1 endonuclease and a fragment having sufficient base pairs to hybridize to a target nucleic acid.
45. 45. The recombinant vector of any one of claims 42 to 44, wherein the crRNA consists of a constant region of 19 or 20 nucleotides and a variable region of 12 to 24 nucleotides, such that the crRNA is at least 40 nucleotides in length, such as 41 nucleotides in length, such as 42 nucleotides in length, such as 43 nucleotides in length, such as 44 nucleotides in length.
46. The recombinant vector of any one of claims 42 to 45, wherein the constant region of the crRNA is as set forth in SEQ ID NO:
5.
47. A cell capable of expressing a mutant Cpf1 or an orthologue thereof according to any one of claims 1 to 39, a polynucleotide according to claim 40 or 41, or a recombinant vector according to any one of claims 42 to 46.
48. A system for expression of the crRNA-Cpf1 complex, comprising: a. the polynucleotide of claim 40 or 41, or the recombinant vector of any one of claims 42 to 46, comprising a polynucleotide encoding a mutant Cpf1 endonuclease or an ortholog thereof; b. A polynucleotide or recombinant vector comprising a polynucleotide encoding a guide RNA (crRNA), optionally operably linked to a promoter. The system comprising:
49. c. A cell for expressing the polynucleotide or recombinant vector of a. and b. above.
49. The system of claim 48, further comprising:
50. 50. The cell of claim 47, or the system of claim 48 or 49, wherein the cell is a prokaryotic or eukaryotic cell.
51. 1. Use of a crRNA-Cpf1 complex to introduce a single-stranded break in a first target nucleic acid, comprising: a. contacting a Cpf1 endonuclease or an orthologue thereof with a guide RNA (crRNA), thereby obtaining a crRNA-Cpf1 complex capable of recognizing a second target nucleic acid, wherein the second target nucleic acid comprises a protospacer adjacent motif (PAM), and wherein the Cpf1 endonuclease or an orthologue thereof is according to any one of claims 1 to 5; b. contacting the crRNA-Cpf1 complex with a first target nucleic acid; This results in a single strand break in the first target sequence. The above uses.
52. 52. The use according to claim 51, wherein the Cpf1 endonuclease or an orthologue thereof is defined in any one of claims 1 to 39 or encoded by a polynucleotide or vector according to any one of claims 40 to 46.
53. The use of claim 52, wherein the crRNA hybridizes to a second target nucleic acid.
54. The use according to any one of claims 51 to 53, wherein hybridization of the crRNA to the second target nucleic acid activates the crRNA-Cpf1 complex.
55. The use according to any one of claims 51 to 54, wherein a single-strand break is made in a specific recognition nucleotide sequence of the first target nucleic acid.
56. The use according to any one of claims 51 to 54, wherein single-strand breaks are made in one or more non-specific nucleotide sequences of the first target nucleic acid.
57. 57. The use according to any one of claims 51 to 56, wherein the crRNA is as defined in any one of claims 44 to 56.
58. 58. Use according to any one of claims 51 to 57, wherein the molar ratio of Cpf1 to crRNA is from 0.5:3.0 to 1.0:1.0, such as 0.7:2.5, for example 0.8:2.0, such as 0.9:1.75, for example 0.95:1.5, such as 1.0:1.4, for example 1.0:1.3, such as 1.1:1.
2.
59. 59. Use according to any one of claims 51 to 58, wherein the molar ratio of Cpf1:crRNA:second target nucleic acid is 0.5:3.0:5.0, such as 0.7:2.5:4.0, for example 0.8:2.0:3.5, such as 0.9:1.75:3.0, for example 0.95:1.5:2.0, such as 1.0:1.4:1.9, for example 1.0:1.3:1.
7.
60. 60. The use according to any one of claims 51 to 59, wherein the second target nucleic acid comprises or consists of a recognition sequence comprising a sequence of at least 15 consecutive nucleotides, such as at least 16 consecutive nucleotides, such as at least 17 consecutive nucleotides, such as at least 18 consecutive nucleotides, such as at least 19 consecutive nucleotides, such as at least 20 consecutive nucleotides, such as at least 21 consecutive nucleotides, such as at least 22 consecutive nucleotides, such as at least 23 consecutive nucleotides, such as at least 24 consecutive nucleotides, such as at least 25 consecutive nucleotides, such as at least 26 consecutive nucleotides, such as at least 27 consecutive nucleotides, with the proviso that the 5'-terminal three nucleic acids consist of a PAM sequence.
61. Use according to any one of claims 51 to 60, wherein the PAM comprises or consists of the sequence 5'-YRN-3' or 5'-RYN-3'.
62. Use according to any one of claims 51 to 61, wherein the PAM comprises or consists of the sequence 5'-TTN-3'.
63. The use according to any one of claims 51 to 62, wherein the first target nucleic acid and the second target nucleic acid are DNA or RNA.
64. The use according to any one of claims 51 to 63, wherein the second target nucleic acid is double-stranded DNA.
65. The use according to any one of claims 51 to 64, wherein the second target nucleic acid is DNA selected from the group consisting of genomic DNA, chromatin, nucleosomes, plasmid DNA, methylated DNA, synthetic DNA, and DNA fragments.
66. The use according to any one of claims 51 to 65, wherein the first target nucleic acid and the second target nucleic acid are identical.
67. 67. The use according to any one of claims 51 to 66, wherein the first target nucleic acid and the second target nucleic acid are the same molecule and the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions 1013 and 1014 of SEQ ID NO:
2.
68. 68. The use according to any one of claims 51 to 67, wherein the first target nucleic acid and the second target nucleic acid are the same molecule and the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:
3.
69. 69. The use according to any one of claims 51 to 68, wherein a single strand break is made at a specific recognition nucleotide sequence of the first target nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions 1013 and 1014 of SEQ ID NO:
2.
70. 70. The use according to any one of claims 51 to 69, wherein Cpf1 specifically or non-specifically effects single strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO:
3.
71. 71. The use according to any one of claims 51 to 70, wherein the first and second target nucleic acids are different and the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions 918 and 1013 of SEQ ID NO:
2.
72. 72. The use according to any one of claims 51 to 71, wherein the first target nucleic acid and the second target nucleic acid are different and the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:
4.
73. 73. The use according to any one of claims 51 to 72, wherein Cpf1 specifically or non-specifically effects single strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions 918 and 1013 of SEQ ID NO:
2.
74. 74. The use according to any one of claims 51 to 73, wherein Cpf1 specifically or non-specifically effects single strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO:
4.
75. 75. The use according to any one of claims 51 to 74, wherein the first target nucleic acid and the second target nucleic acid are different, and wherein the Cpf1 endonuclease or its ortholog comprises two or more amino acid substitutions or deletions at any one of positions 324 to 336 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its ortholog comprises amino acid substitutions at all of the residues corresponding to positions 324 to 336 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its ortholog comprises amino acid deletions at all of the residues corresponding to positions 324 to 336 of SEQ ID NO:
2.
76. 76. The use according to any one of claims 51 to 75, wherein the first target nucleic acid and the second target nucleic acid are different and the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:34 or SEQ ID NO:
36.
77. 77. The use according to any one of claims 51 to 76, wherein Cpf1 specifically or non-specifically effects single-strand breaks in a first target nucleic acid, and wherein the Cpf1 endonuclease, or its orthologue, comprises two or more amino acid substitutions or deletions at any one of positions 324 to 336 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease, or its orthologue, comprises amino acid substitutions at all of the residues corresponding to positions 324 to 336 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease, or its orthologue, comprises amino acid deletions at all of the residues corresponding to positions 324 to 336 of SEQ ID NO:
2.
78. 78. The use according to any one of claims 51 to 77, wherein Cpf1 specifically or non-specifically effects single strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:34 or SEQ ID NO:
36.
79. 79. The use according to any one of claims 51 to 78, wherein the first target nucleic acid and the second target nucleic acid are different, and wherein the Cpf1 endonuclease or its ortholog comprises two or more amino acid substitutions at any one of positions 1006 to 1018 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its ortholog comprises amino acid substitutions at all of the residues corresponding to positions 1006 to 1018 of SEQ ID NO:
2.
80. 80. The use according to any one of claims 51 to 79, wherein the first target nucleic acid and the second target nucleic acid are different and the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:
38.
81. 81. The use according to any one of claims 51 to 80, wherein Cpf1 specifically or non-specifically effects single-strand breaks in a first target nucleic acid, and wherein the Cpf1 endonuclease or its orthologue comprises two or more amino acid substitutions at any one of positions 1006 to 1018 of SEQ ID NO:2, such as wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at all of the residues corresponding to positions 1006 to 1018 of SEQ ID NO:
2.
82. 82. The use according to any one of claims 51 to 81, wherein Cpf1 specifically or non-specifically effects single strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO:
38.
83. 83. The use according to any one of claims 51 to 82, wherein the first target nucleic acid and the second target nucleic acid are identical or different and the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions corresponding to 1025 and 1028 of SEQ ID NO:
2.
84. 84. The use according to any one of claims 51 to 83, wherein the first target nucleic acid and the second target nucleic acid are identical or different and the Cpf1 endonuclease or its orthologue comprises or consists of SEQ ID NO:
32.
85. 85. The use according to any one of claims 51 to 84, wherein Cpf1 specifically effects single-strand breaks in a first target nucleic acid, wherein the target nucleic acid is in double-stranded form, and wherein the Cpf1 endonuclease or an orthologue thereof comprises two or more amino acid substitutions at positions corresponding to 1025 and 1028 of SEQ ID NO:
2.
86. 86. The use according to any one of claims 51 to 85, wherein Cpf1 specifically effects single strand breaks in the first target nucleic acid, and wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO:
32.
87. The use according to any one of claims 51 to 86, which is carried out ex-vivo.
88. 1. A method for introducing a single strand break in a first target nucleic acid, comprising the steps of: a. designing a guide-RNA (crRNA) capable of recognizing a second target nucleic acid containing a protospacer adjacent motif (PAM); b. contacting the crRNA of step a. with a Cpf1 endonuclease or an orthologue thereof, thereby obtaining a crRNA-Cpf1 complex capable of binding to said second target nucleic acid; and c. contacting crRNA and Cpf1 with the first target nucleic acid; Including, thereby introducing one or more single-stranded breaks in the first target nucleic acid. The method.
89. 89. The method of claim 88, wherein the Cpf1 endonuclease or an orthologue thereof is described in any one of claims 1 to 39 or encoded by a polynucleotide or vector according to any one of claims 40 to 46.
90. The method according to any one of claims 88 to 89, wherein steps b. and c. can be performed simultaneously or one after the other.
91. 91. The method of any one of claims 88 to 90, which is carried out in vitro in a cell.
92. 92. The method of any one of claims 88 to 91, wherein a single-stranded break is made at a specific recognition nucleotide sequence of the first target nucleic acid.
93. 93. The method of any one of claims 88 to 92, wherein single-strand breaks are made specifically or non-specifically in the first target nucleic acid.
94. The method according to any one of claims 88 to 93, wherein the first and second target nucleic acids are as defined in any one of claims 1 to 93.
95. 40. An in vitro method for introducing a site-specific double-stranded break in a second target nucleic acid in a mammalian cell, comprising introducing a crRNA-Cpf1 complex into the mammalian cell, wherein Cpf1 is a mutant Cpf1 endonuclease or ortholog of any one of claims 1 to 39, and wherein the crRNA is specific for the second target nucleic acid.
96. 1. A method for detection of a second target nucleic acid in a sample, comprising: a. providing a crRNA-Cpf1 complex, wherein Cpf1 is a Cpf1 endonuclease according to any one of claims 1 to 39 or an orthologue thereof, and wherein the crRNA is specific for a second target nucleic acid; b. providing a labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one dye and at least one quencher; c. contacting the crRNA-Cpf1 complex and the ssDNA with a sample, where the sample contains at least one second target nucleic acid; and d. detecting ssDNA breaks by detecting a fluorescent signal from the fluorophore; Includes thereby detecting the presence of a second target nucleic acid in the sample; wherein step c. optionally comprises activation of the crRNA-Cpf1 complex; The method.
97. 97. The method of claim 96, wherein step c. comprises activation of the crRNA-Cpf1 complex, e.g., activation by single-stranded or double-stranded target DNA.
98. e. Determining the level and / or concentration of a second target nucleic acid Further comprising: wherein the level and / or concentration of the second target nucleic acid correlates with the cleaved ssDNA.
98. The method according to any one of claims 96 to 97.
99. 99. The method of any one of claims 96 to 98, wherein the second target nucleic acid is detectable at a concentration in the sub-nanomolar range, such as in the sub-picomolar range, such as in the sub-femtomolar range, such as in the sub-attomolar range.
100. 100. The method of any one of claims 96 to 99, wherein the ssDNA is labeled at least one base at any position along the strand.
101. 101. The method of any one of claims 96 to 100, wherein at least one dye is a fluorophore.
102. The method of any one of claims 96 to 101, wherein step d. comprises detecting a fluorescent signal resulting from cleavage of the ssDNA.
103. The method of any one of claims 96 to 102, wherein the sample comprises DNA and / or RNA.
104. The method of any one of claims 96 to 103, wherein the sample is suspected of containing a second target nucleic acid.
105. 105. The method of any one of claims 96 to 104, wherein the second target nucleic acid is a nucleic acid fragment of a viral genome, a microbial genome, a gene, or a genome of a pathogen.
106. 1. An in vitro method for diagnosing an infectious disease in a subject, comprising: a. providing a crRNA-Cpf1 complex according to any one of claims 1 to 39, wherein Cpf1 is a Cpf1 endonuclease or an orthologue thereof, and wherein the crRNA is specific for a second target nucleic acid; b. providing a labeled ssDNA, where the ssDNA is labeled with at least one set of interactive labels comprising at least one dye and at least one quencher; c. providing a sample from a subject, wherein said sample contains or is suspected of containing a second target nucleic acid; and d. Determining the level and / or concentration of a second target nucleic acid as defined in any one of claims 51 to 105; wherein the second target nucleic acid is a nucleic acid or a fragment thereof of the genome of a disease-causing infectious pathogen; thereby diagnosing an infectious disease in a subject.
107. 107. The method of claim 106, further comprising treating the infectious disease.
108. 108. The method of claim 107, further comprising treating the infectious disease by administration of a therapeutically effective compound.
109. 109. The method of any one of claims 106 to 108, further comprising the step of comparing the level and / or concentration of the second target nucleic acid to a cut-off value, wherein the cut-off value is determined from a concentration range of the second target nucleic acid in a healthy subject, e.g., a subject not exhibiting an infectious disease; wherein a level and / or concentration higher than the cutoff value indicates the presence of an infectious disease. The method.
110. 110. The method of any one of claims 106 to 109, wherein the infectious disease is caused by an infectious pathogen, infectious pathogens include viruses, viroids, prions, bacteria, nematodes, parasitic roundworms, pinworms, arthropods, fungi, ringworm and macroparasites.
111. The method of any one of claims 106 to 110, wherein the subject is a human.
112. The method of any one of claims 106 to 111, wherein the sample body fluid is selected from the group consisting of blood, whole blood, plasma, serum, urine, saliva, tears, cerebrospinal fluid and semen.
113. 13. The method of any one of claims 106 to 112, wherein the Cpf1 endonuclease or its orthologue comprises amino acid substitutions at positions 918 and 1013 of SEQ ID NO:
2.
114. The method of any one of claims 106 to 113, wherein the Cpf1 endonuclease or an orthologue thereof comprises or consists of SEQ ID NO:4.