Novel, non-naturally occurring CRISPR-CAS nucleases for genome editing

Novel CRISPR-Cas endonucleases BEC85, BEC67, and BEC10, engineered for improved specificity and efficiency, address the limitations of existing systems by enhancing genome editing capabilities and temperature stability, offering a broader range of applications.

JP7763237B2Active Publication Date: 2025-10-31BRAIN BIOTECH AG
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023504539
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-10
Filing Date
2021-07-20
Publication Date
2025-10-31
Estimated Expiration
2041-07-20

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems face limitations in applicability across various genetic backgrounds and can trigger immune responses in organisms like humans, necessitating the identification of novel RNA-guided DNA endonucleases with improved specificity and efficiency for genome editing.

Method used

Development of novel CRISPR-Cas endonucleases, BEC85, BEC67, and BEC10, engineered through protein engineering, which do not require trans-activating crRNA and exhibit distinct molecular mechanisms, enabling efficient homology-directed recombination and broad temperature stability for genome editing.

Benefits of technology

The novel CRISPR-Cas endonucleases demonstrate superior genome editing efficiency and temperature stability, expanding the applicability of CRISPR technology for diverse organisms and applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007763237000005
    Figure 0007763237000005
  • Figure 0007763237000006
    Figure 0007763237000006
  • Figure 0007763237000007
    Figure 0007763237000007
Patent Text Reader

Abstract

The present invention relates to a nucleic acid molecule encoding an RNA-guided DNA endonuclease, wherein the nucleic acid molecule comprises or consists of the amino acid sequence of SEQ ID NO: 29, 1 or 3; (b) a nucleic acid molecule comprises or consists of the nucleotide sequence of SEQ ID NO: 30, 2 or 4; (c) a nucleic acid molecule encoding an RNA-guided DNA endonuclease, the amino acid sequence of which is at least 90%, preferably at least 92%, and most preferably at least 95% identical to the amino acid sequence (a); (d) a nucleic acid molecule comprising or consisting of a nucleotide sequence that is at least 90%, preferably at least 92%, and most preferably at least 95% identical to the nucleotide sequence (b); (e) a nucleic acid molecule degenerate with respect to the nucleic acid molecule (d); or (f) a nucleic acid molecule corresponding to any one of the nucleic acid molecules (a) to (d), in which T is replaced by U.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a nucleic acid molecule encoding an RNA-guided DNA endonuclease, wherein the nucleic acid molecule comprises or consists of the amino acid sequence of SEQ ID NO: 29, 1 or 3; (b) a nucleic acid molecule comprises or consists of the nucleotide sequence of SEQ ID NO: 30, 2 or 4; (c) a nucleic acid molecule encoding an RNA-guided DNA endonuclease, the amino acid sequence of which is at least 90%, preferably at least 92%, and most preferably at least 95% identical to the amino acid sequence (a); (d) a nucleic acid molecule comprising or consisting of a nucleotide sequence that is at least 90%, preferably at least 92%, and most preferably at least 95% identical to the nucleotide sequence (b); (e) a nucleic acid molecule that is degenerate with respect to the nucleic acid molecule (d); or (f) a nucleic acid molecule corresponding to any one of the nucleic acid molecules (a) to (d), in which T is replaced by U.

[0002] Several documents are cited herein, including patent applications and manufacturer's manuals. The disclosures of these documents, although not considered relevant to the patentability of this invention, are incorporated herein by reference in their entirety. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference. [Background technology]

[0003] The CRISPR-Cas system is a widespread adaptive immune system of prokaryotes against the invasion of foreign nucleic acids. To date, over 30 different CRISPR-Cas systems have been identified, differing in the structure of the locus, the number, and the identity of the genes encoding the Cas (CRISPR-associated) proteins.

[0004] A typical feature of CRISPR systems in prokaryotic genomes is the presence of short (30-45 bp) repeat sequences (repeats) interspersed with variable sequences of similar length (spacers). Cas proteins are located either upstream or downstream of the repeat-spacer cluster. Based on differences in the genetic composition and mechanism of these Cas proteins, subtypes are divided into two CRISPR classes (class 1 and class 2). One of the main differences is that class 1 CRISPR systems require a complex of multiple Cas proteins to degrade DNA, whereas class 2 Cas proteins are single, large, multidomain nucleases. The sequence specificity of class 2 Cas proteins can be easily modified by synthetic CRISPR RNA (crRNA) to introduce targeted double-stranded DNA breaks. The most widely known class 2 Cas proteins are Cas9, Cpf1 (Cas12a), and Cms1, which are used for genome editing and have been successfully applied in many eukaryotic organisms, including fungi, plants, and mammalian cells. Cas9 and its orthologs are class 2 type II CRISPR nucleases, whereas Cpf1 (Patent Document 1, Broad Inst.; Patent Document 2, Benson Hill) and Cms1 (Patent Document 3, Benson Hill) belong to class 2 type V nucleases. Cms1 and Cpf1 CRISPR nucleases are a class of CRISPR nucleases that have certain desirable properties compared to other CRISPR nucleases, such as type II nucleases. For example, in contrast to Cas9 nuclease, Cms1 and Cpf1 do not require a trans-activating crRNA (tracrRNA), which is partially complementary to the crRNA precursor (pre-crRNA) (Non-Patent Document 1). Base pairing of the tracrRNA and pre-crRNA forms a Cas9-bound RNA:RNA duplex, which is processed by RNase III and other unidentified nucleases. This mature tracrRNA:crRNA duplex mediates target DNA recognition and cleavage by Cas9.In contrast, type V nucleases can process pre-crRNA without the need for tracrRNA or cellular nucleases (such as RNase III), which significantly simplifies the application of type V nucleases for (multiplexed) genome editing.

[0005] Several new class 2 proteins, such as C2c1 (Cas12b), C2c2 (Cas13a), and C2c3 (Cas12c), have been identified in the genomes of cultured bacteria or in publicly available metagenomic analysis datasets, such as the gut metagenome (Non-Patent Document 2). According to a recent classification of CRISPR-Cas systems, class 2 includes three types and 17 subtypes (Non-Patent Document 3).

[0006] Furthermore, in a recent publication, two new class 2 proteins (CasX (Cas12a) and CasY (Cas12d)) were discovered in uncultured prokaryotes by megagenome sequencing (Non-Patent Document 4), indicating the existence of previously unexploited Cas proteins from as yet uncultured and / or unidentified organisms.

[0007] As discussed, known CRISPR-Cas systems exhibit certain specificities regarding their mode of action. These molecular specificities not only expand the potential for using CRISPR-Cas systems for genome editing across a wide range of genetic backgrounds, but also circumvent issues with applying specific Cas nucleases to specific organisms, such as pre-existing immune responses to Cas9 in humans (Non-Patent Document 5). Therefore, the identification of Cas nucleases from bacterial species with little direct contact with higher eukaryotes, or Cas nucleases of non-natural origin, is particularly important. It is anticipated that uncharacterized CRISPR-Cas systems exist in nature or can be engineered by protein engineering. Thus, although several distinct CRISPR-Cas systems are already known in the prior art, there is a continuing need to identify additional RNA-guided DNA endonucleases. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] International Publication No. 2016 / 205711 [Patent Document 2] International Publication No. 2017 / 141173 [Patent Document 3] International Publication No. 2019 / 030695 [Non-patent literature]

[0009] [Non-Patent Document 1] Deltcheva et al.(2011),Nature,471(7340):602~607 [Non-patent document 2] Shmakov et al. (2015), Mol Cell, 60(3):385~397 [Non-patent document 3] Makarova et al. (2020), Nat Rev Microbiol, 18(2):67~83 [Non-patent document 4] Burstein et al. (2017), Nature, 542: 237~241 [Non-patent document 5] Charlesworth et al. (2019), Nat Med, 25(2):249~254 Summary of the Invention

[0010] Therefore, in a first aspect, the present invention relates to a nucleic acid molecule encoding an RNA-guided DNA endonuclease, wherein: (a) a nucleic acid molecule encoding the RNA-guided DNA endonuclease comprising or consisting of the amino acid sequence of SEQ ID NO: 29, 1 or 3; (b) a nucleic acid molecule comprising or consisting of the nucleotide sequence of SEQ ID NO: 30, 2 or 4; (c) a nucleic acid molecule encoding an RNA-guided DNA endonuclease, the amino acid sequence of which is at least 90%, preferably at least 92%, and most preferably at least 95% identical to the amino acid sequence (a); (d) a nucleic acid molecule comprising or consisting of a nucleotide sequence that is at least 90%, preferably at least 92%, and most preferably at least 95% identical to the nucleotide sequence (b); (e) a nucleic acid molecule degenerate with respect to the nucleic acid molecule of (d); or (f) a nucleic acid molecule corresponding to any one of the nucleic acid molecules of (a) to (d), in which T is replaced by U.

[0011] SEQ ID NOs: 1, 3, and 29 are the amino acid sequences of the novel CRISPR-Cas endonucleases BEC85, BEC67, and BEC10, respectively, where BEC is an abbreviation for BRAIN Engineered Cas. Of the amino acid sequences of SEQ ID NOs: 1, 3, and 29, SEQ ID NO: 29, and therefore the amino acid sequence of BEC10, is preferred. The novel CRISPR-Cas endonucleases BEC85, BEC67, and BEC10 are encoded by the nucleotide sequences of SEQ ID NOs: 2, 4, and 30, respectively. Of the nucleotide sequences of SEQ ID NOs: 2, 4, and 30, SEQ ID NO: 30, and therefore the nucleotide sequence of BEC10, is preferred. As discussed in more detail herein below, the novel CRISPR-Cas endonucleases BEC85, BEC67, and BEC10 do not occur in nature but have been prepared by protein engineering.

[0012] The drawings show: [Brief explanation of the drawings]

[0013] [Figure 1] Schematic diagram visualizing the Ad2 knockout strategy in S. cerevisiae S288c for BEC85, BEC67, and BEC10 in comparison to SpCas9. [Figure 2] An exemplary culture plate showing S. cerevisiae S228c colonies 48 hours after transformation to visualize the different molecular mechanisms of BEC85, BEC67, and BEC10 compared to SpCas9. [Figure 3] An exemplary culture plate showing S. cerevisiae S228c colonies 48 hours after transformation (incubated at 30° C.) to visualize the colony reduction and genome editing efficiency of BEC family nucleases compared to the flanking sequences SuCms1 and SEQ ID NO: 63. Orange colonies, which represent a mixture of edited and non-edited cells, are marked with arrows. [Figure 4]An exemplary culture plate showing S. cerevisiae S228c colonies 48 hours after transformation (incubated at 21°C) to visualize the colony reduction and genome editing efficiency at low temperature (21°C) of BEC family nucleases compared to the flanking sequences SuCms1 and SEQ ID NO: 63. [Figure 5] An exemplary culture plate showing E. coli BW25113 colonies 48 hours after transformation (incubated at 37°C) to visualize the colony depletion efficiency at high temperature (37°C) of BEC family nucleases compared to the flanking sequences SuCms1 and SEQ ID NO: 63.

[0014] According to the present invention, the term "nucleic acid molecule" defines a linear molecular chain of nucleotides. A nucleic acid molecule according to the present invention consists of at least 3327 nucleotides. The group of molecules referred to herein as "nucleic acid molecules" also includes complete genes. The term "nucleic acid molecule" is used interchangeably with the term "polynucleotide" herein.

[0015] The term "nucleic acid molecule" according to the present invention includes DNA, such as cDNA or double- or single-stranded genomic DNA, and RNA. In this context, "DNA" (deoxyribonucleic acid) refers to any strand or sequence of chemical building blocks called nucleotide bases, adenine (A), guanine (G), cytosine (C), and thymine (T), linked together on a deoxyribose sugar backbone. DNA can have a single strand of nucleotide bases or two complementary strands that can form a double helix structure. "RNA" (ribonucleic acid) refers to any strand or sequence of chemical building blocks called nucleotide bases, adenine (A), guanine (G), cytosine (C), and uracil (U), linked together on a ribose sugar backbone. RNA typically has a single strand of nucleotide bases. Also included are single-stranded and double-stranded hybrid molecules, i.e., DNA-DNA, DNA-RNA, and RNA-RNA. Nucleic acid molecules can also be modified by many means known in the art. Non-limiting examples of such modifications include methylation, "capping," substitution of one or more analogs of natural nucleotides, and internucleotide modifications, such as those with uncharged bonds (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.) and those with charged bonds (e.g., phosphorothioates, phosphorodithioates, etc.). Polynucleotides can contain one or more additional covalently linked moieties, such as proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), intercalating agents (e.g., acridine, psoralen, etc.), chelators (e.g., metals, radioactive metals, iron, metal oxides, etc.), and alkylating agents. Polynucleotides can be derivatized by forming a methyl or ethyl phosphotriester bond or an alkyl phosphoramidate bond. Additionally included are nucleic acid mimetic molecules known in the art, such as synthetic or semi-synthetic derivatives of DNA or RNA, and mixed polymers.Such nucleic acid mimic molecules or nucleic acid derivatives according to the present invention include phosphorothioate nucleic acids, phosphoramidate nucleic acids, 2'-O-methoxyethyl ribonucleic acids, morpholino nucleic acids, hexitol nucleic acids (HNA), peptide nucleic acids (PNA), and locked nucleic acids (LNA) (see Braasch and Corey, Chem Biol 2001, 8:1). LNA is an RNA derivative in which the ribose ring is restricted by a methylene bond between the 2'-oxygen and the 4'-carbon. Also included are nucleic acids containing modified bases, such as thiouracil, thioguanine, and fluorouracil. Nucleic acid molecules typically carry genetic information, including information used by the cellular machinery to make proteins and / or polypeptides. Nucleic acid molecules of the present invention may additionally contain promoters, enhancers, response elements, signal sequences, polyadenylation sequences, introns, 5'-noncoding regions, and 3'-noncoding regions, and the like.

[0016] The term "polypeptide," used interchangeably herein with the term "protein," describes a linear molecular chain of amino acids, including a single-chain protein or fragment thereof. A polypeptide / protein according to the present invention contains at least 1108 amino acids. Polypeptides may form oligomers consisting of at least two identical or different molecules. The corresponding higher-order structures of such multimers are equivalently referred to as homodimers or heterodimers, homotrimers or heterotrimers, etc. Polypeptides of the present invention may form heteromultimers or homomultimers, such as heterodimers or homodimers. Furthermore, peptidomimetics of such proteins / polypeptides, in which amino acids and / or peptide bonds are replaced by functional analogs, are also encompassed by the present invention. Such functional analogs include any known amino acid other than the 20 gene-encoded amino acids, such as selenocysteine. The terms "polypeptide" and "protein" also refer to naturally modified polypeptides and proteins, for example, those affected by glycosylation, acetylation, phosphorylation, ubiquitination, and similar modifications well known in the art.

[0017] The terms "RNA-guided DNA endonuclease" or "CRISPR(-Cas) endonuclease" describe enzymes capable of cleaving phosphodiester bonds within deoxyribonucleotide (DNA) strands, thereby generating double-strand breaks (DSBs). BEC85, BEC67, and BEC10 are classified as novel type V class CRISPR nucleases known to introduce staggered cuts with 5' overhangs. Hence, RNA-guided DNA endonucleases contain an endonuclease domain, specifically a RuvC domain. The RuvC domains of BEC85, BEC67, and BEC10 each contain three divergent RuvC motifs (RuvC I-III, SEQ ID NOS: 5-7). RNA-guided DNA endonucleases also contain a domain capable of binding to a crRNA, also known as a guide RNA (gRNA, also referred to herein as a DNA-targeting RNA).

[0018] The cleavage site of the RNA-guided DNA endonuclease is guided by a guide RNA. The gRNA confers target sequence specificity to the RNA-guided DNA endonuclease. Such gRNAs are non-coding short RNA sequences that bind to complementary target DNA sequences. The gRNA first binds to the RNA-guided DNA endonuclease via a binding domain that can interact with the RNA-guided DNA endonuclease. The binding domain that can interact with the RNA-guided DNA endonuclease typically contains a region with a stem-loop structure. This stem-loop preferably has the sequence UCUACN, with base pairing between "UCUAC" and "GUAGA" to form the stem of the stem-loop. 3~5 Contains GUAGAU (SEQ ID NO: 8). 3~5indicates that any base can be present at this position, and 3, 4, or 5 nucleotides can be included at this position. The stem-loop most preferably comprises the stem-loop direct repeat sequences of BEC85 (SEQ ID NO: 9), BEC67 (SEQ ID NO: 10), and BEC10 (SEQ ID NO: 10), respectively, but in RNA form (i.e., T is replaced with U). The gRNA sequence guides a complex (known as a CRISPR ribonucleoprotein (RNP) complex of gRNA and RNA-guided DNA endonuclease) through pairing to a specific position on the DNA strand where the RNA-guided DNA endonuclease exerts its endonuclease activity by cleaving the DNA strand at the target site. The genomic target site for the gRNA can be any approximately 20 (typically 17-26) nucleotide DNA sequence, provided that two conditions are met: (i) the sequence is unique compared to the rest of the genome, and (ii) the target is immediately adjacent to a protospacer adjacent motif (PAM).

[0019] The cleavage site of an RNA-guided DNA endonuclease is therefore further defined by a PAM. The PAM is a short DNA sequence (usually 2-6 base pairs in length) following the DNA region targeted for cleavage by the CRISPR system. The actual sequence depends on the CRISPR endonuclease used. CRISPR endonucleases and their respective PAM sequences are known in the art (see https: / / www.addgene.org / crispr / guide / #pam-table). For example, the PAM recognized by the first identified RNA-guided DNA endonuclease, Cas9, is 5'-NGG-3' (where "N" can be any nucleotide base). The PAM is essential for cleavage by an RNA-guided DNA endonuclease. In Cas9, the PAM is known to be 2-6 nucleotides downstream of the DNA sequence targeted by the guide RNA and 3-6 nucleotides downstream from the cleavage site. In type V systems (including BEC85, BEC67, and BEC10), the PAM is located downstream of both the target sequence and the cleavage site. The complex consisting of an RNA-guided DNA endonuclease and a guide RNA contains a so-called PAM-interacting domain (Andres et al. (2014), Nature, 513(7519):569-573). Therefore, the genomic location that can be targeted for editing by an RNA-guided DNA endonuclease is restricted by the presence and location of the nuclease-specific PAM sequence. Since BEC85, BEC67, and BEC10 belong to the group of type V class 2 CRISPR nucleases, T-rich PAM sites have been predicted, and a TTTA PAM site has been shown to be functional (see Examples).

[0020] The term "percent (%) sequence identity" describes the number of matches ("hits") of identical nucleotides or amino acid residues in two or more aligned nucleic acid or amino acid sequences compared to the number of nucleotides or amino acid residues that make up the entire length of a template nucleic acid or amino acid sequence. In other words, alignment can be used to determine the percentage of amino acid residues or nucleotides that are the same (e.g., 70% identity) for two or more sequences or subsequences by comparing and aligning the (sub)sequences for maximum correspondence over a comparison range or designated region as determined using sequence comparison algorithms known in the art, or by manual alignment and visual inspection. This definition also applies to the complement of any aligned sequences.

[0021] Analysis and alignment of amino acid and nucleotide sequences relevant to the present invention is preferably performed using the NCBI BLAST algorithm (Stephen F. Altschul, Thomas L. Madden, Alejandro A. Schaffer, Jinghui Zhang, Zheng Zhang, Webb Miller, and David J. Lipman (1997), "Gapped BLAST and PSI-BLAST: a new generation of protein database search programs," Nucleic Acids Res. 25:3389-3402). Those skilled in the art will recognize additional suitable programs for aligning nucleic acid sequences.

[0022] As defined herein above, amino acid and nucleotide sequence identities of at least 90% are contemplated by the present invention, and even more preferably, amino acid sequence identities of at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.8%, and at least 99.9% are contemplated by the present invention.

[0023] With respect to these amino acid sequences and the amino acid sequences encoded by these nucleotide sequences, it is preferred to maintain or essentially maintain the RNA-guided DNA endonuclease activity of SEQ ID NOs: 1, 3, and 29 of the present invention. Thus, what is maintained or essentially maintained is the ability to bind to a gRNA to form a complex that can bind to a desired DNA target site, and for the endonuclease activity to induce DSBs.

[0024] The maintenance or substantial maintenance of RNA-guided DNA endonuclease activity can be analyzed in CRISPR-Cas genome editing experiments, for example, as described in Examples 3 to 5. Preferably, the amino acid sequence comprises a RuvC domain and the nucleotide sequence encodes the RuvC domain, as set forth in SEQ ID NOs: 5 to 7. As described, the RuvC domain is an endonuclease domain.

[0025] The term "degenerate" refers to the degeneracy of the genetic code. As is well known, codons encoding an amino acid can differ at any of their three positions, with this difference occurring less frequently at the second or third position. For example, the amino acid glutamic acid is represented by the codons GAA and GAG (different at the third position), the amino acid leucine is represented by the codons UUA, UUG, CUU, CUC, CUA, CUG (different at the first or third position), and the amino acid serine is represented by the codons UCA, UCG, UCC, UCU, AGU, AGC (different at the first, second, or third position).

[0026] As can be seen from the accompanying examples, the novel CRISPR nucleases BEC85, BEC67, and BEC10 of the present invention were generated using protein engineering and in silico approaches. Therefore, the Cas nucleases of the present invention are not simply isolated from bacterial species but are of non-natural origin. More specifically, a number of engineered nuclease sequences were screened, and the activity of the identified sequences was optimized using protein engineering. To the inventors' knowledge, this is the first time that a novel type of Cas nuclease not directly related to a sequence found in nature has been developed.

[0027] Furthermore, experimental results using the novel CRISPR nucleases BEC85, BEC67, and BEC10 in the accompanying Examples of the present application surprisingly demonstrated different molecular mechanisms of the BEC family of CRISPR nucleases compared to conventional CRISPR Cas nucleases. For example, compared to Cas nucleases that support homologous recombination by introducing RNA-induced double-strand breaks, BEC85-, BEC67-, and BEC10-mediated editing results in a strong overall reduction in clones, associated with a significant enrichment of cells that have successfully achieved homologous recombination. For this reason, the novel BEC-type CRISPR nucleases further expand the applicability of CRISPR technology for efficient genome editing.

[0028] As proof of principle, Example 3 demonstrates that BEC85, BEC67, and BEC10 are active CRISPR-Cas endonucleases that can be successfully used for genome editing. In Example 3, the Ade2 gene of Saccharomyces cerevisiae was knocked out using BEC85, BEC67, or BEC10, gRNA, and a homology-directed repair template.

[0029] Similar to the type V CRISPR endonucleases Cms1 and Cpf1, BEC85, BEC67, and BEC10 do not require trans-activating crRNA (tracrRNA). Furthermore, the CRISPR-containing systems identified in this application, BEC85, BEC67, and BEC10, contain CRISPR repeat sequences with an RNA stem-loop at the 3'-end of the repeat that is conserved in the crRNAs of the Cpf1 and Cms protein families, and the "closest neighbors" of BEC85, BEC67, and BEC10 among all known CRISPR-Cas endonucleases are the CMS-like Cas proteins from WO 2017 / 141173, in particular the CMS-like Cas protein SuCms1 (Begemann et al. (2017)) and SEQ ID NO: 63 (WO 2019 / 030695). Interestingly, the activity characteristics of the CMS CRISPR nucleases described in WO 2017 / 141173, WO 2019 / 030695, and Begemann et al. (2017), bioRxiv, are completely different from those of the BEC family of CRISPR nucleases. Example 3 demonstrates that the endonuclease activity of BEC85, BEC67, and BEC10 is based on a novel molecular mechanism that has not been previously described. More specifically, the results of Example 3 using the prior art CRISPR endonuclease SpCas9 and the novel CRISPR endonucleases of the present invention, BEC85, BEC67, and BEC10, are provided. Surprisingly, these results revealed completely different molecular genome editing mechanisms of the three BEC-type CRISPR nucleases compared to conventional CRISPR Cas nucleases. While SpCas9 supports homologous recombination by introducing RNA-directed double-strand breaks, BEC85-, BEC67-, and BEC10-mediated editing results in strong overall clonal reduction associated with significant enrichment of cells that have successfully achieved homologous recombination. Example 3 demonstrates the ability of BEC-type CRISPR nucleases to function as novel genome editing tools through site-specific and highly efficient homology-directed recombination.

[0030] For this reason, BEC85, BEC67, and BEC10 can be classified as novel, non-natural Class 2 Type V nucleases that have no significant overall sequence identity to the known collection of Class 1 and Class 2 CRISPR-Cas endonucleases and have low overall sequence identity to individual Cms1-type endonucleases.

[0031] BEC85, BEC67, and BEC1 are novel CRISPR-Cas endonucleases that differ significantly from the known collection of CRISPR-Cas endonucleases and exhibit a novel mechanism of activity, and BEC85, BEC67, and BEC10 expand the known collection of CRISPR-Cas endonucleases applicable to genome editing, genome regulation, and nucleic acid enrichment / purification in different biotechnology and pharmaceutical sectors. The results described in Example 3 strongly indicate that BEC-type CRISPR nucleases are not only a novel type of effector protein with a distinct locus structure, but also exhibit a novel molecular genome editing mechanism.

[0032] Furthermore, Example 4 demonstrates that genome editing using the novel BEC family nuclease of the present invention results in significantly higher clone reduction numbers and significantly better editing ratios compared to its neighboring element sequence SuCms1 (Begemann et al. (2017), bioRxiv) and SEQ ID NO: 63 (WO 2019 / 030695). The results of Example 4 further demonstrate the general superiority of BEC nucleases for genome editing compared to known CRISPR Cas nucleases.

[0033] Furthermore, Example 5 demonstrates that the novel BEC family nuclease of the present invention demonstrates strong activity at temperatures ranging from 21°C to 37°C and exhibits superior genome editing efficiency and colony reduction rate compared to its flanking element sequence SuCms1 and SEQ ID NO: 63. For example, the genome editing efficiency of SuCms1 nuclease significantly decreases to a level similar to that of the negative control (0.3%) at 21°C, whereas the editing efficiency of BEC10 remains high (65%) even at the relatively low temperature of 21°C. High activity within the temperature range of 21°C to 37°C is of great interest for biotechnology, agriculture, and pharmaceutical applications, because various types of cells are cultured within this temperature range (e.g., various plants and plant cells are cultured at approximately 21°C, various yeast and fungal cells are cultured at approximately 30°C, and various prokaryotic and mammalian cell lines are cultured at 37°C). Novel BEC family nucleases therefore advantageously enable the design of universally applicable CRISPR systems.

[0034] According to a preferred embodiment of the first aspect of the invention, the nucleic acid molecule is operably linked to a promoter that is native or heterologous to the nucleic acid molecule.

[0035] A promoter is a region of DNA that leads to the initiation of transcription of a particular gene. Promoters are generally located upstream (toward the 5' region of the sense strand) of DNA near the transcription start site of a gene. Promoters are typically 100 to 1000 base pairs in length. For transcription to occur, an enzyme that synthesizes RNA, known as RNA polymerase, must bind to DNA near the gene. Promoters contain specific DNA sequences, such as response elements, that provide conserved initiation binding sites for RNA polymerase and for proteins called transcription factors that recruit RNA polymerase. Thus, binding of RNA polymerase and transcription factors to the promoter site ensures transcription of the gene.

[0036] In this context, the term "operably linked" defines a promoter linked to a gene on the same DNA strand, such that upon binding of RNA polymerase and transcription factors, transcription of the gene is initiated. Generally, each gene is operably linked to a promoter in its natural environment in the genome of an organism. This promoter is referred to herein as the "native promoter" or "wild-type promoter." A heterologous promoter is different from a native or wild-type promoter. Thus, a nucleic acid molecule operably linked to a promoter that is heterologous to the nucleic acid molecule does not occur in nature.

[0037] Heterologous promoters that can be used to express a desired gene are known in the art and can be obtained, for example, from EPD (Eukaryotic Promoter Database) or EDPnew (https: / / epd.epfl.ch / / index.php). In this database, eukaryotic promoters can be found, including animal, plant, and yeast promoters.

[0038] The promoter can be, for example, a constitutively active, inducible, tissue-specific, or developmental stage-specific promoter. By using such a promoter, the desired timing and site of expression can be controlled.

[0039] Examples of constitutively active promoters are the AOX1 promoter or GAL1 promoter in yeast, or the CMV (cytomegalovirus), SV40, RSV (Rous sarcoma virus) promoter, chicken beta-actin promoter, CAG promoter (chicken beta-actin promoter and cytomegalovirus immediate early enhancer), gail0 promoter, human elongation factor 1 alpha promoter, CaM-kinase promoter, and Autographa california multiple nuclear polyhedrosis virus (AcMNPV) promoter.

[0040] Examples of inducible promoters include the AdhI promoter, which is inducible by hypoxia or cold stress, the Hsp70 promoter, which is inducible by heat stress, and the PPDK and PEP carboxylase promoters, both of which are inducible by light. Chemically inducible promoters are also useful, including the In2-2 promoter (U.S. Pat. No. 5,364,780), which is induced by herbicide antidotes, the ERE promoter, which is induced by estrogen, and the AxigI promoter, which is induced by auxin and is tapetum-specific but also active in callus (WO 03060123).

[0041] A tissue-specific promoter is a promoter that initiates transcription only in a particular tissue. A developmental stage-specific promoter is a promoter that initiates transcription only in a particular developmental stage.

[0042] In the examples described herein below, Tpil (SEQ ID NO: 12) and the SNR52 promoter (SEQ ID NO: 18) are used. Therefore, the use of Tpil and the SNR52 promoter is preferred.

[0043] According to a further preferred embodiment of the first aspect of the invention, said nucleic acid molecule is linked to a nucleic acid sequence encoding a nuclear localization signal (NLS).

[0044] Further details regarding NLS are provided later in this specification.

[0045] According to another preferred embodiment of the first aspect of the present invention, the nucleic acid molecule is codon optimized for expression in a eukaryotic cell, preferably a yeast cell, a plant cell, or an animal cell.

[0046] As discussed, BEC85, BEC67, and BEC10 are non-naturally occurring CRISPR nucleases, as they were generated by protein engineering.

[0047] The genes encoding the BEC85 polypeptide, the BEC67 polypeptide, and the BEC10 polypeptide may be codon-optimized for expression in target cells and may optionally include sequences encoding peptide tags such as NLSs and / or purification tags. More details regarding tags are provided below.

[0048] Codon optimization is a process used to improve gene expression and increase the transcription efficiency of a gene of interest by accommodating the codon bias of the host cell. A "codon-optimized gene" is therefore a gene whose codon usage frequency is designed to mimic the preferred codon usage frequency of the host cell. A nucleic acid molecule can be fully or partially codon-optimized. Because any single amino acid (except methionine and tryptophan) is coded for by several codons, the sequence of a nucleic acid molecule can be altered without changing the coded amino acid. Codon optimization occurs when one or more codons are changed at the nucleic acid level, thereby increasing expression in a particular host organism without changing the amino acid. Those skilled in the art will recognize that codon tables and other references providing preferred information for a wide range of organisms are available in the art (see, for example, Zhang et al. (1991) Gene 105:61-72; Murray et al. (1989) Nucl. Acids Res. 17:477-508). Methodologies for optimizing nucleotide sequences for expression are provided, for example, in U.S. Patent No. 6,015,891. Programs for codon optimization are available in the art (e.g., OPTIMIZER at genomes.urv.es / OPTIMIZER, OptimumGene™ from GenScript at www.genscript.com / codon_opt.html).

[0049] The eukaryotic cell is preferably a yeast cell, and therefore the codon optimization is preferably for expression in a yeast cell, which is of particular commercial interest because it is one of the most commonly used eukaryotic hosts for industrial-scale production of recombinant proteins.

[0050] In another embodiment, the eukaryotic cell is a mammalian cell, and therefore the codon optimization is preferably for expression in mammalian cells, and mammalian cells (preferably CHO cells and HEK293 cells) are of particular commercial interest, since these cells are the hosts typically used for industrial-scale production of recombinant protein therapeutics.

[0051] Further details regarding suitable eukaryotic cells, including plant and animal cells, are provided below.

[0052] In Example 2, it is described that the nucleotide sequences encoding BEC85, BEC67, or BEC10 are codon optimized for expression in yeast (specifically Saccharomyces cerevisiae) or bacteria (E. coli).

[0053] In a second aspect, the present invention relates to a vector encoding the nucleic acid molecule of the first aspect.

[0054] The definitions and preferred embodiments as set out hereinabove apply mutatis mutandis to the second aspect, where applicable.

[0055] A vector according to the invention is generally preferably capable of controlling the replication and / or expression of a nucleic acid molecule of the invention and / or the expression of a polypeptide encoded thereby.

[0056] Preferably, the vector is a plasmid vector, a cosmid vector, a bacteriophage vector or another vector conventionally used, for example, in genetic engineering.

[0057] Exemplary plasmids and vectors are listed, for example, in Studier and coworkers (Studier, WF; Rosenberg AH; Dunn JJ; Dubendroff JW, 1990, Use of the T7 RNA polymerase to direct expression of cloned genes, Methods Enzymol. 185, 61-89) or in brochures supplied by Novagen, Promega, New England Biolabs, Clontech, and Gibco BRL. Other suitable plasmids and vectors can be found in Glover, DM, 1985, DNA cloning: a practical approach, Vol. I-III, IRL Press Ltd., Oxford; Rodriguez, RL and Denhardt, DT (eds), 1988, Vectors: a survey of molecular cloning vectors and their uses, 179-204, Butterworth, Stoneham; Goedeel, DV, 1990, Systems for heterologous gene expression, Methods Enzymol. 185, 3-7; Sambrook, J.; Russell, DW, 2001, Molecular cloning: a laboratory manual, 3rd ed., Cold Spring Harbor Laboratory Press, New York.

[0058] Particularly preferred vectors are those that can be used for CRISPR genome editing, particularly those that only express the nucleic acid molecule of the present invention that encodes RNA-guided DNA endonuclease, or those that express both the nucleic acid molecule of the present invention that encodes RNA-guided DNA endonuclease and guide RNA (so-called "all-in-one vector").In the former case, a second vector is used to express guide RNA.CRISPR genome editing vectors are commercially available, for example, from OriGene, Vector Builder or ThermoFisher.

[0059] The nucleic acid molecules of the present invention described above can also be inserted into a vector to form a translational fusion with another nucleic acid molecule. For this purpose, overlap extension PCR can be applied (e.g., Wurch, T., Lestienne, F., and Pauwels, PJ, A modified overlap extension PCR method to create chimeric genes in the absence of restriction enzymes, Biotechn. Techn. 12, 9, September 1998, 653-657). The resulting product is called a fusion protein and is further described below. The other nucleic acid molecule can, for example, encode a protein that can increase the solubility and / or facilitate the purification of the protein encoded by the nucleic acid molecule of the present invention. Non-limiting examples include pET32, pET41, and pET43. The vector may also contain an additional expressible nucleic acid encoding one or more chaperones to facilitate correct protein folding. Suitable bacterial expression hosts include, for example, strains derived from BL21 (BL21(DE3), BL21(DE3)PlysS, BL21(DE3)RIL, BL21(DE3)PRARE) or Rosetta®.

[0060] For vector modification techniques, see JF Sambrook and DW Russell, eds., Cold Spring Harbor Laboratory Press, 2001, ISBN-10 0-87969-577-3. Generally, a vector can contain one or more origins of replication (ori) and genetic systems for cloning or expression, one or more markers for selection in the host, e.g., antibiotic resistance, and one or more expression cassettes. Suitable origins of replication include, for example, the Col E1 origin of replication, the SV40 viral origin of replication, and the M13 origin of replication.

[0061] The coding sequence inserted into the vector can be synthesized, for example, by standard methods or isolated from natural sources. Ligation of the coding sequence to transcriptional regulatory elements and / or to other amino acid coding sequences can be performed using established methods. Transcriptional regulatory elements (part of an expression cassette) ensuring expression in prokaryotic or eukaryotic cells are well known to those skilled in the art. These elements include regulatory sequences (e.g., translation initiation codon, transcription termination sequence, promoter, enhancer, and / or insulator) ensuring transcription initiation, an internal ribosome entry site (IRES) (Owens et al., (2001), PNAS.98(4)1471-1476), and, optionally, a poly(A) signal ensuring transcription termination and transcription stabilization. Additional regulatory elements may include transcriptional and translational enhancers, and / or naturally associated or heterologous promoter regions. Regulatory elements may be native to the endonuclease of the present invention or heterologous regulatory elements. Preferably, the nucleic acid molecule of the present invention is operably linked to such expression control sequences to allow expression in prokaryotic or eukaryotic cells. The vector may further comprise a nucleotide sequence encoding a secretion signal as an additional regulatory element. Such sequences are well known to those skilled in the art. Furthermore, depending on the expression system used, a leader sequence capable of directing the expressed polypeptide to a cellular compartment may be added to the coding sequence of the nucleic acid molecule of the present invention. Such leader sequences are well known in the art. Specifically designed vectors allow DNA transfer between different hosts, such as bacteria-fungal cells or bacteria-animal cells.

[0062] Additionally, baculovirus systems or systems based on vaccinia virus or Semliki Forest virus can be used as vectors in eukaryotic expression systems for the nucleic acid molecules of the present invention. Expression vectors derived from viruses such as retroviruses, vaccinia virus, adeno-associated viruses, herpes viruses, or bovine papillomavirus can be used to deliver nucleic acids or vectors to targeted cell populations. Methods well known to those skilled in the art can be used to construct recombinant viral vectors; see, for example, the techniques described in Sambrook and DW Russell, eds., Cold Spring Harbor Laboratory Press, 2001.

[0063] Examples of regulatory elements that allow expression in eukaryotic host cells are promoters, including those described herein above. In addition to elements responsible for transcription initiation, such regulatory elements may also include transcription termination signals, such as the SV40 poly-A site or tk poly-A site downstream of the nucleic acid, or the SV40, laxZ, and AcMNPV polyhedron polyadenylation signals.

[0064] Co-transfection with a selectable marker, such as a kanamycin or ampicillin resistance gene, for culturing in E. coli and other bacteria allows for the identification and isolation of transfected cells. Selectable markers for mammalian cell culture are the dhfr, gpt, neomycin, and hygromycin resistance genes. Transfected nucleic acids can also be amplified to express large amounts of the encoded (poly)peptide. The DHFR (dihydrofolate reductase) marker is useful for developing cell lines carrying hundreds or thousands of copies of the gene of interest. Another useful selectable marker is the enzyme glutamine synthase (GS). Using such a marker, cells are grown in selective medium and the cells with the highest resistance are selected.

[0065] However, the nucleic acid molecules of the invention as described herein above can also be designed to directly introduce phage vectors, or viral vectors (e.g., adenovirus or retrovirus) into cells via liposomes.

[0066] In a third aspect, the present invention relates to a host cell comprising a nucleic acid molecule of the first aspect or transformed, transduced or transfected with a vector of the second aspect.

[0067] The definitions and preferred embodiments as set out herein above apply mutatis mutandis to the third aspect, where applicable.

[0068] Large amounts of RNA-guided DNA endonuclease can be produced by the host cell, in which an isolated nucleotide sequence encoding the RNA-guided DNA endonuclease is inserted into a suitable vector or expression vector which is then introduced into a suitable host cell (preferably one which can be grown in large quantities), and the RNA-guided DNA endonuclease is purified from the host cell or culture medium.

[0069] Host cells can also be used to provide the RNA-guided DNA endonuclease of the present invention without the need for purification of the RNA-guided DNA endonuclease (see Yuan, Y.; Wang, S.; Song, Z.; and Gao, R., Immobilization of an L-aminoacylase-producing strain of Aspergillus oryzae into gelatin pellets and its application in the resolution of D,L-methionine, Biotechnol Appl. Biochem. (2002). 35:107-113). The RNA-guided DNA endonuclease of the present invention can be secreted by the host cell. Those skilled in the field of molecular biology will understand that any of a wide variety of expression systems can be used to provide the RNA-guided DNA endonuclease. The precise host cell used is not critical to the present invention, so long as the host cell produces the RNA-guided DNA endonuclease when grown under appropriate growth conditions.

[0070] Host cells into which vectors containing the nucleic acid molecules of the invention can be cloned are used to replicate and isolate sufficient quantities of the recombinant enzyme. Methods used for this purpose are well known to those skilled in the art (Sambrook and DW Russell, eds., Cold Spring Harbor Laboratory Press, 2001).

[0071] The expression of RNA-guided DNA endonuclease can not only be used to produce RNA-guided DNA endonuclease in host cell, but also its expression can be used to edit the genome of host cell.In this case, host cell also contains guide RNA.The vector that can be used for CRISPR genome editing has been discussed herein above.

[0072] According to a preferred embodiment of the third aspect of the present invention, said host cell is a eukaryotic or prokaryotic cell, preferably a plant cell, a yeast cell or an animal cell.

[0073] The host cell can be a eukaryotic cell, such as a fungal, algal, plant, or animal cell, where the animal can be a bird, reptile, amphibian, fish, cephalopod, crustacean, insect, arachnid, marsupial, or mammal. A gene encoding BEC85, BEC67, or BEC10 that is non-native to the host cell can be operably linked to a regulatory element, such as a promoter. The promoter can be native to the host organism or can be a promoter from another species. Constructs for expressing BEC85, BEC67, or BEC10 in heterologous host cells, such as eukaryotic cells, can further optionally include a transcription terminator. The gene encoding BEC85, BEC67, or BEC10 may optionally be codon-optimized for the host species, optionally include one or more introns, and optionally include one or more peptide tag sequences, one or more nuclear localization sequences (NLS), and / or one or more linkers or engineered cleavage sites (e.g., 2a sequences). In various embodiments, the host cell can comprise any of the engineered BEC85, BEC67, or BEC10 CRISPR systems disclosed above, wherein the nucleic acid sequence encoding the effector is present in the cell prior to introduction of the guide RNA. In other embodiments, the cell engineered to contain a gene for expressing a BEC85, BEC67, or BEC10 polypeptide further comprises a polynucleotide encoding a guide RNA (e.g., a guide RNA) operably linked to a regulatory element.

[0074] The cell or organism can be a prokaryotic cell. Suitable prokaryotic host cells include, for example, E. coli BL21 (e.g., BL21(DE3), BL21(DE3)PlysS, BL21(DE3)RIL, BL21(DE3)PRARE, BL21 Codon Plus, BL21(DE3) Codon Plus), Rosetta®, XL1 Blue, NM522, JM101, JM109, JM105, RR1, DH5α, TOP10, HB101, or MM294. Further suitable bacterial host cells include, but are not limited to, Streptomyces, Pseudomonas such as Pseudomonas putida, Corynebacterium such as C. glutamicum, Lactobacillus such as L. salivarius, Salmonella, or Bacillus such as Bacillus subtilis.

[0075] In general, eukaryotic host cells are preferred over prokaryotic host cells.

[0076] The eukaryotic cell can be a yeast cell, a fungal cell, an amoeba cell, an insect cell, a vertebrate cell (eg, a mammalian cell), or a plant cell.

[0077] The yeast cell may be, for example, a yeast cell of the genus Kluyveromyces, such as Saccharomyces cerevisiae, Ogataea angusta, K. marxianus, or K. lactis, or a yeast cell of the genus Pichia, such as P. pastoris, a yeast cell of the genus Yarrowia, such as Yarroawia lipolytica, a yeast cell of the genus Candida, an insect cell such as a Drosophila S2 cell or a Spodoptera Sf9 cell, a plant cell, or a fungal cell, preferably a fungal cell of the family Trichocomaceae, more preferably the genera Aspergillus, Penicillium, or Trichoderma, or of the family Ustilaginaceae, preferably the genus Ustilago.

[0078] Plant host cells that can be used include monocotyledonous and dicotyledonous plants (ie, monocotyledons and dicotyledons, respectively), such as crop plant cells and tobacco cells.

[0079] Mammalian host cells that can be used include human Hela cells, HEK293 cells, H9 cells, and Jurkat cells, mouse NIH3T3 cells and C127 cells, COS1 cells, COS7 cells, and CV1 cells, quail QC1-QC3 cells, mouse L cells, Bowes melanoma cells, HaCaT cells, BHK cells, HT29 cells, A431 cells, A549 cells, U2OS cells, MDCK cells, HepG2 cells, CaCo-2 cells, and Chinese hamster ovary (CHO) cells.

[0080] In a fourth aspect, the present invention relates to a plant, seed, or part of a plant that is not a single plant cell, or an animal, which comprises a nucleic acid molecule of the first aspect or which has been transformed, transduced or transfected with a vector of the second aspect.

[0081] The definitions and preferred embodiments as set out hereinabove apply mutatis mutandis to the fourth aspect, where applicable.

[0082] The animal is preferably a mammal, and most preferably a non-human mammal, such as a mouse, rat, hamster, cat, dog, horse, pig, cow, monkey, ape, etc.

[0083] By expressing the nucleic acid molecule of the first aspect together with a guide RNA in a plant, seed, or plant part, or animal, the genome of the host can be edited for the purpose of introducing targeted genetic mutations, for example for gene therapy, for creating chromosomal rearrangements, for studying gene function, for generating transgenic organisms, for endogenous gene tagging, or for targeted transgene addition.

[0084] In a fifth aspect, the present invention relates to a method for producing an RNA-guided DNA endonuclease, comprising culturing the host cell of the third aspect and isolating the RNA-guided DNA endonuclease produced.

[0085] The definitions and preferred embodiments as set out hereinabove apply mutatis mutandis to the fifth aspect, where applicable.

[0086] Suitable conditions for culturing prokaryotic or eukaryotic hosts are well known to those skilled in the art. Generally, suitable conditions for culturing bacteria are growing the bacteria under aeration in Luria-Bertani (LB) medium. To increase the yield and solubility of the expression product, the medium can be buffered or suspending with appropriate additives known to enhance or promote both. E. coli can be cultured at temperatures ranging from 4 to about 37°C, with the actual temperature or temperature range depending on the molecule to be overexpressed.

[0087] Generally, Aspergillus can be grown on Sabouraud dextrose agar or potato dextrose agar at temperatures ranging from about 10°C to about 40°C, preferably about 25°C. Suitable conditions for yeast culture are known, for example, from Guthrie and Fink, "Guide to Yeast Genetics and Molecular Cell Biology" (2002); Academic Press Inc. Those skilled in the art will recognize all of these conditions and can further adapt them to the needs of a particular host species and the requirements of the polypeptide to be expressed. When an inducible promoter controls the nucleic acid of the present invention in a vector present in a host cell, expression of the polypeptide can be induced by the addition of an appropriate inducer. Suitable expression protocols and strategies are known to those skilled in the art.

[0088] Depending on the cell type and its specific requirements, mammalian cell culture can be carried out, for example, in RPMI or DMEM medium containing 10% (v / v) FCS, 2 mM L-glutamine, and 100 U / ml penicillin / streptomycin. These cells can be maintained at 37°C, 5% CO2, in a water-saturated atmosphere. Expression protocols suitable for eukaryotic cells are well known to those skilled in the art and can be obtained, for example, from Sambrook, 2001.

[0089] Methods for isolation of the produced RNA-guided DNA endonuclease are well known in the art and include, but are not limited to, method steps such as ion exchange chromatography, gel filtration chromatography (size exclusion chromatography), affinity chromatography, high performance liquid chromatography (HPLC), reverse-phase HPLC, disc gel electrophoresis, or immunoprecipitation; see, e.g., Sambrook, 2001.

[0090] The protein isolation step is preferably a protein purification step. Protein purification according to the present invention refers to a process or series of processes designed to further isolate the polypeptide of the present invention from a complex mixture, preferably to homogeneity. The purification step exploits, for example, differences in protein size, physicochemical properties, and binding affinity. For example, proteins can be purified by isoelectric pointing through a pH gradient gel or an ion exchange column. Furthermore, proteins can be separated by protein size or molecular weight via size exclusion chromatography or SDS-PAGE (sodium dodecyl sulfate polyacrylamide gel electrophoresis) analysis. In the art, proteins are often purified using two-dimensional PAGE and then analyzed by peptide mass fingerprinting to establish the identity of the protein. This is useful for scientific purposes, as the detection limit for proteins is very low, and nanogram amounts of protein are sufficient for analysis. Proteins can also be purified by polarity / hydrophobicity via high-performance liquid chromatography or reverse-phase chromatography. Thus, methods for protein purification are well known to those skilled in the art.

[0091] In a sixth aspect, the present invention relates to an RNA-guided DNA endonuclease encoded by the nucleic acid molecule of the first aspect.

[0092] The definitions and preferred embodiments as set out hereinabove apply mutatis mutandis to the sixth aspect, where applicable.

[0093] The amino acid sequences of SEQ ID NOs: 1, 3, and 29 are particularly preferred examples of the RNA-guided DNA endonuclease of the present invention. An RNA-guided DNA endonuclease comprising or consisting of the amino acid sequence of SEQ ID NO: 29 is most preferred.

[0094] The RNA-guided DNA endonuclease of the sixth aspect of the present invention may also be a fusion protein, in which the amino acid sequence of the RNA-guided DNA endonuclease is fused to a fusion partner. The fusion may be a direct fusion or a fusion via a linker. The linker is preferably a peptide such as a GS-linker.

[0095] The fusion partner can be located at the N-terminus, C-terminus, both termini, or within an internal position of the RNA-guided DNA endonuclease polypeptide, preferably at the N-terminus or C-terminus.

[0096] The fusion partner is preferably a nuclear localization signal (NLS), a cell-transducing domain, a plastid targeting signal, a mitochondrial targeting signal peptide, a signal peptide that targets both plastids and mitochondria, a marker domain, a tag (such as a purification tag), a DNA-modifying enzyme, or a transactivation domain.

[0097] DNA-modifying enzymes can modify DNA by phosphorylating or dephosphorylating blunt-ended DNA, where blunting refers to digesting single-stranded overhangs. Non-limiting examples of dephosphorylating enzymes include shrimp alkaline phosphatase (rSAP), Quick CIP phosphatase, and Antarctic polar phosphatase. Non-limiting examples of phosphorylating enzymes include polynucleotide kinases such as T4 PNK. Non-limiting examples of blunting enzymes include DNA polymerase I large (Klenow) fragment, T4 DNA polymerase, or mung bean nuclease.

[0098] A transactivation domain (or trans-activating domain (TAD)) is a transcription factor scaffolding domain that contains binding sites for other proteins, such as transcriptional coactivators. Non-limiting examples are the 9-amino acid transactivation domain (9aaTAD) and glutamine-rich (Q) TAD.

[0099] Generally, an NLS comprises a chain of basic amino acids. Nuclear localization signals are known in the art. An NLS can be at the N-terminus, C-terminus, or both of the RNA-guided DNA endonuclease polypeptides of the present invention. For example, an RNA-guided DNA endonuclease polypeptide of the present invention can comprise about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the amino-terminus, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the carboxy-terminus, or a combination thereof (e.g., zero or at least one NLS at the amino-terminus and zero or at least one NLS at the carboxy-terminus). When more than one NLS is present, each can be selected independently of the others, such that a single NLS can be present in more than one copy and / or in combination with one or more other NLSs present in one or more copies. In some embodiments, an NLS is considered to be near the N- or C-terminus when the nearest amino acid of the NLS is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N- or C-terminus. The RNA-guided DNA endonuclease polypeptide sequence and the NLS may, in some embodiments, be fused to a linker that is from 1 to about 20 amino acids in length.

[0100] Non-limiting examples of NLSs include the NLS of SV40 virus large T antigen, an NLS derived from nucleoplasmin (e.g., a nucleoplasmin bipartite NLS), c-myc NLS, hRNPAI M9 NLS, the IBB domain from importin-alpha, myo-T protein, p53 protein, c-abl IV protein, or the NLS sequence of influenza virus NS1, the NLS of hepatitis virus delta antigen, Mx1 protein, poly(ADP-ribose) polymerase, and steroid hormone receptor (human) glucocorticoid. Generally, the one or more NLSs are of sufficient length to drive accumulation of the RNA-guided DNA endonuclease polypeptide according to the present invention in detectable amounts in the nucleus of a eukaryotic cell.

[0101] Plastid targeting signal peptide localization signals, mitochondrial targeting signal peptide localization signals, and dual targeting signal peptide localization signals are also known in the art (e.g., Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physiol 6:259; Herrmann and Neupert (2003) IUBMB Life 55:219-225; Soil (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Carrie et al. (2009) FEBS J 276:1187-1195; Silva-Filho (2003) Curr Opin Plant Biol 6:589-595, Peeters and Small (2001) Biochim Biophys Acta 1541:54-63, Murcha et al. (2014) Exp Bot 65:6301-6335, Mackenzie (2005) Trends Cell Biol 15:548-554, Glaser et al. (1998) Plant Mol Biol 38:311-338).

[0102] Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In certain embodiments, the marker domain can be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tag GFP, turbo GFP, EGFP, emerald, thistle green, monomeric thistle green, CopGFP, AceGFP, Zs Green I). Yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhyYFP, ZsYellow), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-Sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyanI, Acropora Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFPI, DsRed-Express, DsRed2, DsRed-monomer, HcRed-tandem, HcRedI, AsRed2, eqFP611, mRaspberry, mStrawberry, J-Red), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange).

[0103] A tag is a short amino acid sequence that allows identification of the RNA-guided DNA endonuclease polypeptide of the present invention in a mixture of polypeptides. Therefore, the tag is preferably a purification tag. Non-limiting examples of purification tags are His tags (e.g., His-6 tags), GST tags, DHFR tags, and CBP tags. A review of known purification tags can be found in Kimple et al. (2015), Curr Protoc Protein Sci. 2013;73:Unit-9.9.

[0104] In a seventh aspect, the present invention relates to a composition comprising the nucleic acid molecule of the first aspect, the vector of the second aspect, the host cell of the third aspect, the plant, seed, part of a cell, or animal of the fourth aspect, the RNA-guided DNA endonuclease of the sixth aspect, or a combination thereof.

[0105] The definitions and preferred embodiments as set out hereinabove apply mutatis mutandis to the seventh aspect, where applicable.

[0106] The term "composition" as used herein refers to a composition comprising at least one of the nucleic acid molecule of the first aspect, the vector of the second aspect, the host cell of the third aspect, the plant, seed, part of a cell, or animal of the fourth aspect, the RNA-guided DNA endonuclease of the sixth aspect, or a combination thereof, which are also hereinafter collectively referred to as compounds.

[0107] According to a preferred embodiment of the seventh aspect, said composition is a pharmaceutical composition or a diagnostic composition.

[0108] According to the present invention, the term "pharmaceutical composition" refers to a composition for administration to a patient, preferably a human patient. The pharmaceutical composition of the present invention comprises at least one of the compounds listed above. The pharmaceutical composition may optionally contain additional molecules that can alter the characteristics of the compound of the present invention, thereby, for example, stabilizing, regulating, and / or activating the function of the compound. The composition may be in solid, liquid, or gaseous form, particularly in the form of a powder, tablet, liquid, or aerosol. The pharmaceutical composition of the present invention may optionally and additionally comprise a pharmaceutically acceptable carrier. Examples of suitable pharmaceutical carriers are well known in the art and include phosphate buffered saline, water, emulsions such as oil / water emulsions, various types of wetting agents, sterile solvents, organic solvents including DMSO, and the like. Compositions containing such carriers can be formulated by conventional methods. These pharmaceutical compositions can be administered to a subject at an appropriate dosage. The dosing regimen is determined by the attending physician and clinical factors. As is well known in the medical field, the dosage for any one patient depends on many factors, including the patient's size, body surface area, age, the specific compound being administered, sex, time and route of administration, health status, and other drugs being administered concomitantly. The therapeutically effective amount for a given situation is readily determined by routine experimentation and is within the skill and judgment of a clinician or physician of ordinary skill. Generally, the dosage regimen for regular administration of the pharmaceutical composition should be in the range of 1 μg to 5 g of active compound per day. However, more preferred dosages may be in the range of 0.01 mg to 100 mg per day, even more preferably 0.01 mg to 50 mg, and most preferably 0.01 mg to 10 mg. The length of treatment needed to observe changes and the interval at which a response occurs after treatment will vary depending on the desired effect. Specific amounts can be determined by conventional testing well known to those skilled in the art.

[0109] The pharmaceutical composition may be used to treat or prevent a pathogenic disease, such as a viral or bacterial disease. For example, the RNA-guided DNA endonuclease of the sixth aspect may be used together with a gRNA that targets the genome of a pathogen, thereby modifying the genome of the pathogen so as to prevent or treat the disease caused by the pathogen.

[0110] The pharmaceutical compositions can also be used to treat or prevent, for example, imbalances in the microflora, which can arise, for example, from antibiotic overuse, which can result in the overgrowth of pathogenic bacteria and yeasts.

[0111] A "diagnostic composition" refers to a composition suitable for detecting both infectious and non-infectious diseases in a subject. Diagnostic compositions may particularly include a marker moiety as described hereinabove in connection with the binding of a fusion protein of the present invention to single-stranded DNA, such that when the RNA-guided DNA endonuclease polypeptide of the present invention cleaves the single-stranded DNA, it activates a reporter, which produces fluorescence or a change in color, thus allowing visual detection of a nuclear marker of a particular disease. Diagnostic compositions may be applied to a body fluid sample, such as blood, urine, or saliva.

[0112] In an eighth aspect, the present invention relates to a nucleic acid molecule of the first aspect, a vector of the second aspect, a host cell of the third aspect, a plant, a seed, a part of a cell or an animal of the fourth aspect, an RNA-guided DNA endonuclease of the sixth aspect, or a combination thereof, for use in treating a disease in a subject or plant by modifying a nucleotide sequence of a target site in the genome of the subject or plant.

[0113] Also described is a method for treating or preventing a disease in a subject or a plant, comprising modifying the nucleotide sequence of a target site in the genome of the subject or plant by a nucleic acid molecule of the first aspect, a vector of the second aspect, a host cell of the third aspect, a plant, seed, part of a cell or animal of the fourth aspect, an RNA-guided DNA endonuclease of the sixth aspect, or a combination thereof.

[0114] The definitions and preferred embodiments as set out hereinabove apply mutatis mutandis to the eighth aspect, where applicable.

[0115] Modification of the nucleotide sequence of a target site in the genome of a subject or plant is, in accordance with the present invention, genome editing by CRISPR technology, in particular by the novel RNA-guided DNA endonucleases described herein, in combination with an appropriate gRNA and, optionally, repair substrates as described below to determine the target site for genome modification.

[0116] Genome editing (also known as genome engineering) is a type of genetic engineering in which a target site, preferably a gene of interest, is inserted, deleted, modified, or replaced in a cell's genome. The target site, preferably a gene of interest, can be within the genome, but can also be within mitochondrial DNA (animal cells) or chloroplast DNA (plant cells). Genome editing can result in loss-of-function or gain-of-function mutations in a cell's genome. Loss-of-function mutations (also called inactivating mutations) result in a reduced or complete loss of function (partial or total inactivation) of the gene of interest. When an allele is completely lost-of-function (total inactivation), this is also referred to herein as a (gene) knockout. Gene knockout can be achieved by inserting, deleting, modifying, or substituting one or more nucleotides of a gene. Gain-of-function mutations (also called activating mutations) alter the gene of interest, thereby making it more effective (enhanced activity) or replacing it with a different (e.g., abnormal) function. Such gain-of-function mutations that introduce new functions or effects are also called gene knock-in.Genome editing can also result in the up-regulation or down-regulation of one or more genes.By targeting the DNA site responsible for the regulation of gene expression (for example, promoter region or transcription factor coding gene), gene expression can be up-regulated or down-regulated by CRISPR technology.More details about the mode of action of CRISPR technology will be described later in this specification.

[0117] Since its discovery, CRISPR technology has been increasingly applied to therapeutic genome editing. The adoption of several viral and non-viral vectors has enabled efficient delivery of CRISPR systems to target cells or tissues in a variety of ways, including mutagenesis, gene integration, epigenetic modulation, chromosomal rearrangement, base editing, and mRNA editing (for a review, see Le and Kim (2019), Hum Genet.;138(6):563-590).

[0118] The modification of the nucleotide sequence of a target site in the genome of a subject is preferably gene therapy, which is based on the principle of genetic manipulation of the nucleotide sequence of a target site to treat and prevent diseases, particularly human diseases.

[0119] In clinical trials of CRISPR technology, scientists are using it to eradicate cancer and blood disorders in humans. In these trials, some cells are removed from the subject to be treated, their DNA is edited, and the edited cells are then returned to the subject's body, where they are equipped to fight the disease being treated.

[0120] In a ninth aspect, the present invention relates to a method for modifying a nucleotide sequence of a target site in the genome of a cell, the method comprising introducing into the cell (i) a DNA-targeting RNA, or a DNA polynucleotide encoding the DNA-targeting RNA, wherein the DNA-targeting RNA comprises (a) a first segment comprising a nucleotide sequence complementary to a sequence in the target DNA and (b) a second segment that interacts with the RNA-guided DNA endonuclease of the sixth aspect, and (ii) the RNA-guided DNA endonuclease of the sixth aspect, or a nucleic acid molecule encoding the RNA-guided DNA endonuclease of the first aspect, or the vector of the second aspect, wherein the RNA-guided DNA endonuclease comprises (a) an RNA-binding portion that interacts with the DNA-targeting RNA and (b) an activity portion that exhibits site-specific enzymatic activity.

[0121] Accordingly, the present invention also relates to a pharmaceutical composition (e.g., a pharmaceutical composition or a diagnostic composition) comprising: (i) a DNA-targeting RNA, or a DNA polynucleotide encoding the DNA-targeting RNA, wherein the DNA-targeting RNA comprises (a) a first segment comprising a nucleotide sequence complementary to a sequence in the target DNA, and (b) a second segment that interacts with the RNA-guided DNA endonuclease of the sixth aspect; and (ii) a nucleic acid molecule encoding the RNA-guided DNA endonuclease of the first aspect, or the vector of the second aspect, wherein the RNA-guided DNA endonuclease comprises (a) an RNA-binding portion that interacts with the DNA-targeting RNA and (b) an activity portion that exhibits site-specific enzymatic activity.

[0122] The definitions and preferred embodiments as set out hereinabove apply mutatis mutandis to the ninth aspect, where applicable.

[0123] The DNA-targeting RNA comprises a first segment comprising a nucleotide sequence complementary to a sequence in the target DNA and a second segment that interacts with the RNA-guided DNA endonuclease. As discussed hereinabove, the nucleotide sequence complementary to a sequence in the target DNA defines the target specificity of the RNA-guided DNA endonuclease. As also discussed hereinabove, the DNA-targeting RNA binds to the RNA-guided DNA endonuclease, thereby forming a complex. The second segment interacts with the RNA-guided DNA endonuclease and is responsible for the formation of the complex. The second segment that interacts with the RNA-guided DNA endonuclease of the sixth aspect preferably comprises or consists of SEQ ID NO: 8, more preferably SEQ ID NO: 9 or 10. SEQ ID NO: 8 is the second segment also known as the 5' handle in type V class 2 CRISPR nucleases. SEQ ID NO: 9 or 10 is the second segment of BEC85, BEC67, or BEC10, respectively.

[0124] The RNA-guided DNA endonuclease comprises a first segment, an RNA-binding portion that interacts with the DNA-targeting RNA, and a second segment, an active portion that exhibits site-specific enzymatic activity. The first segment interacts with the DNA-targeting RNA and is responsible for the formation of the discussed complex. The second segment possesses an endonuclease domain, which preferably comprises a RuvC domain as described herein above (particularly, the RuvC domains of SEQ ID NOS: 5-7).

[0125] As also discussed hereinabove, the DNA-targeting RNA is a guide RNA. The guide RNA can be either directly introduced into a cell or as a DNA polynucleotide encoding the DNA-targeting RNA. In the latter case, the DNA encoding the guide RNA is generally operably linked to one or more promoter sequences for the expression of the guide RNA. For example, the RNA coding sequence can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III) or RNA polymerase II (Pol II). The DNA polynucleotide encoding the DNA-targeting RNA is preferably a vector. Many single gRNA empty vectors (with and without CRISPR endonuclease) are available in the art. In addition, several empty multi-gRNA vectors are available that can be used to express multiple gRNAs from a single plasmid (with or without CRISPR endonuclease expression). The DNA polynucleotide encodes the DNA-targeting RNA in an expressible form.

[0126] Similarly, the RNA-guided DNA endonuclease may either be introduced directly into the cell or as a nucleic acid molecule encoding the RNA-guided DNA endonuclease, the latter being preferably a vector of the second aspect: The DNA polynucleotide encodes the RNA-guided DNA endonuclease in an expressible form.

[0127] As discussed in more detail herein above, the RNA-guided DNA endonuclease and the DNA-targeting RNA can also be encoded by the same DNA polynucleotide, such as an all-in-one CRISPR-cas vector.

[0128] The term "in expressible form" means that the one or more DNA polynucleotides encoding the RNA-guided DNA endonuclease and the DNA-targeting RNA are in a form that ensures that the DNA-targeting RNA is transcribed and that the RNA-guided DNA endonuclease is transcribed and translated into an active enzyme in the cell.

[0129] According to a preferred embodiment of the ninth aspect of the present invention, when the RNA-guided DNA endonuclease and the DNA-targeting RNA are directly introduced into cells, they are introduced in the form of a ribonucleoprotein complex (RNP).

[0130] RNPs can be assembled in vitro and delivered to cells by methods known in the art, such as electroporation or lipofection. RNPs can cleave target sites with similar efficacy as nucleic acid-based (e.g., vector-based) RNA-guided DNA endonucleases (Kim et al. (2014), Genome Research 24(6):1012-1019).

[0131] Means for introducing proteins (or peptides) or RNPs into sex cells are known in the art and include, but are not limited to, microinjection, electroporation, lipofection (using liposomes), nanoparticle-based delivery, and protein transduction. Any one of these methods may be used.

[0132] Liposomes used in lipofection are small vesicles composed of the same material as cell membranes and can be loaded with one or more proteins (e.g., Torchilin VP. (2006), Adv Drug Deliv Rev., 58(14):1532-55). To deliver proteins or RNPs into cells, the lipid bilayer of the liposome fuses with the lipid bilayer of the cell membrane, thereby delivering the contained protein into the cell. Liposomes used according to the present invention are preferably composed of cationic lipids. The cationic liposome strategy has been successfully applied to protein delivery (Zelphati et al. (2001). J. Biol. Chem. 276, 35103-35110). As is known in the art, the actual composition and / or mixture of cationic lipids used can vary depending on the protein of interest and the cell type used (Felgner et al. (1994). J. Biol. Chem. 269, 2550-2561). Nanoparticle-based delivery of Cas9 ribonucleoprotein and donor DNA for induction of homology-directed DNA repair is described, for example, in Lee et al. (2017), Nature Biomedical Engineering, 1:889-90.

[0133] Protein transduction defines the internalization of proteins from the external environment into cells (Ford et al. (2001), Gene Therapy, 8:1-4). This method relies on the unique ability of a small number of proteins and peptides (preferably 10-16 amino acids in length) to penetrate cell membranes. The transduction properties of these molecules can be conferred upon proteins expressed as fusions, thus providing an alternative to gene therapy, for example, for delivering therapeutic proteins into target cells. Commonly used proteins or peptides capable of penetrating cell membranes are, for example, the antennapedia peptide, herpes simplex virus VP22 protein, the HIV TAT protein transduction domain, peptides derived from neurotransmitters or hormones, or the 9xArg tag.

[0134] Microinjection and electroporation are well known in the art, and those skilled in the art know how to perform these methods. Microinjection refers to the process of introducing a substance into a single living cell at the microscopic level or at the macroscopic boundary level. Electroporation is a significant increase in the electrical conductivity and permeability of the cell plasma membrane caused by an externally applied electric field. By increasing the permeability, proteins (or peptides or nucleic acid sequences) can be introduced into living cells.

[0135] The RNA-guided DNA endonuclease can be introduced into a cell either as the active enzyme or as a proenzyme, in which case the RNA-guided DNA endonuclease undergoes a biochemical change within the cell (e.g., by a hydrolysis reaction that exposes an activation site or a conformational change that exposes an activation site), thereby converting the proenzyme into the active enzyme.

[0136] Means and methods for introducing nucleic acid molecules and DNA-targeting RNA into cells are also known in the art, and these methods include transducing or transfecting the cells.

[0137] Transduction is the process of introducing foreign DNA into cells using a virus or viral vector. Transduction is a common tool used by molecular biologists to stably introduce foreign genes into the genome of host cells. Generally, a plasmid is constructed in which the gene to be transferred is flanked by viral sequences that are used by viral proteins to recognize and package the viral genome into viral particles. This plasmid is then inserted (usually by transfection) into producer cells along with other plasmids (DNA constructs) carrying the viral genes required for the formation of infectious virions. In these producer cells, the viral proteins expressed by these packaging constructs bind to sequences on the transferred DNA / RNA (depending on the type of viral vector) and insert it into viral particles. For safety reasons, none of the plasmids used contain all of the sequences required for virus formation, so simultaneous transfection of multiple plasmids is required to obtain infectious virions. Furthermore, only the plasmid carrying the transferred sequence contains a signal that allows the genetic material to be packaged into the virion, so that none of the genes encoding viral proteins are packaged. The viruses collected from these cells are then applied to the cells to be transformed. These initial stages of infection mimic those of natural viruses, leading to the expression of the transferred gene and (in the case of lentiviral / retroviral vectors) the insertion of the transferred DNA into the cell genome. However, because the transferred genetic material does not encode any of the viral genes, these infections do not produce new viruses (the viruses are "replication-deficient"). In the present case, transduction can be used to generate cells containing RNA-guided DNA in an expressible form within the genome.

[0138] Transfection is the process of deliberately introducing naked or purified nucleic acids, or purified proteins, or assembled ribonucleoprotein complexes into cells. Transfection is usually a non-viral-based method.

[0139] Transfection can be chemical. Chemical transfection can be divided into several types: transfection using cyclodextrins, polymers, liposomes, or nanoparticles. One of the least expensive methods uses calcium phosphate. HEPES-buffered saline (HeBS), which contains phosphate ions, is combined with a calcium chloride solution containing the DNA to be transfected. When the two are combined, a fine precipitate of positively charged calcium and negatively charged phosphate is formed, binding the DNA to its surface. The precipitate suspension is then added to the cells to be transfected (usually cell cultures grown in monolayers). Through a process that is not fully understood, the cells take up part of the precipitate, and with it the DNA. This process has been the preferred method for identifying many cancer genes. Another method uses highly branched organic compounds, called dendrimers, to bind DNA and transport it into cells. Another method is the use of cationic polymers such as DEAE-dextran or polyethyleneimine (PEI). Negatively charged DNA binds to polycations, and this complex is taken up by cells via endocytosis. Lipofection (or liposome transfection) is a technique used to inject genetic material into cells using liposomes, which are vesicles that can easily merge with the cell membrane because, as described above, both are made of a phospholipid bilayer. Lipofection generally uses positively charged (cationic) lipids (cationic liposomes or cationic mixtures) to form aggregates with negatively charged (anionic) genetic material. This transfection technique performs the same function in terms of intracellular transfer as other biochemical procedures that utilize polymers, DEAE-dextran, calcium phosphate, and electroporation. The efficiency of lipofection can be improved by treating the transfected cells with a mild heat shock. Fugene is a series of widely used proprietary non-liposomal transfection reagents that can directly transfect a wide variety of cells with high efficiency and low toxicity.

[0140] Transfection can also be achieved by non-chemical methods. Electroporation (gene electrotransfer) is a commonly used method in which a transient increase in cell membrane permeability is achieved when cells are exposed to a short pulse of a strong electric field. Cell constriction allows for the delivery of molecules into cells via membrane deformation. Sonoporation uses high-intensity ultrasound to induce pore formation in cell membranes. This pore formation is primarily due to cavitation of gas bubbles interacting with nearby cell membranes, as pore formation is enhanced by the addition of ultrasound contrast agents, which are the source of cavitation nuclei. Phototransfection is a method in which small (approximately 1 μm diameter) pores are transiently created in the cell plasma membrane with a highly focused laser. Protoplast fusion is a technique in which transformed bacterial cells are treated with lysozyme to remove their cell walls. A fusogenic agent (eg, Sendai virus, PEG, electroporation) is then used to fuse the protoplasts carrying the gene of interest with recipient target cells.

[0141] Finally, transfection can be particle-based. A direct approach for transfection is gene guns, in which DNA is linked to nanoparticles of an inert solid (usually gold), which are then "shot" (or particle bombardment) directly into the nucleus of the target cell. Hence, nucleic acids are usually attached to microprojectiles and delivered through membranes at high speed. Magnetofection, or magnet-assisted transfection, is a transfection method that uses the force of a magnet to deliver DNA into target cells. Impalefection is performed by piercing cells with elongated nanostructures, such as carbon nanofibers or silicon nanowires, functionalized with plasmid DNA, and arrays of such nanostructures.

[0142] The method of the ninth aspect of the present invention relates to a method for editing (i.e., "mutating") the nucleotide sequence of a target site in the genome of a cell using the RNA-guided DNA endonuclease of the present invention. This essentially requires three consecutive pre-processing steps: (1) efficient delivery of a gene encoding the RNA-guided DNA endonuclease or the RNA-guided DNA endonuclease itself into the target cell; (2) efficient expression or presence of CRISPR components (DNA-targeting RNA and the RNA-guided DNA endonuclease of the sixth aspect) in the target cell; and (3) targeting of the desired genomic site by the CRISPR ribonucleoprotein complex and repair of the DNA by the cell's own repair pathway. Step (3) is automatically carried out in the cell upon expression of the CRISPR components in the cell whose genome is to be edited.

[0143] Through genome editing, a target site can be inserted, deleted, modified (including single nucleotide polymorphisms (SNPs)), or replaced in the genome of a cell. The target site can be within the coding region of a gene, within an intron of a gene, within a regulatory region of a gene, within a non-coding region between genes, etc. The gene can be a protein-coding gene or an RNA-coding gene. The gene can be any gene of interest.

[0144] In this context, genome editing uses the cell's own repair pathways, including non-homologous end joining (NHEJ) or homology directed recombination (HDR) pathways. Once DNA is cut by an RNA-guided DNA endonuclease, the cell's own DNA repair mechanisms (NHEJ or HDR) make changes to the DNA by adding or deleting pieces of genetic material, or replacing existing segments with customized DNA sequences. Thus, in the CRISPR-Cas system, CRISPR nucleases make double-strand breaks in DNA at sites determined by short (approximately 20 nucleotides) gRNAs, and this break is then repaired by NHEJ or HDR within the cell. Genome editing preferably uses NHEJ. In different embodiments, genome editing preferably uses HDR.

[0145] NHEJ uses various enzymes to directly join the double-stranded DNA ends. In contrast, in HDR, a homologous sequence is used as a template for regenerating the missing DNA sequence at the break. NHEJ is a canonical homology-independent pathway because it involves the alignment of only one to at most a few complementary bases for religation of the two ends, whereas HDR uses a longer sequence homologous strand to repair the DNA damage.

[0146] The natural properties of these pathways form the very basis of genome editing based on RNA-guided DNA endonucleases. NHEJ has been shown to be error-prone and generate mutations at the repair site. Therefore, if a double-strand break (DSB) can be created in a desired gene in multiple samples, it is highly likely that a mutation will occur at that site during part of the process because errors are created by NHEJ infidelity. On the other hand, HDR's dependence on homologous sequences to repair DSBs can be exploited by inserting a desired sequence within a sequence that is homologous to the flanking sequence of the DSB. When used as a template by the HDR system, this desired sequence can lead to the creation of a desired change in the genomic region of interest. Despite the different mechanisms, the concept of HDR-based gene editing is similar to the concept of homologous recombination-based gene targeting. Therefore, if a DSB can be created at a specific location in the genome based on these principles, the cell's own repair system can be used to create the desired mutation.

[0147] The homologous sequence template for HDR is also referred to herein as a "repair template."

[0148] Thus, by modifying the nucleotide sequence at a target site in the genome of a cell according to the ninth aspect of the invention, a gene can be knocked out (by introducing a premature stop codon) or knocked in (via a repair substrate). Similarly, it is possible to alter the expression of a gene by the method of the ninth aspect of the invention. For example, the target site in the genome can be a promoter region, whereby altering the promoter region can increase or decrease expression of the gene controlled via the target promoter region.

[0149] Hence, according to a preferred embodiment of the ninth aspect, the method further comprises introducing a repair substrate into the cells.

[0150] The design and structure of repair templates suitable for HDR are known in the art. HDR can be error-free if the repair template is identical to the original DNA sequence at the double-strand break (DSB), or it can introduce highly specific mutations into DNA. The three central steps of the HDR pathway are: (1) the 5'-terminal DNA strand is excised at the break, creating a 3' overhang, which serves as both a substrate for proteins required for strand invasion and a primer for DNA repair synthesis; (2) the invading strand can then displace one strand of the homologous DNA duplex and pair with the other, resulting in the formation of a hybrid DNA called a displacement loop (D-loop); and (3) the recombination intermediate can then be degraded to complete the DNA repair process.

[0151] For example, HDR templates used to introduce mutations or insert new nucleotides or nucleotide sequences into genes require a certain amount of homology surrounding the target sequence to be modified. Homologous arms starting at a CRISPR-induced DSB can be used. Generally, the insertion site of the modification should be very close to the DSB, ideally less than 10 bp if possible. One important point to note is that once a DSB is introduced and repaired, the CRISPR enzyme can continue to cut DNA. As long as the gRNA target site / PAM site remains intact, the CRISPR nuclease continues to cut and repair DNA. This repeated editing can be problematic when very specific mutations or sequences are to be introduced into the gene of interest. To address this, repair templates can be designed in a way that ultimately blocks further CRISPR nuclease targeting after the initial DSB is repaired. Two common methods for blocking further editing are to mutate the PAM sequence or the gRNA seed sequence. When designing repair templates, the size of the intended edit should be considered. ssDNA templates (also called ssODNs) are typically used for smaller modifications. Small insertions / edits may require as few as 30–50 bases per homologous arm, and the actual optimal number may vary based on the gene of interest. Homologous arms of 50–80 bases are typically used. For example, Richardson et al. (2016). Nat Biotechnol. 34(3):339-44) found that asymmetric homologous arms (36 bases distal to the PAM and 91 bases proximal to the PAM) support HDR efficiencies of up to 60%. Due to the potential difficulties associated with creating ssODNs longer than 200 bases, it is preferable to use dsDNA plasmid repair templates for larger insertions, such as fluorescent proteins or selection cassettes, into the gene of interest. These templates can have homologous arms of at least 800 bp.To increase the frequency of HDR editing based on a plasmid repair template, a self-cleaving plasmid containing the template and flanking gRNA target sites can be used. In the presence of CRISPR nuclease and the appropriate gRNA, the template is released from the vector. To avoid plasmid cloning, a long dsDNA template generated by PCR can be used. Furthermore, Quadros et al. (2017) Genome Biol. 17;18(1):92) developed Easi-CRISPR, a technology that can perform large mutations and takes advantage of the advantages of ssODNs. To create ssODNs longer than 200 bases, RNA encoding the repair template is transcribed in vitro, followed by the creation of complementary ssDNA using reverse transcriptase. Easi-CRISPR functions well in mouse knock-in models, increasing editing efficiency from 1–10% with dsDNA to 25–50% with ssODNs. While HDR efficiency varies across loci and experimental systems, ssODN templates generally provide the highest frequency of HDR editing.

[0152] According to a preferred embodiment of the ninth aspect, said cell is not the natural host of the gene encoding said RNA-guided DNA endonuclease.

[0153] As discussed herein above, the RNA-guided DNA endonucleases of SEQ ID NOs: 1, 3, and 29 were developed and optimized using various protein engineering strategies, which means that SEQ ID NOs: 1, 3, and 29 are non-naturally occurring sequences that have no natural host.

[0154] Hence, the known cells are not natural hosts for SEQ ID NOs: 1, 3 and 29.

[0155] According to another preferred embodiment of the ninth aspect, the cell is a eukaryotic cell, preferably a yeast cell, a plant cell, or an animal cell.

[0156] Eukaryotic cells, plant cells and animal cells, as well as the eukaryotes, plants and animals from which the cells may be obtained, including preferred examples of these cells, are described herein above in relation to the third and fourth aspects of the invention.

[0157] These cells may also be used in connection with the ninth aspect of the invention.

[0158] According to a more preferred embodiment of the ninth aspect, the method further comprises: culturing plant or animal cells under conditions in which the RNA-guided DNA endonuclease is expressed to cleave the nucleotide sequence at the target site, thereby producing a modified nucleotide sequence, to produce a plant or animal; and and selecting a plant or animal containing the modified nucleotide sequence.

[0159] In this context, the cells into which the components of the CRISPR-Cas system are to be introduced must be totipotent cells capable of developing into complete plants or animals, or germline cells (oocytes and / or sperm) or stem cell populations. Means and methods for the culture of such cells to generate plants or animals are known in the art (see, e.g., https: / / www.stembook.org / node / 720).

[0160] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In case of conflict, the present patent specification, including definitions, will control.

[0161] In a tenth aspect, the present invention relates to modified cells produced by the method according to the ninth aspect of the present invention for use in treating a disease in a subject. The modified cells are preferably modified T lymphocytes, and the disease to be treated is preferably cancer (Stadtmauer et al., Science 28 Feb 2020: Vol. 367, Issue 6481, eaba7365).

[0162] The cells to be modified by the method of the ninth aspect of the invention are preferably obtained from the subject to be treated and the modified cells are then used in accordance with the tenth aspect of the invention.

[0163] In this specification, particularly with respect to embodiments characterized in the claims, each embodiment mentioned in a dependent claim is intended to be combined with each embodiment of each claim (independent or dependent) from which the dependent claim depends. For example, in the case of independent claim 1 reciting alternatives A, B, and C, dependent claim 2 reciting alternatives D, E, and F, and claim 3 dependent on claims 1 and 2 and reciting alternatives G, H, and I, it will be understood that the specification, unless specifically stated otherwise, expressly discloses embodiments corresponding to the combinations A,D,G; A,D,H; A,D,I; A,E,G; A,E,H; A,E,I; A,F,G; A,F,H; A,F,I; B,D,G; B,D,H; B,D,I; B,E,G; B,E,H; B,E,I; B,F,G; B,F,H; B,F,I; C,D,G; C,D,H; C,D,I; C,E,G; C,E,H; C,E,I; C,F,G; C,F,H; C,F,I.

[0164] Similarly, even if an independent and / or dependent claim does not recite alternatives, it is understood that the dependent claim refers to more than one preceding claim, and any combination of the subject matter covered thereby is considered to be explicitly disclosed. For example, in the case of independent claim 1, dependent claim 2 referring to claim 1, and dependent claim 3 referring to both claims 2 and 1, the combination of the subject matter of claim 3 and claim 1 is understood to be clearly and unambiguously disclosed, as is the combination of the subject matter of claims 3, 2, and 1. If there is a further dependent claim 4 pointing to any one of claims 1 to 3, the combination of the subject matter of claims 4 and 1; claims 4, 2, and 1; claims 4, 3, and 1; and claims 4, 3, 2, and 1 is understood to be clearly and unambiguously disclosed. [Example]

[0165] The examples illustrate the invention.

[0166] Example 1: Identification and Engineering of BEC Family Nucleases Metagenomic sequences with the potential to function as novel genome-editing nucleases were identified in silico in various habitats sequenced in the laboratory (Burstein et al., Nature (2017) 542, 237-241). None of these sequences demonstrated sufficient basic DNA targeting efficiency for genome editing, so random shuffling of related sequences was performed (Coco et al., Nat Biotechnol (2001) 19, 354-359). Randomly generated chimeric sequences were further optimized by random mutagenesis (McCullum et al., Methods Mol Biol. (2010) 634, 103-9). In a final step, numerous mutagenized chimeric sequences were screened to assess their DNA targeting activity.

[0167] Using this random and non-rational approach, we successfully identified three sequences (BEC85, BEC67, and BEC10) that exhibit strong DNA targeting activity, potentially sufficient for genome editing approaches. Surprisingly, despite the random approach used, all three identified and designed amino acid sequences share approximately 95% sequence identity with each other. Based on this sequence identity and the unique DNA targeting mechanisms of the three sequences (see Example 3), the three sequences are herein classified as a new family of CRISPR nucleases (BEC family: BRAIN Engineered Cas proteins).

[0168] Example 2: Construction of a functional genome editing system containing BEC family, Cms1 family, and SpCas9 nucleases 2.1 CRISPR / BEC and Cmx1 vector systems for genome editing in S. cerevisiae S288c The genetic elements required for constitutive expression of the novel CRISPR nucleases of the present invention, BEC85, BEC67, and BEC10, and two known Cms1 family CRISPR nucleases, SuCms1 (Begemann et al. (2017), bioRxiv) and SEQ ID NO: 63 (WO 2019 / 030695), as well as the genetic elements required for guide RNA (gRNA) transcription, are provided in all-in-one CRISPR / BEC85, CRISPR / BEC67, CRISPR / BEC10, CRISPR / SuCms1, or CRISPR / SEQ ID NO: 63 vector systems.

[0169] The construction of the CRISPR / BEC10 vector system is described below. The CRISPR / BEC85 and CRISPR / BEC67, CRISPR / SuCms1 and CRISPR / SEQ ID NO: 63 vector systems were constructed in a similar approach.

[0170] (Design of BEC10 protein expression cassette) The 3696-bp synthetic BEC10 nucleotide sequence was codon-optimized for expression in S. cerevisiae S288c (SEQ ID NO: 30) using bioinformatics applications provided by the gene synthesis provider GeneArt (ThermoFisher Scientific, Regensburg, Germany). Additionally, the DNA nuclease coding sequence was 5'-extended by a sequence encoding the SV40 nuclear localization signal (NLS) SEQ ID NO: 11 (Kalderon et al., Cell 39 (1984), 499-509). For protein expression, the resulting 3723-bp synthetic gene was fused to the constitutive S. cerevisiae S288c Tpi1 promoter (SEQ ID NO: 12) and S. cerevisiae S288c Cps1 terminator (SEQ ID NO: 13). The final BEC10 protein expression cassette was inserted by Gibson Assembly Cloning (NEB, Frankuft, Germany) into an E. coli / S. cerevisiae shuttle vector containing all the genetic elements necessary for episomal propagation and selection of recombinant E. coli and recombinant S. cerevisiae cells.

[0171] For propagation and selection of the vector in recombinant E. coli cells, this plasmid contained a high-copy ColE1 origin of replication from pUC under the control of the Em7 synthetic promoter (SEQ ID NO: 14) and the kanMX marker gene, which confers kanamycin resistance. The CEN6 centromere (SEQ ID NO: 15) from S. cerevisiae S288c allowed episomal replication of the shuttle plasmid in S. cerevisiae cells. For selection of transformed S. cerevisiae cells, a bifunctional bacterial / yeast promoter construct upstream to the kanMX marker gene (SEQ ID NO: 16) contained the S. cerevisiae S288c Tef1 promoter sequence (SEQ ID NO: 17).

[0172] (Design of guide RNA (gRNA) expression cassette) Expression of a chimeric gRNA for specific Ade2 gene targeting by BEC10 DNA nuclease was driven by the SNR52 RNA polymerase III promoter (SEQ ID NO: 18) with a SUP4 terminator sequence (SEQ ID NO: 19) (DiCarlo et al., NAR (2013), 41, 4336-4343). The chimeric gRNA consisted of a 19-bp constant BEC family stem-loop sequence (SEQ ID NO: 9 or SEQ ID NO: 10; both stem-loop sequences are interchangeable among all three BEC family nucleases, leading to similar results) fused to an Ade2 target-specific 24-bp spacer sequence (SEQ ID NO: 20). The target spacer sequence was identified in the S. cerevisiae S288c Ade2 gene downstream of the nuclease BEC10-specific PAM-driven 5'-TTTA-3'.

[0173] The complete RNA expression cassette, consisting of the SNR52 RNA polymerase III promoter, the designed chimeric gRNA, and the SUP4 terminator sequence, was obtained as a synthetic gene fragment from GeneArt (Thermo Fisher Scientific, Regensburg, Germany).

[0174] The construction of the all-in-one CRISPR / BEC10 vector system was completed by cloning the synthetic RNA expression cassette into a prepared E. coli / S. cerevisiae shuttle vector containing the BEC10 DNA nuclease expression cassette. The final CRISPR / BEC10 vector system was constructed by Gibson Assembly Cloning (NEB, Frankfurt, Germany).

[0175] The identity of the cloned DNA elements was confirmed by Sanger sequencing at LGC Genomics (Berlin, Germany).

[0176] (CRISPR / BEC10 all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / BEC10 vector system is provided as SEQ ID NO:31.

[0177] (CRISPR / BEC85 all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / BEC85 vector system is provided as SEQ ID NO:21.

[0178] (CRISPR / BEC67 all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / BEC67 vector system is provided as SEQ ID NO:22.

[0179] (CRISPR / SuCms1 all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / SuCms1 vector system is provided as SEQ ID NO:32.

[0180] (CRISPR / SEQ ID NO: 63 all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / SEQ ID NO:63 vector system is provided as SEQ ID NO:33.

[0181] 2.2 Design of Homology-Directed Repair Template (HDR Template) An 838-bp Ade2 BEC85, BEC67, BEC10, SuCms1, and SEQ ID NO: 63 HDR template was designed to generate a 29-bp site-specific deletion in the S. cerevisiae S288c Ade2 chromosomal gene by homologous recombination. Within the HDR template, the introduced Ade2 gene deletion was flanked by 407-bp and 429-bp sequences homologous to the chromosomal target region. Additionally, the HDR fragment created a new recognition sequence for the restriction endonuclease EcoRI at the deleted Ade2 genomic site. Successful recombination events mediated by the designed HDR template eliminated the previously described PAM and protospacer regions (SEQ ID NO: 20) in the chromosomal Ade2 gene, preventing the programmed gRNA / BEC85, BEC67, or BEC10 DNA nuclease complex from retargeting the S. cerevisiae S288c genome. Furthermore, the introduction of gene deletions resulted in Ade2 mutant clones, which were easily recognized by the red color of the colonies because the adenine-depleted mutant cells accumulated red purine precursors in their vacuoles (Ugolini et al., Curr Genet (2006), 485-92).

[0182] The complete sequences of the Ade2 HDR templates for BEC85, BEC67, BEC10, SuCms1, and SEQ ID NO:63 are provided as SEQ ID NO:23.

[0183] 2.3 CRISPR / SpCas9 vector system for genome editing in S. cerevisiae S288c The genetic elements required for constitutive expression of SpCas9 (S. pyogenes Cas9) DNA nuclease and for single guide RNA transcription were provided in an all-in-one CRISPR / SpCas9 vector system.

[0184] (Design of SpCas9 protein expression cassette) DNA synthesis of the SpCas9 coding sequence, codon-optimized for expression in S. cerevisiae S288c, based on the published SpCas9 nucleotide sequence from Streptococcus pyogenes (Deltcheva et al., Nature 471 (2011), 602-607), was ordered from GeneArt (Thermo Fisher Scientific, Regensburg, Germany) (SEQ ID NO: 24). For nuclear translocation, the SpCas9 DNA nuclease coding sequence was 5'-extended by a sequence encoding the SV40 nuclear localization signal (NLS) (SEQ ID NO: 11). The resulting 4134-bp synthetic SpCas9 gene was fused to the constitutive S. cerevisiae S288c Tpi1 promoter (SEQ ID NO: 12) and S. cerevisiae S288c Cps1 terminator (SEQ ID NO: 13), following the protein expression strategy described for the BEC10 DNA nuclease. The final SpCas9 protein expression cassette was inserted by Gibson Assembly Cloning (NEB, Frankfurt, Germany) into an E. coli / S. cerevisiae shuttle vector carrying the same genetic elements for propagation and selection as previously described for the CRISPR / BEC10 vector system.

[0185] (Design of guide RNA expression (gRNA) cassette) Expression of a chimeric gRNA for specific Ade2 gene targeting by SpCas9 DNA nuclease was driven by an SNR52 RNA polymerase III promoter (SEQ ID NO: 18) with a SUP4 terminator sequence (SEQ ID NO: 19). The chimeric guide RNA consisted of an Ade2 target-specific 20-bp spacer sequence (SEQ ID NO: 25) fused to a 76-bp SpCas9-specific sgRNA sequence (SEQ ID NO: 26). The target spacer sequence was identified in the S. cerevisiae S288c Ade2 gene upstream of the nuclease SpCas9-specific PAM-driven 5'-AGG-3'.

[0186] The complete RNA expression cassette, consisting of the SNS52 RNA polymerase III promoter, the designed chimeric guide RNA, and the SUP4 terminator sequence, was obtained as a synthetic DNA fragment from GeneArt (Thermo Fisher Scientific, Regensburg, Germany).

[0187] To generate the final CRISPR / SpCas9 vector system, the synthetic RNA transcription cassette was cloned by Gibson Assembly Cloning (NEB, Frankfurt, Germany) into a prepared E. coli / S. cerevisiae shuttle vector containing the SpCas9 DNA nuclease expression cassette. The identity of all cloned DNA elements was confirmed by Sanger sequencing at LGC Genomics (Berlin, Germany).

[0188] (CRISPR / SpCas9 all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / SpCas9 vector system is provided as SEQ ID NO:27.

[0189] 2.4 Design of homology-directed repair templates An 832-bp synthetic Ade2 SpCas9 HDR template was designed to generate a 26-bp site-specific deletion in the chromosomal S. cerevisiae S288c Ade2 gene by homologous recombination. Within the HDR template, the introduced Ade2 gene deletion was flanked by 402-bp and 428-bp sequences homologous to the chromosomal target region. Successful recombination events mediated by the designed HDR template eliminated the previously described PAM and protospacer regions (SEQ ID NO: 25) in the chromosomal Ade2 gene, preventing the programmed gRNA / SpCas9 DNA nuclease complex from retargeting the S. cerevisiae S288c genome. Furthermore, introduction of the gene deletion resulted in Ade2 mutant clones, which were easily recognized by the red color of the colonies because adenine-depleted mutant cells accumulate red purine precursors in their vacuoles (Ugolini et al., Curr Genet (2006), 485-92).

[0190] The nucleotide sequence of the 832 bp Ade2 HDR template is provided as SEQ ID NO:28.

[0191] 2.5 Cultivation and transformation of Saccharomyces cerevisiae (Transformation of competent S. cerevisiae S288c cells) Preparation and transformation of S. cerevisiae S288c cells were performed as described by Gietz & Schiestl, Nature Protocols (2007), 2, 31-34. Briefly, a single colony of S. cerevisiae S288c was inoculated into 25 ml of 2x YPD medium and incubated at 30°C for 14-16 hours on a horizontal shaker at 200 rpm. An overnight-grown preculture was diluted into 250 ml of fresh 2x YPD medium to an optical density at 600 nm (OD600) of 0.5. The inoculated medium was incubated at 30°C on a horizontal shaker at 200 rpm until the culture reached an OD600 optical density of 2.0-8.0. Cells were transferred into five 50 ml conical tubes and harvested by centrifugation at 3000 x g for 5 minutes. Pelleted cells from a 250 ml culture were resuspended in 125 ml of water and centrifuged at 3000 × g for 5 minutes. The pelleted cells were resuspended in 2.5 ml of water. After an additional centrifugation step at 3000 × g for 5 minutes, the pelleted cells were finally resuspended in 2.5 ml of "frozen competent cell solution" (5% v / v glycerol and 10% v / v DMSO). 50 μl aliquots of competent cells were stored at -80°C until use. For the transformation procedure, aliquots of competent cells were thawed at 37°C for 30 seconds and then centrifuged at 11,600 × g for 2 minutes. The supernatant was removed, and the cell pellet was resuspended in 360 μl of transformation mixture consisting of 1 μg of pScCEN plasmid derivative and 500 ng of HDR template provided in 14 μl of water, 260 μl of 50% w / v PEG3350, 36 μl of 1 M Li-acetate, and 50 μl of single-stranded carrier DNA. The prepared cells were heat-shocked at 42°C for 45 minutes with mixing every 15 minutes. After the heat shock step, the transformed cells were pelleted by centrifugation at 13,000 × g for 30 seconds, and the supernatant was removed. For recovery, the cell pellet was resuspended in 1 ml of YPD. The cell suspension was transferred to a 5 ml tube and incubated at 30°C for 3 hours on a horizontal shaker at 200 rpm.Finally, the transformed cells were plated onto selective agar plates containing 50 μg / ml geneticin (G418) and incubated at 30° C. for at least 2 days.

[0192] 2.6 CRISPR / BEC E. coli and Cms1 E. coli vector systems for genome editing in E. coli BW25113 The genetic elements required for constitutive expression of BEC10, SuCms1, or SEQ ID NO:63 CRISPR nuclease and guide RNA (gRNA) transcription were prepared as an all-in-one CRISPR / BEC10_E. coli (SEQ ID NO:34) vector system, an all-in-one CRISPR / SuCms1_E. coli (SEQ ID NO:35) vector system, or an all-in-one CRISPR / SEQ ID NO:63_E. coli (SEQ ID NO:36) vector system.

[0193] The construction of the CRISPR / BEC10_E. coli vector system is described below. The CRISPR / SuCms1_Coli and CRISPR / SEQ ID NO:63_Coli vector systems were constructed in a similar approach.

[0194] (Design of BEC10_Coli protein expression cassette) The 3696-bp synthetic BEC10 nucleotide sequence was codon-optimized for expression in E. coli BW25113 (SEQ ID NO: 37) using bioinformatics applications provided by the gene synthesis provider GeneArt (Thermo Fisher Scientific, Regensburg, Germany). For protein expression, the resulting synthetic gene was fused to the inducible araBAD promoter (SEQ ID NO: 38) and fdt terminator (SEQ ID NO: 39). The final BEC10_E. coli protein expression cassette was inserted by Gibson Assembly Cloning (NEB, Frankfurt, Germany) into an E. coli shuttle vector containing all the genetic elements necessary for episomal propagation and selection of recombinant E. coli cells.

[0195] (Design of guide RNA (gRNA) expression cassette) Expression of a chimeric gRNA for specific rpoB gene targeting by BEC10 DNA nuclease was driven by a SacB RNA polymerase III promoter (SEQ ID NO: 40) and terminated using an rrnB terminator sequence (SEQ ID NO: 41). The chimeric gRNA consisted of a 19-bp constant BEC family stem-loop sequence (SEQ ID NO: 9 or SEQ ID NO: 10; both stem-loop sequences are interchangeable among all three BEC family nucleases, leading to similar results) fused to an rpoB target-specific 24-bp spacer sequence (SEQ ID NO: 42). The target spacer sequence was identified in the E. coli BW25113 rpoB gene downstream to the nuclease BEC10-specific PAM-driven 5'-TTTA-3'.

[0196] The complete RNA expression cassette, consisting of the SacB RNA polymerase III promoter, the designed chimeric gRNA, and the rrnB terminator sequence, was obtained as a synthetic gene fragment from GeneArt (Thermo Fisher Scientific, Regensburg, Germany).

[0197] The construction of the all-in-one CRISPR / BEC10_E. coli vector system was completed by cloning the synthetic RNA expression cassette into a prepared E. coli shuttle vector containing the BEC10_E. coli DNA nuclease expression cassette. The final CRISPR / BEC10_E. coli vector system was constructed by Gibson Assembly Cloning (NEB, Frankfurt, Germany).

[0198] The identity of all cloned DNA elements was confirmed by Sanger sequencing at LGC Genomics (Berlin, Germany).

[0199] (CRISPR / BEC10_E.coli all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / BEC10_Coli vector system is provided as SEQ ID NO:34.

[0200] (CRISPR / SuCms1_E.coli all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / SuCms1_Coli vector system is provided as SEQ ID NO:35.

[0201] (CRISPR / SEQ ID NO: 63 E. coli all-in-one vector system) The complete nucleotide sequence of the constructed CRISPR / SuCm1_Coli vector system is provided as SEQ ID NO:36.

[0202] 2.7 E. coli Cultivation and Transformation (Transformation of competent E. coli BW25113 cells) Briefly, a single colony of E. coli BW25113 was inoculated into 5 ml of LB-Kan medium and incubated at 37°C for 12–14 hours on a horizontal shaker at 200 rpm. An overnight-grown preculture was diluted into 60 ml of fresh LB medium to an optical density at 600 nm (OD600) of 0.06. The inoculated medium was incubated at 30°C on a horizontal shaker at 200 rpm until the culture reached an optical density of OD600 of 0.2. 20% arabinose was added, and the cells were incubated at 30°C and 200 rpm until the culture reached an optical density of OD600 of 0.5. Cells were transferred into five 50 ml conical tubes and harvested by centrifugation at 4000 × g for 5 minutes at 4°C. Pelleted cells from the 50 ml culture were resuspended in 60 ml of water and centrifuged at 4000 × g for 5 minutes at 4°C.

[0203] A washing step was performed, and the cells were centrifuged at 4000 x g for 5 minutes at 4°C and then resuspended in 30 ml of 10% glycerol. In a second washing step, the cells were centrifuged at 4000 x g for 5 minutes at 4°C and then resuspended in 6 ml of 10% glycerol. In a final step, the cells were resuspended in 150 μl of 10% glycerol. A 25 μl aliquot of competent cells was stored at -80°C until use. For the transformation procedure, an aliquot of competent cells was thawed and 50 ng of plasmid DNA was added. The prepared cells were electroporated using 1800 V, 25 μF, 200 ohms for 5 milliseconds. Then, 975 μL of NEB® 10 Beta / Stable Exgrowth Medium was added, and 100 μL of the suspension was plated onto selective agar plates.

[0204] (2.8 DNA Technology) Plasmid isolation, enzymatic DNA manipulation, and agarose gel electrophoresis were performed according to standard procedures. The Thermo Fisher Scientific Phusion Flash High-Fidelity PCR System (Thermo Fisher Scientific, Regensburg, Germany) was used for PCR amplification. All oligonucleotides used in this work were synthesized by biomers.net (Ulm, Germany) or Eurofins Scientific (Ebersberg, Germany). The DNA Clean and Concentrator Kit and ZymoClean Gel DNA Recovery Kit (Zymo Research, Freiburg, Germany) were used for purification from the agarose and enzymatic reactions. The identity of all cloned DNA fragments was confirmed by Sanger sequencing at LGC Genomics (Berlin, Germany).

[0205] Purified genomic DNA from S. cerevisiae S288c cells was isolated using Zymo Research's YeaStar Genomics DNA Kit (Zymo Research, Freiburg, Germany) according to the manufacturer's instructions. Zymolyase digestion of yeast cell walls was performed at 37°C for 60 minutes, and purified genomic DNA was eluted in 60 μl of 5 mM Tris / HCl pH 8.5.

[0206] Example 3: Functional Characterization of BEC85, BEC67, and BEC10 Compared to SpCas9 in Saccharomyces cerevisiae (S. cerevisiae) 3.1. Experimental Setup In this example, the Ade2 gene was knocked out in S. cerevisiae S288C using the CRISPR / BEC85 (SEQ ID NO: 21), CRISPR / BEC67 (SEQ ID NO: 22), or CRISPR / BEC10 (SEQ ID NO: 31) vector system and the corresponding homology-directed repair template (SEQ ID NO: 23). In comparison to the experiments performed with BEC85, BEC67, or BEC10, similar experiments were performed using the CRISPR / SpCas9 construct (SEQ ID NO: 27) and the corresponding homology-directed repair template (SEQ ID NO: 28) to demonstrate the functionality of the BEC-type CRISPR nuclease.

[0207] Ade2 is a nonessential gene in Saccharomyces cerevisiae, and knockout of this gene results in colonies with a red phenotype due to the accumulation of red purine precursors in the vacuole (Ugolini et al., Curr Genet (2006), 485-92). Because of this easy readout, Ade2 knockouts can be used as a screening system to monitor the ability of CRISPR Cas proteins to function as genome editing tools.

[0208] In this approach, we used CRISPR Cas-directed introduction of a homology-directed repair template, which led to a site-specific deletion that removed the PAM and spacer sequences. Furthermore, a frameshift mutation was introduced by the homology-directed repair template, leading to knockout of the Ade2 gene, to visualize the DNA cleavage activity of BEC85, BEC67, and BEC10 compared to SpCas9, the most commonly used Cas protein in science and medicine.

[0209] The Ade2 knockout strategy used in S. cerevisiae S288c for BEC85, BEC67, BEC10, and SpCas9 is shown schematically in Figure 1.

[0210] Briefly, expression constructs of CRISPR / BEC85, CRISPR / BEC67, CRISPR / BEC10, or CRISPR / SpCas9 and the corresponding homology-directed repair templates were transformed into S. cerevisiae S288c cells and plated as described in Example 2.5.

[0211] In parallel, negative control experiments were performed using CRISPR / BEC85, CRISPR / BEC67, CRISPR / BEC10, or CRISPR / SpCas9 expression constructs that deleted the spacer sequence targeting the Ade2 gene, demonstrating the dependency of Cas proteins to be guided to the target DNA region by the specific spacer.

[0212] After transformation and 48 hours of incubation at 30°C, the culture plates were analyzed by counting the number of colonies that grew and by assessing their phenotype (red or white).

[0213] (3.2 Results) The results are summarized in Table 1 and an exemplary plate is shown in FIG.

[0214] All experiments were performed in five biological replicates, and the results from these replicates were combined to visualize the genome editing efficiency of BEC-type CRISPR nucleases.

[0215] [Table 1]

[0216] (CRISPR / SpCas9) Cells transformed with the negative control construct (CRISPR / SpCas9 (spacer-free) + homology-directed repair template) showed 5831 white colonies and 11 red colonies, demonstrating that the SpCas9 protein did not target the DNA of the Ade2 gene due to the missing spacer sequence. Therefore, 99.8% of the colonies exhibited a wild-type phenotype (white). Furthermore, 11 colonies exhibited a knockout phenotype (red) due to a natural homologous recombination event in which the homology-directed repair template integrated into the Ade2 locus.

[0217] In contrast, the active construct (CRISPR / SpCas9 (containing a spacer targeting the Ade2 gene) + homology-directed repair template) led to 1182 white colonies and 2575 red colonies, thus demonstrating the molecular mechanism and efficacy of SpCas9, which contained 68% of edited colonies compared to the negative control, in which only 0.2% of colonies were edited.

[0218] (CRISPR / BEC85, CRISPR / BEC67, and CRISPR / BEC10) Surprisingly, the same experimental setup using BEC85, BEC67, or BEC10 sequences led to completely different results compared to SpCas9.

[0219] Cells transformed with the negative control construct (CRISPR / BEC10 (spacer-free) + homology-directed repair template) showed 8021 white colonies and 14 red colonies, demonstrating that the BEC10 protein did not target the DNA of the Ade2 gene due to the missing spacer sequence. Therefore, 99.8% of the colonies exhibited a wild-type phenotype (white). Furthermore, 14 colonies exhibited a knockout phenotype (red) due to natural homologous recombination events in which the homology-directed repair template integrated into the Ade2 locus. Similar results were obtained using the BEC85 (6643 wild-type (white) colonies and 14 knockout (red) colonies) and BEC67 (9136 wild-type (white) colonies and 16 knockout (red) colonies) negative control constructs.

[0220] In contrast, the active BEC10 construct (CRISPR / BEC10 (containing a spacer targeting the Ade2 gene) + homology-directed repair template) led to a significant overall reduction in visible colonies (174) compared to the negative control (8035) and also compared to the active SpCas9 approach (3757). However, 155 of these 174 colonies displayed an Ade2 knockout phenotype (red), leading to an editing efficiency of 89%. Similar results were observed using the active BEC85 or BEC67 constructs.

[0221] BEC85: A significant colony reduction to 91 colonies, consisting of 82 red colonies and 9 white colonies, leading to an editing efficiency of 90%

[0222] BEC67: Significant colony reduction to 68 colonies, consisting of 45 red colonies and 21 white colonies, leading to an editing efficiency of 68%

[0223] Taken together, the results obtained using the experimental setup with SpCas9, BEC85, BEC67, and BEC10 surprisingly demonstrated a completely different molecular genome editing mechanism of BEC-type CRISPR nucleases compared to conventional CRISPR Cas nucleases. In contrast to SpCas9, which supports homologous recombination by introducing RNA-directed double-strand breaks, BEC85-, BEC67-, and BEC10-mediated editing resulted in a strong overall clonal reduction associated with a significant enrichment for cells that successfully achieved homologous recombination.

[0224] Although BEC85, BEC67, and BEC10 exhibit novel molecular mechanisms, the results obtained in this example demonstrate the ability of BEC-type CRISPR nucleases to function as novel genome editing tools through site-specific and highly efficient homology-directed recombination.

[0225] Example 4: Evaluation of genome editing activity and efficiency of BEC family nucleases in comparison with flanking sequences SuCms1 and SEQ ID NO: 63 Example 4 demonstrates that the novel BEC family nuclease of the present invention is superior to its closest known member, SuCms1 (Begemann et al. (2017), bioRxiv) and SEQ ID NO: 63 (WO 2019 / 030695), based on comparative experiments.

[0226] 4.1. Experimental Setup In this example, the Ade2 gene was knocked out in S. cerevisiae S288C using the CRISPR / BEC10 (SEQ ID NO: 31), CRISPR / SuCms1 (SEQ ID NO: 32), or CRISPR / SEQ ID NO: 63 (SEQ ID NO: 33) vector systems and the corresponding homology-directed repair template (SEQ ID NO: 23). This example directly compares the genome editing efficiency of BEC family nucleases with the flanking sequences SuCms1 (Begemann et al. (2017), bioRxiv) and SEQ ID NO: 63 (WO 2019 / 030695).

[0227] The experiment was carried out as described in section 3.1 of the Examples above.

[0228] (4.2 Results) The results are summarized in Table 2 and an exemplary plate is shown in FIG.

[0229] All experiments were performed in five biological replicates, and the results from these replicates were combined to visualize the genome editing efficiency of BEC10 in comparison with the prior art nucleases SuCms1 and SEQ ID NO: 63.

[0230] [Table 2]

[0231] (CRISPR / SuCms1) Cells transformed with the active construct (CRISPR / SuCms1 + homology-directed repair template) showed 623 white colonies, 19 red colonies, and 14 orange colonies (orange colonies are marked with arrows in Figure 3), leading to an editing efficiency of 5% (if orange colonies are counted as successfully edited cells). However, further analysis of the orange colonies showed that these clones contained a mixture of successfully edited and unedited (i.e., wild-type) cells, leading to an editing efficiency of only 3% for fully edited colonies.

[0232] (CRISPR / SEQ ID NO: 63) Cells transformed with the active construct (CRISPR / SEQ ID NO: 63 + homology-directed repair template) showed 5231 white colonies and 8 red colonies, leading to an editing efficiency of only 0.2%. The total colony number and editing efficiency were similar to the negative control results shown in Example 3, demonstrating that SEQ ID NO: 63 does not exhibit any nuclease activity.

[0233] (CRISPR / BEC10) Cells transformed with the active construct (CRISPR / BEC10 + homology-directed repair template) showed 11 white and 59 red colonies, leading to a very high editing efficiency of 84%, which is similar to the editing efficiencies obtained in Example 3 for BEC85, BEC65, and BEC10.

[0234] (summary) The results obtained in Example 4 indicate that BEC10 and other BEC family nucleases, as well as BEC85 and BEC67 (with results as described in Example 3), share the same DNA targeting mechanism, and all three have very high and similar editing efficiencies. Furthermore, the BEC family nucleases exhibit significantly stronger colony reduction and significantly better editing efficiency compared to the neighboring sequences SuCms1 and SEQ ID NO: 63. In contrast to the SuCms1 nuclease, which shows an editing efficiency of 5% (also note that of the 33 edited clones, 14 were only partially edited), BEC10 shows an editing efficiency of 84%, BEC85 a 90%, and BEC67 a 68% (see Example 3). Furthermore, SEQ ID NO: 63 does not exhibit any nuclease activity whatsoever.

[0235] Example 5: Evaluation of genome editing efficiency and efficacy of BEC family nucleases in comparison with flanking sequences SuCms1 and SEQ ID NO: 63 at different temperatures (21°C and 37°C) For many biotechnological and pharmaceutical applications, experiments must be performed at specific temperatures to meet the requirements of the organism used and to ensure optimal behavior and reproducible results. The optimum temperature for most organisms used in biotechnological, agricultural, and pharmaceutical applications is between 21°C and 37°C. To demonstrate the behavior of the BEC nuclease of the present invention in this temperature range, experiments were performed using S. cerevisiae (21°C) and E. coli (37°C) in comparison with the flanking sequences SuCms1 (Begemann et al. (2017), bioRxiv) and SEQ ID NO: 63 (WO 2019 / 030695).

[0236] 5.1 Experimental Setup (S. cerevisiae 21°C) In this example, the Ade2 gene was knocked out in S. cerevisiae S 2888C using CRISPR / BEC10 (SEQ ID NO: 31), CRISPR / SuCms1 (SEQ ID NO: 32), or CRISPR / SEQ ID NO: 63 (SEQ ID NO: 33) vector systems and the corresponding homology-directed repair template (SEQ ID NO: 23). Cells were incubated at 21°C to demonstrate the editing efficiency of BEC family nucleases at low temperatures in direct comparison with the flanking sequences SuCms1 and SEQ ID NO: 63.

[0237] Cultivation and transformation of S. cerevisiae was carried out as described in section 2.5 of the Examples, except that the cultivation temperature was varied from 30°C to 21°C.

[0238] The experiments were carried out as described in Section 3.1 of the Examples.

[0239] (5.2 Results) The results are summarized in Table 3 and an exemplary plate is shown in FIG.

[0240] All experiments were performed in five biological replicates, and the results from these replicates were combined to visualize the genome editing efficiency of BEC10 in comparison with the prior art nucleases SuCms1 and SEQ ID NO: 63.

[0241] [Table 3]

[0242] (CRISPR / SuCms1) Cells transformed with the active construct (CRISPR / SuCms1 + homology-directed repair template) showed 8740 white colonies and 28 red colonies, which is just slightly above the editing efficiency of the negative control experiment (0.2%) and is significantly reduced compared to the results obtained at 30°C (Example 4).

[0243] (CRISPR / SEQ ID NO: 63) Cells transformed with the active construct (CRISPR / SEQ ID NO: 63 + homology-directed repair template) showed 10,240 white colonies and 18 red colonies, leading to an editing efficiency of 0.2%. The total colony number and editing efficiency were similar to the negative control results shown in Example 4, demonstrating that SEQ ID NO: 63 does not exhibit any nuclease activity.

[0244] (CRISPR / BEC10) Cells transformed with the active construct (CRISPR / BEC10 + homology-directed repair template) showed 23 white colonies and 42 red colonies, leading to a high editing efficiency of 64%, thereby demonstrating that genome editing with BEC-based CRISPR nucleases, also when used at 21°C, leads to a significant overall reduction in visible colonies and a high genome editing rate.

[0245] (summary) The results obtained in Example 5.2 demonstrate that the BEC10 nuclease from the BEC family exhibits significant overall colony reduction and strong genome editing efficiency (65%) when used at 21°C.

[0246] In contrast, the overall colony reduction and editing efficiency of SuCms1 nuclease was significantly reduced to 0.3% at 21°C, which is just slightly higher than the editing efficiency of the negative control (0.2%) and is unsuitable for functioning as a genome editing tool.

[0247] Furthermore, as already shown at 30°C, SEQ ID NO: 63 does not exhibit any nuclease activity whatsoever.

[0248] 5.3 Experimental Setup (E. coli 37℃) To assess the nuclease activity of BEC family nucleases in comparison with flanking sequences at 37°C, an E. coli assay system was used because of the ideal growth conditions at 37°C.

[0249] To visualize the activity and efficiency of nucleases, a so-called depletion assay is performed, in which the viability of E. coli cells after nuclease targeting is monitored compared to a negative control (lower viability indicates better nuclease activity). Because E. coli cells cannot perform non-homologous end joining (NHEJ), targeting DNA with CRISPR nucleases leads to cell death. Additionally, the essential rpoB gene was targeted in this experimental approach, and knockout of this gene is lethal in E. coli cells.

[0250] For this experimental approach, the CRISPR / BEC10_Coli (SEQ ID NO: 34), CRISPR / SuCms1_Coli (SEQ ID NO: 35), or CRISPR / SEQ ID NO: 63_Coli (SEQ ID NO: 36) vector systems were used to target the rpoB gene in E. coli, and the editing efficiency of BEC family nucleases at high temperature (37°C) was demonstrated in direct comparison with the flanking sequences SuCms1 (Begemann et al. (2017), bioRxiv) and SEQ ID NO: 63 (WO 2019 / 030695).

[0251] In parallel, negative control experiments were performed using expression constructs of CRISPR / BEC10_Coli, CRISPR / SuCms1_Coli, or CRISPR / SEQ ID NO:63_E. coli lacking the spacer sequence targeting the rpoB gene, demonstrating the dependency of the Cas protein to be guided to the target DNA region by a specific spacer.

[0252] After transformation and 48 hours of incubation at 37°C, the culture plates were analyzed by counting the number of colonies that grew.

[0253] (5.4 Results) The results are summarized in Table 4 and an exemplary plate is shown in FIG.

[0254] All experiments were performed in five biological replicates, and the results from these replicates were combined to visualize the genome editing efficiency of BEC10 in comparison with SuCms1 and SEQ ID NO:63.

[0255] [Table 4]

[0256] (CRISPR / SuCms1) Cells transformed with the negative control construct showed 4905 colonies after 48 h of incubation at 37°C, whereas cells transformed with the active construct (CRISPR / SuCms1_Coli) showed 1365 colonies, which leads to a 72% clonal reduction.

[0257] (CRISPR / SEQ ID NO: 63) Cells transformed with the negative control construct showed 5002 colonies after 48 hours of incubation at 37°C, whereas cells transformed with the active construct (CRISPR / SEQ ID NO: 63_Coli) showed 5025 colonies, leading to 0% clonal reduction, demonstrating that SEQ ID NO: 63 does not exhibit any nuclease activity in this experimental approach.

[0258] (CRISPR / BEC10) Cells transformed with the negative control construct showed 4963 colonies after 48 h of incubation at 37°C, whereas cells transformed with the active construct (CRISPR / BEC10_Coli) showed 130 colonies, which translates to a 97% clonal reduction.

[0259] (summary) The results obtained in Example 5.4 demonstrate that when using an E. coli-based depletion assay, BEC10 nuclease showed significant overall colony reduction (97%) at 37°C, thereby demonstrating the very high activity of BEC10 nuclease at higher temperatures. In contrast, SuCms1 nuclease showed a significantly lower reduction in colonies (72%), demonstrating the superior activity of BEC-type nucleases compared to SuCms1 at 37°C.

[0260] Furthermore, SEQ ID NO: 63 does not show any nuclease activity at all with 0% colony reduction compared to the negative control.

[0261] Example 6 - Discussion of Results from Examples 3-5 In summary, the results of Examples 3-5 demonstrate that the sequences of newly identified and developed BEC family nucleases (BEC85, BEC67, and BEC10), which share approximately 95% sequence identity with each other, have similar genome editing efficiencies based on a novel molecular genome editing mechanism compared to Cs9 (Example 3). Furthermore, the results of Example 4 demonstrate that genome editing using BEC family-type nucleases leads to significantly higher clone reduction numbers and significantly better editing ratios compared to the flanking sequences SuCms1 and SEQ ID NO: 63, confirming the overall superiority of BEC-type nucleases for genome editing.

[0262] Most organisms of interest used in biotechnology, agricultural, and pharmaceutical research applications are cultured at temperatures ranging from 21 to 37°C (e.g., various plants and plant cells at approximately 21°C, various yeast and fungal cells at approximately 30°C, and various prokaryotic and mammalian cell lines at approximately 37°C). Therefore, a universally applicable CRISPR system must exhibit strong activity and genome editing efficiency when used within this temperature range. To evaluate the temperature-dependent activity of our newly discovered and developed BEC-type nucleases, experiments using BEC10 nuclease were performed in S. cerevisiae (21°C) and E. coli (37°C) (Example 5) and compared with results obtained using the flanking sequences SuCms1 and SEQ ID NO:63. The results obtained in these experiments demonstrated the robust activity of BEC10 at all temperature levels tested, with superior editing efficiency and colony reduction rates compared to the neighboring sequences SuCms1 (Begemann et al. (2017), bioRxiv) and SEQ ID NO: 63 (WO 2019 / 030695). In addition, the editing efficiency of SuCms1 nuclease significantly decreased at 21°C to a level similar to the negative control (0.3%), whereas BEC10 editing efficiency remained high (65%) even at lower temperatures.

Claims

1. A nucleic acid molecule encoding an RNA-guided DNA endonuclease, (a) a nucleic acid molecule encoding the RNA-guided DNA endonuclease, the nucleic acid molecule comprising or consisting of the amino acid sequence of SEQ ID NO: 29, 1, or 3; (b) a nucleic acid molecule comprising or consisting of the nucleotide sequence of SEQ ID NO: 30, 2 or 4; (c) a nucleic acid molecule encoding an RNA-guided DNA endonuclease, the amino acid sequence of which is at least 93%, most preferably at least 95%, identical to the amino acid sequence of (a); (d) a nucleic acid molecule comprising or consisting of a nucleotide sequence that is at least 93%, most preferably at least 95%, identical to the nucleotide sequence (b); (e) a nucleic acid molecule that is degenerate with respect to the nucleic acid molecule (d); or (f) A nucleic acid molecule corresponding to any one of (a) to (d), wherein T is replaced by U.

2. 10. The nucleic acid molecule of claim 1, wherein the nucleic acid molecule is operably linked to a promoter that is native or heterologous to the nucleic acid molecule.

3. 3. The nucleic acid molecule according to claim 1 or 2, wherein the nucleic acid molecule is codon-optimized for expression in a eukaryotic cell, preferably a plant cell or an animal cell.

4. A vector encoding the nucleic acid molecule of any one of claims 1 to 3.

5. A host cell (excluding totipotent cells or germline cells that develop into a human individual) comprising a nucleic acid molecule according to any one of claims 1 to 3 or a vector according to claim 4 and transformed, transduced or transfected with said vector.

6. The host cell according to claim 5, wherein the host cell is a eukaryotic or prokaryotic cell, preferably a plant or animal cell.

7. 10. A plant, seed or part of a plant that is not a single plant cell, or a non-human animal, which comprises a nucleic acid molecule according to any one of claims 1 to 3, or which comprises a vector according to claim 4 and which is transformed, transduced or transfected with said vector.

8. A method for producing an RNA-guided DNA endonuclease, comprising culturing the host cell of claim 5 or 6 and isolating the RNA-guided DNA endonuclease produced.

9. An RNA-guided DNA endonuclease encoded by the nucleic acid molecule of any one of claims 1 to 3.

10. A composition comprising the nucleic acid molecule of any one of claims 1 to 3, the vector of claim 4, the host cell of claim 5 or 6, the plant, seed, part of a plant, or non-human animal of claim 7, the RNA-guided DNA endonuclease of claim 9, or a combination thereof.

11. The composition of claim 10, which is a pharmaceutical or diagnostic composition.

12. 10. Use of the nucleic acid molecule according to any one of claims 1 to 3, the vector according to claim 4, the host cell according to claim 5 or 6, the plant, seed, plant part or non-human animal according to claim 7, the RNA-guided DNA endonuclease according to claim 9, or a combination thereof, for the manufacture of a composition for treating a disease in a subject or plant by modifying a nucleotide sequence at a target site in the genome of the subject or plant.

13. An agent for modifying a nucleotide sequence of a target site in the genome of a cell, comprising: (i) a DNA-targeting RNA or a DNA polynucleotide encoding a DNA-targeting RNA, wherein the DNA-targeting RNA is: (a) a first segment comprising a nucleotide sequence that is complementary to a sequence in the target DNA; (b) a second segment that interacts with the RNA-guided DNA endonuclease of claim 9; and an RNA targeting the DNA or a DNA polynucleotide encoding the RNA targeting the DNA, comprising: (ii) The RNA-guided DNA endonuclease according to claim 9, or a nucleic acid molecule encoding the RNA-guided DNA endonuclease according to any one of claims 1 to 3, or the vector according to claim 4, wherein the RNA-guided DNA endonuclease is (a) an RNA-binding moiety that interacts with an RNA that targets the DNA; (b) an active moiety that exhibits site-specific enzymatic activity; The RNA-guided DNA endonuclease according to claim 9, or a nucleic acid molecule encoding the RNA-guided DNA endonuclease according to any one of claims 1 to 3, or the vector according to claim 4, a nucleotide sequence modifying agent comprising a combination of:

14. The agent of claim 13 , wherein the cell is not a natural host for the gene encoding the RNA-guided DNA endonuclease.

15. The agent according to claim 13 or 14, wherein when the RNA-guided DNA endonuclease and the DNA-targeting RNA are directly introduced into the cell, they are introduced in the form of a ribonucleoprotein complex (RNP).

Citation Information

Patent Citations

  • Compositions and methods for modifying genomes

    JP2019504649A

  • Novel crispr enzymes and systems

    WO2016205711A1

  • Compositions and methods for modifying genomes

    WO2017141173A2

  • Polypeptides with type v crispr activity and uses thereof

    WO2018191715A2

  • Compositions and methods for modifying genomes

    WO2019030695A1