Novel cas nucleases and polynucleotides encoding the same
By developing novel Cas nucleases, the limitations of targeted genome editing in existing technologies have been addressed, enabling more flexible PAM site compatibility and efficient genome editing, suitable for high-expression and precise gene editing in pharmaceutical methods.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NOVOZYMES AS
- Filing Date
- 2024-12-19
- Publication Date
- 2026-07-24
AI Technical Summary
Existing CRISPR-related nucleases suffer from limited selectivity for targeting bases within the genome and insufficient expression capacity in genome editing, especially due to their dependence on PAM sites.
A novel Cas nuclease was developed with more flexible PAM site compatibility and a smaller size, enabling efficient expression in different organisms and allowing for broader target site selection by binding to a more flexible PAM site.
It improves the flexibility and efficiency of genome editing, expands the number of potential target sites within the genome, and is suitable for high-expression and precise gene editing in pharmaceutical methods.
Smart Images

Figure CN122459451A_ABST
Abstract
Description
References to sequence lists
[0001] This application contains a sequence list in computer-readable form, which is incorporated herein by reference. Background of the Invention Technical Field
[0002] This invention relates to novel CRISPR-associated (Cas) nucleases, variants thereof, and polynucleotides encoding them. The invention also relates to nucleic acid constructs, vectors, and host cells comprising polynucleotides and Cas nucleases, as well as fusion peptides, gene editing methods, methods for producing the Cas nuclease, formulations comprising the Cas nuclease, and uses of the Cas nuclease. Background Technology
[0003] In recent years, genome editing has become a key tool for research and applications. Early methods required complex engineering of nucleases (such as macronucleases, zinc finger fusion proteins, or TALENs) to adapt to each target sequence. This process was time-consuming and costly, posing challenges to scalability and efficiency.
[0004] RNA-guided nucleases, particularly CRISPR-associated (Cas) proteins, have brought about revolutionary breakthroughs. These RNA-guided nucleases can specifically target genetic sequences using guide RNA, simplifying genome editing by eliminating the need for custom-engineered nucleases. RNA-guided nucleases (including CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA)) offer a wide range of genome editing options, from introducing mutations through non-homologous end joining (NHEJ) to precise base editing upon fusion with deaminases.
[0005] Programmable nucleases are the core component of RNA-directed nucleases, binding to and cleaving nucleic acids with sequence specificity. They exhibit activities such as cis-cleavage or nickase activity, directed by specialized RNA molecules. These nucleases can be engineered to reduce catalytic activity while maintaining sequence specificity, thereby expanding their utility.
[0006] CRISPR systems in bacterial and archaea adaptive immunity exhibit distinct characteristics. Differences in size, PAM sites, mid-target activity, and cleavage patterns offer unique advantages for various applications but may also represent limitations (e.g., low frequency or low expression of PAM sites in the target cell genome). Novel Cas nucleases are crucial for meeting the evolving needs of genome engineering. Summary of the Invention
[0007] The inventors of this invention have identified novel Cas nucleases that exhibit nuclease activity in various organisms, including bacterial and fungal species. Surprisingly, these novel nucleases possess several advantages over known nucleases in the prior art. The novel Cas nucleases identified herein are compatible with more flexible PAM sites, thereby generating a greater number of potential target sites in each genome. Furthermore, the Cas nucleases identified herein have a relatively small size, thus showing promising potential for use in pharmaceutical methods. Additionally, the relatively small nuclease size facilitates high expression in recombinant cells, particularly bacterial cells.
[0008] In a first aspect, the present invention relates to Cas nucleases selected from the group consisting of:
[0009] (a) A polypeptide having at least 70% sequence identity with any amino acid sequence of SEQ ID NO: 21, 48, 1, 40, 39, 29, 2-20, 22-28, 30-38, 41-47 or 49-52;
[0010] (b) A polypeptide encoded by a polynucleotide having at least 70% sequence identity with any of the polypeptide coding sequences of SEQ ID No: 73, 100, 53, 92, 91 or 81, or with any of SEQ ID NO: 53-104, or with any of SEQ ID NO: 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549 or 550;
[0011] (c) A polypeptide derived from any one of SEQ ID NO: 21, 48, 1, 40, 39, 29, 2-20, 22-28, 30-38, 41-47 or 49-52 by having 1-30 alterations (e.g., substitution, deletion and / or insertion at one or more positions, e.g., 1 or 2 or 3 or 4 or 4 or 5 or 30 alterations), particularly substitution, e.g., conserved amino acid substitutions;
[0012] (d) A polypeptide having a TM-score of at least 0.80 compared to the three-dimensional structure of a polypeptide of any one of SEQ ID No: 21, 48, 1, 40, 39, 29, 2-20, 22-28, 30-38, 41-47 or 49-52, wherein the three-dimensional structure is calculated using Alphafold.
[0013] (e) A polypeptide derived from (a), (b), (c), or (d), wherein the N-terminus and / or C-terminus have been extended by adding one or more amino acids; and
[0014] Fragments of polypeptides in (f), (a), (b), (c), (d), or (e).
[0015] Some aspects of this disclosure provide Cas nucleases with different PAM specificities. Typically, Cas nucleases, such as Cas9 (spCas9) from *Streptococcus pyogenes*, require a typical “nGG” PAM sequence to bind to a specific nucleotide region. This can limit the ability to target desired bases within the genome. In some embodiments, the Cas nucleases provided herein may need to be placed in precise locations, such as within a 4-base region (e.g., an “edit window”) approximately 15 bases upstream of the PAM. See Komor, AC et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage”, *Nature* 533, 420-424 (2016). Therefore, in some embodiments, any Cas nuclease provided herein may contain a CRISPR nuclease capable of binding nucleotide sequences that do not contain a typical (e.g., nGG) PAM sequence. In one embodiment, the Cas nuclease of the present invention utilizes the “nnAY” PAM sequence, which allows for more flexible selection of target sites compared to the typical “nGG” PAM sequence of spCas9.
[0016] In a second aspect, the present invention relates to a fusion polypeptide comprising the Cas nuclease of the first aspect and one or more second polypeptides.
[0017] In a third aspect, the present invention relates to a non-naturally occurring composition comprising (i) a Cas nuclease of the first aspect and / or a fusion polypeptide of the second aspect, or (ii) a nucleic acid molecule comprising a sequence encoding the Cas nuclease of the first aspect and / or the fusion polypeptide of the second aspect.
[0018] In a fourth aspect, the present invention relates to a method for modifying nucleotide sequences at DNA target sites in the genome of a cell, the method comprising introducing a Cas nuclease according to the first aspect, a fusion polypeptide according to the second aspect, a composition according to the third aspect, a polynucleotide according to the fifth aspect, and / or a nucleic acid construct or expression vector according to the sixth aspect into the cell.
[0019] In a fifth aspect, the present invention relates to polynucleotides encoding a Cas nuclease of the first aspect and / or a fusion polypeptide of the second aspect.
[0020] In a sixth aspect, the present invention relates to nucleic acid constructs or expression vectors comprising a polynucleotide of the fifth aspect, the polynucleotide being operatively linked to one or more control sequences that direct the production of polypeptides in cells.
[0021] In a seventh aspect, the present invention relates to cells comprising the Cas nuclease of the first aspect, the fusion polypeptide of the second aspect, the composition of the third aspect, the polynucleotide of the fifth aspect, or the nucleic acid construct or expression vector of the sixth aspect.
[0022] In an eighth aspect, the present invention relates to a cell comprising a genome modified by the Cas nuclease of the first aspect, the fusion polypeptide of the second aspect, the composition of the third aspect, the method of the fourth aspect, the polynucleotide of the fifth aspect, and / or the nucleic acid construct or expression vector of the sixth aspect.
[0023] In a ninth aspect, the present invention relates to a method for producing a Cas nuclease of the first aspect or a fusion polypeptide of the second aspect, the method comprising culturing a host cell of the seventh aspect under conditions conducive to the production of the Cas nuclease or the fusion polypeptide.
[0024] In a tenth aspect, the present invention relates to the use of Cas nucleases of the first aspect, fusion polypeptides of the second aspect, compositions of the third aspect, methods of the fourth aspect, polynucleotides of the fifth aspect, or nucleic acid constructs or expression vectors of the sixth aspect for modifying target sequences (e.g., target genes) in cells.
[0025] In the eleventh aspect, the present invention relates to the use of Cas nucleases of the first aspect, fusion polypeptides of the second aspect, compositions of the third aspect, methods of the fourth aspect, polynucleotides of the fifth aspect, nucleic acid constructs or expression vectors of the sixth aspect, cells of the seventh aspect, or cells of the eighth aspect for manufacturing medicaments for modifying target sequences (e.g., target genes) in cells.
[0026] In a twelfth aspect, the present invention relates to a formulation comprising (i) a Cas nuclease according to the first aspect, a fusion polypeptide according to the second aspect, a composition according to the third aspect, a polynucleotide according to the fifth aspect, a nucleic acid construct or expression vector according to the sixth aspect, a cell according to the seventh aspect or a cell according to the eighth aspect, and optionally, (ii) one or more of lipids, liposomes, hydrogels, microparticles, nanoparticles or block copolymer micelles. Attached Figure Description
[0027] Figure 1 This paper presents a bioinformatics workflow developed for identifying novel CRISPR-Cas systems in genomic sequences. Genomic sequences from various sources were processed using three state-of-the-art sequence mining tools to identify Cas nuclease genes, CRISPR arrays, and tracrRNA sequences. The potential Cas protein library was further enriched by scanning the genome with a custom HMM and screening based on the presence of domains and residues required for nuclease activity. The three functional elements—Cas, CRISPR arrays, and tracrRNA—are mutually mapped via genomic loci and the complementarity of CRISPR repeats and tracrRNA.
[0028] Figure 2 A sequence homology tree of novel Cas nucleases based on their amino acid sequences (SEQ ID NO: 1-52) is shown.
[0029] Figures 3-6 The ratio of white to black colonies after transformation of Aspergillus niger with CRISPR nuclease is shown.
[0030] Figures 7-10 Insertions / deletions generated by CRISPR nucleases in the Aspergillus niger genome are shown.
[0031] Figures 11-13 The CRISPR plasmid used for Aspergillus niger is shown.
[0032] Figure 14 A schematic diagram of plasmid pTNA665 (PamyL-NZ0076) is shown.
[0033] Figure 15 A schematic diagram of plasmid pTNA666 (Pgrac-NZ0076) is shown.
[0034] Figure 16 A schematic diagram of plasmid pTNA669 (a separate guide RNA of NZ0076) is shown.
[0035] Figure 17 A schematic diagram of plasmid pTNA670 (a single guide RNA of NZ0076) is shown.
[0036] Figure 18 The cytotoxic effect based on nuclease activity in Escherichia coli (E. coli) was demonstrated.
[0037] Figure 19 The gene editing efficiency in Escherichia coli was demonstrated.
[0038] Figure 20The band alignment between the protein structures from nucleases 0076 and 0172 is shown.
[0039] Figure 21 The band alignment between the protein structures from nucleases 0100 and 0172 is shown.
[0040] Figure 22 The band alignment between the protein structures from nucleases 0076 and 0102 is shown.
[0041] Figure 23 The band alignment between the protein structures of nuclease 0076 and Streptococcus pyogenes Cas9 is shown.
[0042] Figure 24 The band alignment between the protein structures from nucleases 0076 and 0100 is shown.
[0043] definition
[0044] Based on this detailed description, the following definitions apply. Note that the singular forms “a / an” and “the” include plural indicators unless the context explicitly indicates otherwise.
[0045] Unless otherwise defined or explicitly indicated by the context, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0046] Base-editing peptides: The term "base-editing peptide" refers to a peptide containing a base editor domain capable of chemically altering a target DNA sequence without introducing double-strand breaks. Non-limiting examples of base-editing peptides include deaminases, such as cytidine deaminase and adenosine deaminase. Base-editing peptides can be fused with the inactivated Cas nuclease of this invention to achieve single nucleotide alterations to specific DNA target sequences. Using base-editing peptides, single nucleotide polymorphisms (SNPs) can be introduced at DNA target sites without generating double-strand breaks.
[0047] Catalytic domain: The term "catalytic domain" refers to the region of an enzyme that contains its catalytic mechanism. The catalytic domain of Cas nucleases contains an HNH domain and one or more RuvC domains.
[0048] Catalytically inactivated nucleases: The terms "catalytically inactivated nuclease," "catalytically inactive nuclease," or "dCas" refer to mutated Cas nucleases whose cleavage activity is reduced, lost, or essentially lost. Catalytically inactivated nucleases exhibit reduced, lost, or essentially lost cleavage activity against both single-stranded and double-stranded DNA. In other words, dCas may not cleave either strand of the target DNA.
[0049] For example, a catalytically inactivated nuclease may comprise an inactivated RuvC domain and an inactivated HNH domain. In some embodiments, the Cas nuclease is a catalytically inactivated nuclease, for example, having substantially no nuclease activity compared to a wild-type Cas nuclease that does not have an inactivated RuvC domain and an inactivated HNH domain, for example having no more than 5% nuclease activity.
[0050] As a non-limiting example, in some cases, dCas has both the D10A and H840A mutations in Streptococcus pyogenes Cas9, or the corresponding mutations in any Cas nuclease of the present invention. Based on this disclosure and knowledge in the art, other suitable nuclease-free dCas will be apparent to those skilled in the art and are within the scope of this disclosure. Such other exemplary suitable dCas include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, for example, Prashant et al., “CAS9 transcriptionalactivators for target specificity screening and paired nickases for cooperative genome engineering” Nature Biotechnology. 2013; 31(9):833-838, the entire contents of which are incorporated herein by reference).
[0051] cDNA: The term "cDNA" refers to a DNA molecule that can be prepared by reverse transcription from mature, spliced mRNA molecules obtained from eukaryotic or prokaryotic cells. cDNA lacks intron sequences that can be present in the corresponding genomic DNA. The initial primary RNA transcript is the precursor of mRNA, which is processed through a series of steps (including splicing) to become mature, spliced mRNA.
[0052] Coding sequence: The term "coding sequence" refers to a polynucleotide that directly specifies the amino acid sequence of a polypeptide. The boundaries of a coding sequence are typically defined by an open reading frame (ORF), which begins with a start codon (such as ATG, GTG, or TTG) and ends with a stop codon (such as TAA, TAG, or TGA). Coding sequences can be genomic DNA, cDNA, synthetic DNA, or a combination thereof.
[0053] Codon-optimized genes: The term "codon-optimized gene" refers to a gene whose codon usage frequency is optimized to match the host cell's preferred codon usage frequency. Nucleic acid alterations used to codon-optimize a gene do not change the amino acid sequence of the polypeptide encoded by the parent gene.
[0054] Control Sequences: The term "control sequence" refers to a nucleic acid sequence involved in regulating the expression of polynucleotides in a particular organism, either in vivo or in vitro. Each control sequence can be native (i.e., from the same gene) or heterologous (i.e., from different genes) for the polynucleotide encoding a polypeptide, and is native or heterologous relative to each other. Such control sequences include, but are not limited to, leader peptides, polyadenylation signals, propeptides, signal peptides, promoters, terminators, enhancers, and transcription or translation initiator and terminator sequences. At a minimum, control sequences include promoters and transcription and translation termination signals. These control sequences may be provided with multiple linkers for the purpose of introducing specific restriction sites that facilitate the linking of control sequences to the coding regions of polynucleotides encoding polypeptides.
[0055] Cas nucleases: The term "Cas nuclease" refers to a CRISPR-associated RNA-guided DNA endonuclease that, when coupled with a guide RNA, can cleave target DNA sequences. Cas nucleases are guided by one or more guide RNAs to recognize and cleave specific target sites in double-stranded DNA within the cellular genome. The CRISPR-Cas system is currently classified into 2 major classes, 6 types, and 33 subtypes (Makarova et al., 2020, Nat Rev Microbiol [Nature Microbiology Reviews] 18: 67-83).
[0056] In one embodiment, the Cas nuclease is a type II_A CRISPR-Cas system employing the Cas nuclease of SEQ ID NO: 1 or a variant thereof (including, for example, a CRISPR nicking enzyme). In one embodiment, the Cas nuclease is a type II_C CRISPR-Cas system employing the Cas nuclease of SEQ ID NO: 21 or a variant thereof (including, for example, a CRISPR nicking enzyme). In one embodiment, the Cas nuclease is a type II_A CRISPR-Cas system employing the Cas nuclease of SEQ ID NO: 40 or a variant thereof (including, for example, a CRISPR nicking enzyme). In one embodiment, the Cas nuclease is a type II_A CRISPR-Cas system employing the Cas nuclease of SEQ ID NO: 39 or a variant thereof (including, for example, a CRISPR nicking enzyme). In one embodiment, the Cas nuclease is a type II_A CRISPR-Cas system employing the Cas nuclease of SEQ ID NO: 48 or a variant thereof (including, for example, a CRISPR nicking enzyme). Typically, type II Cas nucleases contain two nuclease domains: an HNH nuclease domain that cleaves complementary DNA strands and a RuvC-like nuclease domain that cleaves non-complementary DNA strands.
[0057] Target recognition and cleavage by Cas nuclease requires a chimeric RNA, which is a fusion of crRNA (containing a guide sequence and partial unidirectional repeats) and tracrRNA (trans-activating crRNA), along with a short conserved sequence motif downstream of the crRNA-binding region (called a prototypical spacer adjacent motif (PAM)). In one embodiment, target recognition and cleavage by Cas nuclease are performed by separate crRNA and separate tracrRNA molecules, i.e., where the crRNA is not fused with the tracrRNA. In one embodiment, the Cas nuclease (e.g., SEQ ID NO: 21) targets the target DNA immediately adjacent to the 5'-nnGHMA PAM sequence.
[0058] In one embodiment, a Cas nuclease derived from a species of the genus Lactobacillus (e.g., SEQ ID NO: 1) targets the target DNA immediately adjacent to the 5'-nnAY PAM sequence.
[0059] In one embodiment, a Cas nuclease derived from the bacterium Enterococcus asini (e.g., SEQ ID NO: 40) targets the target DNA immediately adjacent to the 5'-nnAMA PAM sequence.
[0060] In one embodiment, a Cas nuclease (e.g., SEQ ID NO: 39) derived from the bacterium Enterococcus hermanniensis targets the target DNA immediately adjacent to the 5'-nnGTA PAM sequence.
[0061] In one embodiment, a Cas nuclease (e.g., SEQ ID NO: 48) derived from the bacterium Vagococcus penaei targets the target DNA immediately adjacent to the 5'-nnRHRD PAM sequence.
[0062] RNA-directed Cas nuclease activity produces site-specific double-strand breaks, which are then repaired via non-homologous end joining (NHEJ) or homology-directed repair (HDR). It should be understood that the term "Cas nuclease" encompasses its variants.
[0063] crRNA: CRISPR RNA (crRNA) serves as molecular guidance for Cas nucleases, providing sequence specificity for Cas nucleases to target and / or edit and / or regulate specific DNA and / or RNA sequences. The crRNA sequence contains spacers (prototype spacers) that recognize different DNA sequences. The crRNA binds to the Cas nuclease and produces a ribonucleoprotein complex called the CRISPR-Cas effector complex. Because the spacers match the complementary target DNA sequence, the Cas nuclease introduces DNA breaks at the target site. crRNA can be reprogrammed, allowing for precise and customizable genome editing.
[0064] DNA target site: The terms “target sequence,” “target site,” or “DNA target site” refer to one or more target DNA (e.g., genomic DNA) or RNA target sequences that may be subject to single-stranded or double-stranded cleavage by Cas nucleases and / or induced or inhibited by Cas nucleases. Typically, the target site (prototype spacer sequence) is at least 15-20 nucleotides in length to allow it to hybridize with the corresponding spacer sequence of the guide RNA. Target sites can be located anywhere in the genome, but are typically located within coding sequences or open reading frames. Non-restrictive examples of target sites include genes, promoters, and other regulatory sequences (e.g., enhancers, silencers, insulators, splice sites, and untranslated regions (UTRs) including the UTR and 5'-UTR of genes).
[0065] Preferably, the target site is flanked by the functional PAM sequence of the selected Cas nuclease.
[0066] Expression: The term “expression” refers to any step involved in peptide production, including but not limited to transcription, post-transcriptional modification, translation, folding the translated peptide into a functional structure, post-translational modification, and secretion.
[0067] Expression vector: An expression vector is a linear or circular DNA construct containing a DNA sequence encoding a polypeptide, with the coding sequence operatively linked to a suitable control sequence that can influence the expression of the DNA in a suitable host. Such control sequences may include promoters that affect transcription, optional operon sequences that control transcription, sequences encoding suitable ribosome binding sites on mRNA, enhancers, and sequences that control the termination of transcription and translation.
[0068] Extension: The term "extension" refers to the addition of one or more amino acids to the amino and / or carboxyl termini of a polypeptide, wherein the "extended" polypeptide has nuclease activity and / or DNA-binding activity.
[0069] Fragment: The term "fragment" refers to a polypeptide, catalytic domain, or DNA binding module having one or more amino acids missing from the amino and / or carboxyl termini of a mature polypeptide, catalytic domain, or binding module, wherein the fragment has nuclease activity or DNA binding activity.
[0070] Fusion polypeptide: The term "fusion polypeptide" is a polypeptide in which one or more polypeptides are fused to the N-terminus and / or C-terminus of the Cas nuclease of the present invention. Fusion polypeptides are generated by fusing a polynucleotide encoding another polypeptide with a polynucleotide of the present invention or by fusing two or more polynucleotides of the present invention together. Techniques for generating fusion polypeptides are known in the art and include linking the coding sequences of the polypeptides such that they conform to reading frames, and that the expression of the fusion polypeptide is under the control of one or more identical promoters and terminators. Fusion polypeptides can also be constructed using intron technology, wherein the fusion polypeptide is generated post-translational (Cooper et al., 1993, EMBO J. [Journal of the European Society for Molecular Biology] 12: 2575-2583; Dawson et al., 1994, Science [Science] 266: 776-779).
[0071] Genomic DNA: As used herein, “genomic DNA” refers to linear and / or chromosomal DNA and / or plasmid or other extrachromosomal DNA sequences present in one or more target cells. In some embodiments, the target cells are eukaryotic cells. In some embodiments, the target cells are prokaryotic cells. In some embodiments, these methods generate double-strand breaks (DSBs) at predetermined target sites in the genomic DNA sequence, resulting in mutations, insertions, and / or deletions of the DNA sequence at the target sites in the genome.
[0072] Guide sequence portion: The “guide sequence portion” or “spacer” of an RNA molecule refers to a nucleotide sequence that can hybridize with a specific target DNA sequence (prototype spacer). For example, the nucleotide sequence of the guide sequence portion is partially or completely complementary to the targeted DNA sequence in length of the guide sequence portion. In some embodiments, the length of the guide sequence portion is 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides, preferably at least 18 nucleotides, for example at least 23 nucleotides, or approximately 17-50, 17-49, 17-48, 17-47, 17-46, 17-45, 17-44, 17-43, or 17-42. 17-41, 17-40, 17-39, 17-38, 17-37, 17-36, 17-35, 17-34, 17-33, 17-31, 17-30, 17-29, 17-28, 17-27, 17-26, 17-25, 17-24, 17-22, 17-21, 18-25, 18-24, 18-23, 18-22, 18-21, 19-25, 19-24, 19-23, 19-22, 19-21, 19-20, 20-22, 18-20, 20-21, 21-22, or 17-20 nucleotides. The entire length of the guide sequence is completely complementary to the length of the targeted DNA sequence in the guide sequence portion. The guide sequence portion can be a part of an RNA molecule capable of forming a complex with a Cas nuclease, wherein the guide sequence portion acts as the DNA targeting portion of the CRISPR complex. When a DNA molecule having a guide sequence portion is present in conjunction with a CRISPR molecule, the RNA molecule is able to target the Cas nuclease to a specific target DNA or RNA sequence. Each possibility represents a separate embodiment. RNA molecules can be custom-designed to target any desired sequence. Therefore, a molecule containing a “guide sequence portion” is a targeting molecule. Throughout this application, the terms “guide molecule,” “RNA guide molecule,” “guide RNA molecule,” and “gRNA molecule” are synonymous with a molecule containing a guide sequence portion, and the term “spacer” is synonymous with “guide sequence portion.”
[0073] In embodiments of the invention, Cas nucleases exhibit maximum cleavage activity when used with RNA molecules containing a guide sequence portion having 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides.
[0074] Single guide RNA (sgRNA) molecules can be used to guide Cas nucleases to desired target sites. A single guide RNA contains a guide sequence portion and a scaffold portion. The scaffold portion interacts with the Cas nuclease and, together with the guide sequence portion, activates the Cas nuclease and targets it to the desired target site. The scaffold portion can be further engineered, for example, to reduce its size.
[0075] In CRISPR-Cas genome editing, the gRNA constitutes the reprogrammable part, which makes the system versatile. In most natural Cas systems, the gRNA is actually a complex of two RNA polynucleotides: a first crRNA (containing about 15-30 nucleotides and defining the specificity of the Cas nuclease) and a tracrRNA (hybridizing with the crRNA to form an RNA complex that interacts with the Cas nuclease) (see Jinek et al., 2012, A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity, Science 337: 816-821).
[0076] Thanks to the discovery of the CRISPR-Cas system, single polynucleotide gRNAs have been developed and successfully applied, proving as effective as natural two-part gRNA complexes.
[0077] The spacer can be part of a targeting guide RNA molecule that can form a complex with Cas nuclease, where the spacer sequence acts as the targeting portion of the CRISPR complex. When a molecule with a spacer sequence is present alongside a CRISPR molecule, the RNA molecule is able to target the Cas nuclease to a specific target sequence. Each possibility represents a separate embodiment. The targeting RNA molecule can be custom-designed to target any desired sequence.
[0078] As used herein, the term "targeted" refers to preferential hybridization of a spacer sequence with a nucleic acid having a target nucleotide sequence (the prototype spacer). It should be understood that the term "targeted" encompasses variable hybridization efficiency, resulting in preferential targeting of nucleic acids having target nucleotide sequences; however, unintentional off-target hybridization may also occur in addition to on-target hybridization. It should be understood that when an RNA molecule targets a sequence, the complex of the RNA molecule and the Cas nuclease molecule targets that sequence to exert nuclease activity.
[0079] In the context of targeting DNA sequences present in multiple cells, it should be understood that targeting encompasses the hybridization of a guide sequence portion of an RNA molecule with a sequence in one or more cells, and also encompasses the hybridization of an RNA molecule with a target sequence in a subset of the multiple cells. Therefore, it should be understood that when an RNA molecule targets a sequence in multiple cells, the complex of the RNA molecule and the Cas nuclease is understood to hybridize with the target sequence in one or more cells, and may also hybridize with the target sequence in a subset of the cells. Therefore, it should be understood that the complex of the RNA molecule and the Cas nuclease introduces double-strand breaks when hybridizing with the target sequence in one or more cells, and may also introduce double-strand breaks when hybridizing with the target sequence in a subset of the cells. As used herein, the terms "modified cell" or "cell containing a genome modified by the Cas nuclease" refer to a cell in which double-strand breaks result from hybridization of the complex of the RNA molecule and the Cas nuclease with the target sequence (i.e., mid-target hybridization).
[0080] Heterogeneous: For host cells, the term "heterogeneous" means that the polypeptide or nucleic acid is not naturally present in the host cell. For polypeptides or nucleic acids, the term "heterogeneous" means that the control sequence (e.g., the promoter of the polypeptide or nucleic acid) is not naturally associated with that polypeptide or nucleic acid; that is, the control sequence comes from a gene other than the gene encoding the mature polypeptide.
[0081] The HNH sequence contains the HNH domain of the Cas nuclease. The HNH domain in the Cas nuclease stands for "histidine-asparagine-histidine". These conserved amino acid residues play a crucial role in the nuclease activity of this domain. The HNH domain is one of the two main types of nuclease domains found in CRISPR-associated (Cas) proteins. In the context of the CRISPR system, the HNH domain is responsible for cleaving the DNA strand complementary to the RNA guide strand, thereby creating a nick in the target DNA. This cleavage is a key step in the CRISPR-Cas gene editing process, enabling precise DNA modification.
[0082] The HNH domain, together with the RuvC domain, creates double-strand breaks in the target DNA, thereby enabling gene editing or modification.
[0083] Host strain or host cell: A "host strain" or "host cell" is an organism in which an expression vector, bacteriophage, virus, or other DNA construct (including a polynucleotide encoding the polypeptide of the present invention) has been introduced. An exemplary host strain is a microbial cell (e.g., bacteria, filamentous fungi, and yeast) capable of expressing a Cas nuclease. The term "host cell" includes protoplasts produced by cells.
[0084] Introduction: In the context of inserting a nucleic acid sequence into a cell, the term “introduction” means “transfection,” “conversion,” or “transduction,” as is known in the art.
[0085] Isolated: The term "isolated" means a polypeptide, nucleic acid, cell, or other specific material or component that has been separated from at least one other material or component (including, but not limited to, other proteins, nucleic acids, cells, etc.). Therefore, the isolated polypeptide, nucleic acid, cell, or other material exists in a form not found in nature. Isolated polypeptides include, but are not limited to, culture media containing secreted polypeptides expressed in host cells.
[0086] Mature peptide: The term “mature peptide” refers to a peptide that has been processed at its N-terminus and / or C-terminus (e.g., removal of the signal peptide) to be in its mature form.
[0087] Mature polypeptide coding sequence: The term "mature polypeptide coding sequence" refers to a polynucleotide that encodes a mature Cas nuclease with nuclease activity and / or DNA binding activity.
[0088] Natural: The term "natural" refers to nucleic acids or polypeptides that are naturally present in host cells.
[0089] Crizoase: The terms “crizoase,” “CRISPR cleavage enzyme,” or “nCas” refer to a nuclease having an inactivated RuvC domain or an inactivated HNH domain. It should be understood that a Cas nuclease does not lose its nuclease activity in cleaving all DNA, but rather may lose the ability to cleave only the target strand of double-stranded DNA or only the non-target strand, thus functioning as a cleavage enzyme (see, Gao et al. (2016) CELL RES. [Cell Research], 26: 901). Therefore, in some embodiments, the Cas nuclease is nCas. In some embodiments, nCas has activity in cleaving the non-complementary strand but essentially lacks activity in cleaving the complementary strand, for example, through mutations in the HNH domain. For example, nCas may have mutations that reduce the function of the HNH domain, such as the H840A mutation in *Streptococcus pyogenes* Cas9 or a corresponding mutation in any of the Cas nucleases of the present invention. In some embodiments, nCas has activity in cleaving the complementary strand but essentially lacks activity in cleaving the non-complementary strand, for example, through mutations in the RuvC domain. For example, nCas can have mutations that reduce the function of the RuvC domain, such as the D10A mutation in Streptococcus pyogenes Cas9 or the corresponding mutation in any Cas nuclease of the present invention.
[0090] Nuclear Localization Sequence: The terms “nuclear localization sequence” and “NLS” are used interchangeably to indicate the amino acid sequence / peptide that guides the transport of its associated protein from the cytoplasm across the nuclear membrane barrier. The term “NLS” is intended to encompass not only the nuclear localization sequence of a specific peptide but also its derivatives capable of guiding the translocation of cytoplasmic peptides across the nuclear membrane barrier. An NLS can guide the nuclear translocation of a Cas nuclease when attached to its N-terminus, C-terminus, or both. Additionally, peptides having an NLS coupled to amino acid side chains randomly distributed along the amino acid sequence of the peptide via its N-terminus or C-terminus will be translocated. Typically, an NLS consists of one or more short sequences of positively charged lysine or arginine residues exposed on the protein surface, but other types of NLS are known. Non-limiting examples of NLS include NLS sequences derived from the following: SV40 virus large T antigen, nucleoplasmic protein, c-myc, hRNPAl M9 NLS, IBB domain from nuclear import protein-α, myoma T protein, human p53, mouse c-abl IV, influenza virus NS1, hepatitis virus δ antigen, mouse Mxl protein, human poly(ADP-ribose) polymerase, and steroid hormone receptor (human) glucocorticoid.
[0091] Nucleic acid: The term "nucleic acid" encompasses DNA, RNA, heteroduplexes, and synthetic molecules capable of encoding polypeptides. Nucleic acids can be single-stranded or double-stranded and can be chemically modified. The terms "nucleic acid" and "polynucleotide" are used interchangeably. Because the genetic code is degenerate, more than one codon can be used to encode a specific amino acid, and the compositions and methods of this invention cover nucleotide sequences encoding specific amino acid sequences. Unless otherwise stated, nucleic acid sequences are presented in a 5' to 3' orientation.
[0092] Nucleic acid constructs: The term “nucleic acid construct” refers to a single-stranded or double-stranded nucleic acid molecule that is isolated from a naturally occurring gene or modified in a way that does not originally exist in nature to contain a segment of nucleic acid or is synthesized and contains one or more control sequences that are operatively linked to the nucleic acid sequence.
[0093] Operationally linked: The term "operationally linked" means that specified components are in a relationship that allows them to function in the intended manner (including, but not limited to, juxtaposition). For example, a regulatory sequence is operationally linked to a coding sequence such that the expression of the coding sequence is under the control of the regulatory sequence.
[0094] PAM: As used herein, the term "PAM" or "prototype spacer adjacent motif" refers to the nucleotide sequence of the target DNA located near the target DNA sequence (prototype spacer) and recognized by the Cas nuclease (i.e., the guide RNA that forms a complex with the Cas nuclease and the target DNA). The PAM sequence can vary depending on the identity of the Cas nuclease. In some cases, the PAM is necessary for the Cas nuclease and guide RNA complex to hybridize with and edit the target sequence. In other cases, the complex does not require the PAM to edit the target sequence.
[0095] The commonly accepted abbreviations for the degeneracy of nucleotide bases representing PAM, as used in the art and herein, include the following: R = G or A; Y = C or T; M = A or C; K = G or T; S = G or C; W = A or T; H = A or C or T; B = G or T or C; V = G or C or A; D = G or A or T; N = A or C or G or T. Non-limiting examples of suitable PAM sequences for the Cas nucleases of the present invention are shown in Table 1.
[0096] Purified: The term "purified" means nucleic acids, peptides, or cells that are substantially free of other components, as determined by analytical techniques well known in the art (e.g., in electrophoretic gels, chromatographic eluates, and / or media subjected to density gradient centrifugation, where purified peptides or nucleic acids form discrete bands). Purified nucleic acids or peptides are at least about 50% pure, and typically at least about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.5%, about 99.6%, about 99.7%, about 99.8%, or more pure (e.g., weight percentage or molar percentage). In a relevant sense, a composition is enriched with the molecule when the concentration of the molecule increases significantly after the application of purification or enrichment techniques. The term “enrichment” refers to the presence of compounds, peptides, cells, nucleic acids, amino acids, or other specified materials or components in a composition at a relative or absolute concentration higher than that of the starting composition.
[0097] In one respect, as used herein, the term "purified" means that the polypeptide or cell is substantially free of components (especially insoluble components) from the producing organism. In other respects, the term "purified" means that the Cas nuclease is substantially free of insoluble components from the natural organism from which it originates. In one respect, the Cas nuclease is separated from some soluble components of the organism from which it is recovered and the culture medium. The polypeptide can be purified (i.e., separated) by one or more of unit operation filtration, precipitation, or chromatography.
[0098] Therefore, Cas nucleases can be purified so that only small amounts of other proteins, particularly other polypeptides, are present. As used herein, the term "purified" can mean the removal of other components, particularly other proteins, and most particularly other enzymes present in the cell from which the polypeptide originated. Cas nucleases can be "substantially pure," meaning they are free from other components from the organism that produced them (e.g., the host organism used to recombinantly produce the polypeptide). In one aspect, the polypeptide is at least 40% pure by weight of the total polypeptide material present in the formulation. In another aspect, the polypeptide is at least 50%, 60%, 70%, 80%, or 90% pure by weight of the total polypeptide material present in the formulation. As used herein, "substantially pure polypeptide" can mean a Cas nuclease formulation containing up to 10%, preferably up to 8%, more preferably up to 6%, more preferably up to 5%, more preferably up to 4%, more preferably up to 3%, even more preferably up to 2%, most preferably up to 1%, and even most preferably up to 0.5% by weight of other polypeptide material (Cas nuclease associated with its natural or recombinant forms).
[0099] Therefore, it is preferred that, based on the weight of the total polypeptide material present in the formulation, the substantially pure Cas nuclease or fusion polypeptide is at least 92% pure, preferably at least 94% pure, more preferably at least 95% pure, more preferably at least 96% pure, more preferably at least 97% pure, more preferably at least 98% pure, even more preferably at least 99% pure, and most preferably at least 99.5% pure. The polypeptides of the present invention are preferably in a substantially pure form (i.e., the formulation is substantially free of other polypeptide material associated with its natural or recombinant form). For example, this can be achieved by preparing the polypeptide using well-known recombinant methods or classical purification methods.
[0100] Recombination: The term "recombination," used in its conventional sense, refers to the manipulation (e.g., cutting and rejoining) of nucleic acid sequences to form a sequence group different from that found in nature. The term recombination refers to cells, nucleic acids, polypeptides, or vectors that have been modified from their natural state. Thus, for example, recombinant cells express genes not found in their natural (non-recombinant) forms, or express natural genes at different levels or under different conditions compared to those found in nature. The term "recombination" is synonymous with "genetically modified" and "transgenic."
[0101] Recovery: The term "recovery" refers to the removal of peptides from at least one fermentation broth component selected from a list of cells, nucleic acids, or other specified materials, for example, by means of: harvesting peptides from whole or cell-free fermentation broths by means of peptide crystallization, by means of filtration (e.g., deep filtration (using filter aids or packed filter media, cloth filtration in a box filter, rotary drum filtration, drum filtration, rotary vacuum drum filtration, candle filter, horizontal leaf filter, or the like, sheet or pad filtration in a frame or modular apparatus) or membrane filtration (using plate filtration, modular filtration, candle filtration, microfiltration, crossflow, dynamic crossflow, or ultrafiltration in dead-end operation)), or by means of centrifugation (using a horizontal centrifuge, disc stack centrifuge, hyrdo cyclone, or the like), or by means of precipitating peptides and using relevant solid-liquid separation methods to harvest peptides from broth media by means of particle size fractionation. Recovery encompasses the isolation and / or purification of peptides.
[0102] Reverse transcriptase: The term “reverse transcriptase” or “RT” describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require primers to synthesize DNA transcripts from an RNA template. Historically, reverse transcriptases have primarily been used to transcribe mRNA into cDNA, which is then cloned into vectors for further manipulation. Avian myeloblastomavirus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta [Chinese Journal of Biochemistry and Biophysics] 473:1 (1977)). The enzyme possesses 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a continuous 5' and 3' ribonuclease that is specific for the RNA strand of RNA-DNA hybrids (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons [Wiley & Sons Publishing Co.] (1984)). Errors in transcription cannot be corrected by reverse transcriptase because known viral reverse transcriptases lack the 3'-5' exonuclease activity necessary for proofreading (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). Berger et al., Biochemistry 22:2365-2372 (1983) provided a detailed study of the activity of AMV reverse transcriptase and its associated RNase H activity. Another reverse transcriptase widely used in molecular biology is that derived from Moloney mouse leukemia virus (M-MLV). See, for example, Gerard, GR, DNA 5:271-279 (1986) and Kotewicz, ML, et al., Gene 35:249-258 (1985). M-MLV reverse transcriptases that are essentially lacking RNase H activity have also been described. See, for example, U.S. Patent No. 5,244,797.
[0103] When fused with Cas nucleases, reverse transcriptases provide a versatile and precise method for gene editing. They allow the conversion of RNA targets into DNA, enhancing the specificity and accuracy of Cas-mediated gene editing, while simultaneously manipulating both DNA and RNA molecules, enabling a wide range of applications in genetics and molecular biology.
[0104] This disclosure covers any wild-type reverse transcriptase obtained from any naturally occurring organism or virus, or from commercial or non-commercial sources. Additionally, the reverse transcriptase that can be used for the fusion polypeptides disclosed herein may include any naturally occurring mutant RT, engineered mutant RT, or other variant RT, including truncated variants that retain function. RT may also be engineered to contain specific amino acid substitutions, such as those specifically disclosed herein.
[0105] RT is a multifunctional enzyme, typically possessing three enzymatic activities: RNA- and DNA-dependent DNA polymerization activity, and RNase H activity, which catalyzes the cleavage of RNA in RNA-DNA hybrids. Some mutants of RT partially inactivate RNase H to prevent accidental damage to mRNA. These enzymes, which use mRNA as a template to synthesize complementary DNA (cDNA), were initially identified in RNA viruses.
[0106] Exemplary enzymes that can be used with the fusion peptides disclosed herein may include, but are not limited to, M-MLV reverse transcriptase and RSV reverse transcriptase. Enzymes with RT activity are commercially available. Based on various embodiments of this disclosure, several exemplary reverse transcriptases that can be fused with CRISPR nucleases or provided as standalone proteins are provided below.
[0107] Those skilled in the art will recognize that wild-type reverse transcriptases (RTs) including, but not limited to, Moloney murine leukosis virus (M-MLV), human immunodeficiency virus (HIV) reverse transcriptase, and avian sarcoma leukosis virus (ASLV) reverse transcriptase, including but not limited to: Laure's sarcoma virus (RSV) reverse transcriptase, avian myeloblastoma virus (AMV) reverse transcriptase, avian erythroblastosis virus (AEV) helper virus MCAV reverse transcriptase, avian myelomavirus MC29 helper virus MCAV reverse transcriptase, avian reticuloendothelioma virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Laure's-associated virus (RAV) reverse transcriptase, and myeloblastoma-associated virus (MAV) reverse transcriptase) are suitable for use in the subject methods and compositions described herein. In some embodiments, the RT can be any RT described in WO 2020 / 191248 (the contents of which are incorporated herein by reference).
[0108] In some embodiments, a suitable reverse transcriptase may be any reverse transcriptase described in WO 2020191233, WO 2020191233, WO2020191243, WO 2020191246, WO 2020191245, WO 2020191234, WO 2020191233, WO2020191241, US20200085066, US 20200109398, US 20200109398, WO 2020191239, WO2020191245 and WO 2020191248 (the contents of each of which are incorporated herein by reference in their entirety).
[0109] RuvC sequence: In the context of the CRISPR-Cas system, the RuvC sequence contains or is composed of a "RuvC" or "RuvC-like" domain. RuvC stands for "Hollywood linker dissociative nuclease" and is responsible for cleaving the DNA strand opposite to the RNA guide strand. One or more RuvC domains, together with the HNH domain, create double-strand breaks in the target DNA, thereby achieving gene editing or modification. In some embodiments, the Cas nuclease contains three RuvC domains, for example, RuvC I, RuvC II, and RuvC III domains.
[0110] Sequence identity: The degree of association between two amino acid sequences or two nucleotide sequences is described by the parameter "sequence identity".
[0111] For the purposes of this invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. [Journal of Molecular Biology] 48: 443-453) is used to determine the sequence identity between two amino acid sequences as the output of "longest identity". This algorithm is implemented in the Niedel program of the EMBOSS software package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. [Trends in Genetics] 16: 276-277) (preferably version 6.6.0 or later). The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. For the Niedel program to report the longest identity, the non-simplification option must be specified in the command line. The Niedel-marked "longest identity" output is calculated as follows:
[0112] (identical residues × 100) / (alignment length - total number of vacancies in the alignment)
[0113] For the purposes of this invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, ibid.) is used to determine the sequence identity between two polynucleotide sequences as the output of "longest identity," as implemented by the Niedle program in the EMBOSS software package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, ibid.) (preferably version 6.6.0 or later). The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EDNAFULL (EMBOSS version of NCBINUC 4.4) substitution matrix. For the Niedle program to report the longest identity, the non-simplified option must be specified on the command line. The Niedle-marked "longest identity" output is calculated as follows:
[0114] (Identical deoxyribonucleotides × 100) / (Alignment length - Total number of vacancies in the alignment)
[0115] Signal peptide: A signal peptide is an amino acid sequence that attaches to the N-terminal portion of a protein and promotes its secretion outside the cell. The mature form of extracellular proteins lacks a signal peptide, which is cleaved during the secretion process.
[0116] Subsequence: The term "subsequence" refers to a polynucleotide that omits one or more nucleotides from the 5' and / or 3' end of the coding sequence of a mature polypeptide, wherein the subsequence encodes a fragment with nuclease activity.
[0117] tracrRNA: Trans-activating CRISPR RNA (tracrRNA) is a class of RNA molecules that form part of the CRISPR-Cas system. tracrRNA acts as a scaffold or scaffold-like molecule that facilitates the binding of Cas nucleases to CRISPR RNA (crRNA) molecules. In this complex, tracrRNA interacts with Cas nucleases to form a ribonucleoprotein complex that recognizes and binds to target DNA or RNA sequences, guiding the Cas nucleases to precise sites for cleavage or editing. Non-restrictive examples of tracrRNA coding sequences are listed in column 3 of Table 4.
[0118] Variant: The term "variant" refers to a Cas nuclease that has nuclease activity and / or DNA-binding activity and contains artificial mutations (i.e., substitutions, insertions (including extensions), and / or deletions (e.g., truncations)) at one or more positions. Substitution means replacing an amino acid occupying a position with a different amino acid; deletion means removing an amino acid occupying a position; and insertion means adding an amino acid adjacent to and immediately following the amino acid occupying a position (e.g., 1-5 amino acids, 1-3 amino acids, or particularly 1 amino acid).
[0119] Wild-type: When referring to an amino acid sequence or nucleic acid sequence, the term "wild-type" means that the amino acid sequence or nucleic acid sequence is natural or naturally occurring. As used herein, the term "naturally occurring" means any substance found in nature (e.g., protein, amino acid, or nucleic acid sequence). Conversely, the term "non-naturally occurring" means any substance not found in nature (e.g., compositions produced in a laboratory or during manufacturing, and / or recombinant nucleic acid and protein sequences produced in a laboratory, or modifications of wild-type sequences). In embodiments of the invention, the engineered Cas nuclease is a variant Cas nuclease containing at least one amino acid modification (e.g., substitution, deletion, and / or insertion) compared to any Cas nuclease shown in column 1 of Table 4. Detailed Implementation
[0120] Cas nuclease
[0121] In a first aspect, the present invention relates to Cas nucleases selected from the group consisting of:
[0122] (a) A polypeptide having at least 70% sequence identity with any amino acid sequence of SEQ ID NO: 21, 48, 1, 40, 39, 29, 2-20, 22-28, 30-38, 41-47 or 49-52;
[0123] (b) A polypeptide encoded by a polynucleotide having at least 70% sequence identity with any of the polypeptide coding sequences of SEQ ID NO: 73, 100, 53, 92, 91 or 81, or with any of SEQ ID NO: 53-104, or with any of SEQ ID NO: 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549 or 550;
[0124] (c) A polypeptide derived from any one of SEQ ID NO: 21, 48, 1, 40, 39 or 29 or any one of SEQ ID NO: 1-52 by having 1-30 alterations (e.g., substitution, deletion and / or insertion at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations), particularly by substitution, e.g., conserved amino acid substitution;
[0125] (d) A polypeptide having a TM-score of at least 0.80 compared to the three-dimensional structure of a polypeptide of any one of SEQ ID No: 21, 48, 1, 40, 39 or 29 or any one of SEQ ID No: 1-52, wherein the three-dimensional structure is calculated using Alphafold.
[0126] (e) A polypeptide derived from (a), (b), (c), or (d), wherein the N-terminus and / or C-terminus have been extended by adding one or more amino acids; and
[0127] Fragments of polypeptides in (f), (a), (b), (c), (d), or (e).
[0128] In one embodiment, the Cas nuclease is selected from the group consisting of:
[0129] (a) Polypeptides corresponding to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13. SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26. SEQ ID NO: 27. SEQ ID NO: 28. SEQ ID NO: 29. SEQ ID NO: 30. SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: 52 have an amino acid sequence identity of at least 70%, for example, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100%;
[0130] (b) A polypeptide encoded by a polynucleotide, wherein the polynucleotide is associated with SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, SEQ ID NO: 89, SEQ ID NO: 90, SEQ ID NO: 91, SEQ ID NO: 92, SEQ ID NO: 93, SEQ ID NO: 94, SEQ ID NO: 95. SEQ ID NO: 96, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103, SEQ ID NO: 104, or SEQ ID NO: The polypeptide coding sequence of any one of 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549 or 550 has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity;
[0131] (c) A polypeptide derived from SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12 ...13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 30, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 19, SEQ ID NO: 10, SEQ ID NO: 19, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO 17. SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30. SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43. SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48. SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: 52;
[0132] (d) A polypeptide having a TM-score of at least 0.80, for example, at least 0.85, at least 0.90, at least 0.91, at least 0.92, at least 0.93, at least 0.94, at least 0.95, at least 0.96, at least 0.97, at least 0.98, at least 0.99 or even 1.0 compared to the three-dimensional structure of the polypeptide of any one of SEQ ID NO:1 or SEQ ID NO:1-52, wherein the three-dimensional structure is calculated using Alphafold;
[0133] (e) A polypeptide derived from (a), (b), (c), or (d), wherein the N-terminus and / or C-terminus have been extended by adding one or more amino acids; and
[0134] Fragments of polypeptides in (f), (a), (b), (c), (d), or (e).
[0135] In one embodiment, the nuclease has nuclease activity.
[0136] In one embodiment, the nuclease has DNA-binding activity.
[0137] In one embodiment, the nuclease has both nuclease activity and DNA-binding activity.
[0138] In one embodiment, the nuclease comprises or consists of an amino acid sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39, SEQ ID NO: 48, or SEQ ID NO: 29.
[0139] In one embodiment, the nuclease comprises, is substantially composed of, or is composed of SEQ ID NO:1, SEQ ID NO:21, SEQ ID NO:40, SEQ ID NO:39, SEQ ID NO:48, or SEQ ID NO:29.
[0140] In one embodiment, the nuclease is SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13. SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26. SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ Fragments of SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, or SEQ ID NO: 52, wherein the fragment preferably contains at least 600 amino acid residues (e.g.,Amino acids 9 to 640 of SEQ ID NO: 1, 13 to 628 of SEQ ID NO: 39, 16 to 637 of SEQ ID NO: 40, 10 to 637 of SEQ ID NO: 41, 10 to 639 of SEQ ID NO: 42, 10 to 636 of SEQ ID NO: 43, 10 to 635 of SEQ ID NO: 44, 9 to 640 of SEQ ID NO: 45, 10 to 637 of SEQ ID NO: 46, 10 to 633 of SEQ ID NO: 47, 12 to 632 of SEQ ID NO: 48, 8 to 620 of SEQ ID NO: 21, 9 to 640 of SEQ ID NO: 51, or 9 to 640 of SEQ ID NO: 52.
[0141] In one embodiment, the nuclease comprises, substantially comprises, or comprises the following: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30 ...26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:29, SEQ ID NO:29, SEQ ID NO:29, SEQ ID NO:29, SEQ ID NO:29, SEQ ID NO:29, SEQ ID NO:2 ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: 52.
[0142] In one embodiment, the nuclease is encoded by a polynucleotide, which is associated with SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 7 ... ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, SEQ ID NO: 89, SEQ ID NO: 90, SEQ ID NO: 91, SEQ ID NO: 92, SEQ ID NO: 93, SEQ ID NO: 94, SEQ ID NO: 95, SEQ ID NO: 96, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103 or SEQ ID NO: 104, or SEQ ID NO: The mature polypeptide coding sequence of any one of 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549 or 550 has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity.
[0143] In one embodiment, the nuclease comprises an N-terminal extension and / or a C-terminal extension of 1-10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids, preferably an extension of 1-10 amino acid residues in the N-terminus and / or 1-10 amino acids, such as 1-5 amino acids, in the C-terminus.
[0144] In one embodiment, the nuclease is combined with SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13. SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26. SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ The polypeptides of SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: 52 have a sequence difference of up to 10%, up to 9%, up to 8%, up to 7%, up to 6%, up to 5%, up to 4%, up to 3%, up to 2% or up to 1%.
[0145] In one embodiment, the nuclease is combined with SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13. SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26. SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ The polypeptides of SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, or SEQ ID NO: 52 differ by up to 15 amino acids, for example, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids.
[0146] In one embodiment, the nuclease comprises one or more functional RuvC domains.
[0147] In one embodiment, the nuclease comprises two or more functional RuvC domains.
[0148] In one embodiment, the nuclease comprises three or more functional RuvC domains.
[0149] In one embodiment, the nuclease comprises one or more functional HNH domains.
[0150] In one embodiment, the nuclease comprises one or more domains selected from the group consisting of:
[0151] (a) The RuvC domain having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 105-143 or 313-318;
[0152] (b) An HNH domain having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 144-156 or 319-320;
[0153] (c) The RuvC domain, which is derived from any one of SEQ ID NO: 105-143 or 313-318 by substitution, deletion or addition of one or more amino acids of SEQ ID NO: 105-143 or 313-318;
[0154] (d) An HNH domain derived from any one of SEQ ID NO: 144-156 or 319-320 by substitution, deletion, or addition of one or more amino acids of SEQ ID NO: 144-156 or 319-320; and
[0155] Fragments of the catalytic domains of (e), (a), (b), (c), or (d);
[0156] Preferably, the nuclease has nuclease activity, or the nuclease has nicking enzyme activity.
[0157] In one embodiment, the HNH domain has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 144-156 or 319-320.
[0158] In one embodiment, the HNH domain comprises or consists of an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 144, 146, 145, 154, 319, or 320.
[0159] In one embodiment, the HNH domain is a variant of any one of SEQ ID NO: 144-156 or 319-320, which includes substitutions, such as conserved amino acid substitutions, deletions, and / or insertions, at one or more positions.
[0160] In one embodiment, the HNH domain differs from any of SEQ ID NO: 144-156 or 319-320 by up to 15 amino acids, for example, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids.
[0161] In one embodiment, the HNH domain is a fragment of any one of SEQ ID NO: 144-156 or 319-320, wherein the fragment preferably contains at least 20 amino acid residues (e.g., amino acids 613 to 640 of SEQ ID NO: 1) or at least 27 amino acid residues (e.g., amino acids 613 to 640 of SEQ ID NO: 1).
[0162] In one embodiment, the HNH domain comprises, is substantially composed of, or is composed of any one of SEQ ID NO: 144-156 or 319-320.
[0163] In one embodiment, the RuvC domain has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 105-143 or 313-318.
[0164] In one embodiment, the RuvC domain comprises or consists of an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 105-107, 111-113, 108-110, 135-137, or 313-318.
[0165] In one embodiment, the RuvC domain is a variant of any one of SEQ ID NO: 105-143 or 313-318, which includes substitutions, such as conserved amino acid substitutions, deletions, and / or insertions, at one or more positions.
[0166] In one embodiment, the RuvC domain differs from any of SEQ ID NO: 105-143 or 313-318 by up to 15 amino acids, for example, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids.
[0167] In one embodiment, the RuvC domain is a fragment of any one of SEQ ID NO: 105-143 or 313-318, wherein the fragment preferably contains at least 10 amino acid residues (e.g., amino acids 5 to 20 of SEQ ID NO: 1).
[0168] In one embodiment, the RuvC domain comprises, is substantially composed of, or is composed of any one of SEQ ID NO: 105-143 or 313-318.
[0169] In one embodiment, the nuclease has double-strand break activity against DNA target sites.
[0170] In one embodiment, sequence identity is determined by the method described in the definition section of "Sequence Identity".
[0171] In one embodiment, the polynucleotide encoding a nuclease is codon-optimized for expression in eukaryotic cells.
[0172] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in mammalian cells, such as non-human mammalian cells.
[0173] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Escherichia coli cells.
[0174] In one embodiment, the polynucleotide encoding a nuclease is codon-optimized for expression in Escherichia coli cells, wherein the polynucleotide comprises or consists of a sequence having at least 80%, for example, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the nucleotide sequences in SEQ ID NO: 512-520.
[0175] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Bacillus cells.
[0176] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Bacillus subtilis cells.
[0177] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Bacillus licheniformis cells.
[0178] In one embodiment, the polynucleotide encoding a nuclease is codon-optimized for expression in Bacillus licheniformis cells, wherein the polynucleotide comprises or consists of a sequence having at least 80%, for example, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the nucleotide sequences of SEQ ID NO: 405, 416, 417, 434, 449, 465, or 466.
[0179] In one embodiment, the polynucleotide encoding a nuclease is codon-optimized for expression in Bacillus subtilis cells, wherein the polynucleotide comprises or consists of a sequence having at least 80%, for example, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the nucleotide sequence of any one of SEQ ID NO: 528, 549, or 550.
[0180] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Lactobacillus paracasei (Lb. paracasei, Lacticaseibacillus paracasei, or Lactobacillus paracasei) cells.
[0181] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Streptococcus thermophilus cells.
[0182] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in filamentous fungal cells.
[0183] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Aspergillus niger cells.
[0184] In one embodiment, the polynucleotide encoding a nuclease is codon-optimized for expression in Aspergillus niger cells, wherein the polynucleotide comprises or consists of a sequence having at least 80%, for example, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the nucleotide sequences in SEQ ID NO: 347, 349, 351, or 353.
[0185] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Pichia pastoris cells.
[0186] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Aspergillus oryzae cells.
[0187] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Trichoderma reesei cells.
[0188] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Lactobacillus cells.
[0189] In one embodiment, the polynucleotide encoding a nuclease is codon-optimized for expression in probiotic cells.
[0190] In one embodiment, the polynucleotide encoding the nuclease is codon-optimized for expression in Saccharomyces cerevisiae cells.
[0191] In one embodiment, the nuclease is a Class 2 Cas nuclease.
[0192] In one embodiment, the nuclease is a type II Cas nuclease.
[0193] In one embodiment, the nuclease is a type II-A Cas nuclease.
[0194] In one embodiment, the nuclease is a type II-B Cas nuclease.
[0195] In one embodiment, the nuclease is a type II-C Cas nuclease.
[0196] In one embodiment, the nuclease utilizes the prototype spacer adjacent motif (PAM) sequence provided for the nuclease in Table 1.
[0197] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “nnAY”.
[0198] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “nnGHMA”.
[0199] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “nnGTA”.
[0200] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “nnAMA”.
[0201] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “nnRHRD”.
[0202] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “ATGTCA”.
[0203] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “CCATA”.
[0204] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “TTACA”.
[0205] In one embodiment, the nuclease utilizes a prototype spacer adjacent motif (PAM) sequence with the sequence “TTACAA”.
[0206] In one embodiment, the nuclease is not naturally occurring, for example, wherein the nuclease is engineered and contains non-natural or synthetic amino acids.
[0207] In one embodiment, the nuclease is naturally occurring.
[0208] On the one hand, Cas nucleases are isolated.
[0209] On the other hand, Cas nuclease is purified.
[0210] Prototype spacer adjacent motif (PAM) sequence
[0211] The Cas nuclease disclosed herein can cleave, split, or bind to a target nucleic acid within or near a prototype spacer adjacent motif (PAM) sequence. In some embodiments, the target nucleic acid is a double-stranded nucleic acid comprising a target strand and a non-target strand. In some embodiments, cleavage occurs within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides from the 5' or 3' end of the PAM sequence. In some embodiments, the effector Cas nuclease described herein recognizes the PAM sequence. In some embodiments, recognizing the PAM sequence includes interaction with a sequence adjacent to the PAM. In some embodiments, the target nucleic acid comprises a target sequence adjacent to the PAM sequence. In some embodiments, the Cas nuclease does not require a PAM to bind to and / or cleave the target nucleic acid. Examples of identified PAM sequences are shown in Table 1.
[0212] In some embodiments, the target nucleic acid is a single-stranded target nucleic acid containing a target sequence. Therefore, in some embodiments, the single-stranded target nucleic acid contains the PAM sequence described herein, which is adjacent to (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides) or directly adjacent to the target sequence. In some embodiments, a complex containing a Cas nuclease and a guide RNA cleaves the single-stranded target nucleic acid.
[0213] In some embodiments, the target nucleic acid is a double-stranded nucleic acid comprising a target strand and a non-target strand, wherein the target strand contains a target sequence. In some embodiments, the PAM sequence is located on the target strand. In some embodiments, the PAM sequence is located on the non-target strand. In some embodiments, the PAM sequence described herein is adjacent to the target sequence on the target strand or the non-target strand (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides). In some embodiments, such a PAM is directly adjacent to the target sequence on the target strand or the non-target strand. In some embodiments, a complex comprising a Cas nuclease and a guide RNA cleaves the target strand or the non-target strand. In some embodiments, the complex cleaves both the target strand and the non-target strand. In some embodiments, the complex recognizes the PAM sequence and hybridizes with the target sequence of the target nucleic acid. In some embodiments, the complex cleaves the target nucleic acid, wherein the complex has recognized the PAM sequence and hybridized with the target sequence.
[0214] In some embodiments, the Cas nucleases described herein, or their multimeric complexes, recognize PAMs on target nucleic acids. In some embodiments, multiple Cas nucleases of the multimeric complex recognize PAMs on target nucleic acids. In some embodiments, at least two of the multiple Cas nucleases recognize the same PAM sequence. In some embodiments, at least two of the multiple Cas nucleases recognize different PAM sequences. In some embodiments, only one Cas nuclease of the multimeric complex recognizes PAMs on target nucleic acids.
[0215] The Cas nuclease or its multimeric complex disclosed herein can cleave or split a target nucleic acid within or near the prototype spacer adjacent motif (PAM) sequence. In some embodiments, the cleavage occurs within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides at the 5' or 3' end of the PAM sequence.
[0216] In some embodiments, the compositions and methods described herein do not include a PAM sequence. In some embodiments, the Cas nuclease does not recognize a PAM sequence. In some embodiments, the compositions and methods described herein include a prototype spacer flanking site (PFS) sequence. The PFS sequence can be used to detect and / or modify RNA.
[0217] Table 1. PAM sequence of novel Cas nuclease.
[0218]
[0219] Remove or reduce Cas nuclease activity
[0220] The present invention also relates to methods for generating Cas nuclease variants, which include mutating, for example deleting, inserting or substituting a polynucleotide encoding the Cas nuclease of the present invention, thereby generating variants having reduced nuclease activity (e.g., nick enzyme activity only) or no nuclease activity (e.g., Cas nuclease without catalytic activity).
[0221] Variants can be constructed by mutating polynucleotides using methods well known in the art, such as inserting one or more nucleotides, replacing one or more nucleotides, or deleting one or more nucleotides.
[0222] In one embodiment, the nuclease comprises amino acid substitutions, insertions, or deletions in one or more RuvC domains.
[0223] In one embodiment, the nuclease comprises amino acid substitutions, insertions, or deletions in one or more HNH domains.
[0224] In one embodiment, the nuclease has single-strand breakage activity against DNA target sites.
[0225] In one embodiment, the nuclease is a catalytically inactivated nuclease, for example, due to the inactivation / mutation of at least one RuvC domain and at least one HNH domain.
[0226] In one embodiment, the catalytically inactivated nuclease comprises one or more inactivated RuvC domains and one or more inactivated HNH domains.
[0227] In one embodiment, a catalytically inactivated nuclease comprising one or more inactivated RuvC domains and one or more inactivated HNH domains is generated by substitution, deletion, or insertion of one or more amino acids at positions provided for the nuclease in column 3 of Table 2 or column 3 of Table 3, respectively.
[0228] In one embodiment, the RuvC domain of SEQ ID NO:1 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D9 at position 9 of SEQ ID NO:1.
[0229] In one embodiment, the RuvC domain of SEQ ID NO:21 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D8 at position 8 of SEQ ID NO:21.
[0230] In one embodiment, the RuvC domain of SEQ ID NO:40 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D16 at position 16 of SEQ ID NO:40.
[0231] In one embodiment, the RuvC domain of SEQ ID NO:39 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D13 at position 13 of SEQ ID NO:39.
[0232] In one embodiment, the RuvC domain of SEQ ID NO:48 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D12 at position 12 of SEQ ID NO:48.
[0233] In one embodiment, the HNH domain of SEQ ID NO:1 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D613 at position 613 of SEQ ID NO:1.
[0234] In one embodiment, the HNH domain of SEQ ID NO:1 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid H614 at position 614 of SEQ ID NO:1.
[0235] In one embodiment, the HNH domain of SEQ ID NO:1 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid N637 at position 637 of SEQ ID NO:1.
[0236] In one embodiment, the HNH domain of SEQ ID NO:1 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid K640 at position 640 of SEQ ID NO:1.
[0237] In one embodiment, the HNH domain of SEQ ID NO:21 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D593 at position 593 of SEQ ID NO:21.
[0238] In one embodiment, the HNH domain of SEQ ID NO:21 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid H594 at position 594 of SEQ ID NO:21.
[0239] In one embodiment, the HNH domain of SEQ ID NO:21 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid N617 at position 617 of SEQ ID NO:21.
[0240] In one embodiment, the HNH domain of SEQ ID NO:21 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid K620 at position 620 of SEQ ID NO:21.
[0241] In one embodiment, the HNH domain of SEQ ID NO:40 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D610 at position 610 of SEQ ID NO:40.
[0242] In one embodiment, the HNH domain of SEQ ID NO:40 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid H611 at position 611 of SEQ ID NO:40.
[0243] In one embodiment, the HNH domain of SEQ ID NO:40 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid N634 at position 634 of SEQ ID NO:40.
[0244] In one embodiment, the HNH domain of SEQ ID NO:40 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid K637 at position 637 of SEQ ID NO:40.
[0245] In one embodiment, the HNH domain of SEQ ID NO:39 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D601 at position 601 of SEQ ID NO:39.
[0246] In one embodiment, the HNH domain of SEQ ID NO:39 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid H602 at position 602 of SEQ ID NO:39.
[0247] In one embodiment, the HNH domain of SEQ ID NO:39 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid N625 at position 625 of SEQ ID NO:39.
[0248] In one embodiment, the HNH domain of SEQ ID NO:39 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid K628 at position 628 of SEQ ID NO:39.
[0249] In one embodiment, the HNH domain of SEQ ID NO:48 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid D605 at position 605 of SEQ ID NO:48.
[0250] In one embodiment, the HNH domain of SEQ ID NO:48 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid H606 at position 606 of SEQ ID NO:48.
[0251] In one embodiment, the HNH domain of SEQ ID NO:48 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid N629 at position 629 of SEQ ID NO:48.
[0252] In one embodiment, the HNH domain of SEQ ID NO:48 is inactivated by a mutation (e.g., deletion, insertion, or substitution) of amino acid K632 at position 632 of SEQ ID NO:48.
[0253] In one aspect, the polypeptide is derived from SEQ ID NO:1-52 by substitution, deletion, or addition of one or more amino acids. In some embodiments, the polypeptide is a variant of SEQ ID NO:1-52 containing substitutions, deletions, and / or insertions at one or more positions. In one aspect, the number of amino acid substitutions, deletions, and / or insertions introduced into the polypeptide of SEQ ID NO:1-52 is up to 15, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. The amino acid changes can be minor, i.e., conserved amino acid substitutions or insertions that do not significantly affect the folding of the Cas nuclease; typically small deletions of 1-30 amino acids; small N-terminal or C-terminal extensions, such as methionine residues at the N-terminus; small linker peptides of up to 20-25 residues; or small extensions that facilitate purification by altering net charge or another function, such as a polyhistidine fragment, an antigenic epitope, or a binding module.
[0254] In one embodiment, the nuclease is a nicking enzyme having one or more inactive RuvC domains generated by substitution, insertion, or deletion of amino acids at the positions provided for the nuclease in column 3 of Table 2.
[0255] In some embodiments, the RuvC domain is derived from the amino acid sequences provided in column 2 of Table 2 and / or the amino acid sequences provided in column 3 of Table 2 at one or more positions by substitution, deletion, or addition of one or more amino acids. In one aspect, the number of amino acid substitutions, deletions, and / or insertions introduced into the RuvC domain can be up to 15, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. Amino acid changes can be minor, i.e., conserved amino acid substitutions or insertions that do not significantly affect protein folding; typically small deletions of 1–30 amino acids; small N-terminal or C-terminal extensions, such as methionine residues at the N-terminus; small linker peptides of up to 20–25 residues; or small extensions that facilitate purification by altering net charge or another function, such as a polyhistidine tag, antigenic epitope, or binding module.
[0256] Table 2. RuvC domain of novel Cas nucleases.
[0257]
[0258] In one embodiment, the nuclease is a nicking enzyme having one or more inactive HNH domains generated by substitution, insertion, or deletion of amino acids at the positions provided for the nuclease in column 3 of Table 3.
[0259] Table 3. HNH domain of novel Cas nucleases
[0260]
[0261] In some embodiments, the HNH domain is derived from the amino acid sequences provided in column 2 of Table 3 and / or the amino acid sequences provided in column 3 of Table 3 at one or more positions by substitution, deletion, or addition of one or more amino acids. In one aspect, the number of amino acid substitutions, deletions, and / or insertions introduced into the HNH domain can be up to 15, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. Amino acid changes can be minor, i.e., conserved amino acid substitutions or insertions that do not significantly affect protein folding; typically small deletions of 1–30 amino acids; small N-terminal or C-terminal extensions, such as methionine residues at the N-terminus; small linker peptides of up to 20–25 residues; or small extensions that facilitate purification by altering net charge or another function, such as a polyhistidine segment, an antigenic epitope, or a binding module.
[0262] Essential amino acids in peptides can be identified using procedures known in the art, such as site-directed mutagenesis or alanine scanning mutagenesis (Cunningham and Wells, 1989, Science 244: 1081-1085). In the latter technique, a single alanine mutation is introduced at each residue in the molecule, and the nuclease activity of the resulting molecule is tested to identify amino acid residues critical to the molecule's activity. See also Hilton et al., 1996, J. Biol. Chem. 271: 4699-4708. Active sites of enzymes or other biological interactions can also be determined by physical analysis of the structure, such as by techniques like nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, along with mutagenesis of the amino acids at the putative contact sites. See, for example, de Vos et al., 1992, Science 255: 306-312; Smith et al., 1992, J. Mol. Biol. 224: 899-904; Wlodaver et al., 1992, FEBS Lett. 309: 59-64. The identity of essential amino acids can also be inferred from alignment with related peptides, and / or from sequence homology and conserved catalytic mechanisms with related peptides or peptide / protein families from a common ancestor (typically possessing similar three-dimensional structures, functions, and significant sequence similarities). Alternatively or concurrently, protein structure prediction tools can be used for protein structure modeling to identify essential amino acids and / or active sites of peptides. See, for example, Jumper et al., 2021, “Highly accurate protein structure prediction with AlphaFold”, Nature 596: 583-589.
[0263] Using known mutagenesis, recombination, and / or hybridization methods, followed by relevant screening procedures, single or multiple amino acid substitutions, deletions, and / or insertions can be made and tested, such as those disclosed by Reidhaar-Olson and Sauer, 1988, Science 241: 53-57; Bowie and Sauer, 1989, Proc. Natl. Acad. Sci. USA 86: 2152-2156; WO 95 / 17413; or WO 95 / 22625. Other methods that can be used include error-prone PCR, phage display (e.g., Lowman et al., 1991, Biochemistry 30: 10832-10837; US 5,223,409; WO 92 / 06204), and region-directed mutagenesis (Derbyshire et al., 1986, Gene 46: 145; Ner et al., 1988, DNA 7:127).
[0264] In a second aspect, the present invention relates to a fusion polypeptide comprising any one of the Cas nucleases of the first aspect and one or more second polypeptides.
[0265] In one embodiment, one or more second polypeptides comprise polypeptides located in one or more subcellular organelles.
[0266] In one embodiment, one or more second polypeptides are nuclear localization sequences (NLS), cell-penetrating peptides, and / or affinity tags.
[0267] In one embodiment, the fusion polypeptide comprises 1-10 or more NLS at or near the amino terminus, 1-10 or more NLS at or near the carboxyl terminus, or a combination of 1-10 or more NLS at or near the amino terminus and 1-10 or more NLS at or near the carboxyl terminus.
[0268] In one embodiment, the fusion peptide contains 1-4 NLS.
[0269] In one embodiment, the fusion peptide comprises an NLS.
[0270] In one embodiment, one or more NLS are located within the open reading frame (ORF) of the nuclease.
[0271] In one embodiment, one or more NLSs are repeated in series.
[0272] In one embodiment, the fusion peptide comprises a first NLS and a second NLS.
[0273] In one embodiment, the fusion peptide comprises a linker sequence between the first NLS and the second NLS.
[0274] In one embodiment, the linker between the first NLS and the second NLS contains at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 amino acids.
[0275] In one embodiment, one or more second polypeptides comprise base-editing polypeptides.
[0276] In one embodiment, the base-editing polypeptide includes a base editor domain.
[0277] In one embodiment, the fusion polypeptide includes a linker between the Cas nuclease and the base-editing polypeptide.
[0278] In one embodiment, the base-editing polypeptide contains a deaminase, such as cytidine deaminase, such as APOBEC3A deaminase or adenosine deaminase.
[0279] In one embodiment, one or more second polypeptides contain a reverse transcriptase, which preferably contains a reverse transcriptase domain.
[0280] In one embodiment, the nuclease is fused with one or more NLS of sufficient strength to drive the accumulation of a CRISPR complex containing the Cas nuclease in the nucleus of a eukaryotic cell to a detectable amount.
[0281] In one embodiment, sequence identity is determined by the method described in the definition section of "Sequence Identity".
[0282] On the one hand, the fusion peptides are separate.
[0283] On the other hand, the fusion peptides are purified.
[0284] Sources of Cas nuclease
[0285] The Cas nucleases of this invention can be obtained from microorganisms of any genus. For the purposes of this invention, the term "obtained from," as used herein in conjunction with a given source, should mean that the polypeptide encoded by the polynucleotide is produced by that source or by a strain that has inserted the polynucleotide of this invention. In one aspect, polypeptides obtained from a given source are not secreted extracellularly.
[0286] In one respect, Cas nucleases are obtained from bacterial cells.
[0287] In one embodiment, the Cas nuclease is a polypeptide obtained from Streptococcus cells. In one embodiment, the Streptococcus cells are Streptococcus equinus, Streptococcus mutans, Streptococcus species, Streptococcus henryi DSM19005, Streptococcus species CCH8-G7, Streptococcus pacificus, Streptococcus orisratti DSM 15617, Streptococcus salivarius, or Streptococcus ruminantium cells.
[0288] In one embodiment, the Cas nuclease is a polypeptide obtained from Bacillus cells (e.g., Bacillus species-63030 cells).
[0289] In one embodiment, the Cas nuclease is a polypeptide obtained from Turicibacter cells (e.g., Turicibacter species cells).
[0290] In one embodiment, the Cas nuclease is a polypeptide obtained from Ureibacillus cells (e.g., Ureibacillus thermophilus cells).
[0291] In one embodiment, the Cas nuclease is a polypeptide obtained from Lentihominibacter cells (e.g., human Lentihominibacter hominis cells).
[0292] In one embodiment, the Cas nuclease is a polypeptide obtained from Clostridia cells.
[0293] In one embodiment, the Cas nuclease is a polypeptide obtained from Ruminococcus cells (e.g., Ruminococcus species cells).
[0294] In one embodiment, the Cas nuclease is a polypeptide obtained from Alicyclobacillus cells (e.g., Alicyclobacillus sacchari cells).
[0295] In one embodiment, the Cas nuclease is a polypeptide obtained from Enterococcus cells (e.g., Enterococcus gilvus, Enterococcus hermannii, or Enterococcus spp. cells).
[0296] In one embodiment, the Cas nuclease is a polypeptide obtained from Companilacbtobacillus cells (e.g., Companilacbtobacillus zhachilii, Companilacbtobacillus halodurans, Companilacbtobacillus keshanensis, Companilacbtobacillus suantsaicola, or Companilacbtobacillus hulinensis cells).
[0297] In one embodiment, the Cas nuclease is a polypeptide obtained from Bombilactobacillus cells (e.g., Bombilactobacillus apium cells).
[0298] In one embodiment, the Cas nuclease is a polypeptide obtained from Vagococcus cells (e.g., Vagococcus penaei cells).
[0299] In one embodiment, the Cas nuclease is a polypeptide obtained from Lactobacillus cells. In one embodiment, the Lactobacillus cells are Lactobacillus species, Lactobacillus farciminis (DSM 20184), Lactobacillus murinus, Lactobacillus ruminis, Lactobacillus salivarius, Lactobacillus jensenii, Lactobacillus hamster, Lactobacillus delbrueckii, Lactobacillus johnsonii, Lactobacillus plantarum, Lactobacillus rhamnosus, or Lactobacillus gallinarum cells.
[0300] In one embodiment, the Cas nuclease is obtained from or can be obtained from Streptococcus cells (e.g., Streptococcus equi, Streptococcus mutans, Streptococcus species, Streptococcus henryi DSM 19005, Streptococcus species CCH8-G7, Streptococcus paclitaxel, Streptococcus mutans DSM 15617, Streptococcus salivarius, or ruminant streptococcus cells), Bacillus cells (e.g., Bacillus species-63030 cells), Zurich bacillus cells (e.g., Zurich bacillus species cells), Ureaplasma cells (e.g., Ureaplasma thermophilus cells), Chlorobacterium cells (e.g., Chlorobacterium humanis cells, Clostridium cells), Ruminococcus cells (e.g., Ruminococcus species cells), and Cyclocycline Bacillus cells (e.g., Cyclocycline Bacillus cells). The cells may contain *Bacillus* cells, *Enterococcus* cells (e.g., *Enterococcus faecalis*, *Enterococcus harzianum*, or *Enterococcus asper* cells), *Lactobacillus* cells (e.g., *Lactobacillus aspera*, *Lactobacillus halophilus*, *Lactobacillus keshanensis*, *Lactobacillus sauerkraut*, or *Lactobacillus hulinensis* cells), *Lactobacillus bumblebee* cells (e.g., *Lactobacillus bumblebee* cells), or *Zygococcus* cells (e.g., *Zygococcus prawni* cells), preferably *Enterococcus asper* cells, *Enterococcus harzianum* cells, *Zygococcus prawni* cells, or *H. cirrhosum* cells.
[0301] In one embodiment, the Cas nuclease is obtained from or can be obtained from Lactobacillus cells, such as Lactobacillus species, Lactobacillus sausageii (DSM 20184), Lactobacillus sausageii, Lactobacillus murineis, Lactobacillus rumenii, Lactobacillus salivarius, Lactobacillus janniae, Lactobacillus hamsterii, Lactobacillus delbrueckii, Lactobacillus johnsonii, Lactobacillus plantarum, Lactobacillus rhamnosus, or Lactobacillus chrysogenus cells.
[0302] It should be understood that, for the aforementioned species, this invention covers complete and incomplete stages, as well as other taxonomic equivalents, such as asexual forms, regardless of their known species names. Those skilled in the art will readily identify the appropriate equivalents.
[0303] The probes described above can be used to identify and obtain Cas nucleases from other sources, including microorganisms isolated from nature (e.g., soil, compost, water, etc.) or DNA samples obtained directly from natural materials (e.g., soil, compost, water, etc.). Techniques for directly isolating microorganisms and DNA from natural habitats are well known in the art. The polynucleotide encoding the Cas nuclease can then be obtained by similarly screening a library of genomic DNA or cDNA from another microorganism or a mixed DNA sample. Once the polynucleotide encoding the Cas nuclease has been detected with one or more probes, it can be isolated or cloned using techniques known to those skilled in the art (see, for example, Davis et al., 2012, Basic Methods in Molecular Biology, Elsevier).
[0304] AlphaFold structural prediction
[0305] AlphaFold is a computational method for predicting the three-dimensional structure of peptides based on their amino acid sequences (Jumper et al., Highly accurate protein structure prediction with AlphaFold. Nature, 2021). The predicted structures of millions of peptides stored in the UniProt database are already available in the AlphaFold protein structure database, using the AlphaFold monomer v2.0 model (Varadi et al., AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Research, 2021). In the AlphaFold protein structure database, the three-dimensional structure of a peptide can be obtained by searching for its UniProt accession number.
[0306] In addition to many publicly available 3D structures, the code can be used to reproduce and predict the structures of novel peptides in source code repositories, such as notebooks / AlphaFold.ipynb at deepmind / alphafold / on Github.com using Alphafold v2.3.1 or later. Alternatively, it can be found at sokrypton / ColabFold on Github.com using v1.5.2 or later (using AlphaFold2.ipynb). For technical details, see Jumper et al. (see above).
[0307] AlphaFold generates per-residue estimates of its confidence level on a scale of 0 to 100. This confidence level metric is called pLDDT and corresponds to the model's prediction score on the lDDT-Cα index. It is stored in the B-factor field of the downloadable mmCIF and PDB files (although, unlike the B-factor, a higher pLDDT score is better). Regions with pLDDT scores above 90 are expected to be modeled with high accuracy. These should be suitable for any application that benefits from high accuracy (e.g., characterization of binding sites). Regions with pLDDT scores between 70 and 90 are expected to be well-modeled, corresponding to generally good main-chain predictions.
[0308] Structural similarity
[0309] The correlation between two amino acid sequences is typically described by the parameter "sequence identity". However, since the biological function of a polypeptide is defined by its three-dimensional structure rather than its amino acid sequence, a better approach to assess the functional relationship between polypeptides is to compare their three-dimensional structures. Therefore, for the purposes of this invention, the correlation between the three-dimensional structures of two polypeptides is described by the parameter "structural similarity".
[0310] The three-dimensional structure of any polypeptide can be obtained experimentally, for example by X-ray crystallography or using computational methods such as AlphaFold (see above). The structural similarity between the three-dimensional structures can then be determined by the TM-score, which is calculated using the following general formula (Zhang and Skolnick, Proteins 57:702-710, 2004):
[0311]
[0312] Where L N It is the length of the natural structure, L T It is the length of the residues compared with the template structure, d i d is the distance between residues in the i-th pair, and d0 is the scale of the normalized matching difference. “Max” represents the maximum value after optimal spatial superposition.
[0313] For the purposes of this invention, L N Always referencing the protein length indicates the use of a fixed reference length L to prevent artificially large TM-scores in substructure alignment:
[0314]
[0315] Before the TM-score can be calculated, structural alignment of the three-dimensional structures of the two peptides is necessary. This is achieved by optimizing the algorithm for structural overlap, and several methods are available, such as CEalign (Shindyalov and Bourne, ProteinEng., 11, 739-747, 1998), DALI (Holm and Sander, Trends Biochem., 20, 478-480, 1995), or TM-align (Nucleic Acids Res. 33:2302-2309, 2005).
[0316] For the purposes of this invention, TM-align is applied. For convenience, the TM-score is integrated into the TM-align software, which is available from the authors' website. The version of TM-align is preferably the latest version released on August 22, 2019, or later, and the TM-score between the reference protein and the query protein is determined by running the following command:
[0317] TMalign<query.pdb><reference.pdb> -L<length of reference>
[0318] in<query.pdb> It is the name of the PDB file containing the coordinates of the queried peptide.<reference.pdb> This is the name of the PDB file containing the coordinates of the reference peptide. The TM-score is calculated and reported in the output, along with several other parameters from the alignment.
[0319] Composition
[0320] In a third aspect, the invention also relates to a non-naturally occurring composition comprising (i) a Cas nuclease of the first aspect or a fusion polypeptide of the second aspect, and / or (ii) a nucleic acid molecule comprising a sequence encoding the Cas nuclease of the first aspect or the fusion polypeptide of the second aspect.
[0321] In one embodiment, the nucleic acid molecule is a chemically modified nucleic acid molecule.
[0322] In one embodiment, the nucleic acid molecule is DNA.
[0323] In one embodiment, the nucleic acid molecule is RNA.
[0324] In one embodiment, the RNA is mRNA that includes one or more of the following: a 5' untranslated region (UTR), an open reading frame (ORF) encoding a Cas nuclease or a fusion polypeptide, a 3' UTR, and a polyadenylated (polyA) tail.
[0325] In one embodiment, the ORF is composed of a nucleoside selected from adenosine, modified adenosine, uridine, modified uridine, guanosine, modified guanosine, cytidine, and modified cytidine.
[0326] In one embodiment, the ORF consists of a nucleoside selected from adenosine, uridine, modified uridine, guanosine, and cytidine.
[0327] In one embodiment, nucleic acid molecules are linear.
[0328] In one embodiment, the nucleic acid molecule is circular.
[0329] In one embodiment, the composition further comprises one or more RNA molecules, or DNA polynucleotides encoding one or more of one or more RNA molecules, wherein the one or more RNA molecules and the Cas nuclease or fusion polypeptide do not naturally coexist, and the one or more RNA molecules are configured to form a complex with the Cas nuclease or fusion polypeptide and / or target the complex to a target site.
[0330] In one embodiment, one or more RNA molecules contain guide RNA (gRNA) that includes CRISPR RNA (crRNA) and trans-activating RNA (tracrRNA).
[0331] In one embodiment, one or more RNA molecules are single-molecule RNAs (sgRNAs), for example, where crRNA and tracrRNA are part of the same RNA molecule.
[0332] In another embodiment, one or more RNA molecules are bimolecular RNAs, for example, where crRNA and tracrRNA are separate RNA molecules.
[0333] In one embodiment, the composition further comprises a donor template for homologous directional repair (HDR).
[0334] In one embodiment, the sequence encoding the Cas nuclease or fusion polypeptide comprises sequences corresponding to SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82. SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, SEQ ID NO: 89, SEQ ID NO: 90, SEQ ID NO: 91, SEQ ID NO: 92, SEQ ID NO: 93, SEQ ID NO: 94, SEQ ID NO: 95. SEQ ID NO: 96, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103 or SEQ ID NO: 104, or SEQ ID NO: The polynucleotide sequence of any one of 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549 or 550 has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity.
[0335] In one embodiment, the sequence encoding the Cas nuclease or fusion polypeptide comprises a polynucleotide sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the polynucleotide sequences in SEQ ID NO: 53, SEQ ID NO: 73, SEQ ID NO: 92, SEQ ID NO: 91, SEQ ID NO: 100, or SEQ ID NO: 81.
[0336] In one embodiment, the sequence encoding the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 53.
[0337] In one embodiment, the sequence encoding the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 73.
[0338] In one embodiment, the sequence encoding the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 92.
[0339] In one embodiment, the sequence encoding the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 91.
[0340] In one embodiment, the sequence encoding the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 100.
[0341] In one embodiment, the sequence encoding the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 81.
[0342] In one embodiment, one or more RNA molecules comprise trans-activating RNA (tracrRNA) encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of SEQ ID NO: 157-208.
[0343] In one embodiment, one or more RNA molecules comprise a trans-activating RNA (tracrRNA) encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of SEQ NO: 157, SEQ ID NO: 177, SEQ ID NO: 196, SEQ ID NO: 195, SEQ ID NO: 204, or SEQ ID NO: 185.
[0344] In one embodiment, at least one of one or more RNA molecules comprises a CRISPR RNA (crRNA) molecule containing a guide sequence portion and a sequence encoded by a polynucleotide, the polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of SEQ ID No: 209-260.
[0345] In one embodiment, at least one of one or more RNA molecules comprises a CRISPR RNA (crRNA) molecule containing a guide sequence portion and a sequence encoded by a polynucleotide, the polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of SEQ ID NO: 209, SEQ ID NO: 229, SEQ ID NO: 248, SEQ ID NO: 247, SEQ ID NO: 256, or SEQ ID NO: 237.
[0346] In one embodiment, at least one of one or more RNA molecules comprises or is composed of an RNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of SEQ ID NO:261-312.
[0347] In one embodiment, at least one of one or more RNA molecules comprises or is composed of an RNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide, the polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of SEQ ID NO:261, SEQ ID NO:281, SEQ ID NO:300, SEQ ID NO:299, SEQ ID NO:308, or SEQ ID NO:289.
[0348] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any amino acid sequence in column 4 of Table 4 (e.g., any one of SEQ ID NO: 261-312), having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence in column 4 of Table 4 (e.g., any one of SEQ ID NO: 261-312).
[0349] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 1, and at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 209.
[0350] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 21, and at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 229.
[0351] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 40, and at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 248.
[0352] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 39, and at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 247.
[0353] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 48, and at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 256.
[0354] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 237, and at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 237.
[0355] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the amino acid sequences in column 1 of Table 4 (e.g., any of SEQ ID NO: 209-260), having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the polynucleotide sequences in column 2 of Table 4 (e.g., any of SEQ ID NO: 209-260).
[0356] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 157, and at least one RNA molecule comprises a tracrRNA molecule containing a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 157.
[0357] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 21, and at least one RNA molecule comprises a tracrRNA molecule containing a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 177.
[0358] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 40, and at least one RNA molecule comprises a tracrRNA molecule containing a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 196.
[0359] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 39, and at least one RNA molecule comprises a tracrRNA molecule containing a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 195.
[0360] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 48, and at least one RNA molecule comprises a tracrRNA molecule containing a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 204.
[0361] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 29, and at least one RNA molecule comprises a tracrRNA molecule containing a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 185.
[0362] In one embodiment, the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the polynucleotide sequences in column 3 of Table 4 (e.g., any of SEQ ID NO: 157-208), having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the polynucleotide sequences in column 3 of Table 4 (e.g., any of SEQ ID NO: 157-208).
[0363] In one embodiment, the composition further comprises a base editor enzyme.
[0364] In one embodiment, the base editor enzyme is adenosine deaminase or cytidine deaminase.
[0365] In one embodiment, the composition further comprises reverse transcriptase.
[0366] Table 4 discloses which crRNA and tracrRNA coding sequences are associated with each novel Cas nuclease. For example, Cas nuclease 0076, having SEQ ID NO: 1, utilizes a crRNA sequence encoded by SEQ ID NO: 209 and a tracrRNA sequence encoded by SEQ ID NO: 157. Additionally and / or alternatively, Cas nucleases having SEQ ID NO: 1 utilize a gRNA or sgRNA sequence encoded by SEQ ID NO: 261, which comprises both the crRNA sequence encoded by SEQ ID NO: 209 and the tracrRNA sequence encoded by SEQ ID NO: 157.
[0367] Table 4. Coding sequences of crRNA, tracrRNA, and guide RNA of novel Cas nucleases
[0368]
[0369] Methods for modifying DNA target sites
[0370] In a fourth aspect, the present invention also relates to a method for modifying nucleotide sequences at DNA target sites in the genome of a cell, the method comprising introducing a Cas nuclease of the first aspect or a fusion polypeptide of the second aspect, a polynucleotide encoding a Cas nuclease of the first aspect or a fusion polypeptide of the second aspect, and / or a composition of the third aspect into the cell.
[0371] In one embodiment, the method includes introducing DNA breaks at a DNA target site.
[0372] In one embodiment, the DNA break is a single-strand break.
[0373] In one embodiment, the DNA break is a double-strand break.
[0374] In one embodiment, the method is performed under conditions that allow non-homologous end-joining (NHEJ) and homologous directional repair (HDR).
[0375] In one embodiment, the method is performed under conditions that allow non-homologous end connections (NHEJ).
[0376] In one embodiment, the method is performed under conditions that allow for homologous directional inpainting (HDR).
[0377] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in a DNA strand adjacent to the PAM sequence (e.g., adjacent to the PAM sequence “nnAY”, or adjacent to any of the PAM sequences mentioned in Table 1).
[0378] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in the DNA strand adjacent to the PAM sequence “nnGHMA” (e.g., “nnGHCA”, “nnGGCA”, or “nnGCMA”).
[0379] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in the DNA strand adjacent to the PAM sequence “nnGTA” (e.g., an Aspergillus DNA strand).
[0380] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in the DNA strand adjacent to the PAM sequence “nnAMA” (e.g., “nnAAA”).
[0381] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in the DNA strand adjacent to the PAM sequence “nnRHRD” (e.g., “nnACAG”, “nnACAR”, “nnATGT”, or “nnAYRD”).
[0382] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in a DNA strand (e.g., Bacillus subtilis DNA) adjacent to the PAM sequence “ATGTCA”, “CCATA”, “TTACA”, or “TTACAA”.
[0383] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in the Aspergillus DNA strand adjacent to the PAM sequence “nnGHCA”, “nnGTA”, or “nnACAG”.
[0384] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in the Bacillus DNA strand adjacent to the PAM sequence “nnGHMA”, “nnGGCA”, “nnACAR”, “nnATGT”, or “nnAMA”.
[0385] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in the Escherichia coli DNA strand adjacent to the PAM sequence “nnGCMA”, “nnAAA”, or “nnAYRD”.
[0386] In one embodiment, the Cas nuclease or fusion polypeptide achieves DNA breakage in a DNA strand adjacent to a sequence complementary to the PAM sequence.
[0387] In one embodiment, the target site is located within the coding region of the protein.
[0388] In one embodiment, the target site is located in the non-coding region of the protein.
[0389] In one embodiment, the target site is located within the regulatory region of the protein (e.g., the promoter).
[0390] In one embodiment, the cell is a eukaryotic cell.
[0391] In one embodiment, the cell is a prokaryotic cell.
[0392] In one embodiment, the cell is a eukaryotic cell, such as a mammalian cell, a human cell, or a non-human mammalian cell, such as a BHK cell, a CHO cell, a mouse cell, a hamster cell, or a rat cell.
[0393] In a preferred embodiment, the cell is a fungal cell, such as a filamentous fungal cell or a yeast cell.
[0394] In one embodiment, the fungal cell is a Pichia pastoris cell, for example, a Pichia pastoris cell.
[0395] In one embodiment, the cell is a yeast cell, such as a cell of the genera *Candida*, *Hansenula*, *Kluyveromyces*, *Pichia*, *Saccharomyces*, *Schizosaccharomyces*, or *Yarrowia*, such as *Kluyveromyces lactis*, *Saccharomyces carlsbergensis*, *Saccharomyces cerevisiae*, *Saccharomyces diastaticus*, *Saccharomyces douglasii*, *Saccharomyces kluyveri*, *Saccharomyces norbensis*, *Saccharomyces oviformis*, or *Yarrowia lipolytica*.
[0396] In one embodiment, the cells are filamentous fungal cells, such as those from the genera *Acremonium*, *Aspergillus*, *Aureobasidium*, *Bjerkandera*, *Ceriporiopsis*, *Chrysosporium*, *Coprinus*, *Coriolus*, *Cryptococcus*, *Filibasidium*, *Fusarium*, *Humicola*, *Magnaporthe*, *Mucor*, *Myceliophthora*, and *Neo*. Cells of the genera *Calmastix*, *Neurospora*, *Paecilomyces*, *Penicillium*, *Phanerochaete*, *Phlebia*, *Piromyces*, *Pleurotus*, *Schizophyllum*, *Talaromyces*, *Thermoascus*, *Thielavia*, *Tolypocladium*, *Trametes*, or *Trichoderma*, especially *Aspergillus*. Aspergillus awamori, Aspergillus foetidus, Aspergillus fumigatus, Aspergillus japonicus, Aspergillus nidulans, Aspergillus niger, Aspergillus oryzae, Bjerkandera adusta, Ceriporiopsisaneirina, Ceriporiopsis caregiea, Ceriporiopsis gilvescens, Ceriporiopsis pannocinta, Ceriporiopsis rivulosa, Ceriporiopsis subrufa, Ceriporiopsis subvermispora, ChrysosporiumThe following fungi are listed: *Chrysosporium keratinophilum*, *Chrysosporium lucknowense*, *Chrysosporium merdarium*, *Chrysosporium pannicola*, *Chrysosporium queenslandicum*, *Chrysosporium tropicum*, *Chrysosporium zonatum*, *Coprinus cinereus*, *Coriolushirsutus*, *Fusarium bactridioides*, *Fusarium cerealis*, *Fusarium crookwellense*, *Fusarium culmorum*, *Fusarium graminearum*, and *Fusarium graminearum*. Fusarium graminum, Fusarium heterosporum, Fusarium negundi, Fusarium oxysporum, Fusarium reticulatum, Fusarium roseum, Fusarium sambucinum, Fusarium sarcochroum, Fusarium sporotrichioides, Fusarium sulphureum, Fusarium torulosum, Fusarium trichothecioides, Fusarium venenatum, Humicola insolens, Humicola lanuginosa, Mucor miehei, Myceliophthora thermophila, Neurospora *Penicillium purpurogenum*, *Phanerochaete chrysosporium*, *Phlebia*The cells of *Trichoderma radiata*, *Pleurotus eryngii*, *Talaromyces emersonii*, *Thielavia terrestris*, *Trametes villosa*, *Trametes versicolor*, *Trichoderma harzianum*, *Trichoderma koningii*, *Trichoderma longibrachiatum*, *Trichoderma reesei*, or *Trichoderma viride*.
[0397] In one embodiment, the cells are Trichoderma cells.
[0398] In one embodiment, the cells are Trichoderma reesei cells.
[0399] In one embodiment, the cells are Aspergillus cells.
[0400] In one embodiment, the cells are Aspergillus niger cells.
[0401] In one embodiment, the cells are Aspergillus oryzae cells.
[0402] In one embodiment, the cell is a plant cell.
[0403] In one embodiment, the plant cell is one or more of the following: corn, rice, sorghum, rye, barley, wheat, millet, oats, sugarcane, turfgrass, switchgrass, soybean, canola, alfalfa, sunflower, cotton, tobacco, peanut, potato, tobacco, Arabidopsis, vegetable, or safflower cells.
[0404] In a preferred embodiment, the cells are prokaryotic cells, for example, Gram-positive cells selected from the group consisting of: Bacillus, Clostridium, Corynebacterium, Enterococcus, Geobacillus, Lactobacillus, Lacticaseibacillus, Lactiplantibacillus, Levilactobacillus, Ligilactobacillus, Limosilactobacillus, Lactococcus, Oceanobacillus, and Staphylococcus. Cells of Staphylococcus, Streptococcus, or Streptomyces, or Gram-negative bacteria selected from the group consisting of: Campylobacter, Escherichia coli, Flavobacterium, Fusobacterium, Helicobacter, Ilyobacter, Neisseria, Pseudomonas, Salmonella, and Ureaplasma, such as Lactobacillus casei, Lactobacillus paracasei, and Lactobacillus rhamnosus. Lactobacillus rhamnosus, Lactiplantibacillus plantarum, Lactobacillus brevis, Lactobacillus salivarius, Lactobacillus fermentum, Lactobacillus reuteri, Lactobacillus acidophilus, Lactobacillus bulgaricus, Lactobacillus crispatus, Lactobacillus gasseri, Lactobacillus johnsonii, Lactobacillus helveticus*Helveticus*, *Corynebacterium glutamicum*, *Bacillus alkalophilus*, *Bacillus amyloliquefaciens*, *Bacillus brevis*, *Bacillus circulans*, *Bacillus clausii*, *Bacillus scoagulans*, *Bacillus firmus*, *Bacillus lautus*, *Bacillus lentus*, *Bacillus licheniformis*, *Bacillus megaterium*, *Bacillus pumilus*, *Bacillus stearothermophilus*, *Bacillus subtilis*, *Bacillus thuringiensis* Streptococcus thuringiensis), Streptococcus equisimilis, Streptococcus pyogenes, Streptococcus uberis, Streptococcus equi subsp. Zooepidemicus, Strepttomyces achromogenes, Strepttomyces avermitilis, Strepttomyces coelicolor, Strepttomyces griseus, and Strepttomyces lividans cells.
[0405] In one embodiment, the cell is a Bacillus cell.
[0406] In one embodiment, the cell is a Bacillus subtilis cell.
[0407] In one embodiment, the cell is a Bacillus licheniformis cell.
[0408] In one embodiment, the cells are Lactobacillus paracasei cells.
[0409] In one embodiment, the cell is a thermophilic streptococcal cell.
[0410] In one embodiment, the cells are Escherichia coli cells.
[0411] DNA repair via non-homologous end joining
[0412] Following target recognition, Cas nucleases induce double-strand breaks in the target sequence, which can lead to frameshift mutations and gene knockdown when repaired via non-homologous end joining (NHEJ). Frameshift mutations caused by error-prone NHEJ can include nucleotide insertions or deletions (insertions or deletions). Alternatively, homology-directed repair (HDR) at the double-strand break site can allow the insertion of the desired sequence.
[0413] DNA repair through homologous recombination
[0414] The term "homology-directed repair" or "HDR" refers to a mechanism used to repair DNA damage in cells, such as during the repair of double-strand and single-strand breaks in DNA. HDR requires nucleotide sequence homology and uses a "nucleic acid template" (the terms "nucleic acid template" or "donor template" are used interchangeably herein) to repair the sequence where double-strand or single-strand breaks occur (e.g., the DNA target sequence). This results in the transfer of genetic information from, for example, the nucleic acid template to the DNA target sequence. If the nucleic acid template sequence differs from the DNA target sequence, and part or all of the nucleic acid template polynucleotide or oligonucleotide is incorporated into the DNA target sequence, HDR may result in alterations to the DNA target sequence (e.g., insertions, deletions, mutations). In some embodiments, the entire nucleic acid template polynucleotide, a portion of the nucleic acid template polynucleotide, or a copy of the nucleic acid template is integrated at a site in the DNA target sequence.
[0415] The terms "nucleic acid template" and "donor" refer to nucleotide sequences that are inserted into or copied into the genome. A nucleic acid template comprises, for example, a nucleotide sequence of one or more nucleotides, which is added to a target nucleic acid, used as a template for changes to the target nucleic acid, or can be used to modify the target sequence. The nucleic acid template sequence can be of any length, for example, between 2 and 10,000 nucleotides (or any integer value between or greater than this), preferably between about 100 and 1,000 nucleotides (or any integer between this), and more preferably between about 200 and 500 nucleotides. The nucleic acid template can be a single-stranded nucleic acid or a double-stranded nucleic acid. In some embodiments, the nucleic acid template comprises, for example, a nucleotide sequence of one or more nucleotides corresponding to the wild-type sequence of a target nucleic acid at, for example, a target location. In some embodiments, the nucleic acid template comprises, for example, a ribonucleotide sequence of one or more ribonucleotides corresponding to the wild-type sequence of a target nucleic acid at, for example, a target location. In some embodiments, the nucleic acid template comprises modified ribonucleotides.
[0416] The insertion of exogenous sequences (also known as "donor sequences," "donor templates," or "donors") can also be performed, for example, to correct mutated genes or to increase the expression of wild-type genes. Obviously, the donor sequence is typically different from the genome sequence to which it is placed. The donor sequence may contain non-homologous sequences flanked by two homologous regions to allow for efficient HDR at the target location. Additionally, the donor sequence may comprise a vector molecule containing a sequence disproportionate to the target region in the cellular chromatin. The donor molecule may contain several discontinuous regions homologous to the cellular chromatin. For example, for targeted insertion of sequences not normally present in the target region, the sequence may be present in the donor nucleic acid molecule and flanked by regions homologous to sequences in the target region.
[0417] The donor polynucleotide can be single-stranded and / or double-stranded DNA or RNA and can be introduced into the cell in a linear or circular form. See, for example, U.S. Patent Publications 2010 / 0047805; 2011 / 0281361; 2011 / 0207221; and 2019 / 0330620. If introduced in a linear form, the ends of the donor sequence can be protected by methods known to those skilled in the art (e.g., to prevent exonuclease degradation). For example, one or more dideoxynucleotide residues can be added to the 3' end of a linear molecule and / or a self-complementary oligonucleotide can be attached to one or both ends. See, for example, Chang and Wilson, Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] (1987); Nehls et al., Science [Science] (1996). Other methods to protect exogenous polynucleotides from degradation include, but are not limited to, adding terminal amino groups and using modified internucleotide bonds, such as thiophosphates, aminophosphates, and O-methylribose or deoxyribose residues.
[0418] Therefore, embodiments of the present invention that use donor templates for repair can use DNA or RNA, single-stranded and / or double-stranded donor templates, which can be introduced into cells in a linear or circular form.
[0419] In embodiments of the present invention, the gene editing composition comprises: (1) an RNA molecule containing a guide sequence that affects double-strand breaks in a gene prior to repair, and (2) a donor RNA template for repair, wherein the RNA molecule containing the guide sequence is a first RNA molecule and the donor RNA template is a second RNA molecule. In some embodiments, the guide RNA molecule and the template RNA molecule are linked as part of a single molecule.
[0420] Donor sequences can also be oligonucleotides and used for gene correction or targeted alteration of endogenous sequences. Oligonucleotides can be introduced into cells via vectors, electroporated into cells, or introduced by other methods known in the art. Oligonucleotides can be used to correct mutated sequences in endogenous genes (e.g., sickle mutations in β-globin) or to insert sequences with the desired purpose into endogenous gene loci.
[0421] Polynucleotides can be introduced into cells as part of a vector molecule that has additional sequences, such as origin of replication, promoter, and genes encoding antibiotic resistance. Furthermore, donor polynucleotides can be introduced as naked nucleic acids, as nucleic acids complexed with agents such as liposomes or poloxamer, or delivered via recombinant viruses (e.g., adenoviruses, AAVs, herpesviruses, retroviruses, lentiviruses, and integrase-deficient lentiviruses (IDLVs)).
[0422] Donors are typically inserted in such a manner that their expression is driven by an endogenous promoter at the integration site—that is, a promoter that drives the expression of the endogenous gene into which the donor is inserted. However, it is evident that the donor may contain promoters and / or enhancers, such as constitutive promoters or inducible or tissue-specific promoters.
[0423] Donor molecules can be inserted into endogenous genes to render the endogenous gene wholly, partially, or completely unexpressed. For example, a transgene as described herein can be inserted into an endogenous locus to render a portion (the N-terminus and / or C-terminus of the transgene) of the endogenous sequence expressed or unexpressed, for example, as a fusion with the transgene. In other embodiments, the transgene (e.g., having or not having an additional coding sequence, such as an endogenous gene) is integrated into any endogenous locus (e.g., a safe harbor locus, such as the CCR5 gene, CXCR4 gene, PPP1R12c (also known as AAVS1) gene, albumin gene, or Rosa gene). See, for example, U.S. Patent Nos. 7,951,925 and 8,110,379; U.S. Publications Nos. 2008 / 0159996, 20100 / 0218264, 2010 / 0291048, 2012 / 0017290, 2011 / 0265198, 2013 / 0137104, 2013 / 0122591, 2013 / 0177983 and 2013 / 0177960; and U.S. Provisional Application No. 61 / 823,689.
[0424] When an endogenous sequence (either endogenous or part of a transgene) is expressed together with a transgene, the endogenous sequence can be a full-length sequence (wild-type or mutant) or a partial sequence. Preferably, the endogenous sequence is functional. Non-limiting examples of the function of these full-length or partial sequences include increasing the serum half-life of a polypeptide expressed by a transgene (e.g., a therapeutic gene) and / or serving as a vector.
[0425] In addition, although not essential for expression, exogenous sequences may also include transcriptional or translational regulatory sequences, such as promoters, enhancers, insulators, internal ribosome entry sites, sequences encoding 2A peptides, and / or polyadenylation signals.
[0426] In some embodiments, the donor molecule comprises sequences selected from the group consisting of: genes encoding proteins (e.g., sequences encoding proteins lacking in cells or individuals or alternative forms of genes encoding proteins), regulatory sequences, and / or sequences encoding structural nucleic acids (such as microRNAs or siRNAs).
[0427] Polynucleotides encoding Cas nuclease
[0428] In a fifth aspect, the invention also relates to polynucleotides encoding the Cas nuclease of the first aspect and / or the fusion polypeptide of the second aspect.
[0429] In one embodiment, the polynucleotide comprises or consists of the following: SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, SEQ ID NO: 89, SEQ ID NO: 90, SEQ ID NO: 91, SEQ ID NO: 92, SEQ ID NO: 93, SEQ ID NO: 94, SEQ ID NO: 95, SEQ ID NO: 96, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103, SEQ ID NO: 104, or SEQ ID NO: The polypeptide coding sequence of any one of 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549 or 550 has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity.
[0430] In one embodiment, the polynucleotide comprises or consists of the following: a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any polypeptide coding sequence of SEQ ID NO: 53, SEQ ID NO: 73, SEQ ID NO: 92, SEQ ID NO: 91, SEQ ID NO: 100, or SEQ ID NO: 81.
[0431] In one embodiment, the polynucleotide is a chemically modified nucleic acid molecule.
[0432] In one embodiment, the polynucleotide is DNA.
[0433] In one embodiment, the polynucleotide is RNA.
[0434] In one embodiment, the RNA is mRNA that includes one or more of the following: a 5' untranslated region (UTR), an open reading frame (ORF) encoding a Cas nuclease or a fusion polypeptide, a 3' UTR, and a polyadenylated (polyA) tail.
[0435] In one embodiment, the ORF is composed of a nucleoside selected from adenosine, modified adenosine, uridine, modified uridine, guanosine, modified guanosine, cytidine, and modified cytidine.
[0436] In one embodiment, the ORF consists of a nucleoside selected from adenosine, uridine, modified uridine, guanosine, and cytidine.
[0437] In one embodiment, the polynucleotide is linear.
[0438] In one embodiment, the polynucleotide is cyclic.
[0439] In one embodiment, the polyA sequence comprises a non-adenine nucleotide.
[0440] In one embodiment, the polyA sequence contains 100-400 nucleotides.
[0441] In one embodiment, the polynucleotide is operatively linked to one or more heterologous control sequences.
[0442] In one embodiment, the heterogeneous control sequence is a heterogeneous promoter.
[0443] In one embodiment, the polynucleotides are isolated.
[0444] In one embodiment, the polynucleotide is purified.
[0445] The polynucleotide can be genomic DNA, cDNA, synthetic DNA, synthetic RNA, mRNA, or a combination thereof. The polynucleotide can be cloned from strains or related organisms of *Lactobacillus*, *Bacillus*, *Zurichobacterium*, *Ureaplasma*, *Hypertropha*, *Clostridium*, *Ruminococcus*, *Bacillus cyclophosphamide*, *Enterococcus*, *Lactobacillus symbioticus*, *Lactobacillus bumblebee*, *Lactococcus*, or *Lactococcus*, and therefore, for example, can be a polynucleotide sequence encoding a variant of the Cas nuclease of the present invention.
[0446] In one embodiment, the polynucleotide is obtained from Streptococcus cells. In one embodiment, the Streptococcus cells are Streptococcus equi, Streptococcus mutans, Streptococcus species, Streptococcus henryi DSM 19005, Streptococcus species CCH8-G7, Streptococcus paclitaxum, Streptococcus mutans DSM 15617, Streptococcus salivans, or ruminant streptococcus cells.
[0447] In one embodiment, the polynucleotide is obtained from Bacillus cells, for example, Bacillus species 63030 cells.
[0448] In one embodiment, the polynucleotide is obtained from Zurich bacillus cells, for example, Zurich bacillus species cells.
[0449] In one embodiment, the polynucleotide is obtained from Ureaplasma cells, such as Ureaplasma thermophila cells.
[0450] In one embodiment, the polynucleotide is obtained from *Cytobacter* cells, such as *Cytobacter humanis* cells.
[0451] In one embodiment, the polynucleotide is obtained from Clostridium cells.
[0452] In one embodiment, the polynucleotide is obtained from Ruminococcus cells, for example, Ruminococcus species cells.
[0453] In one embodiment, the polynucleotide is obtained from Bacillus cycloalisate cells, such as Bacillus glycocycloalisate cells.
[0454] In one embodiment, the polynucleotide is obtained from Enterococcus cells, such as Enterococcus faecalis, Enterococcus hermannii, or Enterococcus asperoides cells.
[0455] In one embodiment, the polynucleotide is obtained from Lactobacillus spp. cells, such as Lactobacillus zaguangjiao, Lactobacillus halophilus, Lactobacillus keshanensis, Lactobacillus sauerkraut, or Lactobacillus hulinensis cells.
[0456] In one embodiment, the polynucleotide is obtained from *Lactobacillus* cells, for example, *Lactobacillus bumblebee* cells.
[0457] In one embodiment, the polynucleotide is obtained from *Vibrio spp.* cells, such as *Vibrio spp.* shrimp cells.
[0458] In one embodiment, the polynucleotide is obtained from Lactobacillus cells. In one embodiment, the Lactobacillus cells are Lactobacillus species, Lactobacillus sausageii (DSM 20184), Lactobacillus sausageii, Lactobacillus murineis, Lactobacillus rumenii, Lactobacillus salivarius, Lactobacillus janniae, Lactobacillus hamsterii, Lactobacillus delbrueckii, Lactobacillus johnsonii, Lactobacillus plantarum, Lactobacillus rhamnosus, or Lactobacillus chickenii cells.
[0459] In a preferred embodiment, the polynucleotide encoding the Cas nuclease was isolated from Lactobacillus cells.
[0460] In the embodiments, the polynucleotide is a subsequence encoding a fragment having the Cas nuclease activity and / or DNA binding activity of the present invention.
[0461] These polynucleotides can also be constructed by introducing nucleotide substitutions that do not cause a change in the amino acid sequence of the polypeptide but correspond to the codons used by the host organism intended to produce the enzyme, or by introducing nucleotide substitutions that may produce different amino acid sequences. For a general description of nucleotide substitutions, see, for example, Ford et al., 1991, Protein Expression and Purification 2: 95-107.
[0462] Nucleic acid constructs
[0463] In a sixth aspect, the present invention relates to nucleic acid constructs or expression vectors comprising a polynucleotide according to a fifth aspect of the invention, the polynucleotide being operatively linked to one or more control sequences that direct the production of nucleases or fusion polypeptides in cells.
[0464] The present invention also relates to nucleic acid constructs or expression vectors comprising the polynucleotide of the present invention, wherein the polynucleotide is operatively linked to one or more control sequences, which, under conditions compatible with the control sequences, guide the expression of the coding sequence in a suitable host cell.
[0465] Polynucleotides can be manipulated in various ways to provide polypeptide expression. Depending on the expression vector, manipulating the polynucleotide before insertion into the vector may be desirable or necessary. Techniques for modifying polynucleotides using recombinant DNA methods are well known in the art.
[0466] promoter
[0467] The control sequence can be a promoter, i.e., a polynucleotide recognized by the host cell for expressing a polynucleotide encoding a Cas nuclease. The promoter contains a transcriptional control sequence that mediates the expression of the Cas nuclease. The promoter can be any polynucleotide that exhibits transcriptional activity in the host cell, including mutant promoters, truncated promoters, and heterozygous promoters, and can acquire a gene that encodes an extracellular or intracellular polypeptide that is homologous or heterologous to that of the host cell.
[0468] Examples of suitable promoters for guiding the transcription of polynucleotides in bacterial host cells are described in Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Lab, New York; Davis et al., 2012, ibid.; and Song et al., 2016, PLOS One 11(7): e0158447.
[0469] Examples of suitable promoters for guiding the transcription of polynucleotides in filamentous fungal host cells in this invention are promoters obtained from Aspergillus, Fusarium, Rhizomucor, and Trichoderma cells, such as those described in Mukherjee et al., 2013, “Trichoderma: Biology and Applications” and Schmoll and Dattenböck, 2016, “Gene Expression Systems in Fungi: Advancements and Applications”, Fungal Biology.
[0470] Examples of useful promoters for expression in yeast hosts are described in: Smolke et al., 2018, “Synthetic Biology: Parts, Devices and Applications” (Chapter 6: Constitutive and Regulated Promoters in Yeast: How to Design and Make Use of Promoters in S. cerevisiae) and Schmoll and Dattenböck, 2016, “Gene Expression Systems in Fungi: Advancements and Applications”, Fungal Biology.
[0471] Termination
[0472] The control sequence can also be a transcription terminator that is recognized by the host cell to terminate transcription. The terminator is operatively linked to the 3' end of a polynucleotide encoding a Cas nuclease. Any terminator that is functional in the host cell can be used in this invention.
[0473] Preferred terminators for bacterial host cells can be obtained from genes of Bacillus clausti alkaline protease (aprH), Bacillus licheniformis α-amylase (amyL), and Escherichia coli ribosomal RNA (rrnB).
[0474] Preferred terminators for filamentous fungal host cells can be obtained from Aspergillus or Trichoderma species, such as genes obtained from Aspergillus niger glucosidase, Trichoderma reesei β-glucosidase, Trichoderma reesei cellobiase I, and Trichoderma reesei endoglucanase I, such as terminators described in Mukherjee et al., 2013, “Trichoderma: Biology and Applications” and Schmoll and Dattenböck, 2016, “Gene Expression Systems in Fungi: Advancements and Applications”, Fungal Biology.
[0475] Preferred terminators for yeast host cells can be obtained from the genes of *Saccharomyces cerevisiae* enolase, *Saccharomyces cerevisiae* cytochrome C (CYC1), and *Saccharomyces cerevisiae* glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are described by Romanos et al., 1992, Yeast [Yeast] 8: 423-488.
[0476] mRNA stabilizers
[0477] Control sequences can also be mRNA stabilizing regions downstream of the promoter and upstream of the coding sequence of a gene, which increase the expression of genes encoding Cas nucleases.
[0478] Examples of suitable mRNA stabilizer regions were obtained from the Bacillus thuringiensis cryIIIA gene (WO 94 / 25612) and the Bacillus subtilis SP82 gene (Hue et al., 1995, J. Bacteriol. [Journal of Bacteriology] 177:3465-3471).
[0479] Examples of mRNA stabilizer regions in fungal cells are described in Geisberg et al., 2014, Cell [Cell] 156(4): 812-824 and Morozov et al., 2006, Eukaryotic Cell [Eukaryotic Cell] 5(11): 1838-1846.
[0480] Leader sequence
[0481] The control sequence can also be a leader sequence, i.e., an untranslated region of mRNA that is important for translation in the host cell. The leader sequence is operatively linked to the 5' end of the polynucleotide encoding the Cas nuclease. Any leader sequence that is functional in the host cell can be used.
[0482] The appropriate leader sequence for bacterial host cells is described by Hambraeus et al., 2000, Microbiology 146(12): 3051-3059 and Kaberdin and Bläsi, 2006, FEMS Microbiol. Rev. 30(6): 967-979.
[0483] Preferred leader sequences for filamentous fungal host cells can be obtained from the genes of Aspergillus oryzae TAKA amylase and Aspergillus nidulans triphosphate isomerase.
[0484] Suitable leader sequences for yeast host cells can be obtained from the following genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).
[0485] polyadenylation sequence
[0486] The control sequence can also be a polyadenylated sequence, which is a sequence operatively linked to the 3' end of a polynucleotide encoding a Cas nuclease. This sequence is recognized by the host cell during transcription as a signal to add polyadenylated residues to the transcribed mRNA. Any polyadenylated sequence that is functional in the host cell can be used.
[0487] Preferred polyadenylated sequences for filamentous fungal host cells are obtained from the genes of Aspergillus nidulans anthranilate synthase, Aspergillus niger glucosyl amylase, Aspergillus niger α-glucosidase, Aspergillus oryzae TAKA amylase, and Fusarium oxysporum trypsin-like protease.
[0488] Useful polyadenylation sequences in yeast host cells are described by Guo and Sherman, 1995, Mol. Cellular Biol. [Molecular Cell Biology] 15: 5983-5990.
[0489] signal peptide
[0490] The control sequence can also be a signal peptide coding region that encodes a signal peptide linked to the N-terminus of a Cas nuclease and guides the Cas nuclease into the cell's secretory pathway. The 5' end of the polynucleotide coding sequence may inherently contain a signal peptide coding sequence naturally linked to a segment of the coding sequence of the coding polypeptide within the translation reading frame. Alternatively, the 5' end of the coding sequence may contain a signal peptide coding sequence that is heterologous to the coding sequence. In cases where the coding sequence does not naturally contain a signal peptide coding sequence, a heterologous signal peptide coding sequence may be required. Alternatively, a heterologous signal peptide coding sequence may simply replace the natural signal peptide coding sequence to enhance polypeptide secretion. Any signal peptide coding sequence that guides the secretory pathway of the expressed polypeptide into the host cell can be used.
[0491] The effective signal peptide coding sequences of bacterial host cells are obtained from the signal peptide coding sequences of the following genes: maltose amylase produced by Bacillus NCIB 11837, Bacillus subtilis protease, Bacillus β-lactamase, Bacillus stearothermophilus α-amylase, Bacillus stearothermophilus neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA. Other signal peptides are described by Freudl, 2018, Microbial Cell Factories 17: 52.
[0492] The effective signal peptide coding sequences for filamentous fungal host cells are obtained from the following genes: Aspergillus niger neutral amylase, Aspergillus niger glucosylase, Aspergillus oryzae TAKA amylase, Specific Humus cellulase, Specific Humus endoglucanase V, Specific Humus lipase, and Rhizomucor miehei aspartic protease, as described by Xu et al., 2018, Biotechnology Letters [Biotechnology Letters] 40: 949-955.
[0493] Useful signal peptides in yeast host cells are obtained from the genes of *Saccharomyces cerevisiae* α-factor and *Saccharomyces cerevisiae* invertase. Sequences encoding other useful signal peptides are described above by Romanos et al., 1992.
[0494] propeptide
[0495] The control sequence can also be a propeptide-coding sequence encoding the propeptide located at the N-terminus of the Cas nuclease. The resulting polypeptide is called a proenzyme or propeptide progenitor (or, in some cases, a zymogen). The propeptide progenitor is usually inactive and can be converted into an active polypeptide by catalytic cleavage or autocatalytic cleavage of the propeptide progenitor. Propeptide-coding sequences can be obtained from the genes of Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Thermophilus laccase (WO 95 / 33836), Rhizopus oryzae aspartic protease, and Saccharomyces cerevisiae α-factor.
[0496] When both a signal peptide sequence and a propeptide sequence are present, the propeptide sequence is located immediately adjacent to the N-terminus of the polypeptide, and the signal peptide sequence is located immediately adjacent to the N-terminus of the propeptide sequence. Alternatively, when both a signal peptide sequence and a propeptide sequence are present, the polypeptide may contain only a portion of the signal peptide sequence and / or only a portion of the propeptide sequence. Alternatively, the final or isolated polypeptide may comprise a mixture of a mature polypeptide and a polypeptide containing partial or full-length propeptide sequences and / or signal peptide sequences.
[0497] Regulation sequence
[0498] It may also be desirable to add regulatory sequences that modulate the expression of Cas nucleases relative to the growth of the host cell. Examples of regulatory sequences are those that cause gene expression to turn on or off in response to chemical or physical stimuli, including the presence of regulatory compounds. Regulatory sequences in prokaryotic systems include the lac, tac, and trp operon systems. In yeast, the ADH2 or GAL1 system can be used. In filamentous fungi, the *Aspergillus niger* glucosylamylase promoter, the *Aspergillus oryzae* TAKAα-amylase promoter and *Aspergillus oryzae* glucosylamylase promoter, the *Trichoderma reesei* cellobiose hydrolase I promoter and the *Trichoderma reesei* cellobiose hydrolase II promoter can be used. Other examples of regulatory sequences are those that allow gene amplification. In fungal systems, these regulatory sequences include dihydrofolate reductase genes amplified in the presence of methotrexate and metallothionein genes amplified with heavy metals.
[0499] expression carrier
[0500] The present invention also relates to recombinant expression vectors comprising the polynucleotide, promoter, and transcription and translation termination signals of the present invention. Various nucleotides and control sequences can be linked together to produce a recombinant expression vector, which may include one or more suitable restriction sites to allow insertion or substitution of the polynucleotide encoding a Cas nuclease at such sites. Alternatively, the polynucleotide can be expressed by inserting the polynucleotide or a nucleic acid construct containing the polynucleotide into a suitable vector for expression. In producing the expression vector, the coding sequence is located in the vector such that the coding sequence is operatively linked to a suitable control sequence for expression.
[0501] Recombinant expression vectors can be any vector (e.g., plasmids or viruses) that can readily undergo recombinant DNA procedures and induce polynucleotide expression. The choice of vector will typically depend on its compatibility with the host cell to which it will be introduced. Vectors can be linear or closed circular plasmids.
[0502] The vector can be a self-replicating vector, that is, a vector that exists as an extrachromosomal entity and replicates independently of chromosome replication, such as a plasmid, extrachromosomal element, microchromosome, or artificial chromosome. The vector can contain any means to ensure self-replication. Alternatively, the vector can be one that integrates into the genome when introduced into a host cell and replicates along with the chromosome in which it has been integrated. Furthermore, a single vector or plasmid, or two or more vectors or plasmids collectively containing the total DNA of the host cell genome to be introduced, or transposons can be used.
[0503] The vector preferably contains one or more selective markers that allow for convenient selection of cells such as transformed cells, transfected cells, and transduced cells. A selective marker is a gene whose product provides resistance to biocides or viruses, resistance to heavy metals, or prototrophic auxotrophic traits, etc.
[0504] The vector preferably contains at least one element that allows the vector to integrate into the genome of the host cell or to replicate autonomously in the cell independently of the genome.
[0505] In order to integrate into the host cell genome, the vector may depend on a polynucleotide sequence encoding a polypeptide or any other element of the vector used for integration into the genome via homologous recombination (such as homologous directed repair (HDR)) or non-homologous recombination (such as non-homologous end joining (NHEJ)).
[0506] For autonomous replication, the vector may further include an origin of replication, which enables the vector to replicate autonomously in the host cell discussed. The origin of replication can be any plasmid replicon that mediates autonomous replication and functions within the cell. The terms "origin of replication" or "plasmid replicon" refer to the polynucleotide that enables a plasmid or vector to replicate in vivo.
[0507] More than one copy of the polynucleotide of the present invention can be inserted into a host cell to increase the production of Cas nuclease. For example, two, three, four, five or more copies can be inserted into a host cell. The increased copy number of the polynucleotide can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selective marker gene along with the polynucleotide, wherein cells containing the amplified copy of the selective marker gene and thus additional copies of the polynucleotide can be selected by culturing cells in the presence of a suitable selective reagent.
[0508] host cells
[0509] In a seventh aspect, the present invention relates to cells comprising the Cas nuclease of the first aspect, the fusion polypeptide of the second aspect, the composition of the third aspect, the polynucleotide of the fifth aspect, and / or the nucleic acid construct or expression vector of the sixth aspect.
[0510] In one embodiment, the cell is a recombinant cell.
[0511] In a preferred embodiment, the Cas nuclease is heterologous to the cell.
[0512] In one embodiment, the cell contains at least two copies, such as three, four, or five or more copies of a fifth-side polynucleotide or a sixth-side vector or construct.
[0513] In one embodiment, the cell's genome contains a polynucleotide encoding a first-sided Cas nuclease or a second-sided fusion polypeptide, a fifth-sided polynucleotide, or a sixth-sided nucleic acid construct or expression vector.
[0514] In one embodiment, the genome of the recombinant cell contains at least two, such as three, four, or five or more, copies of a polynucleotide encoding a first-sided Cas nuclease or a second-sided fusion polypeptide, a fifth-sided polynucleotide, or a sixth-sided nucleic acid construct or expression vector.
[0515] In an eighth aspect, the present invention relates to cells comprising a genome modified by the Cas nuclease of the first aspect, the fusion polypeptide of the second aspect, the composition of the third aspect, the method of the fourth aspect, the polynucleotide of the fifth aspect, and / or the nucleic acid construct or expression vector of the sixth aspect.
[0516] In one embodiment, the cell is a recombinant cell.
[0517] In one embodiment, the cell is selected from the group consisting of: archaea cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, bird cells, mammalian cells, non-human mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells.
[0518] In one embodiment, the cell is a eukaryotic cell.
[0519] In one embodiment, the cell is a prokaryotic cell.
[0520] In one embodiment, the cell is a eukaryotic cell, such as a mammalian cell, a human cell, or a non-human mammalian cell, such as a BHK cell, a CHO cell, a mouse cell, a hamster cell, or a rat cell.
[0521] In one embodiment, the cell is a fungal cell, such as a filamentous fungal cell or a yeast cell.
[0522] In one embodiment, the cells are Pichia pastoris cells.
[0523] In one embodiment, the cell is a brewer's yeast cell.
[0524] In one embodiment, the cell is a yeast cell, such as cells of the genera *Candida*, *Hansenula*, *Kluyveromyces*, *Pichia*, *Saccharomyces*, *Saccharomyces*, or *Yersinia*, such as *Kluyveromyces lactis*, *Kalvatia*, *Saccharomyces cerevisiae*, *Saccharomyces sacchariformis*, *Saccharomyces davidiana*, *Douglas*, *Kluyveromyces rufi*, *Nordic*, *Ovoyces*, or *Yersinia lipolytica*.
[0525] In one embodiment, the cells are filamentous fungal cells, such as those from the genera *Cladosporium*, *Aspergillus*, *Briefomus*, *Cirsium*, *Pseudomonas*, *Aureospora*, *Coprinus*, *Pseudomonas*, *Cryptococcus*, *Ustilago*, *Fusarium*, *Pyrophyllus*, *Gastrophyllus*, *Mucor*, *Hydrophyllus*, *Neurospora*, *Penicillium*, *Penicillium*, *Pseudomonas ... Cells of the genera *Pleurotus*, *Schizophyllum*, *Basilaria*, *Thermophila*, *Fusporium*, *Cyclophorus*, *Vallisneria*, or *Trichoderma*, particularly *Aspergillus pumilus*, *Aspergillus sulphureus*, *Aspergillus fumigatus*, *Aspergillus japonicus*, *Aspergillus nidus*, *Aspergillus niger*, *Aspergillus oryzae*, *Cyclophorus nidus*, *Cyclophorus carnegieii*, *Cyclophorus pachytyphae*, *Cyclophorus pannohita*, *Cyclophorus circumsus*, *Cyclophorus microcarne*, *Cyclophorus wormina*, and *Aureospora narrowensis*. *Aureobasidium keratophyllum*, *Aureobasidium rucinosa*, *Aureobasidium fecalithum*, *Aureobasidium monnieri*, *Aureobasidium queenslandense*, *Aureobasidium tropicalis*, *Aureobasidium bandedense*, *Coprinus comatus*, *Aureobasidium triflorum*, *Fusarium moniliforme*, *Fusarium cerealis*, *Fusarium graminearum*, *Fusarium graminearum*, *Fusarium heterosporum*, *Fusarium hygrosvenorii*, *Fusarium oxysporum ... Fusarium, Fusarium pseudocranioides, Fusarium sulfideum, Fusarium rotundum, Fusarium pseudofilariae, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Mucor, Fusarium variegatum, Neurospora crassa, Penicillium purpureum, Pleurotus eryngii, Pleurotus eryngii, Pleurotus emarginatus, Clostridium perfringens, Clostridium perfringens, Clostridium chrysogenum, Clostridium perfringens, Clostridium chrysogenum, Trichoderma harzianum, Trichoderma cornigrin, Trichoderma longibranchii, Trichoderma reesei, or Trichoderma viride cells.
[0526] In one embodiment, the cells are Trichoderma cells.
[0527] In one embodiment, the cells are Trichoderma reesei cells.
[0528] In one embodiment, the cells are Aspergillus cells.
[0529] In one embodiment, the cells are Aspergillus niger cells.
[0530] In one embodiment, the cells are Aspergillus oryzae cells.
[0531] In one embodiment, the cell is a plant cell.
[0532] In one embodiment, the cell is one or more of the following: corn, rice, sorghum, rye, barley, wheat, millet, oats, sugarcane, turfgrass, switchgrass, soybean, canola, alfalfa, sunflower, cotton, tobacco, peanut, potato, tobacco, Arabidopsis, vegetable, or safflower cells.
[0533] In one embodiment, the cell is a prokaryotic cell, for example, Gram-positive cells selected from the group consisting of: *Bacillus*, *Clostridium*, *Corynebacterium*, *Enterococcus*, *Bacillus terrestris*, *Lactobacillus*, *C. casei*, *L. lactis*, *L. proliferative lactobacillus*, *L. assemblica*, *L. mucolyticus*, *L. lactococcus*, *B. oceanicus*, *Staphylococcus*, *Streptococcus*, or *Streptomyces* cells, or Gram-negative bacteria selected from the group consisting of: *Campylobacter*, *Escherichia coli*, *Flavobacterium*, *Clostridium*, *Helicobacter*, *Staphylococcus*, *Neisseria*, *Pseudomonas*, *Salmonella*, and *Ureaplasma* cells, such as *C. casei*, *C. paracasei*, *C. rhamnosus*, *L. plantarum*, *L. proliferative lactobacillus ... Lactobacillus, Lactobacillus salivarius, Lactobacillus fermentum, Lactobacillus reuteri, Lactobacillus acidophilus, Lactobacillus bulgaricus, Lactobacillus curvatureus, Lactobacillus gasseri, Lactobacillus johnsonii, Lactobacillus helveticus, Corynebacterium glutamicum, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus croceus, Bacillus coagulans, Bacillus sclerotiorum, Bacillus splenicus, Bacillus scintillans, Bacillus licheniformis, Bacillus megaterium, Bacillus brevis, Bacillus thermophilus, Bacillus subtilis, Bacillus thuringiensis, Streptococcus equine, Streptococcus pyogenes, Streptococcus lactis and Streptococcus equine subsp. porphyria, Streptococcus nonchromogenic, Streptococcus pyogenes, Streptococcus aureus, Streptococcus gray, and Streptococcus purpureus cells.
[0534] In a preferred embodiment, the cells are Bacillus cells.
[0535] In one embodiment, the cell is a Bacillus subtilis cell.
[0536] In one embodiment, the cell is a Bacillus licheniformis cell.
[0537] In one embodiment, the cells are Lactobacillus paracasei cells.
[0538] In one embodiment, the cell is a thermophilic streptococcal cell.
[0539] In one embodiment, the cells are Escherichia coli cells.
[0540] In one embodiment, the cells are isolated.
[0541] In one embodiment, the cells are purified.
[0542] The present invention also relates to recombinant host cells containing polynucleotides of the present invention operatively linked to one or more control sequences that direct the production of Cas nucleases.
[0543] A construct or vector containing a polynucleotide is introduced into a host cell, such that the construct or vector is maintained as a chromosomal integrase or as a self-replicating extrachromosomal vector, as previously described. The choice of host cell will depend largely on the gene encoding the polypeptide and its origin. The Cas nuclease for the recombinant host cell can be native or heterologous. Furthermore, at least one of one or more control sequences for the polynucleotide encoding the Cas nuclease can be heterologous. The recombinant host cell may contain a single copy or at least two copies, such as three, four, five or more copies, of the polynucleotide of the present invention.
[0544] For the purposes of this invention, the class / genus / species of Bacillus should be defined as described in Patel and Gupta, 2020, Int. J.Syst.Evol.Microbiol. [International Journal of Systematic and Evolutionary Microbiology] 70: 406-438.
[0545] The bacterial host cell can also be any streptococcus cell, including but not limited to Streptococcus equisimilis, Streptococcus pyogenes, Streptococcus uberis, and Streptococcus equi subsp. Zooepidemicus.
[0546] The bacterial host cell can also be any Streptomyces cell, including but not limited to Streptomyces achromogenes, Streptomyces avermitilis, Streptomyces coelicolor, Streptomyces griseus, and Streptomyces lividans.
[0547] Methods for introducing DNA into prokaryotic host cells are well known in the art, and any suitable method can be used, including but not limited to protoplast transformation, competent cell transformation, electroporation, conjugation, and transduction, wherein the DNA is introduced as a linearized or circular polynucleotide. Those skilled in the art will be able to readily determine, for example, a suitable method for introducing DNA into a given prokaryotic cell based on the genus. Methods for introducing DNA into prokaryotic host cells are described, for example, in Heinze et al., 2018, BMC Microbiology 18:56; Burke et al., 2001, Proc. Natl. Acad. Sci. USA 98: 6289-6294; Choi et al., 2006, J. Microbiol. Methods 64: 391-397; and Donald et al., 2013, J. Bacteriol. 195(11): 2612-2620.
[0548] The host cell can be a fungal cell. As used herein, “fungus” includes Ascomycota, Basidiomycota, Chytridiomycota, Zygomycota, Oomycota, and all mitotic fungi (as defined by Hawksworth et al. in Ainsworth and Bisby’s Dictionary of The Fungi, 8th edition, 1995, CAB International, University Press, Cambridge, UK).
[0549] Fungal cells can be transformed via processes involving protoplast-mediated transformation, Agrobacterium-mediated transformation, electroporation, gene gun methods, and shock wave-mediated transformation (as reviewed in Li et al., 2017, Microbial Cell Factories, 16: 168), as well as procedures described in EP 238023, Yelton et al., 1984, Proc. Natl. Acad. Sci. USA, 81: 1470-1474; Christensen et al., 1988, Bio / Technology, 6: 1419-1422; and Lubertozzi and Keasling, 2009, Biotechn. Advances, 27: 53-75. However, any method known in the art for introducing DNA into fungal host cells can be used, and the DNA can be introduced as linearized or circular polynucleotides.
[0550] The host cell for fungi can be a yeast cell. As used herein, “yeast” includes ascosporogenous yeast (Endomycetales), basidiosporogenous yeast, and yeasts belonging to the class Fungi Imperfecti (Blastomycetes). For the purposes of this invention, yeast should be defined as described in Biology and Activities of Yeast (edited by Skinner, Passmore, and Davenport, Soc. App. Bacteriol. Symposium Series No. 9, 1980).
[0551] In a preferred embodiment, the yeast host cell is a cell of the genera *Pichia* or *Komagataella*, such as *Pichia pastoris* cells (*Komagataella phaffii*).
[0552] The host cell for fungi can be a filamentous fungal cell. "Filamentous fungi" includes all filamentous forms within the phylum Eumycota and the subphylum Oomycetes (as defined by Hawksworth et al., 1995, ibid.). Filamentous fungi are generally characterized by a hyphal wall composed of chitin, cellulose, glucan, chitosan, mannan, and other complex polysaccharides. Vegetative growth occurs through hyphal elongation, and carbon metabolism is obligate aerobic. In contrast, the vegetative growth of yeasts (such as Saccharomyces cerevisiae) occurs through budding of single-celled cells, and carbon metabolism can be fermentative.
[0553] The host cells of filamentous fungi can be cells from genera such as *Cladosporium*, *Aspergillus*, *Briefomus*, *Cirsium*, *Pseudomonas*, *Aureosporium*, *Coprinus*, *Pseudomonas*, *Cryptococcus*, *Filibasidium*, *Fusarium*, *Pyrophyllus*, *Gastrophyllus*, *Mucor*, *Pyrophyllus*, *Pseudomonas ...
[0554] In a preferred embodiment, the filamentous fungal host cell is a cell of the genus Aspergillus, Trichoderma, or Fusarium.
[0555] In another preferred embodiment, the filamentous fungal host cell is Aspergillus niger, Aspergillus oryzae, Trichoderma reesei, or Fusarium moniliforme cells.
[0556] In an eighth aspect, the present invention relates to plant cells comprising the Cas nuclease of the first aspect, the fusion polypeptide of the second aspect, the composition of the third aspect, the polynucleotide of the fifth aspect, and / or the nucleic acid construct or expression vector of the sixth aspect.
[0557] In one embodiment, the plant cell is one or more of the following: corn, rice, sorghum, rye, barley, wheat, millet, oats, sugarcane, turfgrass, switchgrass, soybean, canola, alfalfa, sunflower, cotton, tobacco, peanut, potato, tobacco, Arabidopsis, vegetable, or safflower cells.
[0558] Generation method
[0559] In a ninth aspect, the invention also relates to a method for generating the Cas nuclease of the first aspect or the fusion polypeptide of the second aspect, the method comprising culturing the host cell of the seventh aspect under conditions conducive to generating the Cas nuclease or the fusion polypeptide; and optionally, (b) recovering the Cas nuclease and / or the fusion polypeptide.
[0560] In one respect, the cell is a Bacillus cell. In another respect, the cell is a Bacillus subtilis cell. In yet another respect, the cell is a Bacillus licheniformis cell.
[0561] In one respect, the cell is an Aspergillus cell. In another respect, the cell is an Aspergillus niger cell. In yet another respect, the cell is an Aspergillus oryzae cell.
[0562] In one respect, the cells are Trichoderma cells. In another respect, the cells are Trichoderma reesei cells.
[0563] In one embodiment, the cells are Pichia pastoris cells.
[0564] In one embodiment, the cell is a brewer's yeast cell.
[0565] In one embodiment, the cells are Lactobacillus paracasei cells.
[0566] In one embodiment, the cell is a thermophilic streptococcal cell.
[0567] In one embodiment, the cells are Escherichia coli cells.
[0568] The host cells are cultured in a nutrient medium suitable for producing peptides using methods known in the art. For example, cells can be cultured in a suitable medium and under conditions that allow for peptide expression and / or isolation by shake-flask culture or by small-scale or large-scale fermentation (including continuous, batch, fed-batch, or solid-state and / or microcarrier-based fermentation) in a laboratory or industrial fermenter. Suitable media are available from commercial suppliers or can be prepared according to publicly available compositions (e.g., in the catalogue of the U.S. Center for Type Culture Collection). If the peptide is secreted into the nutrient medium, it can be recovered directly from that medium. If the peptide is not secreted, it can be recovered from cell lysates.
[0569] Peptides can be detected using methods known in the art that are specific to peptides, including but not limited to the use of specific antibodies, enzyme product formation, enzyme substrate disappearance, or determination of the relative or specific activity of the peptide.
[0570] Peptides can be recovered from culture media using methods known in the art, including but not limited to collection, centrifugation, filtration, extraction, spray drying, evaporation, or precipitation. In one aspect, whole fermentation broth containing peptides is recovered. In another aspect, cell-free fermentation broth containing peptides is recovered.
[0571] Peptides can be purified using a variety of procedures known in the art to obtain substantially pure peptides and / or peptide fragments (see, for example, Wingfield, 2015, Current Protocols in Protein Science; 80(1): 6.1.1-6.1.35; Labrou, 2014, Protein Downstream Processing, 1129: 3-10).
[0572] In terms of alternatives, peptides are not recycled.
[0573] Uses of Cas nuclease
[0574] In a tenth aspect, the present invention relates to the use of Cas nucleases of the first aspect, fusion polypeptides of the second aspect, compositions of the third aspect, methods of the fourth aspect, polynucleotides of the fifth aspect, or nucleic acid constructs or expression vectors of the sixth aspect for modifying target sequences (e.g., target genes) in cells.
[0575] In the eleventh aspect, the present invention relates to the use of Cas nucleases of the first aspect, fusion polypeptides of the second aspect, compositions of the third aspect, methods of the fourth aspect, polynucleotides of the fifth aspect, nucleic acid constructs or expression vectors of the sixth aspect, cells of the seventh aspect, or cells of the eighth aspect for manufacturing medicaments for modifying target sequences (e.g., target genes) in cells.
[0576] In one embodiment, the target cell is selected from the group consisting of: archaea cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, non-human animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, bird cells, mammalian cells, non-human mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells.
[0577] Preparations
[0578] In a twelfth aspect, the invention also relates to formulations comprising (i) a Cas nuclease according to the first aspect, a fusion polypeptide according to the second aspect, a composition according to the third aspect, a polynucleotide according to the fifth aspect, a nucleic acid construct or expression vector according to the sixth aspect, a cell according to the seventh aspect or a cell according to the eighth aspect, and optionally, (ii) one or more of lipids, liposomes, hydrogels, microparticles, nanoparticles or block copolymer micelles.
[0579] In one embodiment, the lipid is a lipid nanoparticle.
[0580] In one embodiment, the Cas nuclease or fusion polypeptide is in the form of a lyophilized preparation.
[0581] In one embodiment, the Cas nuclease or fusion polypeptide is in the form of a liquid formulation.
[0582] In one embodiment, the Cas nuclease or fusion polypeptide is a formulation that is substantially free of endotoxins.
[0583] deliver
[0584] The Cas nuclease or CRISPR composition described herein can be used for the delivery of proteins, DNA molecules, RNA molecules, ribonucleoproteins (RNPs), nucleic acid vectors, or any combination thereof. In some embodiments, the RNA molecule contains chemical modifications. Non-limiting examples of suitable chemical modifications include 2'-O-methyl (M), 2'-O-methyl, 3'-thiophosphate (MS), or 2'-O-methyl, 3'-thiophosphate (MSP), pseudouridine, and 1-methylpseudouridine. Each possibility represents a separate embodiment of the invention.
[0585] The Cas nucleases and / or the polynucleotides encoding them described herein, as well as optionally additional proteins (e.g., ZFP, TALEN, transcription factors, restriction enzymes) and / or nucleotide molecules (e.g., guide RNA), can be delivered to target cells by any suitable means. Target cells can be any type of cell, such as eukaryotic or prokaryotic cells, in any environment, whether isolated or not, maintained in culture, in vitro, ex vivo, in vivo, or in plants.
[0586] In some embodiments, the composition to be delivered includes the mRNA of a nuclease and a guide RNA. In some embodiments, the composition to be delivered includes the mRNA of a nuclease, a guide RNA, and a donor template. In some embodiments, the composition to be delivered includes a Cas nuclease and a guide RNA. In some embodiments, the composition to be delivered includes a Cas nuclease, a guide RNA, and a donor template for gene editing via, for example, homology-directed repair (HDR). In some embodiments, the composition to be delivered includes the mRNA of a nuclease, a DNA-targeting RNA, and tracrRNA. In some embodiments, the composition to be delivered includes the mRNA of a nuclease, a DNA-targeting RNA, tracrRNA, and a donor template. In some embodiments, the composition to be delivered includes a Cas nuclease, a DNA-targeting RNA, and tracrRNA. In some embodiments, the composition to be delivered includes a Cas nuclease, a DNA-targeting RNA, tracrRNA, and a donor template for gene editing via, for example, homology-directed repair.
[0587] With respect to the foregoing embodiments, each embodiment disclosed herein is contemplated to be applicable to each other disclosed embodiment. For example, it should be understood that any RNA molecule or composition of the present invention can be used in any method of the present invention.
[0588] As used herein, all headings are for organizational purposes only and are not intended to limit this disclosure in any way. The content of any single section may apply equally to all sections.
[0589] The invention is further described by the following examples, which should not be construed as limiting the scope of the invention. Example
[0590] Example 1: Identification of a novel Cas nuclease
[0591] The inventors of this invention used their developed bioinformatics workflow to identify the Cas nucleases SEQ ID NO: 1-52 by mining bacterial genomes. Figure 1 As shown, the workflow comprises several task-specific modules, including the identification of Cas nuclease genes, CRISPR arrays, and tracrRNAs, as well as the matching and sequencing of these features. To identify Cas nuclease genes, the workflow employs a state-of-the-art tool with a large set of Hidden Markov Models (HMMs) and scoring schemes to predict Cas nuclease isotypes. Additionally, a new HMM was constructed based on proprietary data. Identified Cas enzymes are screened to determine the presence of conserved domains and essential catalytic residues. To identify CRISPR arrays, the workflow uses a combination of searching for repetitive sequences and aligning them with known repetitive sequences. A screening process is applied to exclude repetitive sequences that do not meet CRISPR-specific criteria. A kmer-based machine learning method (extreme gradient boosting tree) is applied to predict CRISPR array isotypes based on shared repetitive sequences. To identify tracrRNAs, the workflow uses a combination of scanning for antirepetitive sequences using shared CRISPR repetitive sequences from identified CRISPR arrays as queries and a sequence-structure covariance model derived from aligning sequences with experimentally validated tracrRNA tail structures. The complementarity of CRISPR direct repeats and tracrRNA was assessed by comparison using a custom scoring system.
[0592] Table 5 provides an overview of the sequences (SEQ ID NO.) of the identified novel Cas nucleases, nuclease domains, tracrRNA coding sequences, and crRNA coding sequences.
[0593] Table 5. Sequence Overview.
[0594]
[0595] Figure 2 A phylogenetic tree of the Cas nucleases of the present invention based on their amino acid sequences (SEQ ID NO: 1-52) is shown. Full-length amino acid sequences were aligned using Clustal Omega, and the phylogenetic tree was calculated using FastTree (using the Whelan-And-Goldman model). The scale bar shows 0.5 substitutions at each site, indicating the evolutionary distance between Cas nucleases. Most of the novel nucleases belong to class 2 type II Cas nucleases, including the nuclease of SEQ ID NO: 21, and nuclease 0076 (SEQ ID NO: 1) and its closely related homologs (SEQ ID NO: 52, SEQ ID NO: 50, SEQ ID NO: 51, and SEQ ID NO: 45), all of which cluster very closely with SEQ ID NO: 1, such as... Figure 2 (As shown).
[0596] For example Figure 2 As shown, the novel nucleases of SEQ ID NO: 39, SEQ ID NO: 40 and SEQ ID NO: 48 are closely clustered together.
[0597] Example 2: Activity of CRISPR nucleases 0076, 0100, 0105, and 0149 in Aspergillus niger.
[0598] To confirm that CRISPR nucleases 0076 (SEQ ID NO:1), 0100 (SEQ ID NO:21), 0105 (SEQ ID NO:39), and 0149 (SEQ ID NO:48) can be used to specifically induce insertions and deletions at target regions for targeted gene editing, the *Aspergillus niger* fwnA(wA) knockout / spore color assay was used. Using this assay, plasmids with different nucleases and spacer sequences targeting fwnA(wA) were transformed into *Aspergillus niger* Mbin118, and their effectiveness in inducing white phenotype colonies was screened, indicating the presence of insertions / deletions in fwnA(wA). As a result, one or more white phenotype transformants were found for all four nucleases, indicating that all four nucleases were capable of gene editing in this fungal host. Sequencing of the corresponding target prototype spacer regions revealed the presence of insertions and deletions, indicating nuclease-induced fwnA(wA) knockout. These results also indicate that, similar to Cas9, the PAM of all nucleases is located in the 3' region of the prototype spacer. This study confirms that 076, 0100, 0105, and 0149, along with their predicted gRNA scaffolds, can be used for targeted gene editing in Aspergillus niger.
[0599] White spore assay
[0600] To verify the potential of nucleases for targeted gene editing in Aspergillus niger, fwnA (wA) knockout-induced spore color change assays were used.
[0601] In *Aspergillus niger*, knockout of the polyketide synthase fwnA (wA) results in a white / light yellowish-brown spore color phenotype, thought to be due to inactivation of PptA-dependent lysine biosynthesis / side carrier biosynthesis (Jorgensen et al., 2015, *Fungal Genetics and Biology*, 48(5), pp. 544-553). This spore color assay can be used as an indicator of CRISPR nuclease activity. In the absence of DNA repair, DNA strand breaks in eukaryotes typically result in insertions and deletions due to error-prone NHEJ DNA repair. Therefore, CRISPR nuclease-induced DNA breaks in fwnA (wA) are expected to lead to insertions and deletions, gene knockout, and subsequently white spore color. Thus, white colonies resulting from transformation with a putative CRISPR system targeting fwnA indicate targeted nuclease activity.
[0602] Design of control plasmids and screening plasmids
[0603] To ensure proper system function, two controls were used for each nuclease screening: as a positive control (+ctrl), the *Streptococcus pyogenes* CRISPR / Cas9 system was used to demonstrate that insertion-deletion-induced knockout of fwnA (wA) resulted in a detectable color change under the tested experimental conditions. As a negative control (-ctrl), a plasmid containing the nuclease and spacerless sgRNA was used to confirm that the system itself did not induce untargeted insertions or deletions in fwnA (wA).
[0604] For screening, a 21 bp spacer sequence with the expected activity targeting *Aspergillus niger* fwnA was designed based on the hypothesis of corresponding nucleases (PAM). The screening plasmid was designed to contain two expression cassettes to express sgRNAs (each targeting a different region of fwnA) and nucleases. The sgRNAs were derived from a tandem of spacers, direct repeats, and tracrRNAs, while the nuclease gene sequences were codon-optimized for *Aspergillus niger*. Except for the nuclease (0076) in SEQ ID NO: 1 and the nuclease (0105) in SEQ ID NO: 39, all expression cassettes used constitutive expression, and nuclease expression utilized a thiamine-induced expression system to fine-tune expression levels to prevent cytotoxicity (Shoji et al., 2005, FEMS microbiology letters, 244(1), pp. 41-46). This system utilized the transcriptionally / posttranscriptionally regulated promoter PthiA, which downregulates gene expression in the presence of thiamine.
[0605] Transformation of control and screening plasmids
[0606] The control plasmid (+ctrl and -ctrl) and the selection plasmid were transformed into Aspergillus niger Mbin118, and 12 random colonies of the target bacteria were isolated for each transformation according to the method recorded in "Protoplast Generation and Transformation". The ratio of white to black colonies was then observed for each transformation.
[0607] As expected, all -ctrl conversions did not induce the white spore phenotype for all nucleases. Figures 3 to 6 Transformations 0100 and 0149 produced only black colonies, while transformations 0076 and 0105 showed a general lack of growth on the transformation plates. In contrast, the *Streptococcus pyogenes* Cas9 + ctrl system produced almost only white spore colonies. Figures 3 to 6 The results from these controls confirm the effectiveness of this assay in validating fwnA knockout.
[0608] For the primary screening of nuclease 0100 (SEQ ID NO: 21), we initially performed transformation and culture at 30°C and 34°C to assess the effect of temperature on the white colony ratio. Figure 3 It shows that at 30°C ( Figure 3 A) and 34°C ( Figure 3 B) The number of colonies obtained after transformation of 0100 and the ratio of white to black phenotypes.
[0609] White colonies were observed in 10 of the 12 spacers tested at both temperatures. Figure 3 A and Figure 3 B). The percentage of white colonies was 17%–100% at 30°C and 42%–100% at 34°C. Since a higher percentage of white spores was observed at 34°C for all positive 0100 transformations, transformation and culture for nucleases 0076 (SEQ ID NO:1), 0105 (SEQ ID NO:39), and 0149 (SEQ ID NO:48) were continued at 34°C.
[0610] The nuclease 076 of SEQ ID NO: 1 was transformed using a thiamine-adjustable system, in which three different thiamine concentrations (0.02, 0.4 and 5 μM) were tested. Figure 4 The colony counts and white / black phenotype ratios obtained after transformation of 076 with supplemental concentrations of 0.02 μM (A), 0.4 μM (B), and 5 μM (C) of thiamine are shown.
[0611] When agar was supplemented with 0.02 μM and 0.4 μM thiamine, white colonies were observed in spacer 76-T2 (Table 7). Figure 4 A and Figure 4 B). The percentages of white colonies were 9% and 8%, respectively. Figure 4 A and Figure 4 B). As expected, no white colonies were observed in all transformations supplemented with 5 μM thiamine. Figure 4 C) indicates that PthiA inhibition leads to downregulation of nuclease expression.
[0612] Nuclease 0105 (SEQ ID NO: 39) was converted using a similar thiamine-regulated system. Figure 5 The colony counts and white / black phenotype ratios obtained after transformation of 0105 with supplementation of 0.02 μM (A), 0.4 μM (B), and 5 μM (C) thiamine are shown.
[0613] Here, white colonies of spacer 105-T4 (Table 6) were observed at all thiamine concentrations. Figure 5 At 0.02 μM and 0.4 μM thiamine, the percentage of white colonies was 92%, but at 5 μM thiamine, the percentage of white colonies decreased to 67%. Similarly, these results indicate that nuclease expression is downregulated when 5 μM thiamine is supplemented.
[0614] For nuclease 0149 (SEQ ID NO: 48), a constitutive expression system was used for transformation. Figure 6 The table shows the number of colonies and the ratio of white to black phenotypes obtained after transformation of 0149. White colonies were observed in spacer 149-T4 (Table 6) (11 out of 12 selected colonies, 92%), while the other three selection spacers produced only black colonies. Figure 6 ).
[0615] In summary, the conversion results of nucleases 0076, 0100, 0105, and 0149 showed that insertion and deletion activities were observed at the target sites for all four nucleases.
[0616] Spore PCR / Sanger sequencing of selected transformants
[0617] To further confirm that the white spore phenotype is a result of targeted insertion / deletion formation in fwnA, spore PCR / sequencing was performed on 1, 6, 7, and 8 white colonies transformed from 0076, 0100, 0105, and 0149, according to the methods described in "Spore PCR, DNA Purification, and Sanger Sequencing". Sequencing results showed that insertions / deletions of varying lengths and deletions (…) were present in all sequenced samples. Figures 7-9Furthermore, these insertions and deletions are all located near the target prototype spacer region. This indicates that DNA strand breaks are induced in the target region by the expressed nucleases, leading to insertions / deletions and subsequent fwnA (wA) knockout. Since targeted insertion / deletion activity was observed for all four nucleases, this data also suggests that their PAMs are located at the 3' relative to their prototype spacers, and the putative PAM sequences disclosed herein can be used for targeted gene editing.
[0618] Using Aspergillus niger fwnA (wA) knockout / spore color assay, four putative nucleases, 0076, 0100, 0105, and 0149, induced a white spore phenotype. Subsequent spore PCR and sequencing confirmed that the white spore colonies were caused by fwnA knockout induced by nuclease-induced insertions or deletions in the target region.
[0619] Figure 7 The sequenced prototype spacer region (SEQ ID NO: 322) of white spore colonies transformed by plasmid targeting 76-T2 under supplemental 0.02 μM thiamine is shown, compared with the wild-type fwnA target sequence (SEQ ID NO: 321).
[0620] Figure 8 The sequenced prototype spacer regions of white spore colonies transformed by plasmids targeting spacers 100-T2 (A) and 100-T5 (B) are shown, compared with wild-type fwnA target sequences (above, SEQ ID NO: 323 and SEQ ID NO: 328). Figure 8 The insertions / deletions in A are shown as SEQ ID NO: 324-327. Figure 8 The insertions / deletions in B are shown as SEQ ID NO: 329-330.
[0621] Figure 9 The sequenced prototype spacer region of white spore colonies transformed with a plasmid targeting spacer 105-T4 under supplemental 0.02 μM thiamine is shown, compared with the original fwnA target sequence (above, SEQ ID NO: 331). Figure 9 The insertions / deletions in the data are shown as SEQ ID NO: 331-338.
[0622] Figure 10 The sequenced prototype spacer region of white spore colonies transformed at 34C by a plasmid targeting 149-T4 is shown, compared with the original fwnA target sequence (above, SEQ ID NO: 339). Figure 10 The insertions / deletions in the figure are shown as SEQ ID NO: 340-346. A sequence with a 346 bp insertion is omitted from the figure.
[0623] from Figures 7-10 As can be seen, (multi) nucleotide insertions or deletions occurred at each target locus.
[0624] In summary, these results confirm that nucleases 0076, 0100, 0105, and 0149, along with their corresponding homologous repeats and gRNA scaffolds, can be used for targeted gene editing in fungal hosts.
[0625] expression plasmid construction
[0626] Construction of positive control plasmid
[0627] First, the intermediate plasmid pHUda2351 containing Streptococcus pyogenes Cas9 was... Figure 11 A) The linearized vector was digested with PmeI and purified by silica gel column chromatography (Nucleobond Xtra Midi, Takara Bio Inc.). To meet the PAM requirements of Cas9, a 20 bp gRNA spacer sequence was designed to target fwnA (positions 37-56 bp, sense strand). To facilitate plasmid assembly, an additional 25 bp region homologous to the plasmid backbone insertion site was added to the 5' and 3' ends. The spacer sequence with flanking ends was synthesized using a commercial oligonucleotide synthesis service. The spacer was then ligated into the linearized vector using HiFi DNA assembly (New England Biolabs, USA) and named pBKHM0015-32. Figure 11 B).
[0628] Figure 11The plasmid vectors used for the Streptococcus pyogenes CRISPR / Cas9 positive control 1 are shown. (A) The intermediate plasmid vector pHUda2351 containing the Streptococcus pyogenes Cas9 system for the construction of the control plasmid and (B) the positive control 1 containing Cas9 and a spacer targeting the region of spacer 32 or spacer 112. ampR, Escherichia coli ampicillin resistance gene; AMA1, Aspergillus nidulans AMA1 replication origin; P. oryzae RNAPIII U6-2 promoter, Magnaporthe grisea RNAPIII U6-2 promoter; Af tRNA gly, Aspergillus fumigatus glycine tRNA; gRNA backbone, Cas9 scaffold sequence; spacer 32 or 112, spacer targeting fwnA; P. oryzae U6-2 terminator (long), Magnaporthe grisea U6-2 long terminator; Ptef (nid), Aspergillus nidulans tef promoter; Cas9, Streptococcus pyogenes Cas9; Ttef (nid), Aspergillus nidulans tef terminator; Ptef1, Aspergillus niger tef1 promoter; NATr, norsinolate resistance gene; TniaD, niaD terminator.
[0629] Construction of intermediate plasmid (negative control plasmid)
[0630] To construct an intermediate plasmid for screening (which also serves as a negative control plasmid), the plasmid vector pBKHM0001 was first digested with AscI and SbfI. Figure 12 The plasmid backbone fragment with a length of 11097 bp was recovered by gel extraction.
[0631] Figure 12 The plasmid vector pBKHM0001 used as the backbone sequence is shown. *Aspergillus oryzae* U6 promoter; *Af tRNA gly*, *Aspergillus fumigatus* glycine tRNA; *Aspergillus oryzae* U6 terminator; *AnPtef1*, *Aspergillus niger* tef1 promoter; CRISPR-CasΦ, CasΦ nuclease gene; Ttef(nid), *Aspergillus niger* tef1 terminator; Ptef1, *Aspergillus niger* tef1 promoter; NATr, norsinolate resistance gene; TniaD, niaD terminator; pUC, pUC plasmid backbone sequence; ampR, *Escherichia coli* ampicillin resistance gene; AMA1, *Aspergillus niger* AMA1 origin of replication.
[0632] Before constructing the inserts, codon optimization was performed on the gene sequences encoding nucleases 0076, 0100, 0105, and 0149 from *Aspergillus niger*, and nucleoplasmic proteins and SV40 nuclear localization signals were added to the 5' and 3' ends, respectively. In the absence of codon optimization, their respective spacerless (unidirectional repeat) hypothetical scaffold sequences were used. The nuclease sequences used and their corresponding spacerless sgRNAs can be found in Table 6.
[0633] Table 6. Sequences of the nuclease genes and corresponding sgRNAs used.
[0634]
[0635] The inserts 0100 and 0149 contain constitutive expression cassettes of their respective gRNAs and nucleases. For 0075 and 0105, the inserts contain constitutive expression cassettes encoding their gRNAs and thiamine-inducible cassettes encoding their nucleases. A 25 bp region homologous to the plasmid backbone insertion site was added to the 5' and 3' flankings to aid in future plasmid construction, and the resulting fragments were synthesized using a commercial gene synthesis service.
[0636] Using HiFi DNA Assembly (New England Biolabs, USA), plasmid backbones and inserts were ligated to form their corresponding intermediate plasmids. These intermediate plasmids were named pBKHM-076-thiA-int, pBKHM-0100-int, pBKHM-0105-thiA-int, and pBKHM-0149-int. Figure 13 A- Figure 13 D). These plasmids were used directly as negative controls.
[0637] Figure 13The intermediate plasmid / negative control vectors shown are (A) pBKHM-076-thiA-int, (B) pBKHM-0100-int, (C) pBKHM-0105-thiA-int, and (D) pBKHM-0149-int. Aspergillus oryzae U6 promoter; Af tRNA gly, Aspergillus fumigatus glycine tRNA; Aspergillus oryzae U6 terminator; DR, direct repeat; tracr RNA, tracr RNA; sgRNA, tandem of the corresponding direct repeat and tracrRNA sequences; AnPtef1, Aspergillus niger tef1 promoter; Ttef(nid), Aspergillus niger tef1 terminator; Ptef1, Aspergillus niger tef1 promoter; NATr, norsinolate resistance gene; TniaD, niaD terminator; pUC, pUC plasmid backbone sequence; ampR, Escherichia coli ampicillin resistance gene; AMA1, Aspergillus niger AMA1 origin of replication.
[0638] Plasmid construction
[0639] Before constructing the screening plasmids, the four plasmids pBKHM-076-thiA-int, pBKHM-0100-int, pBKHM-0105-thiA-int, and pBKHM-0149-int were digested with AsciI, and the linearized vectors were purified and recovered by silica gel column purification (Nucleobond Xtra Midi, Takara Bio Inc.).
[0640] To construct the selection plasmid, various 21 bp gRNA spacer oligonucleotide sequences targeting different regions of fwnA were used (Table 7). A 20 bp flanking region homologous to the plasmid backbone insertion site was added, and the oligonucleotides were synthesized using a commercial oligonucleotide synthesis service. Subsequently, the spacer sequences were ligated into four AscI-cleaved intermediate plasmids using HiFi DNA Assembly (New England Biolabs, USA) to form the final selection plasmid.
[0641] Table 7. Target sequences of various enzymes
[0642]
[0643] fungal strains
[0644] All procedures used Aspergillus niger strain MBin118 (NN049549).
[0645] culture medium
[0646] The following culture media were used in this study:
[0647] COVE-N-glyX: 218 g / L xylitol, 10 g / L glycerol, 2.02 g / L KNO3, 50 ml / L COVE salt solution, 25 g / L agar BA10, pH 5.3
[0648] YPG: 4 g / L yeast extract, 1 g / L KH2PO4, 0.5 g / L MgSO4.7aq, 15 g / L glucose, pH 6.0
[0649] Cove salt solution: 26 g KCl, 26 g MgSO4.7aq, 76 g KH2PO4, 50 ml Cove trace metal / L
[0650] COVE trace metals: 0.04 g NaB4O7.10aq, 0.4 g CuSO4.5aq, 1.2 g FeSO4.7aq, 0.7 gMnSO4.aq, 0.8 g Na2MoO2.2aq, 10 g ZnSO4.7aq / L
[0651] COVE-N top agar solution: 342.3 g / L sucrose, 20 ml / L COVE salt solution, 3 g / L NaNO2, 10 g / L Nippon Gene L-type low-melting-point agarose, 6 drops / L 5N NaOH
[0652] STC: 0.8 M sorbitol, 50 mM Tris (pH 8), 50 mM CaCl2
[0653] STPC: 40% PEG4000 in STC buffer.
[0654] Tween water: 1 g / L polyoxyethylene (20) dehydrated sorbitan monolaurate (Tween 20)
[0655] LB: 10 g / L Bacto tryptone, 10 g / L NaCl, 5 g / L Bacto yeast extract, pH 7.0
[0656] Protoplast formation and transformation
[0657] Protoplast formation
[0658] Inoculate MBin118 spores onto an agar slant (COVE-N-glyX) and allow the strain to grow at 30°C until complete spore formation. Add 9 ml of 0.1% Tween 20 water to the slant and manually suspend the spores. Transfer the spore suspension to a shake flask (500 ml) containing 100 ml of YPG medium with a baffle. Incubate the flask at 30°C or 32°C for 15–20 hours (60–80 rpm). Collect the mycelium by filtration through Mira cloth. Wash the mycelium 2–3 times with 0.6 M KCl or 0.7 M KCl + 10 mM CaCl2. Resuspend the mycelium in 20–30 ml of 0.6 M KCl or 0.7 M KCl + 10 mM CaCl2 containing 20–48 mg / ml Glucanex and 1.2 mg / ml BSA in a 50 ml centrifuge tube. Incubate the samples at 30°C or 32°C and 80 rpm for 1–1.5 hours, monitoring protoplast formation frequently under a microscope. After observing protoplast formation, filter the solution through Mira cloth into a 25 ml universal container (Nunc 364211). Then centrifuge the solution slowly at 2000 rpm for 10 minutes. Discard the supernatant and wash the precipitate with 5–15 ml of STC buffer, followed by slow centrifugation at 2000 rpm for 10 minutes. Resuspend the protoplasts in protoplast solution (STC / STPC / DMSO = 8:2:0.1) to a concentration of approximately 2 x 10⁻⁶. 7 One protoplast per ml. Gently mix and resuspend the pellet using pipettes.
[0659] Transformation
[0660] For each transformation, add the transforming DNA to 100 µl of protoplasts in a 14 ml Falcon tube, mix gently, and incubate on ice for more than 30 minutes. Add 1 ml of SPTC buffer, mix gently, and incubate at 37°C for 20 minutes. Add 10–15 ml of COVE-N top agar solution containing 50 µg / ml norocin to the solution, mix, and pour into the transformation plate. After the agar has hardened, incubate the plate at 34°C until colonies are clearly visible. The volume of transforming DNA should be less than 10 µl (1–10 µg).
[0661] strain isolation
[0662] Colonies were selected for each transformation and isolated onto COVE-N-glyX agar. The colonies were then incubated at 30°C for one week to allow them to form spores.
[0663] Spore PCR, DNA purification and Sanger sequencing
[0664] Reagents from the Phire Plant Direct PCR Kit (Thermofisher) were used. Spores from each fungal strain were first scraped from the fungal colony and immersed in 10 µl of dilution buffer. Spores were then picked using a 1 µl inoculation loop. The samples were briefly vortexed and incubated at room temperature for 5 min, followed by brief centrifugation. The samples were diluted 10-fold with sterile water and used as templates in subsequent PCR. A 20 µL PCR reaction was established using 0.5 µL of the prepared template, according to the manufacturer's instructions. PCR amplification was performed according to the manufacturer's instructions.
[0665] DNA was purified using standard molecular biology procedures, including gel electrophoresis, excision of DNA with the correct band length, and purification using a silica gel column (QIAquick Gel Extraction Kit, Qiagen). The purified DNA samples were then sequenced using a commercial Sanger sequencing service.
[0666] Example 3: Activity of a novel CRISPR nuclease in Bacillus licheniformis
[0667] Design of prototype spacers targeting the DsRED gene in Bacillus licheniformis
[0668] To evaluate the editing activity of the CRISPR nuclease of the present invention in Bacillus licheniformis, a knockout experiment was designed as follows. Because the Bacillus licheniformis strain MDT545 (WO 2021 / 183622) possesses a DsRED expression cassette (SEQ ID NO: 383) at the amyL locus on its chromosome, it was used as the host strain. Since MDT545 also possesses a GFP expression cassette at the xylA locus on its chromosome, DsRED knockout resulted in reduced red fluorescence and the resulting cells produced only green fluorescence; therefore, edited cells could be screened by observing the visible phenotypic changes. A deletion cassette (SEQ ID NO: 384) of the DsRED gene was designed to delete 67-bp DNA in the middle of the DsRED coding sequence. The prototype spacer sequences used to disrupt the DsRED gene are shown in SEQ ID NO: 385-404. The prototype spacer was designed as a 21-bp long DNA segment within the region to be deleted after disruption.
[0669] Construction of plasmids for expressing CRISPR components targeting the DsRED gene in Bacillus licheniformis
[0670] The complete construction of plasmid DNA is accomplished in several consecutive steps through the following sequential DNA manipulation.
[0671] Plasmids for expressing CRISPR components were assembled by PCR amplification of synthetic DNA. Purified PCR products were used in subsequent PCR reactions to generate single plasmids using SOE PCR as described in Materials and Methods. The PCR amplification reaction mixture contained 50 ng of each gel-purified PCR product, and plasmids were assembled and amplified using a thermal cycler. The resulting SOE products were directly used to transform the Bacillus subtilis host PP3724 to construct plasmids, which were subsequently used as a medium for transferring the constructed plasmids to the Bacillus licheniformis host MDT545. The synthetic DNA used for plasmid construction is shown in SEQ ID NO: 405-484. In short, a synthetic DNA fragment encoding a given CRISPR nuclease is integrated into a region of the mobile plasmid vector pBC16, flanked upstream by the Bacillus clausti promoter PamyL4199 (US Patent No. 6,100,063) and downstream by its aprH transcription terminator. This vector is tagged with tetracycline resistance (Bernhard et al. (1978), J Bacteriol. [Journal of Bacteriology] 133, 897-903). In the cases of 0076 nuclease (SEQ ID NO:1) and 0102 nuclease (SEQ ID NO:40), the IPTG-inducible promoter Pgrac is used instead of PamyL4199 for nuclease expression. On the other hand, synthetic DNA fragments encoding guide RNA expression cassettes were integrated upstream of the DsRED deletion cassette in a plasmid vector based on the temperature-sensitive pAMβ1-derived plasmid pWT, which was tagged with erythromycin resistance (Bidnenko et al. (1998), 28, 1005-1016) and contained the transfer origin oriT of plasmid pUB110, for transfer via conjugation (Selinger et al. (1990), J. Bacteriol. [Journal of Bacteriology], 172, 3290-3297). DNA diagrams representing typical nuclease vectors and guide RNA expression vectors are shown in the diagram. Figures 14-17 If necessary, the synthesized DNA was amplified by PCR using the oligonucleotide DNA (SEQ ID NO: 485-507) shown in Table 8, and purified by gel electrophoresis before SOE PCR and subsequent transformation into the Bacillus subtilis host PP3724. Finally, the transformants were isolated by tetracycline or erythromycin resistance.
[0672] Table 8. Oligonucleotide DNA used for amplifying synthetic DNA fragments
[0673]
[0674] Activity of novel CRISPR nuclease in Bacillus licheniformis
[0675] The aim of this experiment was to demonstrate that the novel nuclease could efficiently introduce the desired modifications into the chromosome of Bacillus licheniformis. As described above, the DsRED knockout phenotype (and green fluorescence) was used to visualize the targeted nuclease activity.
[0676] First, the guide RNA expression vector was transferred to the Bacillus licheniformis host MDT545 by conjugation with the Bacillus subtilis PP3724 host containing the plasmid. Since the Bacillus subtilis PP3724 host requires the addition of D-alanine to the growth medium, the desired Bacillus licheniformis conjugate could be isolated in an erythromycin-containing medium without D-alanine. Second, the corresponding nuclease vector was transferred to the Bacillus licheniformis host MDT545 by conjugation together with the guide RNA vector. Finally, the conjugate grown on a medium containing tetracycline and erythromycin was the conjugate containing the vector required for editing. The editing efficiency of the nuclease and guide RNA expression vector for a given group was calculated as follows: Editing efficiency (%) = G / N, where G represents the number of colonies showing only green fluorescence on a dual-resistance medium (agar plate), and N represents the total number of colonies grown on said medium. A summary of the editing efficiency of the nuclease and guide RNA for a given group is shown in Table 2. In the "Guide RNA Backbone" column, "Single" indicates that crRNA and tracrRNA are expressed separately, while "Single Guide RNA_a" or "Single Guide RNA_b" indicates that crRNA and tracrRNA are tandemly linked in different ways. In the "Target" column, "Empty" indicates that there is no prototype spacer sequence (0-bp) designed to target the DsRED gene, while prototype spacer-1 to -4 indicate different prototype spacer sequences (21-bp) designed.
[0677] Table 9. Summary of editing efficiency in Bacillus licheniformis host MDT545
[0678]
[0679] Nuclease 0076 was expressed by an IPTG-inducible Pgrac promoter. Green fluorescent colonies were observed in dual-resistance medium without IPTG (EXP_08 in Table 9), indicating that leak expression from Pgrac was sufficient to introduce the desired modification into the chromosome.
[0680] Nuclease 0102 was expressed by an IPTG-inducible Pgrac promoter. Total colony counts were performed on a dual-resistance medium without IPTG, where all colonies showed red fluorescence. Up to eight colonies per plate were then isolated into a medium containing IPTG, where phenotypic changes in the isolates were examined. Therefore, editing efficiency was calculated using the number of isolates rather than the total number of colonies, as shown in parentheses in Table 9.
[0681] Next, to confirm that the desired modification (a 67-bp deletion in the DsRED gene) was introduced into the chromosomes of green colonies, genomic PCR and Sanger sequencing were performed on selected colonies from EXP_08,_18,_20,_22,_27,_33,_37,_38,_43,_48,_49,_53,_55,_58, and_61, respectively. The oligonucleotide DNA used for genomic region amplification and Sanger sequencing is shown in Table 10 (SEQ ID NO: 508-511). Finally, it was demonstrated that all tested green colonies had the desired deletion within the DsRED gene in their chromosomes. Given that even with empty prototypical spacers (EXP_61 in Table 9, efficiency 1%), a small number of green colonies may still appear due to spontaneous homologous recombination, 1% should be used as a criterion for determining the efficiency of a given nuclease in increasing the desired modification in the chromosomes of Bacillus licheniformis. In summary, novel CRISPR nucleases 0076, 0100, 0102, and 0149 have shown promise as tools for genome engineering in Bacillus licheniformis hosts.
[0682] Table 10. Oligonucleotide DNA used to amplify the DsRED region of the genome.
[0683]
[0684] Materials and Methods
[0685] Material
[0686] The chemicals used as buffers and substrates are at least reagent-grade commercial products.
[0687] PCR amplification was performed using standard textbook procedures, a commercial thermal cycler, and either PrimeStar GXL polymerase (Takara Bio Inc., Japan) or KOD One polymerase (Toyobo Co., Ltd., Japan).
[0688] Use the following culture medium for bacterial growth.
[0689] LB: See EP 0 506 780.
[0690] TY: See WO 94 / 14968, page 16.
[0691] To select for erythromycin resistance, supplement agar and liquid medium with 5 μg / ml erythromycin. To select for tetracycline resistance, supplement agar and liquid medium with 15 μg / ml tetracycline. If necessary, add IPTG (isopropyl β-d-1-thiogalactoside) to a final concentration of 1 mM.
[0692] Oligonucleotide primers were obtained from Macrogen, a South Korean company. DNA manipulation (plasmid and genomic DNA preparation, restriction digestion, purification, ligation, and DNA sequencing) was performed using commercially available kits and reagents following standard textbook procedures.
[0693] DNA was introduced into naturally competent Bacillus subtilis using either a two-step procedure (Yasbin et al., 1975, J. Bacteriol. [Journal of Bacteriology] 121: 296-304.) or a one-step procedure, in which cell material from agar plates was resuspended in Spizisen 1 medium (12 ml) (WO 2014 / 052630) and shaken at 200 rpm at 37°C for approximately 4 hours. DNA was then added to 400 μL aliquots, and these aliquots were shaken at 150 rpm for 1 hour at the desired temperature before being plated on selective agar plates.
[0694] Using a modified Bacillus subtilis donor strain PP3724 containing pLS20, essentially as previously described (EP 2029732 B1), DNA was introduced into Bacillus licheniformis via conjugation from Bacillus subtilis, wherein the methyltransferase gene M.bli1904II (US 20130177942) was expressed by a triple promoter at the amyE locus, the pBC16-derived orf β and Bacillus subtilis comS genes (and kanamycin resistance genes) were expressed by a triple promoter at the alar locus (making the strain require D-alanine), and the Bacillus subtilis comS genes (and cat genes) were expressed by a triple promoter at the pel locus.
[0695] All constructs described in the examples were assembled from synthetic DNA fragments ordered from TWIST Bioscience, Inc., USA. As described in the examples, the fragments were assembled using sequence overlap extension (SOE).
[0696] strain
[0697] PP3724: This strain is a Bacillus subtilis derivative containing pLS20, in which the methyltransferase gene M.bli1904II (US 20130177942) is expressed by a triple promoter at the amyE locus, the pBC16-derived orf β and Bacillus subtilis comS genes (and kanamycin resistance genes) are expressed by a triple promoter at the alar locus (making the strain require D-alanine), and the Bacillus subtilis comS genes (and cat genes) are expressed by a triple promoter at the pel locus.
[0698] SJ1904: This strain is a Bacillus licheniformis derivative described in WO 2008 / 066931.
[0699] MDT545: This strain is the SJ1904 derivative described in WO 2021 / 183622.
[0700] plasmid
[0701] pC194: A plasmid isolated from Staphylococcus aureus (Horinouchi and Weisblum, 1982).
[0702] pE194: A plasmid isolated from Staphylococcus aureus (Horinouchi and Weisblum, 1982).
[0703] pUB110: The isolated plasmid was derived from (McKenzie et al., 1986).
[0704] Example 4: The novel nuclease is structurally highly different from Streptococcus pyogenes Cas9.
[0705] Table 11 shows the TM values for comparing the three-dimensional structures of the selected novel nucleases with those of *Streptococcus pyogenes* Cas9. When compared with the three-dimensional structure of *Streptococcus pyogenes* Cas9, the analyzed novel nucleases showed TM values of up to 0.62, indicating that these novel nucleases are structurally very different. Furthermore, Table 11 further highlights the close relationship between the nucleases of SEQ ID NO: 21 and SEQ ID NO: 29, with a TM of 0.98.
[0706] The close structural similarity among the novel nucleases 0076, 0100, and 0172 is also evident. Figures 20-22 and Figure 24 In the band comparison.
[0707] Figure 20 The band alignment between the protein structures of nucleases 0076 (black) and 0172 is shown. Figure 21The band alignment between the protein structures from nucleases 0100 (black) and 0172 is shown. Figure 22 The band alignment between the protein structures from nucleases 0076 (black) and 0102 is shown. Figure 24 The band alignment between the protein structures from nucleases 0076 (black) and 0100 is shown.
[0708] Figure 23 The band alignment between the protein structures of nuclease 0076 (black) and Streptococcus pyogenes Cas9 is shown, further demonstrating the significant structural differences between the novel nucleases.
[0709] Table 11. TM scores of novel Cas nucleases.
[0710]
[0711] Example 5: Novel CRISPR nucleases with nuclease activity in Escherichia coli
[0712] The cytotoxic effect of the novel CRISPR nuclease was confirmed.
[0713] In this experiment, the nuclease activities and efficiencies of novel CRISPR nucleases 0100 and 0149 were tested in Escherichia coli BL21 (DE3). Kill assays were used to evaluate the activity of the nucleases. The killing effect of CRISPR nucleases on Escherichia coli involves the precise targeting and cleavage of bacterial DNA, leading to cell death due to the inability to repair critical genomic damage [Bikard et al., (2014), Nature biotechnology, 32(11), 1146-1150].
[0714] Unlike eukaryotic cells, bacteria have limited mechanisms for repairing double-strand breaks (DSBs). The primary repair mechanism, non-homologous end joining (NHEJ), is error-prone and unavailable in *Escherichia coli* BL21 (DE3) strains. Without providing the cells with repair gene fragments, they struggle to survive after DSBs. Therefore, cytotoxic effects have been used as a primary method for assessing and identifying the efficacy of specific gRNA sequences targeting the adhE and tam loci in BL21 strains.
[0715] Neither adhE nor tam is an essential gene in BL21 strain, and therefore they can be knocked out or knocked in for different purposes. adhE is involved in metabolic processes related to fermentation and energy production, while tam is associated with stress response mechanisms that enhance bacterial survival under adverse conditions [Membrillo-Hernández et al. (2000), Journal of Biological Chemistry, 275(43), 33869-33875; and Ferla, M. and Patrick, W. (2014), Microbiology, 160(8), 1571-1584].
[0716] In this example, a dual-plasmid system was designed for CRISPR. Genes encoding spCas9, the novel nuclease 0100 (SEQ ID NO: 21), and the novel nuclease 0149 (SEQ ID NO: 48) were constructed in separate low-copy plasmids (CRISPR plasmids), and the gRNA sequence of the target gene, along with a CRISPR gRNA scaffold based on the specific nuclease, were constructed in another high-copy plasmid (guide plasmid). The gRNA and PAM sequences are illustrated in Table 12. Codon optimization of the sequence encoding the 0100 nuclease was performed for experiments (SEQ ID NO: 513). The wild-type (WT) coding sequence for the 0149 nuclease (SEQ ID NO: 100) and the codon-optimized sequence (OPT, SEQ ID NO: 518) were ordered and used in this experiment. The amino acid sequence of the nuclease remained unchanged for both codon-optimized sequences. An overview of the codon-optimized sequences used for expression in *Escherichia coli* is shown below. Figure 13 The CRISPR plasmid and guide plasmid were sequentially introduced into strain BL21 and induced according to the procedure described in "Materials and Methods" below. After transformation and recovery, the desired constructs were plated on agar plates containing antibiotics (with or without 0.4 μg / ml tetracycline adeoxycycline) and incubated at 30°C for at least 1–2 days. The killing effect was validated by comparing the number of colonies on the plates under induced and non-induced conditions.
[0717] Table 12. gRNA sequences and PAMs of spCas9, 0100, and 0149.
[0718]
[0719] Table 13. Codon-optimized nuclease coding sequences for expression in Escherichia coli.
[0720]
[0721] In this experiment, cytotoxic effects of 0100, 0149 WT, and 0149 OPT were observed at both the adhE and tam loci. Figure 18 A- Figure 18 As shown in F. Figure 18 The colony-forming comparisons of the cytotoxic effects of different constructs under induced and uninduced conditions are shown. The left-side plate was induced with 0.4 μg / ml adeoxytetracycline. The right-side plate was not induced. Figure 18 A shows 0100 with a gRNA that targets adhE. Figure 18 B shows the 0149 OPT with a gRNA that targets adhE. Figure 18 C shows the 0149 WT with a gRNA targeting adhE. Figure 18 D indicates 0100 with a gRNA targeting tam. Figure 18 E shows 0149 OPT with a gRNA targeting tam. Figure 18 F shows 0149 WT with gRNA targeting tam. Figure 18 G shows spCas9 with gRNA targeting adhE (positive control).
[0722] SpCas9 was used as a positive control, such as Figure 18 G is shown. This indicates some constructs (e.g., 0149 WT). Figure 18 C and Figure 18 F) and 0149 OPT ( Figure 18 B and Figure 18 E)) Even under uninduced conditions, fewer colonies were observed overall, likely due to leaky expression of CRISPR nucleases. This result suggests that these nucleases can be highly active even under leaky expression (where there are small amounts of nucleases in the cell), leading to cell death during the process.
[0723] Editing efficiency verification in the KIKO experiment
[0724] In Escherichia coli, genetic modification involves introducing specific DNA sequences into the bacterial chromosome using homologous recombination. This recombination event can be facilitated by the expression of recombinant proteins, such as those derived from λ phage (e.g., the λ-Red recombinase system), which enhances efficiency [Sabri et al. (2013), Microbial Cell Factories, 12(1), and Zhang et al. (2017), Current Microbiology, 74(8), 961-964]. Therefore, in the construction of CRISPR plasmids, the pSIM6 vector is used as a backbone to provide the λ-Red recombination system for efficient homologous recombination [Datta et al. (2006), Gene, 379, 109-115]. This method requires a DNA fragment containing the desired genetic sequence, flanked by regions homologous to the target site in the chromosome. The size of the homologous arms is crucial because it facilitates the integration of DNA into the chromosome via homologous recombination.
[0725] In this experiment, donor P J23104 The -mCherry DNA fragment was constructed with a very short homologous arm (100 bp), which significantly reduced recombination efficiency in the absence of a selection marker. The goal of this experiment was to test the efficiency of the novel CRISPR system as a reverse selection method, and whether it could provide better editing results compared to traditional methods using only the λ-Red recombinase system for genetic modification. Therefore, donor fragment P with a homologous arm from the adhE locus was used... J23104 -mCherry was introduced as a repair template to perform knock-in / knock-out (KIKO) at the adhE locus in strain BL21. No additional antibiotic selection markers were provided on the repair template.
[0726] The test constructs, including the required CRISPR and guide plasmid, were first thermally activated to induce the expression of λ-Red recombinase to facilitate the recombination of the donor fragment. The donor fragment was then introduced via electroporation. After transformation and recovery, the test candidates were induced with 0.4 μg / ml tetracycline and incubated overnight at 30°C in shake flasks. The induced cultures were then plated the following day to select individual colonies for characterization. 0149 Ctrl, lacking a guide plasmid, was considered a negative control along with the pSIM6 plasmid. Both constructs contained only the λ-Red recombinase system for genetic modification. Uninduced controls (0149 WT Ctrl and pSIM6 Ctrl) were plated immediately after recovery on agar plates with the appropriate antibiotics and incubated overnight at 30°C. Colony PCR was performed to confirm the insertion of the donor fragment under all conditions.
[0727] Figure 19 Editing efficiency of each CRISPR system is shown compared to conventional methods using only the λ-Red recombinase system. Donor DNA fragments were provided for all four conditions. Expression of nuclease 0149 (both OPT and WT) was induced overnight with 0.4 μg / ml adehydrotetracycline. The 0149 WT Ctrl contains the 0149 nuclease (encoded by the WT sequence) but without any guide plasmid. The pSIM6 Ctrl is a control containing only the λ-Red recombinase system and contains no CRISPR nuclease. Both conditions (pSIM6 and 0149 WT Ctrl) were used as negative controls for the experiments.
[0728] like Figure 19 As shown, the 0149 OPT sequence encoded by codon-optimized sequences exhibited 100% editing efficiency, with 79 out of 79 clones showing positive results for the desired insertion. The 0149 sequence encoded by WT sequences demonstrated 62.2% editing efficiency, with 28 out of 45 clones showing positive results for the desired insertion. Sequence validation of the insertions from selected clones was performed (data not shown). However, the 0149 WT Ctrl and pSIM6 controls (considered conventional editing methods) showed only 2% (1 out of 48 clones) and 3% (1 out of 33 clones), respectively. Compared to classical homologous recombination methods, the novel CRISPR nuclease 0149 exhibits significantly improved editing efficiency. Furthermore, codon optimization of the 0149 nuclease further improves its editing efficiency.
[0729] The results show that nucleases 0100 and 0149 exhibit nuclease activity in Escherichia coli. The codon-optimized 0149 nuclease showed 100% editing efficiency at the adhE locus, while the 0149 encoded by the wild-type DNA sequence showed 62.2% efficiency.
[0730] Materials and methods
[0731] Strains and plasmids construction
[0732] This study used Escherichia coli strain BL21(DE3) (Novagen). The CRISPR plasmid was constructed on the pSIM6 vector. Datta et al. (2006), Gene [Gene], 379, 109-115], this vector possesses inducible P that controls the expression of different CRISPR nucleases. tet Promoter. The CRISPR plasmid carries an ampicillin resistance marker and is generated by... Supplier synthesis. Guide plasmids containing gRNA sequences and corresponding gRNA scaffolds from different nucleases were synthesized on pTwist Kan high-copy v2. The gRNA sequences were synthesized under the control of the promoter. The guide plasmids carried a kanamycin resistance marker. The gRNA and PAM sequences are shown in Table 1. The gRNA sequence for the spCas9 adhE locus was ordered from Shukal et al. [Shukal et al. (2022), Microbial Cell Factories, 21(1), 19]. The donor DNA fragment P was synthesized. J23104 -mCherry (containing a 100 bp homologous flanking region from the adhE locus) was further amplified by PCR. It was then purified using NucleoSpin gel and a PCR Clean Mini Kit (Macherey-Nagel) for further experiments.
[0733] Culture medium and culture conditions
[0734] All cells were grown in 2x yeast extract tryptone medium. In short, cells from overnight cultures or 2xYT plates were seeded into 20 mL of fresh medium in shake flasks. Cells were incubated at 30°C for 16 h or longer before harvest. The medium was supplemented with appropriate antibiotics (100 mg / L ampicillin and 30 mg / L kanamycin) to maintain the appropriate plasmids.
[0735] CRISPR-Cas-mediated kill assay
[0736] To construct BL21 carrying the CRISPR plasmid, BL21 electrocompetent cells were prepared for transformation with the corresponding plasmid. For electroporation, 10 ng of plasmid was mixed with 60 μl of competent cells in a 1 mm Gene Pulser cuvette (Bio-Rad) and electroporated at 1.8 kV. The cells were recovered in 600 μl of SOC medium at 30°C for 1 h at 120 rpm, then spread onto 2xYT agar plates containing ampicillin (100 μg / ml) and incubated overnight at 30°C. This method was modified from that of Shukal et al. [Shukal et al. (2022), Microbial Cell Factories, 21(1), 19]. The guide plasmid was then introduced into the corresponding BL21-CRISPR-carrying strain using the same electroporation setup. After recovery, 100 μl of culture was spread on 2xYT agar plates containing antibiotics (100 μg / ml ampicillin and 30 μg / ml kanamycin) (with or without 0.4 μg / ml dehydrated tetracycline) and incubated at 30°C for at least 1–2 days to observe the cytotoxic effect.
[0737] CRISPR-Cas mediated gene knock-in / knockout (KIKO)
[0738] Single colonies containing both CRISPR and the guide plasmid from the kill assay were picked and inoculated into 20 ml of 2xYT medium containing 100 μg / ml ampicillin and 30 μg / ml kanamycin, and incubated overnight at 30°C with shaking at 120 rpm. Cells were induced at 42°C for 15 min for λ-Red recombinase expression, and then electroporated into competent cells [Datta et al. (2006), Gene, 379, 109-115]. Then, following the same electroporation protocol as described above, 300-600 ng of donor DNA fragment P was introduced. J23104 -mCherry.
[0739] Cells were recovered in 600 μl of SOC medium at 30°C and 120 rpm for 1 hour. 100 μl of the recovered culture was serially diluted and plated onto a 2xYT agar plate containing 100 μg / ml ampicillin and 30 μg / ml kanamycin, and incubated overnight at 30°C as a non-induction control. Separately, 200 μl of the recovered culture was transferred to 20 ml of 2xYT medium containing 0.4 μg / ml tetracycline, 100 μg / ml ampicillin, and 30 μg / ml kanamycin (to maintain the plasmid), and shaken overnight at 30°C and 120 rpm. The next day, the cell culture was serially diluted to obtain single colonies for characterization. This method was modified from publications [Datta et al. (2006), Gene, 379, 109-115; and Shukal et al. (2022), Microbial Cell Factories, 21(1), 19]. For the pSIM6 control, the procedure was similar to that described by Datta et al. (2006), Gene, 379, 109-115), with the same amount of donor fragment added as described above. After recovery, the recovered culture was plated directly onto selective agar plates and incubated overnight at 30°C.
[0740] To characterize the clones, individual colonies were selected and colony PCR was performed to identify successful recombination events based on the size of the insert.
[0741] Example 6: Novel nucleases for gene editing in Bacillus subtilis
[0742] As described by Sachla et al. (2021) in “A simplified method for CRISPR-Cas9 engineering of Bacillus subtilis” (Microbiol Spectr 9:e00754-21), the pJOE8999 plasmid was obtained from the Bacillus Genetic Collection (BGSC No. ECE358). This plasmid expresses the Cas9 nuclease under a mannose-inducible promoter and includes the Pvan promoter for constitutive expression when directing sgRNA cloning into a vector.
[0743] Plasmids were constructed using Golden Gate assembly.
[0744] The following pJOE899 (SEQ ID NO: 542) derived plasmids (including pJOE_Cas9_001, pJOE_NZ0149_002, and pJOE_NZ0149_empty) were generated by assembling two or three DNA fragments using BsaI restriction enzyme and T4 DNA ligase with standard Golden Gate assembly. The Golden Gate assemblies were then transformed into Escherichia coli TOP10 and selected with kanamycin.
[0745] The DNA fragments were generated by PCR or ordered as synthetic genes from Topway Biotechnology Co., Ltd. DNA containing the NZ0149 nuclease gene was codon-optimized for Bacillus subtilis expression and contained no BsaI restriction enzyme site, and ordered as synthetic DNA from Topway Biotechnology Co., Ltd. (SEQ ID NO: 528). DNA fragments containing guide RNAs of the CRISPR endonuclease Cas9 and the novel nuclease 0149, sgRNA_Cas9_001 (SEQ ID NO: 530) and sgRNA_NZ0149_002 (SEQ ID NO: 531), respectively, were ordered as synthetic DNA from Topway Biotechnology Co., Ltd. Using pJOE8999 plasmid as template DNA, DNA fragments containing kanamycin (KAN) resistance and the Cas9 gene were amplified by PCR using oligonucleotides.
[0746] The assembly of DNA fragments and oligonucleotides and corresponding template DNA used for PCR reactions is shown in Table 14 below.
[0747] Bacillus subtilis strains with the dsRed gene integrated into their genome
[0748] A Bacillus subtilis strain containing the full-length dsRed gene (SEQ ID NO: 544) integrated into the Pel locus was constructed via homologous recombination. This strain was then prepared into competent cells and used to transform the pJOE plasmid and repair DNA.
[0749] Preparation of repair DNA containing a 130 bp deletion of the dsRed gene.
[0750] Repair DNA (SEQ ID NO:18) containing a 130 bp deletion of the dsRed gene (SEQ ID NO:543) was ordered as synthetic DNA from TopVest Biotechnology Co., Ltd.
[0751] The synthesized repair DNA (SEQ ID NO: 545) was amplified by PCR using forward and reverse oligonucleotides identical to the 5' and 3' sequences. The PCR amplification products were then co-transformed with the pJOE plasmid into Bacillus subtilis.
[0752] Table 14. DNA fragments used for Golden Gate assembly
[0753]
[0754] Plasmid preparation and transformation of Bacillus subtilis containing the dsRed gene integrated into the Pel locus:
[0755] Each plasmid was sequence verified in Escherichia coli TOP10 and transformed into Escherichia coli TG1 to produce multimeric plasmid DNA as described by Sachla et al. Small-scale plasmid preparations were then prepared, and competent cells of Bacillus subtilis strains with dsRED integrated into their genome were co-transformed with 2 µg plasmid DNA and 2 µg repair template DNA (SEQ ID NO: 545) (provided as PCR product). The transformants were plated on Luria medium (LB) containing 0.5% mannose (M2069; SIGMA) and 15 µg / ml kanamycin and incubated at 30°C for 3 days. White and red colonies were counted. Results are shown in Table 15.
[0756] Table 15. Colony count results of Bacillus subtilis on LB+Kan+mannose plates
[0757]
[0758] White and red Bacillus subtilis colonies were selected and streaked onto new LB+KAN+mannose plates and incubated at 30°C for 2 days. Colony PCR was then performed using oligonucleotides (SEQ ID NO:540 and SEQ ID NO:541) to amplify the dsRed gene and flanking regions from the Pel locus. The results showed that white colonies contained smaller fragments corresponding to the deleted dsRed gene, while red colonies contained larger PCR fragments corresponding to the full-length dsRed gene.
[0759] These findings, along with sequencing results, confirm that the functional dsRed gene at the Pel locus has been replaced by a non-functional dsRed gene derived from DNA repair due to the action of the Cas9 or O149 nuclease. Therefore, nuclease O149 can be used for gene editing in the Bacillus subtilis host.
[0760] Plasmid construction via POE-PCR
[0761] The following plasmids derived from pJOE899 (SEQ ID NO:542) (including pJOE_NZ0100_004, pJOE_NZ0100_empty, pJOE_NZ0102_003, pJOE_NZ0102_004, and pJOE_NZ0102_empty) were generated by POE-PCR (POE-PCR described in WO 24133344) by assembling synthetic DNA (ordered from Topsys Biotechnology, containing codon-optimized DNA containing the 0100 and 0102 nuclease genes for Bacillus subtilis expression) and plasmid elements. See Table 16 for fragment assembly and SEQ IDs.
[0762] Table 16. DNA fragments used for POE
[0763]
[0764] Table 17 shows additional sequence information used in the Bacillus subtilis experiments.
[0765] Table 17. CRISPR nucleases and their corresponding PAM motifs and PAM sequences, spacer sequences, and sgRNA sequences
[0766]
[0767] POE-PCR was used to transform Bacillus subtilis containing the dsRed gene integrated into the Pel locus:
[0768] POE-PCR was used to transform Bacillus subtilis strains with dsRED integrated into their genome. Transformations were plated in Luria medium (LB) containing 0.5% mannose (M2069; SIGMA) and 15 µg / ml kanamycin and incubated at 30°C for 3 days. As previously described, the efficiency of recombination was assessed by counting the number of red and white colonies. While only red colonies were obtained on plates transformed with empty plasmids pJOE_NZ0102_empty and pJOE_NZ0100_empty, plates transformed with pJOE_NZ0100_004, pJOE_NZ0102_003, and pJOE_102_004 contained both white and red colonies.
[0769] As previously mentioned, PCR and sequencing analysis of red and white colonies confirmed that the deletion in dsRED had been introduced as expected. Therefore, the novel nucleases 0100 and 0102, along with their corresponding sgRNAs, can be used for gene editing in the Bacillus subtilis host.
[0770] The invention described and claimed herein is not limited to the specific aspects disclosed herein, as these aspects are intended to illustrate several aspects of the invention. Any equivalent aspects are intended to be within the scope of the invention. In fact, various modifications to the invention, in addition to those shown and described herein, will become apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims. In case of conflict, this disclosure (including definitions) shall prevail.
[0771] The invention is further defined by the following numbered paragraphs:
[0772] 1. A Cas nuclease selected from the group consisting of:
[0773] (a) Polypeptides corresponding to SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13. SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26. SEQ ID NO: 27. SEQ ID NO: 28. SEQ ID NO: 29. SEQ ID NO: 30. SEQ ID NO: 31. SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44. SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: The amino acid sequence of SEQ ID NO: 52 has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity, preferably with the amino acid sequence of SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39, SEQ ID NO: 48, or SEQ ID NO: 29 having at least 60%, for example,At least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity;
[0774] (b) A polypeptide encoded by a polynucleotide, wherein the polynucleotide is associated with SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, SEQ ID NO: 89, SEQ ID NO: 90, SEQ ID NO: 91, SEQ ID NO: 92, SEQ ID NO: 93, SEQ ID NO: 94, SEQ ID NO: 95. SEQ ID NO: 96, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103 or SEQ ID NO: The polypeptide coding sequence of 104 has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity, preferably the polypeptide is identical to any one of SEQ ID NO: 53, SEQ ID NO: 73, SEQ ID NO: 92, SEQ ID NO: 91, SEQ ID NO: 100, or SEQ ID NO: 81.The polypeptide coding sequence of any one of SEQ ID NO: 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549, or 550 has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity;
[0775] (c) A polypeptide derived from SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12 ...13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 30, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 19, SEQ ID NO: 10, SEQ ID NO: 19, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO 17. SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30. SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43. SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: 52, preferably derived from SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39, SEQ ID NO: 48 or SEQ ID NO: 29;
[0776] (d) A polypeptide having a TM-score of at least 0.80, for example, at least 0.85, at least 0.90, at least 0.91, at least 0.92, at least 0.93, at least 0.94, at least 0.95, at least 0.96, at least 0.97, at least 0.98, at least 0.99 or even 1.0, compared to the three-dimensional structure of the polypeptide of any one of SEQ ID NO: 1-52, preferably SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39, SEQ ID NO: 48 or SEQ ID NO: 29, wherein the three-dimensional structure is calculated using Alphafold;
[0777] (e) A polypeptide derived from (a), (b), (c), or (d), wherein the N-terminus and / or C-terminus have been extended by adding one or more amino acids; and
[0778] Fragments of polypeptides in (f), (a), (b), (c), (d), or (e).
[0779] 2. The nuclease described in paragraph 1, which has nuclease activity and / or DNA binding activity.
[0780] 3. The nuclease according to any one of paragraphs 1-2, wherein the nuclease comprises or is composed of the following: an amino acid sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39, SEQ ID NO: 48, or SEQ ID NO: 29.
[0781] 4. The nuclease according to any one of paragraphs 1-3, wherein the nuclease comprises, is substantially composed of, or is composed of: SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39, SEQ ID NO: 48 or SEQ ID NO: 29.
[0782] 5. The nuclease according to any one of paragraphs 1-4, wherein the nuclease is SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30. SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43. SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or fragments of SEQ ID NO: 52, preferably SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39, SEQ ID NO: 48 or SEQ ID NO: 29 fragments, wherein the fragment preferably contains at least 600 amino acid residues (e.g.,Amino acids 9 to 640 of SEQ ID NO: 1, 13 to 628 of SEQ ID NO: 39, 16 to 637 of SEQ ID NO: 40, 10 to 637 of SEQ ID NO: 41, 10 to 639 of SEQ ID NO: 42, 10 to 636 of SEQ ID NO: 43, 10 to 635 of SEQ ID NO: 44, 9 to 640 of SEQ ID NO: 45, 10 to 637 of SEQ ID NO: 46, 10 to 633 of SEQ ID NO: 47, 12 to 632 of SEQ ID NO: 48, 9 to 640 of SEQ ID NO: 51, 8 to 620 of SEQ ID NO: 21, or 9 to 640 of SEQ ID NO: 52.
[0783] 6. The nuclease according to any one of paragraphs 1-5, comprising, substantially comprising, or comprising of: SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29. SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42. SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: 52, preferably SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40. SEQ ID NO: 39, SEQ ID NO: 48 or SEQ ID NO: 29.
[0784] 7. The nuclease according to any one of paragraphs 1-6, wherein the nuclease is encoded by a polynucleotide, the polynucleotide being associated with SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80. SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, SEQ ID NO: 89, SEQ ID NO: 90, SEQ ID NO: 91, SEQ ID NO: 92, SEQ ID NO: 93. SEQ ID NO: 94, SEQ ID NO: 95, SEQ ID NO: 96, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103 or SEQ ID NO: 104, or SEQ ID NO: Any one of 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549, or 550.Preferably, the mature polypeptide coding sequences of SEQ ID NO: 53, SEQ ID NO: 73, SEQ ID NO: 92, SEQ ID NO: 91, SEQ ID NO: 100, or SEQ ID NO: 81 have at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity.
[0785] 8. The nuclease according to any one of paragraphs 1-7, the nuclease comprising an N-terminal extension and / or a C-terminal extension of 1-10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids, preferably an extension of 1-10 amino acid residues in the N-terminus and / or 1-10 amino acids, such as 1-5 amino acids, in the C-terminus.
[0786] 9. The nuclease according to any one of paragraphs 1-8, wherein the nuclease is associated with SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 1. ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: 52, preferably SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39, SEQ ID NO: 48 or SEQ ID NO: The 29-peptide has a sequence difference of up to 10%, up to 9%, up to 8%, up to 7%, up to 6%, up to 5%, up to 4%, up to 3%, up to 2%, or up to 1%.
[0787] 10. The nuclease according to any one of paragraphs 1-9, wherein the nuclease is associated with SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30. SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43. SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51 or SEQ ID NO: 52, preferably SEQ ID NO: 1, SEQ ID NO: 21, SEQ ID NO: 40, SEQ ID NO: 39. SEQ ID NO: 48 or SEQ ID NO: The polypeptides of number 29 differ by up to 15 amino acids, for example, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids.
[0788] 11. The nuclease according to any one of paragraphs 1-10, wherein the nuclease is obtained from or can be obtained from Streptococcus cells (e.g., Streptococcus equi, Streptococcus mutans, Streptococcus species, Streptococcus henryi DSM 19005, Streptococcus species CCH8-G7, Streptococcus paclitaxel, Streptococcus mutans DSM 15617, Streptococcus salivarius, or Streptococcus ruminant cells), Bacillus cells (e.g., Bacillus species-63030 cells), Zurich bacillus cells (e.g., Zurich bacillus species cells), Ureaplasma cells (e.g., Ureaplasma thermophilus cells), Chlorobacterium cells (e.g., Chlorobacterium humanis cells, Clostridium cells), Ruminococcus cells (e.g., Ruminococcus species cells), and Cyclocycline Bacillus cells (e.g., Cyclocycline Bacillus cells). The cells may contain *Bacillus* cells, *Enterococcus* cells (e.g., *Enterococcus faecalis*, *Enterococcus harzianum*, or *Enterococcus asper* cells), *Lactobacillus* cells (e.g., *Lactobacillus aspera*, *Lactobacillus halophilus*, *Lactobacillus keshanensis*, *Lactobacillus sauerkraut*, or *Lactobacillus hulinensis* cells), *Lactobacillus bumblebee* cells (e.g., *Lactobacillus bumblebee* cells), or *Zygococcus* cells (e.g., *Zygococcus prawni* cells), preferably *Enterococcus asper* cells, *Enterococcus harzianum* cells, *Zygococcus prawni* cells, or *H. cirrhosum* cells.
[0789] 12. The nuclease according to any one of paragraphs 1-10, wherein the nuclease is obtained from or can be obtained from Lactobacillus cells, for example, Lactobacillus species, Lactobacillus sausageii (DSM 20184), Lactobacillus sausageii, Lactobacillus murineis, Lactobacillus rumenii, Lactobacillus salivarius, Lactobacillus janniae, Lactobacillus hamsterii, Lactobacillus delbrueckii, Lactobacillus johnsonii, Lactobacillus plantarum, Lactobacillus rhamnosus, or Lactobacillus chrysogenus cells.
[0790] 13. The nuclease according to any one of paragraphs 1-12, wherein the nuclease comprises one or more functional RuvC domains.
[0791] 14. The nuclease according to any one of paragraphs 1-13, wherein the nuclease comprises one or more functional HNH domains.
[0792] 15. The nuclease according to any one of paragraphs 1-14, wherein the nuclease comprises one or more domains selected from the group consisting of:
[0793] (a) A RuvC domain having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 105-143 or 313-318, preferably SEQ ID NO: 105-107, 111-113, 108-110, 135-137 or 313-318;
[0794] (b) An HNH domain having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 144-156 or 319-320, preferably any of SEQ ID NO: 144, 146, 145, 154, 319, or 320;
[0795] (c) The RuvC domain, which is derived from SEQ ID NO: 105-143 or 313-318 by substitution, deletion or addition of one or more amino acids of SEQ ID NO: 105-107, 111-113, 108-110, 135-137 or 313-318, preferably any one of SEQ ID NO: 105-107, 111-113, 108-110, 135-137 or 313-318;
[0796] (d) An HNH domain derived from SEQ ID NO: 144-156 or 319-320, preferably any one of SEQ ID NO: 144-156, 145-154, 319 or 320, by substitution, deletion or addition of one or more amino acids of SEQ ID NO: 144-156, preferably SEQ ID NO: 144, 146, 145, 154, 319 or 320; and
[0797] Fragments of the catalytic domains of (e), (a), (b), (c), or (d);
[0798] Preferably, the nuclease has nuclease activity, or the nuclease has nicking enzyme activity.
[0799] 16. The nuclease according to any one of paragraphs 1-15, wherein the HNH domain has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 144-156 or 319-320, preferably any one of SEQ ID NO: 144, 146, 145, 154, 319, or 320.
[0800] 17. The nuclease according to any one of paragraphs 1-16, wherein the HNH domain comprises or consists of an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 144, 146, 145, 154, 319, or 320.
[0801] 18. The nuclease according to any one of paragraphs 1-17, wherein the HNH domain is a variant of SEQ ID NO: 144-156 or 319-320, preferably a variant of any one of SEQ ID NO: 144, 146, 145, 154, 319 or 320, the variant comprising substitutions, such as conserved amino acid substitutions, deletions and / or insertions, at one or more positions.
[0802] 19. The nuclease according to any one of paragraphs 1-18, wherein the HNH domain differs from SEQ ID NO: 144-156 or 319-320, preferably any one of SEQ ID NO: 144, 146, 145, 154, 319 or 320 by up to 15 amino acids, for example up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids.
[0803] 20. The nuclease according to any one of paragraphs 1-19, wherein the HNH domain is a fragment of SEQ ID NO: 144-156 or 319-320, preferably any one of SEQ ID NO: 144, 146, 145, 154, 319 or 320, wherein the fragment preferably contains at least 20 amino acid residues (e.g., amino acids 613 to 640 of SEQ ID NO: 1) or at least 27 amino acid residues (e.g., amino acids 613 to 640 of SEQ ID NO: 1).
[0804] 21. The nuclease according to any one of paragraphs 1-20, wherein the HNH domain comprises, is substantially composed of, or is composed of: SEQ ID NO: 144-156 or 319-320, preferably any one of SEQ ID NO: 144, 146, 145, 154, 319 or 320.
[0805] 22. The nuclease according to any one of paragraphs 1-21, wherein the RuvC domain has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO: 105-143 or 313-318, preferably SEQ ID No: 105-107, 111-113, 108-110, 135-137, or 313-318.
[0806] 23. The nuclease according to any one of paragraphs 1-22, wherein the RuvC domain comprises or consists of the following: an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the amino acid sequence of SEQ ID NO: 105-143 or 313-318, preferably SEQ ID NO: 105-107, 111-113, 108-110, 135-137, or 313-318.
[0807] 24. The nuclease according to any one of paragraphs 1-23, wherein the RuvC domain is a variant of SEQ ID NO: 105-143 or 313-318, preferably a variant of any one of SEQ ID No: 105-107, 111-113, 108-110, 135-137 or 313-318, the variant comprising substitutions, such as conserved amino acid substitutions, deletions and / or insertions, at one or more positions.
[0808] 25. The nuclease according to any one of paragraphs 1-24, wherein the RuvC domain differs from any one of SEQ ID NO: 105-143 or 313-318, preferably SEQ ID No: 105-107, 111-113, 108-110, 135-137 or 313-318 by up to 15 amino acids, for example, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 amino acids.
[0809] 26. The nuclease according to any one of paragraphs 1-25, wherein the RuvC domain is a fragment of SEQ ID NO: 105-143 or 313-318, preferably any one of SEQ ID No: 105-107, 111-113, 108-110, 135-137 or 313-318, wherein the fragment preferably contains at least 10 amino acid residues (e.g., amino acid 5 to 20 of SEQ ID NO: 1).
[0810] 27. The nuclease according to any one of paragraphs 1-26, wherein the RuvC domain comprises, is substantially composed of, or is composed of: SEQ ID NO: 105-143 or 313-318, preferably any one of SEQ ID No: 105-107, 111-113, 108-110, 135-137 or 313-318.
[0811] 28. The nuclease according to any one of paragraphs 1-27, wherein the nuclease has double-strand break activity against a DNA target site.
[0812] 29. The nuclease according to any one of paragraphs 1-28, wherein the nuclease comprises an amino acid substitution, insertion, or deletion in one or more RuvC domains.
[0813] 30. The nuclease according to any one of paragraphs 1-29, wherein the nuclease comprises an amino acid substitution, insertion, or deletion in one or more HNH domains.
[0814] 31. The nuclease according to any one of paragraphs 1-30, wherein the nuclease is a nicking enzyme having one or more inactive RuvC domains generated by substitution, insertion or deletion of amino acids at the positions provided for the nuclease in column 3 of Table 2.
[0815] 32. The nuclease according to any one of paragraphs 1-31, wherein the nuclease is a nicking enzyme having one or more inactive HNH domains generated by substitution, insertion or deletion of amino acids at the positions provided for the nuclease in column 3 of Table 3.
[0816] 33. The nuclease according to any one of paragraphs 1-32, wherein the nuclease has single-strand breakage activity against a DNA target site.
[0817] 34. The nuclease according to any one of paragraphs 1-32, wherein the nuclease is a catalytically inactivated nuclease.
[0818] 35. The nuclease according to paragraph 34, wherein the catalytically inactivated nuclease comprises one or more inactivated RuvC domains and one or more inactivated HNH domains.
[0819] 36. The nuclease according to any one of paragraphs 34-35, wherein the catalytically inactivated nuclease comprising one or more inactivated RuvC domains and one or more inactivated HNH domains is produced by substitution, deletion or insertion of one or more amino acids at the positions provided for the nuclease in column 3 of Table 2 or column 3 of Table 3.
[0820] 37. The nuclease according to any one of paragraphs 1-36, wherein sequence identity is determined by the method described in the definition section of “Sequence Identity”.
[0821] 38. The nuclease according to any one of paragraphs 1-37, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in eukaryotic cells.
[0822] 39. The nuclease according to any one of paragraphs 1-38, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in mammalian cells, such as non-human mammalian cells.
[0823] 40. The nuclease according to any one of paragraphs 1-37, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Escherichia coli cells.
[0824] 41. The nuclease according to any one of paragraphs 1-37, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Bacillus cells.
[0825] 42. The nuclease according to any one of paragraphs 1-37, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Bacillus subtilis cells.
[0826] 43. The nuclease according to any one of paragraphs 1-37, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Bacillus licheniformis cells.
[0827] 44. The nuclease according to any one of paragraphs 1-38, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in filamentous fungal cells.
[0828] 45. The nuclease according to any one of paragraphs 1-38, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Aspergillus niger cells.
[0829] 46. The nuclease according to any one of paragraphs 1-38, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Aspergillus oryzae cells.
[0830] 47. The nuclease according to any one of paragraphs 1-38, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Trichoderma reesei cells.
[0831] 48. The nuclease according to any one of paragraphs 1-37, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Lactobacillus cells.
[0832] 49. The nuclease according to any one of paragraphs 1-37, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in probiotic cells.
[0833] The nuclease according to any one of paragraphs 1-38, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Saccharomyces cerevisiae cells.
[0834] 50. The nuclease according to any one of the preceding paragraphs, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Pichia pastoris.
[0835] 50b. The nuclease according to any one of the preceding paragraphs, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Lactobacillus paracasei (Lb. paracasei, Lacticaseibacillus paracasei or Lactobacillus paracasei).
[0836] 50c. The nuclease according to any one of the preceding paragraphs, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Streptococcus thermophilus.
[0837] 50d. The nuclease according to any one of the preceding paragraphs, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Escherichia coli cells, wherein the polynucleotide comprises or consists of a sequence having at least 80%, for example, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of the nucleotide sequences in SEQ ID NO: 512-520.
[0838] 50e. The nuclease according to any one of the preceding paragraphs, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Bacillus licheniformis cells, wherein the polynucleotide comprises or consists of a sequence having at least 80%, for example, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the nucleotide sequence of any one of SEQ ID NO: 405, 416, 417, 434, 449, 465, or 466.
[0839] 50f. The nuclease according to any one of the preceding paragraphs, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Aspergillus niger cells, wherein the polynucleotide comprises or consists of a sequence having at least 80%, for example, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 347, 349, 351, or 353.
[0840] 50g. The nuclease according to any one of the preceding paragraphs, wherein the polynucleotide encoding the nuclease is codon-optimized for expression in Bacillus subtilis cells, wherein the polynucleotide comprises or consists of a sequence having at least 80%, for example, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the nucleotide sequence of any one of SEQ ID NO: 528, 549, or 550.
[0841] 51. The nuclease according to any one of paragraphs 1-50g, wherein the nuclease is a class 2 Cas nuclease.
[0842] 52. The nuclease according to any one of paragraphs 1-51, wherein the nuclease is a type II Cas nuclease.
[0843] 53. The nuclease according to any one of paragraphs 1-52, wherein the nuclease is a type II-A Cas nuclease.
[0844] 54. The nuclease according to any one of paragraphs 1-52, wherein the nuclease is a type II-B Cas nuclease.
[0845] 55. The nuclease according to any one of paragraphs 1-52, wherein the nuclease is a type II-C Cas nuclease.
[0846] 56. The nuclease according to any one of paragraphs 1-55, wherein the nuclease utilizes the prototype spacer adjacent motif (PAM) sequence provided for the nuclease in Table 1.
[0847] 57. The nuclease according to any one of paragraphs 1-56, wherein the nuclease is not naturally occurring, for example, wherein the nuclease is engineered and contains non-natural or synthetic amino acids.
[0848] 58. The nuclease according to any one of paragraphs 1-57, wherein the nuclease is naturally occurring.
[0849] 59. A fusion polypeptide comprising a Cas nuclease according to any one of paragraphs 1-58, and one or more second polypeptides.
[0850] 60. The fusion polypeptide according to paragraph 59, wherein the one or more second polypeptides comprises polypeptides localized to one or more subcellular organelles.
[0851] 61. The fusion polypeptide according to any one of paragraphs 59-60, wherein the one or more second polypeptides is a nuclear localization sequence (NLS), a cell-penetrating peptide, and / or an affinity tag.
[0852] 62. The fusion polypeptide according to any one of paragraphs 59-61, wherein the fusion polypeptide comprises 1-10 or more NLS at or near an amino terminus, 1-10 or more NLS at or near a carboxyl terminus, or a combination of 1-10 or more NLS at or near an amino terminus and 1-10 or more NLS at or near a carboxyl terminus.
[0853] 63. The fusion polypeptide according to any one of paragraphs 59-62, wherein the fusion polypeptide comprises 1-4 NLS.
[0854] 64. The fusion polypeptide according to any one of paragraphs 59-63, wherein the fusion polypeptide comprises an NLS.
[0855] 65. The fusion polypeptide according to any one of paragraphs 59-64, wherein the one or more NLS are located within the open reading frame (ORF) of the nuclease.
[0856] 66. The fusion polypeptide according to any one of paragraphs 59-65, wherein the one or more NLS are tandem repeats.
[0857] 67. The fusion polypeptide according to any one of paragraphs 59-66, wherein the fusion polypeptide comprises a first NLS and a second NLS.
[0858] 68. The fusion polypeptide according to paragraph 67, wherein the fusion polypeptide comprises a linker sequence between the first NLS and the second NLS.
[0859] 69. The fusion polypeptide according to paragraph 68, wherein the linker between the first NLS and the second NLS comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 amino acids.
[0860] 70. The fusion polypeptide according to any one of paragraphs 59-69, wherein the one or more second polypeptides comprises a base-editing polypeptide.
[0861] 71. The fusion polypeptide according to any one of paragraphs 59-70, wherein the base-editing polypeptide comprises a base editor domain.
[0862] 72. The fusion polypeptide according to any one of paragraphs 59-71, wherein the fusion polypeptide comprises a linker between the Cas nuclease and the base-editing polypeptide.
[0863] 73. The fusion polypeptide according to any one of paragraphs 59-72, wherein the base-editing polypeptide comprises a deaminase, for example, cytidine deaminase, such as APOBEC3A deaminase or adenosine deaminase.
[0864] 74. The fusion polypeptide according to any one of paragraphs 59-73, wherein the one or more second polypeptides comprises a reverse transcriptase, the reverse transcriptase preferably comprising a reverse transcriptase domain.
[0865] 75. The fusion polypeptide according to any one of paragraphs 59-74, wherein the nuclease is fused with one or more NLS of sufficient strength to drive the accumulation of a CRISPR complex containing the Cas nuclease in the nucleus of a eukaryotic cell to a detectable amount.
[0866] 76. The nuclease or fusion polypeptide according to any one of paragraphs 1-75, wherein the nuclease or fusion polypeptide is isolated.
[0867] 77. The nuclease or fusion polypeptide according to any one of paragraphs 1-76, wherein the nuclease or fusion polypeptide is purified.
[0868] 78. The nuclease or fusion polypeptide according to any one of paragraphs 1-77, wherein sequence identity is determined by the method described in the definition section of “Sequence Identity”.
[0869] 79. A non-naturally occurring composition comprising (i) a nuclease or fusion polypeptide according to any one of paragraphs 1-78, and / or (ii) a nucleic acid molecule comprising a sequence encoding the nuclease or fusion polypeptide according to any one of paragraphs 1-78.
[0870] 80. The composition according to paragraph 79, wherein the nucleic acid molecule is a chemically modified nucleic acid molecule.
[0871] 81. The composition according to any one of paragraphs 79-80, wherein the nucleic acid molecule is DNA.
[0872] 82. The composition according to any one of paragraphs 79-81, wherein the nucleic acid molecule is RNA.
[0873] 83. The composition according to any one of paragraphs 79-82, wherein the RNA is mRNA comprising one or more of the following: a 5' untranslated region (UTR), an open reading frame (ORF) encoding a Cas nuclease or a fusion polypeptide, a 3' UTR, and a polyadenylated (polyA) tail.
[0874] 84. The composition according to any one of paragraphs 79-83, wherein the ORF is composed of a nucleoside selected from adenosine, modified adenosine, uridine, modified uridine, guanosine, modified guanosine, cytidine and modified cytidine.
[0875] 85. The composition according to any one of paragraphs 79-84, wherein the ORF is composed of a nucleoside selected from adenosine, uridine, modified uridine, guanosine and cytidine.
[0876] 86. The composition according to any one of paragraphs 79-86, wherein the nucleic acid molecule is linear.
[0877] 87. The composition according to any one of paragraphs 79-85, wherein the nucleic acid molecule is circular.
[0878] 88. The composition according to any one of paragraphs 79-87, further comprising one or more RNA molecules, or DNA polynucleotides encoding one or more of the one or more RNA molecules, wherein the one or more RNA molecules and the Cas nuclease or fusion polypeptide do not naturally coexist, and the one or more RNA molecules are configured to form a complex with the Cas nuclease or fusion polypeptide and / or to target the complex to a target site.
[0879] 89. The composition according to any one of paragraphs 79-88, wherein the one or more RNA molecules comprise guide RNA (gRNA), the gRNA comprising CRISPR RNA (crRNA) and trans-activating RNA (tracrRNA).
[0880] 90. The composition according to any one of paragraphs 79-89, wherein the one or more RNA molecules are single-molecule RNAs (sgRNAs), for example, wherein the crRNA and the tracrRNA are part of the same RNA molecule.
[0881] 91. The composition according to any one of paragraphs 79-89, wherein the one or more RNA molecules are bimolecular RNAs, for example, wherein the crRNA and the tracrRNA are separate RNA molecules.
[0882] 92. The composition according to any one of paragraphs 79-91, wherein the composition further comprises a donor template for homologous directional repair (HDR).
[0883] 93. The composition according to any one of paragraphs 79-92, wherein the sequence encoding the Cas nuclease or fusion polypeptide comprises the sequences of SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 65, SEQ ID NO: 66, SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 79, SEQ ID NO: 80. SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, SEQ ID NO: 89, SEQ ID NO: 90, SEQ ID NO: 91, SEQ ID NO: 92, SEQ ID NO: 93. SEQ ID NO: 94, SEQ ID NO: 95, SEQ ID NO: 96, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103 or SEQ ID NO: 104, preferably SEQ ID NO: 53. SEQ ID NO: 73, SEQ ID NO: 92, SEQ ID NO: 91, SEQ ID NO: 100 or SEQ ID NO: 81, or any one of SEQ ID NO: 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549, or 550, has a polynucleotide sequence containing at least 60%, for example,A sequence with at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity.
[0884] 94. The composition according to any one of paragraphs 79-93, wherein the one or more RNA molecules comprise a trans-activating RNA (tracrRNA) sequence encoded by a polynucleotide, the polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of any one of SEQ ID NO: 157-208, preferably SEQ NO: 157, SEQ ID NO: 177, SEQ ID NO: 196, SEQ ID NO: 195, SEQ ID NO: 204, or SEQ ID NO: 185.
[0885] 95. The composition according to any one of paragraphs 79-93, wherein at least one of the one or more RNA molecules comprises a CRISPR RNA (crRNA) molecule containing a guide sequence portion and a sequence encoded by a polynucleotide, the polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of any one of SEQ ID NO: 209-260, preferably SEQ ID NO: 209, SEQ ID NO: 229, SEQ ID NO: 248, SEQ ID NO: 247, SEQ ID NO: 256, or SEQ ID NO: 237.
[0886] 96. The composition according to any one of paragraphs 79-95, wherein at least one of the one or more RNA molecules comprises or is composed of an RNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide, the polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any polynucleotide sequence of any one of SEQ ID NO: 261-312, preferably SEQ ID NO: 261, SEQ ID NO: 281, SEQ ID NO: 300, SEQ ID NO: 299, SEQ ID NO: 308, or SEQ ID NO: 289.
[0887] 97. The composition according to any one of paragraphs 79-96, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any amino acid sequence in column 1 of Table 4, and the at least one RNA molecule is an RNA molecule comprising a guide sequence portion and a sequence encoded by a polynucleotide, the polynucleotide being any polynucleotide sequence in column 4 of Table 4, for example, any one of SEQ ID NO: 261-312, preferably SEQ ID NO: 261, SEQ ID NO: 281, SEQ ID NO: 300, SEQ ID NO: 299, SEQ ID NO: 308, or SEQ ID NO: Any one of 289 has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity.
[0888] 98. The composition according to any one of paragraphs 79-97, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 1, and the at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 209.
[0889] 98a. The composition according to any one of paragraphs 79-97, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 21, and the at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 229.
[0890] 98b. The composition according to any one of paragraphs 79-97, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 40, and the at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 248.
[0891] 98c. The composition according to any one of paragraphs 79-97, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 39, and the at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 247.
[0892] 98d. The composition according to any one of paragraphs 79-97, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 48, and the at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 256.
[0893] 98e. The composition according to any one of paragraphs 79-97, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 29, and the at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 237.
[0894] 99. The composition according to any one of paragraphs 79-98e, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any amino acid sequence in column 1 of Table 4, and the at least one RNA molecule comprises a crRNA molecule containing a guide sequence portion and a sequence encoded by a polynucleotide, the polynucleotide being any one of the polynucleotide sequences in column 2 of Table 4, for example, any one of SEQ ID NO: 209-260, preferably SEQ ID NO: 209, SEQ ID NO: 229, SEQ ID NO: 248, SEQ ID NO: 247, SEQ ID NO: 256 or SEQ ID NO: Any one of 237 has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity.
[0895] 100. The composition according to any one of paragraphs 79-99, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 1, and the at least one RNA molecule comprises a tracrRNA molecule containing a sequence encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the polynucleotide sequence of SEQ ID NO: 157.
[0896] 100a. The composition according to any one of paragraphs 79-99, wherein the Cas nuclease or fusion polypeptide comprises a sequence having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at le...
Claims
1. A Cas nuclease selected from the group consisting of: (a) A polypeptide having at least 70% sequence identity with any amino acid sequence of SEQ ID NO: 21, 48, 1, 40, 39, 29, 2-20, 22-28, 30-38, 41-47 or 49-52; (b) A polypeptide encoded by a polynucleotide having at least 70% sequence identity with any of the polypeptide coding sequences of SEQ ID NO: 73, 100, 53, 92, 91 or 81, or with any of SEQ ID NO: 53-104, 347, 349, 351, 353, 405, 416, 417, 434, 449, 465, 466, 512-520, 528, 549 or 550; (c) A polypeptide derived from any one of SEQ ID NO: 21, 48, 1, 40, 39, 29, 2-20, 22-28, 30-38, 41-47 or 49-52 by having 1-30 alterations (e.g., substitution, deletion and / or insertion at one or more positions, e.g., 1 or 2 or 3 or 4 or 4 or 5 or 30 alterations), particularly substitution, e.g., conserved amino acid substitutions; (d) A polypeptide having a TM-score of at least 0.80 compared to the three-dimensional structure of a polypeptide of any one of SEQ ID No: 21, 48, 1, 40, 39, 29, 2-20, 22-28, 30-38, 41-47 or 49-52, wherein the three-dimensional structure is calculated using Alphafold. (e) A polypeptide derived from (a), (b), (c), or (d) wherein the N-terminus and / or C-terminus have been extended by adding one or more amino acids; and Fragments of polypeptides in (f), (a), (b), (c), (d), or (e).
2. The nuclease according to claim 1, wherein the nuclease comprises one or more domains selected from the group consisting of: (a) The RuvC domain, which has at least 80% sequence identity with any one of the amino acid sequences of SEQ ID NO: 105-143 or 313-318; (b) An HNH domain that has at least 80% sequence identity with any one of the amino acid sequences of SEQ ID NO: 144-156 or 319-320; (c) The RuvC domain, which is derived from any one of SEQ ID NO: 105-143 or 313-318 by substitution, deletion or addition of one or more amino acids of SEQ ID NO: 105-143 or 313-318; (d) An HNH domain derived from any one of SEQ ID NO: 144-156 or 319-320 by substitution, deletion, or addition of one or more amino acids of SEQ ID NO: 144-156; and Fragments of the catalytic domains of (e), (a), (b), (c), or (d).
3. The nuclease according to any one of claims 1-2, wherein the nuclease is a nicking enzyme having one or more inactive RuvC domains generated by substitution, insertion or deletion of amino acids at the positions provided for the nuclease in column 3 of Table 2.
4. The nuclease according to any one of claims 1-3, wherein the nuclease is a nicking enzyme having one or more inactive HNH domains generated by substitution, insertion or deletion of amino acids at the positions provided for the nuclease in column 3 of Table 3.
5. The nuclease according to any one of claims 1-4, wherein the nuclease is a nicking enzyme and has single-strand breakage activity against DNA target sites.
6. The nuclease according to any one of claims 1-2, wherein the nuclease is a catalytically inactivated nuclease.
7. The nuclease according to any one of claims 1-6, wherein the nuclease is a type II Cas nuclease.
8. The nuclease according to any one of claims 1-7, wherein the nuclease utilizes the prototype spacer adjacent motif (PAM) sequence provided for the nuclease in Table 1.
9. A non-naturally occurring composition comprising (i) a Cas nuclease according to any one of claims 1-8, or (ii) a nucleic acid molecule comprising a sequence encoding a Cas nuclease according to any one of claims 1-8.
10. The composition of claim 9, further comprising one or more RNA molecules, or DNA polynucleotides encoding one or more of the one or more RNA molecules, wherein the one or more RNA molecules and the Cas nuclease do not naturally coexist, and the one or more RNA molecules are configured to form a complex with the Cas nuclease and / or target the complex to a target site.
11. The composition according to any one of claims 9-10, wherein the one or more RNA molecules comprise guide RNA (gRNA), the gRNA comprising CRISPR RNA (crRNA) and trans-activating RNA (tracrRNA).
12. A method for modifying a nucleotide sequence at a DNA target site in the genome of a cell, the method comprising introducing a Cas nuclease according to any one of claims 1-8, a polynucleotide encoding a Cas nuclease according to any one of claims 1-8, and / or a composition according to any one of claims 9-11 into the cell.
13. A polynucleotide encoding a Cas nuclease according to any one of claims 1-8.
14. A nucleic acid construct or expression vector comprising the polynucleotide of claim 13, the polynucleotide being operatively linked to one or more control sequences that direct the production of the Cas nuclease in cells.
15. A cell comprising a Cas nuclease according to any one of claims 1-8, a polynucleotide according to claim 13, a nucleic acid construct or expression vector according to claim 14, or a composition according to any one of claims 9-11.