CRISPR nuclease polypeptides and gene editing systems containing them

Engineered CRISPR nuclease polypeptides with mutations improve gene editing efficiency and precision by enhancing enzymatic activity and binding to guide RNAs, addressing inefficiencies in existing CRISPR-Cas systems.

JP2026516573APending Publication Date: 2026-05-26ARBOR BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ARBOR BIOTECHNOLOGIES INC
Filing Date
2024-03-29
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems for gene editing lack efficiency and precision, particularly in terms of nuclease activity and binding to guide RNAs.

Method used

Development of CRISPR nuclease polypeptides derived from reference nucleases A, K, and M, with engineered mutations such as arginine and lysine substitutions, nickase mutations, and N-terminal cleavages, enhancing enzymatic activity and binding to guide RNAs.

Benefits of technology

The engineered CRISPR nuclease polypeptides exhibit improved gene editing efficiency and precision, with enhanced nickase activity and higher binding to guide RNAs, leading to better target site modification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516573000001_ABST
    Figure 2026516573000001_ABST
Patent Text Reader

Abstract

A nuclease polypeptide derived from a reference nuclease, for example, a CRISPR nuclease polypeptide, wherein the reference nuclease may be nuclease A, nuclease K, or nuclease M, and the nuclease polypeptide comprises a RuvC nuclease domain and an HNH nuclease domain. Also provided herein are a gene editing system comprising such a nuclease polypeptide and a gene editing method using the gene editing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of the filing dates of U.S. Patent Provisional Application No. 63 / 493,360, filed Mar. 31, 2023; U.S. Patent Provisional Application No. 63 / 516,246, filed Jul. 28, 2023; U.S. Patent Provisional Application No. 63 / 566,661, filed Mar. 18, 2024; U.S. Patent Provisional Application No. 63 / 493,355, filed Mar. 31, 2023; U.S. Patent Provisional Application No. 63 / 562,132, filed Mar. 6, 2024; and U.S. Patent Provisional Application No. 63 / 493,363, filed Mar. 31, 2023. Each of the priority applications is hereby incorporated by reference in its entirety.

[0002] Sequence Listing This application includes a sequence listing that is electronically filed in XML format, which is hereby incorporated by reference in its entirety. The XML copy created on Mar. 27, 2024 is named 063586 - 514001WO_SeqList_ST26.xml and is 137,220 bytes in size.

Background Art

[0003] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR - associated (Cas) genes are collectively known as the CRISPR - Cas or CRISPR / Cas system, which is an adaptive immune system in archaea and bacteria that defends certain species against foreign genetic elements.

[0004] The CRISPR - Cas system typically includes a CRISPR nuclease and one or more RNA components that direct the CRISPR nuclease to a target genomic site for gene editing. The goal is to develop efficient CRISPR nucleases to improve gene - editing efficiency.

Summary of the Invention

[0005] This disclosure provides CRISPR nucleases exhibiting advantageous enzymatic activity (e.g., nickase activity, high indel activity, high binding activity to congeneral guide RNA scaffolds, and / or high DNA cleavage activity). Therefore, the CRISPR nucleases provided herein are expected to exhibit superior efficacy when used in gene editing.

[0006] Accordingly, CRISPR nuclease polypeptides derived from a reference CRISPR nuclease are provided herein, wherein the reference CRISPR nuclease comprises nuclease A (SEQ ID NO: 1), nuclease K (SEQ ID NO: 65), or nuclease M (SEQ ID NO: 80). Such CRISPR nucleases comprise a RuvC nuclease domain and an HNH nuclease domain. In some cases, the CRISPR nuclease polypeptides provided herein share at least 85% (e.g., at least 90%) the same amino acid sequence as the reference CRISPR nucleases provided herein. In some cases, the CRISPR nuclease polypeptides disclosed herein comprise a variant CRISPR nuclease (also known as an engineered CRISPR nuclease), wherein the variant CRISPR nuclease comprises at least one mutation relative to the reference CRISPR nuclease.

[0007] In some embodiments, the disclosure provides a gene editing system comprising a CRISPR nuclease polypeptide derived from reference nuclease A, a guide RNA targeting a genomic site of interest, and the use of the gene editing system for modifying a genomic site of interest in a host cell.

[0008] In some embodiments, the CRISPR nuclease polypeptide is an engineered variant of reference nuclease A (SEQ ID NO: 1), comprising at least one mutation relative to nuclease A. In some cases, the at least one mutation may comprise (a) one or more arginine and / or lysine substitutions (e.g., arginine substitutions) relative to nuclease A, (b) one or more nickase mutations in the HNH nuclease domain or RuvC nuclease domain of nuclease A, (c) an N-terminal cleavage relative to nuclease A, or any combination of (a), (b), and (c).

[0009] Similar to nuclease A, CRISPR nuclease polypeptides derived from nuclease A may contain a bridge helix (BH) domain, a phosphate-locked loop (PLL) domain, a wedge-type (WED) domain, and a PAM interaction (PID) domain. In some cases, variants of nuclease A provided herein may contain one or more arginine and / or lysine substitutions (e.g., arginine substitutions). In some cases, one or more arginine and / or lysine substitutions (e.g., arginine substitutions) are located in the BH domain, PLL domain, WED domain, PID domain, or a combination thereof. In some cases, one or more arginine and / or lysine substitutions (e.g., arginine substitutions) may be located in one or more of the positions D56, E59, G60, E63, I67, G71, S564, D568, W593, T605, E655, E694, and E706 in Sequence ID No. 1. In some specific examples, the manipulated variant of nuclease A contains arginine and / or lysine substitutions (e.g., arginine substitutions) at the following positions relative to SEQ ID NO: a) I67, D568, and E706; b) I67, D56, and D568; c) I67, T605, and D568; d) I67, D56, and E706; e) I67, T605, and E706; f) I67, D568, and E694; g) I67, D56, E694, and E706; h) I67, D56, T605, and E706; or i) I67, D568, W593, and E706. In one specific example, the manipulated variant of nuclease A includes arginine and / or lysine substitutions at I67, D568, and E706.

[0010] In some examples, the manipulated variant nuclease polypeptide of nuclease A contains (a) arginine substitutions of I67R, D568R, and E706R. In some examples, the manipulated CRISPR nuclease polypeptide contains (b) arginine substitutions of I67R, D56R, and D568R. In some examples, the manipulated CRISPR nuclease polypeptide contains (c) arginine substitutions of I67R, T605R, and D568R. In some examples, the manipulated CRISPR nuclease polypeptide contains (d) arginine substitutions of I67R, D56R, and E706R. In some examples, the manipulated CRISPR nuclease polypeptide contains (e) arginine substitutions of I67R, T605R, and E706R. In some examples, the manipulated CRISPR nuclease polypeptide contains (f) arginine substitutions of I67R, D568R, and E694R. In some examples, the manipulated CRISPR nuclease polypeptide contains (g) arginine substitutions of I67R, D56R, E694R, and E706R. In some examples, the manipulated CRISPR nuclease polypeptide contains (h) arginine substitutions of I67R, D56R, T605R, and E706R. In some examples, the manipulated CRISPR nuclease polypeptide contains (i) arginine substitutions of I67R, D568R, W593R, and E706R. In some specific examples, the manipulated CRISPR nuclease polypeptide contains arginine substitutions of I67R, D568R, and E706R.

[0011] Any of the manipulated variants of nuclease A disclosed herein may contain up to 20 arginine and / or lysine substitutions (e.g., arginine substitutions), for example, up to 15 arginine and / or lysine substitutions (e.g., arginine substitutions).

[0012] Alternatively, or in addition, manipulated variants of nuclease A may have an N-terminal cleavage relative to nuclease A (SEQ ID NO: 1). In some examples, the N-terminal cleavage is a deletion within residues 1-15 of SEQ ID NO: 1. In some specific examples, the N-terminal cleavage is a deletion within residues 1-14. In another specific example, the N-terminal cleavage is a deletion within residues 1-15 of SEQ ID NO: 1.

[0013] In addition, the CRISPR nuclease polypeptides disclosed herein may be nickase variants of nuclease A that exhibit nickase activity. Such nickase variants may include one or more mutations at positions H397, D396, N420, D24, E337, and / or D524 of SEQ ID NO: 1. In some examples, the nickase variant may include a mutation at position H397. In some specific examples, the mutation is an amino acid substitution of H397A or H397L.

[0014] In some embodiments, the manipulated variants of nuclease A disclosed herein may include any of the mutations disclosed herein, for example, a combination of one or more arginine / lysine substitutions, one or more nickase mutations, and / or N-terminal cleavages. For example, a variant may include (a) one or more nickase mutations in the HNH nuclease domain (e.g., at positions D396, H397, and / or N420 relative to SEQ ID NO: 1) and one or more arginine and / or lysine substitutions. In a specific example, the nickase mutation is at position H397 in SEQ ID NO: 1, and the arginine / lysine substitutions are at positions I67, D568, and E706.

[0015] In another example, an engineered variant of nuclease A may include (a) one or more nickase mutations in the HNH nuclease domain (e.g., at positions D396, H397, and / or N420 relative to SEQ ID NO: 1), (b) one or more arginine and / or lysine substitutions (e.g., at positions I67, D568, and E706), and (c) an N-terminal cleavage within residues 1-15 of SEQ ID NO: 1 (e.g., deletion of residues 1-14 or 1-15 of SEQ ID NO: 1). Alternatively or in addition, an engineered variant of nuclease A may include a C-terminal cleavage relative to nuclease A (e.g., within the C-terminal residues 1-5 of SEQ ID NO: 1).

[0016] In another example, variants of nuclease A include (a) a nickase mutation in H397A, (b) arginine substitutions in I67R, D568R, and E706R, and (c) an N-terminal cleavage involving deletions of residues 1-14 or 1-15 of SEQ ID NO: 1.

[0017] Any of the manipulated variants of nuclease A disclosed herein may contain an amino acid sequence that is at least 95% identical to SEQ ID NO: 1. In some examples, the manipulated variant may contain an amino acid sequence that is at least 98% identical to SEQ ID NO: 1.

[0018] Exemplary CRISPR nuclease polypeptides derived from nuclease A provided herein are listed in Table 1, Table 11, or Table 14, each of which is within the scope of this disclosure.

[0019] Any of the CRISPR nuclease polypeptides derived from nuclease A disclosed herein may be fusion polypeptides, which may further comprise one or more functional fragments. In some embodiments, one or more functional fragments may comprise one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof. One or more NLS may be located at the N-terminus, C-terminus, or both.

[0020] Furthermore, nucleic acids comprising a nucleotide sequence encoding one of the CRISPR nuclease polypeptides derived from nuclease A disclosed herein are also provided herein. In some cases, the nucleic acid is an expression vector, and the nucleotide sequence encoding the CRISPR nuclease polypeptide is operably bound to a promoter. Moreover, this disclosure provides a host cell comprising a nucleic acid encoding the CRISPR nuclease polypeptide disclosed herein.

[0021] In some embodiments, the Disclosure features a gene editing system comprising (a) a CRISPR nuclease polypeptide derived from any of the nuclease A disclosed herein or a first nucleic acid encoding a CRISPR nuclease polypeptide, and (b) a guide RNA (gRNA) or a second nucleic acid encoding a gRNA. The gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease polypeptide and a spacer sequence specific to a target sequence at a genomic site of interest. The target sequence is adjacent to a protospacer-adjacent motif (PAM). In some embodiments, the target sequence is upstream (5'-) of the 5'-NGG-3' protospacer-adjacent motif (PAM), where N represents any nucleotide.

[0022] In some cases, the scaffold sequences that can be utilized by the CRISPR nuclease polypeptides derived from nuclease A disclosed herein may contain a nucleotide sequence that is at least 85% identical to SEQ ID NO: 2. In some cases, the scaffold sequences may contain one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to SEQ ID NO: 2.

[0023] In some cases, the scaffold sequence is a cleavage variant of SEQ ID NO: 2. Such a cleavage variant may be approximately 110–140 nt in length. In some cases, the cleavage variant may have a 3' cleavage relative to SEQ ID NO: 2. For example, the 3' cleavage may include a deletion within residues 143–202 of SEQ ID NO: 2. In other cases, the cleavage variant of SEQ ID NO: 2 may have an internal cleavage relative to SEQ ID NO: 2. For example, the variant may include one or more deletions within residues 10–40 of SEQ ID NO: 2. In a specific example, the deletions may include residues 14–20 and / or 25–32 of SEQ ID NO: 2. In other cases, the cleavage variant of SEQ ID NO: 2 may have both a 3' cleavage (e.g., as disclosed herein) and an internal cleavage (e.g., as disclosed herein). Alternatively or in addition, the cleavage variant may further include one or more mutations relative to SEQ ID NO: 2, e.g., deletions and / or substitutions within residues 81–85 of SEQ ID NO: 2. In one specific example, the scaffold sequence contains (e.g., consists of) the nucleotide sequence of sequence number 27. In another specific example, the scaffold sequence contains (e.g., consists of) sequence number 28. Such scaffold sequences can be approximately 115–130 nt in length.

[0024] Furthermore, guide RNAs comprising a spacer sequence as defined herein and one of the variant scaffold sequences for sequence number 2 are also within the scope of this disclosure.

[0025] In some cases, the CRISPR nuclease polypeptide contained within the gene editing system disclosed herein is the reference CRISPR nuclease of SEQ ID NO: 1, and the scaffold sequence contains the nucleotide sequence of SEQ ID NO: 27. Alternatively, the CRISPR nuclease polypeptide in the gene editing system may be a variant of SEQ ID NO: 1, for example, the variant may contain mutations at positions I67, D568, and E706 (e.g., arginine substitutions I67R, D568R, and E706R), and the scaffold sequence may contain the nucleotide sequence of SEQ ID NO: 28.

[0026] In other aspects, provided herein is a CRISPR nuclease polypeptide derived from CRISPR nuclease K (SEQ ID NO: 65), comprising a RuvC nuclease domain and a HNH nuclease domain. The CRISPR nuclease polypeptide derived from nuclease K can comprise an amino acid sequence that is at least 90% identical to SEQ ID NO: 65. In addition to the RuvC and HNH nuclease domains, the CRISPR nuclease polypeptide can also comprise a bridge helix (BH) domain, a wedge (WED) domain, and a PAM interaction (PID) domain.

[0027] In some embodiments, the CRISPR nuclease polypeptide can be a variant of nuclease K that comprises at least one mutation relative to SEQ ID NO: 65. Such variant CRISPR nuclease polypeptides can have enhanced enzymatic activity as compared to the reference CRISPR nuclease K. In some cases, the at least one mutation can comprise (a) one or more arginine and / or lysine substitutions, optionally one or more arginine substitutions, relative to SEQ ID NO: 65, (b) one or more nickase mutations in the HNH nuclease domain or the RuvC nuclease domain of SEQ ID NO: 65, or (c) a combination of (a) and (b).

[0028] In some examples, variants of nuclease K may include one or more arginine and / or lysine substitutions (e.g., arginine substitutions). Any of the variant CRISPR nuclease polypeptides disclosed herein may contain up to 20 arginine and / or lysine substitutions (e.g., up to 20 arginine substitutions), e.g., up to 15 arginine and / or lysine substitutions (e.g., up to 15 arginine substitutions). In some cases, one or more arginine and / or lysine substitutions are located in the BH domain, the WED domain, the PID domain, or a combination thereof. In a specific example, one or more arginine and / or lysine substitutions (e.g., arginine substitutions) are located in one or more of positions G42, D46, I53, F83, E128, E541, F570, E576, D582, C630, I631, E683, S719, and T734 in SEQ ID NO:65.

[0029] In some cases, variants of nuclease K may include a combination of arginine and / or lysine substitutions. For example, a variant may include arginine and / or lysine substitutions at positions G42 and D582 of SEQ ID NO:65 (e.g., G42R and D582R). In another example, a variant may include arginine and / or lysine substitutions at positions I53, D582, and T734 of SEQ ID NO:65 (e.g., I53R, D582R, and T734R).

[0030] In some embodiments, the engineered CRISPR nuclease polypeptides derived from nuclease K provided herein may contain one or more mutations that result in nickase activity, e.g., one or more mutations at positions D374, H375, and / or N398 in SEQ ID NO: 65. In some examples, the nickase variant may contain a mutation at position H375, and optionally, the mutation is an amino acid substitution of H375A. In specific examples, the engineered CRISPR nuclease polypeptide may contain arginine and / or lysine substitutions (e.g., arginine substitutions) at positions I53, D582, and T734 in SEQ ID NO: 65, and may further contain an amino acid substitution of H375A.

[0031] In some embodiments, the engineered CRISPR nuclease polypeptide derived from nuclease K disclosed herein may have an amino acid sequence identical to at least 90% of SEQ ID NO: 65. In some examples, the engineered CRISPR nuclease derived from nuclease K may have an amino acid sequence identical to at least 95% of SEQ ID NO: 65. In specific examples, the engineered CRISPR nuclease derived from nuclease K may have an amino acid sequence identical to at least 98% of SEQ ID NO: 65.

[0032] In some embodiments, the manipulated CRISPR nuclease polypeptide derived from nuclease K disclosed herein may be a fusion polypeptide comprising a CRISPR nuclease and one or more additional fragments, the one or more additional fragments being heterogeneous to the CRISPR nuclease. In some embodiments, the one or more functional fragments may comprise one or more NLSs, one or more peptide linkers, or a combination thereof. The one or more NLSs may be located at the N-terminus, C-terminus, or both.

[0033] Furthermore, nucleic acids are provided herein that comprise a nucleotide sequence encoding one of the engineered CRISPR nuclease polypeptides derived from nuclease K disclosed herein. In some cases, the nucleic acid is an expression vector, and the nucleotide sequence encoding the engineered CRISPR nuclease polypeptide is operably bound to a promoter. Moreover, this disclosure provides a host cell comprising a nucleic acid encoding one of the engineered CRISPR nuclease polypeptides disclosed herein.

[0034] In another embodiment, the Disclosure features a gene editing system comprising (a) a first nucleic acid encoding any of the engineered CRISPR nuclease polypeptides derived from nuclease K disclosed herein or a CRISPR nuclease polypeptide, and (b) a guide RNA (gRNA) or a second nucleic acid encoding a gRNA, wherein the gRNA comprises a scaffold sequence recognizable by the engineered CRISPR nuclease polypeptide and a spacer sequence specific to a target sequence at a genomic site of interest. The target sequence is adjacent to a protospacer-adjacent motif (PAM). In some embodiments, the target sequence is upstream (5'-) of the 5'-NGG-3' protospacer-adjacent motif (PAM), where N represents any nucleotide.

[0035] In some embodiments, a scaffold that can be utilized by a CRISPR nuclease polypeptide derived from nuclease A disclosed herein may contain a nucleotide sequence that is at least 85% identical to SEQ ID NO: 79. Alternatively, or in addition, the scaffold may be a fragment of SEQ ID NO: 79. In some cases, the scaffold may contain one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to SEQ ID NO: 79.

[0036] In another embodiment, the disclosure provides a nuclease polypeptide derived from reference nuclease M (SEQ ID NO: 80), comprising a RuvC nuclease domain and an HNH nuclease domain. Such a nuclease polypeptide may have an amino acid sequence at least 90% identical to SEQ ID NO: 80 (referring to reference nuclease M disclosed herein). In some embodiments, a nuclease polypeptide derived from nuclease M may be a variant of nuclease M comprising at least one mutation relative to nuclease M. Such a variant of nuclease M may have enhanced enzymatic activity compared to nuclease M.

[0037] In some cases, at least one mutation in a variant of nuclease M may include (a) one or more arginine and / or lysine substitutions (e.g., arginine substitutions) to nuclease M (SEQ ID NO: 80), (b) one or more nickase mutations in the HNH nuclease domain or RuvC nuclease domain of SEQ ID NO: 80, or (c) a combination of (a) and (b).

[0038] In some cases, a variant of nuclease M may contain one or more arginine and / or lysine substitutions (e.g., arginine substitutions), and one or more arginine and / or lysine substitutions may be located in one or more of the following: bridge helix (BH) domain, nucleic acid recognition (REC) domain, phosphate-locked loop (PLL) domain, wedge-type (WED) domain, PAM interaction (PID) domain, nuclease domain, or a combination thereof. In some cases, a variant nuclease polypeptide of nuclease M disclosed herein may contain up to 20 arginine and / or lysine substitutions (e.g., arginine substitutions) relative to reference nuclease M. For example, a variant nuclease polypeptide may contain up to 15 arginine and / or lysine substitutions (e.g., up to 12 arginine / lysine substitutions), e.g., up to 15 arginine substitutions or up to 12 arginine substitutions relative to reference nuclease M.

[0039] In a specific example, one or more arginine and / or lysine substitutions (e.g., arginine substitutions) may be located at positions E88, S95, L92, E401, E83, N371, P481, and / or A373 in SEQ ID NO: 80.

[0040] Alternatively, or in addition, the variant nuclease polypeptide of nuclease M may contain one or more mutations that result in nickase activity. In some embodiments, one or more nickase mutations are located at one or more of the positions D58, E189, D341, H243, H244, H267, R329, and / or H338 in SEQ ID NO: 80.

[0041] Any of the variants of nuclease M provided herein may have an amino acid sequence that is at least 90% identical to SEQ ID NO: 80. In some examples, a nuclease polypeptide derived from nuclease M may have an amino acid sequence that is at least 95% identical to SEQ ID NO: 80. In other examples, a nuclease polypeptide derived from nuclease M may have an amino acid sequence that is at least 98% identical to SEQ ID NO: 80.

[0042] In some embodiments, the nuclease polypeptide derived from nuclease M may be a fusion polypeptide, which may further comprise one or more functional fragments. In some cases, one or more functional fragments may be heterogeneous to the nuclease moiety in the fusion polypeptide. In some cases, one or more functional fragments may comprise one or more NLSs, one or more peptide linkers, or a combination thereof, which may be located at the N-terminus, C-terminus, or both.

[0043] Furthermore, nucleic acids are provided herein, comprising a nucleotide sequence encoding one of the nuclease polypeptides derived from nuclease M disclosed herein. In some cases, the nucleic acid is an expression vector, and the nucleotide sequence encoding the nuclease polypeptide is operably bound to a promoter. Alternatively, the nucleic acid may be a messenger RNA (mRNA) molecule. Moreover, this disclosure provides a host cell comprising a nucleic acid encoding a nuclease polypeptide derived from nuclease M disclosed herein.

[0044] In addition, the Disclosure features a gene editing system comprising (a) one of the nuclease polypeptides derived from nuclease M disclosed herein or a first nucleic acid encoding one thereof, and (b) a guide RNA (gRNA) or a second nucleic acid encoding a gRNA, wherein the gRNA comprises a scaffold sequence recognizable by the nuclease polypeptide derived from nuclease M and a spacer specific to a target sequence in a genomic region, the target sequence being adjacent to a protospacer-adjacent motif (PAM), and the gRNA or the second nucleic acid encoding a gRNA. In some embodiments, the PAM is 5'-WTAAH-3', where W is A or T and H is A, C, or T. In one example, the PAM is 5'-TTAAA-3'.

[0045] In some embodiments, a scaffold sequence that can be utilized by a nuclease polypeptide derived from nuclease M disclosed herein may contain a nucleotide sequence that is at least 85% identical to SEQ ID NO: 94. Alternatively, or in addition, the scaffold sequence may be a fragment of SEQ ID NO: 94. In some cases, the scaffold sequence may contain one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to SEQ ID NO: 94.

[0046] Any of the gene editing systems disclosed herein may further comprise one or more lipid excipients associated with elements (a) and / or (b) of the gene editing system. In some examples, one or more lipid excipients form lipid nanoparticles, which are associated with or encapsulate elements (a) and / or (b) of the gene editing system.

[0047] Alternatively, any of the gene editing systems disclosed herein may include a viral vector, such as an adeno-associated virus (AAV) vector, the viral vector containing coding sequences for both the CRISPR nuclease polypeptide and gRNA within the gene editing system.

[0048] Furthermore, gene editing methods are provided herein, comprising delivering one of the gene editing systems disclosed herein to a host cell to edit a genomic site targeted by the gRNA of the gene editing system. In some cases, the host cell is cultured in vitro. In other cases, the host cell is located at a target requiring gene editing.

[0049] Details of one or more embodiments of the present invention are shown in the following description. Other features or advantages of the present invention will become apparent from the following drawings and detailed descriptions of some embodiments, and from the appended claims.

[0050] The following drawings form part of this specification and are included to further demonstrate certain aspects of the present disclosure, which can be better understood by referring to the drawings together with a detailed description of the specific embodiments presented herein. [Brief explanation of the drawing]

[0051] [Figure 1] This figure shows the percentage of NGS reads containing indels at six loci, shown in the presence or absence of CRISPR nuclease for Sequence ID No. 1 (nuclease A). [Figure 2A] Gel images showing in vitro cleavage of the target or non-target strand of the target DNA substrate, and bar graphs showing quantification of nuclease activity by nuclease A and its variants. Gel images captured using an 800 nm channel showing in vitro cleavage of the target strand of the target DNA substrate (labeled with IR800 dye at the 5' end) by a reference CRISPR nuclease, putative HNH knockout nickasase, or putative RuvC knockout nickasase. [Figure 2B]Gel images showing in vitro cleavage of the target or non-target strand of the target DNA substrate, and bar graphs showing quantification of nuclease activity by nuclease A and its variants. Gel images captured using a 700 nm channel showing in vitro cleavage of the non-target strand of the target DNA substrate (labeled with IR700 dye at the 5' end) by a reference CRISPR nuclease, putative HNH knockout nickasase, or putative RuvC knockout nickasase. [Figure 2C] These are gel images showing in vitro cleavage of the target or non-target strand of the target DNA substrate, and bar graphs showing quantification of nuclease activity by nuclease A and its variants. Figures 2A and 2B are overlay images captured using the 800nm ​​and 700nm channels, respectively. [Figure 2D] Gel images showing in vitro cleavage of the target or non-target strand of the target DNA substrate, and bar graphs showing quantification of nuclease activity by nuclease A and its variants. Bar graphs showing quantification of the percentage of cleaved target and non-target DNA produced by the tested reference CRISPR nuclease, putative HNH knockout nickasase, and putative RuvC knockout nickasase. [Figure 3A] This is a schematic diagram showing the predicted secondary structure of the scaffolding arrangement. Reference scaffolding (Sequence ID 2). [Figure 3B] This is a schematic diagram showing the predicted secondary structure of the scaffolding arrangement. It depicts a sectional reference scaffolding (sequence number 27). [Figure 3C] This is a schematic diagram showing the predicted secondary structure of the scaffolding arrangement. Scaffolding 2 (Sequence ID 28). [Figure 4] This figure shows the percentage of NGS reads containing indels at six loci, with or without the CRISPR nuclease for Sequence ID No. 65 (nuclease K). [Figure 5] This figure shows the percentage of NGS reads containing indels at six gene loci, shown in the presence or absence of nuclease SEQ ID NO: 80 (nuclease M). [Modes for carrying out the invention]

[0052] Nuclease polypeptides derived from reference nuclease A (SEQ ID NO: 1), nuclease K (SEQ ID NO: 65), or nuclease M (SEQ ID NO: 80), such as CRISPR nuclease polypeptides, are provided herein. Such CRISPR nuclease polypeptides may include the reference nucleases disclosed herein or variants thereof. In some cases, the CRISPR nuclease polypeptides provided herein may be variant nucleases of a reference nuclease, comprising at least one mutation relative to the reference nuclease.

[0053] In some embodiments, a variant CRISPR nuclease polypeptide may contain one or more mutations (e.g., arginine substitutions, lysine substitutions, or combinations thereof) relative to a reference CRISPR nuclease. Alternatively or in addition, a variant CRISPR nuclease polypeptide may contain one or more mutations in either the RuvC nuclease domain or the HNH nuclease domain. Such mutations (i.e., nickase mutations) may reduce or eliminate the nuclease activity of either the RuvC or HNH nuclease domain, resulting in a variant exhibiting nickase activity. As used herein, the term “nickase” refers to an enzyme that cleaves one strand of double-stranded DNA at a specific recognition nucleotide sequence (e.g., a target sequence disclosed herein). A nickase may interact with one strand of a DNA double helix to produce a DNA molecule that is cleaved (also known as nicked) on one strand. In some embodiments, a nickase is a variant of a CRISPR nuclease containing an inactivated HNH domain. In some embodiments, nickase is a variant of a CRISPR nuclease containing an inactivated RuvC domain.

[0054] Any of the variant CRISPR nuclease polypeptides provided herein may share high sequence homology (e.g., at least 85% sequence identity, e.g., at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or higher sequence identity) with reference CRISPR nuclease A (SEQ ID NO: 1), nuclease K (SEQ ID NO: 65), or nuclease M (SEQ ID NO: 80).

[0055] The variant CRISPR nuclease polypeptides provided herein are expected to possess advantageous characteristics compared to reference CRISPR nucleases, such as increased binding to congeneral guide RNAs and higher nuclease activity. Therefore, the variant CRISPR nuclease polypeptides disclosed herein are expected to exhibit better activity in gene editing, such as nickase activity, and higher efficiency and precision in gene editing, including strand substitution, compared to reference CRISPR nucleases.

[0056] Alternatively or in addition, a variant CRISPR nuclease polypeptide may be a fusion polypeptide comprising a CRISPR nuclease (e.g., reference nuclease A, nuclease K, or nuclease M, or a variant of the reference nuclease provided herein) (nuclease portion) and one or more additional functional fragments, e.g., those described herein (e.g., NLS and / or peptide linkers). In addition to the above-mentioned advantageous features, the fusion polypeptide may have additional functionality that can be attributed to its fusion partner.

[0057] Accordingly, this disclosure provides CRISPR nuclease polypeptides derived from reference CRISPR nucleases (nuclease A, nuclease K, or nuclease M), gene editing systems containing them, and gene editing methods using them.

[0058] I. CRISPR nuclease polypeptide As used herein, the term “CRISPR nuclease” refers to an RNA-guided effector capable of binding to nucleic acids and introducing single-strand or double-strand breaks. CRISPR nucleases typically comprise multiple functional domains, such as nuclease domains (e.g., RuvC and / or HNH), PLMP domains, bridge helix (BH) domains, nucleic acid recognition (REC) domains, phosphate-locked loop (PLL) domains, wedge-type domains (WED), PAM interaction domains (PID), or combinations thereof. As used herein, the term “domain” refers to a characteristic functional and / or structural unit of a polypeptide. In some cases, functional domains may be linear. In other cases, functional domains may be discontinuous and conformal. In some embodiments, domains may comprise conserved amino acid sequences.

[0059] As used herein, the term “variant CRISPR nuclease polypeptide” refers to a CRISPR nuclease polypeptide that, compared to a reference CRISPR nuclease (nuclease A, nuclease K, or nuclease M), contains alterations, e.g., substitutions, insertions, deletions, and / or fusions at one or more residue positions.

[0060] The variant CRISPR nuclease polypeptides provided herein are expected to exhibit one or more regulated activities (e.g., enhanced or reduced) relative to a reference CRISPR nuclease. As used herein, the term “activity” refers to biological activity. In some embodiments, activity includes enzymatic activity, e.g., catalytic ability of an effector. For example, activity may include nuclease activity. In some embodiments, activity includes binding activity, e.g., binding of an effector (e.g., a CRISPR nuclease) to an RNA guide and / or target nucleic acid. In some examples, the variant CRISPR nuclease polypeptides disclosed herein have enhanced binding to a congeneral guide RNA (gRNA) compared to a reference CRISPR nuclease, for example, having binding activity that is at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2x, 2x, 5x, 10x, or more greater than that of a reference CRISPR nuclease. Congeneral gRNAs refer to gRNAs that have a scaffold recognizable by CRISPR nucleases.

[0061] In some cases, the variant CRISPR nuclease polypeptides disclosed herein have enhanced enzymatic activity relative to a reference CRISPR nuclease, for example, having at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2x, 2x, 5x, 10x, or more greater enzymatic activity than that of the reference CRISPR nuclease. In other cases, the variant CRISPR nuclease polypeptides disclosed herein have reduced enzymatic activity relative to a reference CRISPR nuclease, for example, having at least 20%, 30%, 40%, 50%, 60%, or 70% lower enzymatic activity than that of the reference CRISPR nuclease. In some cases, the reduced enzymatic activity is achieved by reducing or decreasing the nuclease activity of the RuvC domain. In some cases, the reduced enzymatic activity is achieved by reducing or decreasing the nuclease activity of the HNH domain.

[0062] In some cases, the variant CRISPR nuclease polypeptides disclosed herein have enhanced indel activity relative to the reference CRISPR nuclease. As used herein, the term “indel activity” refers to the ability of a CRISPR nuclease to introduce indels (insertions / deletions) into a sequence (e.g., a genomic target).

[0063] In some embodiments, the variant CRISPR nuclease polypeptides provided herein share high sequence homology with a reference CRISPR nuclease. For example, a variant CRISPR nuclease polypeptide may contain at least 70% (e.g., at least 80%, 85%, 90%, 95%, or more) identical amino acid sequences to SEQ ID NO: 1, SEQ ID NO: 65, or SEQ ID NO: 80. In some cases, a variant CRISPR nuclease polypeptide may contain at least 90% identical amino acid sequences to SEQ ID NO: 1, SEQ ID NO: 65, or SEQ ID NO: 80. In some cases, a variant CRISPR nuclease polypeptide may contain at least 95% identical amino acid sequences to SEQ ID NO: 1, SEQ ID NO: 65, or SEQ ID NO: 80. In other cases, a variant CRISPR nuclease polypeptide may contain at least 97% (e.g., 98%, 99%, 99.5%, or more) identical amino acid sequences to SEQ ID NO: 1, SEQ ID NO: 65, or SEQ ID NO: 80.

[0064] The "percent identity" (also known as sequence identity) of two nucleic acids or two amino acid sequences is determined using the algorithm of Karlin and Altschul Proc.Natl.Acad.Sci.USA 87:2264-68, 1990, as modified in Karlin and Altschul Proc.Natl.Acad.Sci.USA 90:5873-77, 1993. Such an algorithm is incorporated into the NBLAST and XBLAST programs (version 2.0) of Altschul, et al. J.Mol.Biol.215:403-10, 1990. BLAST nucleotide search can be performed using the NBLAST program, score = 100, word length = 12, to obtain nucleotide sequences homologous to the nucleic acid molecule of the present invention. BLAST protein search can be performed using the XBLAST program, score = 50, word length = 3, to obtain amino acid sequences homologous to the protein molecule of the present invention. If a gap exists between two sequences, gap BLAST may be used as described in Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997. When utilizing the BLAST and gap BLAST programs, the default parameters of each program (e.g., XBLAST and NBLAST) may be used.

[0065] In some cases, the variant CRISPR nuclease polypeptides disclosed herein may include one or more arginine and / or lysine substitutions relative to the reference nuclease A, nuclease K, or nuclease M provided herein. "Arginine substitution" and / or "lysine substitution" means replacing a non-arginine or non-lysine residue in SEQ ID NO: 1 (nuclease A), SEQ ID NO: 65 (nuclease K), or SEQ ID NO: 80 (nuclease M) with an arginine or lysine residue.

[0066] In some cases, one or more of the substituted arginine residues may be replaced by conserved amino acid residues, such as lysine or histidine. In some embodiments, the variant CRISPR nuclease polypeptides provided herein may comprise one or more arginine substitutions, one or more lysine substitutions, or a combination thereof.

[0067] In some cases, the variant CRISPR nuclease polypeptides provided herein may contain one or more conserved amino acid residue substitutions, either alone or in combination with other types of mutations disclosed herein.

[0068] As used herein, “conservative amino acid substitution” refers to an amino acid substitution that does not alter the relative charge or size properties of the protein being substituted. Variants can be prepared according to methods for modifying polypeptide sequences known to those skilled in the art. For example, such methods can be found in the references Molecular Cloning: A Laboratory Manual, J. Sambrook, et al., eds., Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989, or Current Protocols in Molecular Biology, FMAusubel, et al., eds., John Wiley & Sons, Inc., New York. Conservative amino acid substitutions include substitutions made with amino acids in the following groups: (a) M, I, L, V; (b) F, Y, W; (c) K, R, H; (d) A, G; (e) S, T; (f) Q, N; and (g) E, D.

[0069] (A) Nuclease A and its modified variants In some embodiments, the CRISPR nuclease polypeptide is derived from nuclease A, where nuclease A includes both wild-type nuclease A (SEQ ID NO: 1), its variants, e.g., those disclosed herein, or a fusion polypeptide comprising the same. The variant CRISPR nuclease polypeptide of nuclease A may be produced by introducing one or more mutations into a reference CRISPR nuclease to modulate (e.g., enhance or reduce) one or more activities of the nuclease.

[0070] The reference CRISPR nuclease of nuclease A (SEQ ID NO: 1) (see Table 1 below) is a CRISPR nuclease containing both a RuvC nuclease domain (located at residues 15-53, 308-338, and 472-566 of SEQ ID NO: 1) and an HNH domain (located at residues 339-471 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate DNA strand cleavage adjacent to the 5'-NGG-3'PAM motif, where N represents any nucleotide. Positions D24, E335, and D524 are considered active sites in the RuvC domain, while positions D396, H397, and N420 are considered active sites in the HNH domain. R506 and H521 may also be important for the nuclease activity of the RuvC domain. In addition to the nuclease domain, the reference CRISPR nuclease of SEQ ID NO: 1 also contains a BH domain (residues 54-89 of SEQ ID NO: 1), a REC domain (residues 90-307 of SEQ ID NO: 1), a PLL domain (residues 567-580 of SEQ ID NO: 1), a WED domain (residues 581-672 of SEQ ID NO: 1), and a PID domain (residues 673-776 of SEQ ID NO: 1).

[0071] Compared to the CRISPR-Cas9 nuclease from Streptococcus pyogenes, the reference CRISPR nuclease of SEQ ID NO: 1 disclosed herein is smaller. The scaffolds utilized by the reference CRISPR nuclease of SEQ ID NO: 1 and its variants can be miniaturized. This scaffold contains a characteristic structure compared to the SpCas9 nuclease scaffold. The characteristic scaffold allows for a reduced size of several domains (e.g., the REC domain), and is therefore expected to contribute to a smaller CRISPR nuclease size. These features are beneficial for delivery. Arginine and / or lysine substitutions (e.g., arginine substitutions) can be introduced into the CRISPR nuclease of SEQ ID NO: 1 to increase indel activity.

[0072] Furthermore, the RuvC domain of nuclease A (SEQ ID NO: 1) cleaves the non-target strand of the target nucleic acid within 4-8 nucleotides upstream of the PAM, while the HNH domain cleaves the target strand within 3-4 nucleotides upstream of the PAM, each resulting in a cleavage site with a 0-5 nucleotide overhang (most commonly, a 3-5 nucleotide overhang). In contrast, the RuvC domain of SpCas9 cleaves the non-target strand of the target nucleic acid within 3-5 nucleotides upstream of a congeneral PAM (5'-NGG-3', where N is any nucleotide in the sequence), while the HNH domain cleaves the target strand within 3-4 nucleotides upstream of the PAM, each resulting in a 0-3 nucleotide overhang (most commonly, a blunt-end cleavage).

[0073] Additionally, the use of gene editing systems containing nuclease A (SEQ ID NO: 1) or its variants can result in the introduction of larger indels into target nucleic acids than those that can be introduced by SpCas9. For example, insertions induced by the use of nuclease A or its variants can range from about 1 to about 7 nucleotides (most commonly about 4 nucleotides). Deletions induced by the use of nuclease A or its variants can range from about 1 to about 25 nucleotides or more (most commonly about 11 nucleotides). Furthermore, since the reference CRISPR nuclease of SEQ ID NO: 1 contains a RuvC domain and an HNH domain, nickasase variants can be manipulated, for example, by disrupting the nuclease activity of either the RuvC domain or the HNH domain.

[0074] The CRISPR nuclease polypeptide variants of nuclease A provided herein may contain one or more alterations to nuclease A (SEQ ID NO: 1), such as one or more amino acid residue substitutions, one or more deletions, one or more insertions, one or more fusions, or a combination thereof. In some cases, the alterations may be introduced within the BH domain, PLL domain, WED domain, PID domain, or a combination thereof. In some cases, the alterations are not introduced within the RuvC and / or HNH nuclease domain, or within the active site and / or site involved in activity in these domains provided herein. Alternatively, conservative amino acid substitutions may be introduced within SEQ ID NO: 1, including within the RuvC and / or HNH nuclease domain.

[0075] In some embodiments, the variant CRISPR nuclease polypeptide of nuclease A may contain one or more arginine substitutions, one or more lysine substitutions, or combinations thereof, relative to SEQ ID NO: 1. In some examples, the variant CRISPR nuclease polypeptide may contain up to 20 arginine and / or lysine substitutions (e.g., up to 20 arginine substitutions, up to 20 lysine substitutions, or combinations thereof), for example, up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine substitutions, lysine substitutions, or combinations thereof. In specific examples, the variant CRISPR nuclease polypeptide may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine substitutions, lysine substitutions, or combinations thereof. In some examples, the variant CRISPR nuclease polypeptides provided herein contain arginine substitutions.

[0076] In some cases, arginine and / or lysine substitutions may be located in the BH domain, PLL domain, WED domain, PID domain, or a combination thereof of nuclease A. In some cases, variant CRISPR nuclease polypeptides derived from nuclease A may contain one or more arginine and / or lysine substitutions (e.g., arginine substitutions) at one or more of the following positions in Sequence ID No. 1: D56, E59, G60, E63, I67, G71, S564, D568, T605, E655, E694, and E706. In some cases, variant CRISPR nuclease polypeptides derived from nuclease A contain one or more of the following arginine substitutions relative to SEQ ID NO: D56R, E59R, G60R, E63R, I67R, G71R, S564R, D568R, T605R, E655R, E694R, and E706R.

[0077] In some cases, arginine and / or lysine substitutions may be located within the RuvC and / or HNH nuclease domains of nuclease A. In some cases, arginine and / or lysine substitutions may be located within the RuvC nuclease domain of nuclease A and may reduce or inactivate the RuvC domain (e.g., at positions D24, E335, R506, H521, and / or D524 in the RuvC domain of SEQ ID NO: 1). In other cases, arginine and / or lysine substitutions may be located within the HNH domain of nuclease A, for example, at positions D396, H397, and / or N420 in the HNH domain of SEQ ID NO: 1. In other cases, arginine and / or lysine substitutions may be located within both the RuvC and HNH domains of nuclease A and may reduce or decrease nuclease enzyme activity.

[0078] Alternatively, arginine and / or lysine substitutions do not have to be at the active site and / or site involved in activity within the RuvC and / or HNH nuclease domain of nuclease A (for example, not at positions D24, E335, R506, H521, and / or D524 in the RuvC domain of SEQ ID NO: 1, and / or at positions D396, H397, and / or N420 in the HNH domain). In some examples, arginine and / or lysine substitutions do not have to be at the RuvC and / or HNH domain of nuclease A.

[0079] Arginine substitutions at the following positions in SEQ ID NO: 1 result in reduced or no nuclease activity, indicating that these positions cannot tolerate mutations with respect to nuclease activity, as reported herein. L21, G22, I23, D24, G26, G27, T30, G31, L32, A33, V34, V35, V42, V48, M50, L85, V107, Y108, C111, G115, M169, V182, F186, I270, C286, H289, S309, L310, V31 7, V321, M336, I343, V334, E337, S338, N339, F341, T347, Y358, T377, V382, Y383, C 384, T389, A393, D396, H397, I398, F399, P400, I406, N411, V413, A414, C415, C416, N420, K423, K450, L452, A455, I466, M469, S470, A472, S473, I474, G475, L483, G49 9, T502, W509, F511, H521, L523, D524, A525, V526, I527, L528, A529, P581, V595, T 596, I606, Y611, L615, L630, A635, F640, Y641, L651, G657, L658, G659, Q662, M663, V664, K672, T673, N674, V675, Y683, L700, V725, I736, L739, P743, L745, and L765.

[0080] In some embodiments, variants of CRISPR nuclease A provided herein may exhibit nuclease activity and may not have mutations at the above-mentioned positions. Alternatively, the variant CRISPR nuclease may be an inactive nuclease (for example, for use in base editing) having mutations at one or more of these positions.

[0081] In some cases, variant CRISPR nucleases include arginine and / or lysine substitutions in combinations of two or more (e.g., three or four) positions in sequence number 1, e.g., combinations of D56, I67, D568, T605, E694, and E706. Examples include, but are not limited to, (a) I67, D568, and E706, (b) D56, I67, and D568, (c) I67, D568, and T605, (d) D56, I67, and E706, (e) I67, T605, and E706, (f) I67, D568, and E694, (g) D56, I67, E694, and E706, (h) D56, I67, T605, and E706, and (i) I67, D568, W593, and E706. In some specific examples, the manipulated CRISPR nuclease polypeptides include arginine substitutions of I67, D568, and E706. In some examples, the variant CRISPR nuclease includes arginine substitutions at the combined positions listed herein.

[0082] Additional exemplary combinations of arginine and / or lysine substitutions (e.g., arginine substitutions) (for SEQ ID NO: 1) (i) at least two positions of D56, E59, G60, E63, I67, G71, S564, D568, T605, E655, E694, and E706, (ii) at least two positions of D56, I67, T605, E694, and E706, ( iii) at least two of the positions D56, I67, D568, T605, E694, and E706; (iv) at least two of the positions D56, E63, G71, D568, T605, E655, E694, and E706; and (v) at least two of the positions D56, I67, G71, D568, T605, E655, E694, and E706.

[0083] Alternatively or in addition, the manipulated variant CRISPR nuclease polypeptides disclosed herein may have an N-terminal cleavage relative to nuclease A of SEQ ID NO: 1. Such manipulated variant CRISPR nuclease polypeptides have a deleted fragment at the N-terminus of nuclease A. In some cases, the deleted N-terminal fragment has up to 80 amino acid residues, e.g., up to 70 amino acid residues, up to 60 amino acid residues, up to 50 amino acid residues, up to 40 amino acid residues, up to 30 amino acid residues, or up to 20 amino acid residues. In some cases, the N-terminal cleavage is a deletion within residues 1-15 of SEQ ID NO: 1. In one specific example, the N-terminal cleavage is a deletion within residues 1-14 of SEQ ID NO: 1. In another specific example, the N-terminal cleavage is a deletion within residues 1-15 of SEQ ID NO: 1.

[0084] Alternatively or in addition, the engineered CRISPR nuclease polypeptides disclosed herein may have a C-terminal cleavage relative to the reference CRISPR nuclease A of SEQ ID NO: 1. Such engineered CRISPR nuclease polypeptides have a deleted fragment at the C-terminus of the reference CRISPR nuclease. In some cases, the deleted C-terminal fragment has up to 10 amino acid residues, e.g., up to 5 amino acid residues, up to 4 amino acid residues, up to 3 amino acid residues, up to 2 amino acid residues, or up to 1 amino acid residue. In some cases, the C-terminal cleavage is the deletion of the last residue, the last two residues, the last three residues, the last four residues, or the last five residues of SEQ ID NO: 1. In some cases, the C-terminal cleavage variant may further include one or more of the mutations disclosed herein, e.g., one or more arginine / lysine substitutions, one or more mutations resulting in nickase activity, or a combination thereof. In some specific examples, the C-terminal cleavage variant further includes arginine substitutions I67R, D568R, and E706R relative to SEQ ID NO: 1. In other specific examples, the C-terminal cleavage variant further includes arginine substitutions I67R, D568R, W593R, and E706R relative to SEQ ID NO: 1. In some examples, the C-terminal cleavage variant further includes an N-terminal cleavage. In some examples, the C-terminal cleavage variant further includes deletions of residues 1-14 or 1-15 of SEQ ID NO: 1.

[0085] In some cases, the cleavage variant may further include one or more of the mutations disclosed herein, e.g., one or more arginine / lysine substitutions, one or more mutations resulting in nickase activity (nickase mutations), or combinations thereof. In some specific examples, the cleavage variant may further include arginine / lysine substitutions at positions I67, D568, and E706 of SEQ ID NO: 1 (e.g., arginine substitutions I67R, D568R, and E706R). In other specific examples, the cleavage variant may further include arginine / lysine substitutions at positions I67, D568, W593, and E706 of SEQ ID NO: 1 (e.g., arginine substitutions at I67R, D568R, W593R, and E706R).

[0086] Alternatively, or in addition, any of the CRISPR nuclease polypeptide variants of nuclease A provided herein may contain one or more nickase mutations in either the RuvC or HNH nuclease domain to reduce or eliminate the nuclease activity of a target domain, thereby producing a variant with nickase activity. Such mutations may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation in either the RuvC or HNH nuclease domain is an amino acid substitution, and the substituted amino acid residue of the amino acid substitution is not a conserved substitution of the native amino acid residue at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue except K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue except R. Bases for conserved amino acid residue substitutions are provided herein.

[0087] Positions D24, E337, and D524 are identified as putative catalytic residues in the RuvC domain of nuclease A of SEQ ID NO: 1, and positions D396, H397, and N420 are identified as putative catalytic residues in the HNH domain of nuclease A. In some cases, one or more mutations may be located within the RuvC nuclease domain and reduce or inactivate the RuvC domain (e.g., at positions D24, E337, and / or D524 in the RuvC domain). In some specific examples, variant CRISPR nuclease polypeptides derived from nuclease A may contain substitutions of D24A, E337A, and / or D524A. In other examples, any of D24, E337, and D524 may be substituted with amino acid residues similar to A, e.g., G, S, or L.

[0088] In some cases, one or more mutations may be located within the HNH domain and may reduce or inactivate the HNH domain (e.g., at positions D396, H397, and / or N420 in the HNH domain). In some specific cases, the variant CRISPR nuclease polypeptide may contain substitutions of H397A, D396A, and / or N420A. Alternatively, any of H397, D396, and N420 may be replaced with an amino acid residue similar to A, e.g., G, L, or S. In one specific case, the variant CRISPR nuclease polypeptide may contain substitutions of H397A or H397L.

[0089] In other examples, one or more mutations may be located within both the RuvC domain and the HNH domain of nuclease A, and may reduce or decrease the nuclease enzyme activity.

[0090] In some embodiments, the variant CRISPR nuclease polypeptides provided herein may include both arginine / lysine substitutions (e.g., arginine substitutions), e.g., arginine / lysine substitutions (e.g., arginine substitutions) at one or more positions provided herein, and nickase mutations. For example, the variant CRISPR nuclease polypeptide may include arginine / lysine substitutions (e.g., arginine substitutions) at positions I67, D568, and / or E706 in SEQ ID NO: 1, and may further include an amino acid substitution at H397 (e.g., H1397A or H397L). In some cases, such variants may further include N-terminal cleavage as disclosed herein, e.g., deletions within residues 1-15 of SEQ ID NO: 1, e.g., deletions between 1-14aa or 1-15aa of SEQ ID NO: 1. In some cases, such variants may further include C-terminal cleavage as disclosed herein, e.g., deletions within the last five residues of SEQ ID NO: 1.

[0091] In some examples, the variant CRISPR nuclease polypeptides provided herein may share at least 90% (e.g., 95%, 97%, 98%, 99%, 99.5%, or more) sequence identity with SEQ ID NO: 1. Exemplary manipulated variant CRISPR nuclease polypeptides of nuclease A are listed in Table 1, Table 11, or Table 14, each of which is within the scope of this disclosure. In specific examples, the variant CRISPR nuclease polypeptide may contain (e.g., consist of) any one of SEQ ID NOs: 6-13, 34, 36, 38, 40, 42, 44, 46, 58, 62, and 64 (e.g., SEQ ID NO: 6 or SEQ ID NO: 36).

[0092] (B) Nuclease K and its modified variants In some embodiments, the CRISPR nuclease polypeptide is derived from nuclease K, and nuclease K includes both wild-type nuclease K (SEQ ID NO: 65), its variants, e.g., those disclosed herein, or a fusion polypeptide containing them. The nuclease K variant CRISPR nuclease polypeptide may be produced by introducing one or more mutations into a reference CRISPR nuclease to modulate (e.g., enhance or reduce) one or more activities of the nuclease.

[0093] Reference CRISPR nuclease K (SEQ ID NO: 65; see Table 15 below) is a CRISPR nuclease containing both a RuvC nuclease domain (located at residues 1-39, 286-320, and 448-543 of SEQ ID NO: 1) and an HNH domain (located at residues 321-447 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate DNA strand cleavage adjacent to the 5'-NGG-3'PAM motif, where N represents any nucleotide. Positions D10, E315, and D500 are considered active sites in the RuvC domain, while positions D374, H375, and N398 are considered active sites in the HNH domain. R482 and H497 may also be important for the nuclease activity of the RuvC domain. In addition to the nuclease domain, the reference CRISPR nuclease of SEQ ID NO: 1 also contains a BH domain (residues 40-77 of SEQ ID NO: 1), an REC domain (residues 78-285 of SEQ ID NO: 1), a PLL domain (residues 544-556), a WED domain (residues 557-648), and a PID domain (residues 649-747).

[0094] Compared to the CRISPR-Cas9 nuclease from Streptococcus pyogenes, the reference CRISPR nuclease K of SEQ ID NO: 65 disclosed herein is smaller. The scaffolds utilized by the reference CRISPR nuclease of SEQ ID NO: 65 and its variants can be miniaturized. This scaffold contains a characteristic structure compared to the SpCas9 nuclease scaffold. The characteristic scaffold allows for a reduced size of several domains (e.g., the REC domain), and is therefore expected to contribute to a smaller CRISPR nuclease size. These features are beneficial for delivery. Arginine and / or lysine substitutions can be introduced into the CRISPR nuclease of SEQ ID NO: 65 to increase indel activity.

[0095] Furthermore, the RuvC domain of nuclease K cleaves the non-target strand within 4-6 nucleotides upstream of the PAM, and the HNH domain cleaves the target strand within 3-4 nucleotides upstream of the PAM, each resulting in a cleavage site with an overhang of 0-3 nucleotides (most commonly, 3 nucleotides). This cleavage pattern differs from that of SpCas9 described above. In addition, the use of gene editing systems containing nuclease K or its variants can result in the introduction of larger indels into the target nucleic acid than those that can be introduced by SpCas9. For example, insertions induced by the use of nuclease K or its variants can range from about 1 to about 10 nucleotides (most commonly, about 4 nucleotides). Deletions induced by the use of nuclease A or its variants can range from about 1 to about 25 nucleotides or more (most commonly, about 11 nucleotides). Furthermore, since nuclease K contains both a RuvC domain and an HNH domain, nickasase variants can be manipulated, for example, by disrupting the nuclease activity of either the RuvC domain or the HNH domain.

[0096] The variants of the nuclease K polypeptide provided herein may contain one or more mutations of nuclease K (SEQ ID NO: 65), such as one or more amino acid residue substitutions, one or more deletions, one or more insertions, one or more fusions, or a combination thereof. In some cases, the changes may be introduced within the BH domain, WED, PID, or a combination thereof of nuclease K. In some cases, the changes are not introduced within the RuvC and / or HNH nuclease domain of nuclease K, or within the active site and / or site involved in activity in these domains provided herein. Alternatively, a conservative amino acid substitution may be introduced within SEQ ID NO: 65, including within the RuvC and / or HNH nuclease domain of nuclease K.

[0097] In some embodiments, the variant CRISPR nuclease polypeptide derived from nuclease K provided herein may contain one or more arginine substitutions, one or more lysine substitutions, or combinations thereof with respect to SEQ ID NO: 65. In some examples, the variant CRISPR nuclease polypeptide derived from nuclease K may contain up to 20 arginine and / or lysine substitutions (e.g., up to 20 arginine substitutions, up to 20 lysine substitutions, or combinations thereof), for example, up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine substitutions, lysine substitutions, or combinations thereof. In specific examples, the variant CRISPR nuclease polypeptide derived from nuclease K may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine substitutions, lysine substitutions, or combinations thereof. In some examples, the variant CRISPR nuclease polypeptides derived from nuclease K provided herein contain arginine substitutions.

[0098] In some cases, arginine and / or lysine substitutions may be located in the BH domain, WED domain, PID domain, or a combination thereof of nuclease K. For example, arginine and / or lysine substitutions may be introduced in one or more of the following positions in SEQ ID NO: 65: G42, D46, I53, F83, E128, E541, F570, E576, D582, C630, I631, E683, S719, and T734. In some cases, arginine and / or lysine substitutions may be located at positions G42 and D582 in SEQ ID NO: 65. In other cases, arginine and / or lysine substitutions may be located at positions I53, D582, and T734 in SEQ ID NO: 65.

[0099] In some cases, variant CRISPR nuclease polypeptides derived from nuclease K (SEQ ID NO: 65) may contain one or more of the following arginine substitutions: G42R, D46R, I53R, E128R, D582R, F570R, E576R, C630R, I631R, E683R, T734R, and S719R. In other cases, variant CRISPR nuclease polypeptides derived from nuclease K may contain one or more of the following arginine substitutions: G42R, D46R, I53R, F83R, E541R, F570R, D582R, E683R, and S719R. In some specific examples, variant CRISPR nuclease polypeptides derived from nuclease K may contain arginine substitutions of D582R and G42R. In other specific examples, variant CRISPR nuclease polypeptides may contain arginine substitutions at D582R, I53R, and T734R.

[0100] In some cases, arginine and / or lysine substitutions may be located within the RuvC and / or HNH nuclease domains of nuclease K. In some cases, arginine and / or lysine substitutions may be located within the RuvC nuclease domain of nuclease K and may reduce or inactivate the RuvC domain (e.g., at positions D10, E315, R482, H497, and / or D500 in the RuvC domain of SEQ ID NO: 65). In other cases, arginine and / or lysine substitutions may be located within the HNH domain of nuclease K, for example, at positions D374, H375, and / or N398 in the HNH domain of SEQ ID NO: 65. In other cases, arginine and / or lysine substitutions may be located within both the RuvC and HNH domains of nuclease K and may reduce or decrease nuclease enzyme activity.

[0101] Alternatively, arginine and / or lysine substitutions may not be located in the active site and / or activity-involved site of the RuvC and / or HNH nuclease domain of nuclease K (e.g., not at positions D10, E315, R482, H497, and / or D500 in the RuvC domain, and / or not at positions D374, H375, and / or N398 in the HNH domain). In some examples, arginine and / or lysine substitutions may not be located in the RuvC and / or HNH domain of nuclease K.

[0102] Alternatively, or in addition, variant CRISPR nuclease polypeptides derived from nuclease K provided herein may contain one or more mutations (i.e., nickase mutations) within either the RuvC or HNH nuclease domain to reduce or eliminate the nuclease activity of a target domain, thereby producing variants with nickase activity. Such mutations may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation within either the RuvC or HNH nuclease domain of nuclease K is an amino acid substitution, and the substituted amino acid residue of the amino acid substitution is not a conserved substitution of the native amino acid residue at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue other than K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue other than R. Bases for conserved amino acid residue substitutions are provided herein.

[0103] Positions D10, E315, and D500 are considered to be active sites in the RuvC domain of nuclease K (SEQ ID NO: 65), and positions D374, H375, and N398 are considered to be active sites in the HNH domain. R482 and H497 may also be important for the nuclease activity of the RuvC domain of nuclease K. In some cases, one or more mutations may be located within the RuvC nuclease domain and may reduce or inactivate the RuvC domain (e.g., at positions D10, E315, R482, H497, and / or D500 in the RuvC domain). In some specific examples, variant CRISPR nuclease polypeptides may contain substitutions of D10A, E315A, and / or D500A. In other examples, one of D10, E315, and D500 may be substituted with an amino acid residue similar to A, for example, G, S, or L.

[0104] In some cases, one or more mutations may occur within the HNH domain (e.g., at positions H375, D374, and / or N398 in the HNH domain of SEQ ID NO: 65). In some specific cases, variant CRISPR nuclease polypeptides derived from nuclease K may contain substitutions of H375A, D374A, and / or N398A. Alternatively, any of H375, D374, and N398 may be replaced with an amino acid residue similar to A, e.g., G, L, or S. In one specific case, variant CRISPR nuclease polypeptides derived from nuclease K may contain substitutions of H375A or H375L.

[0105] In other examples, one or more mutations may be located within both the RuvC domain and the HNH domain of nuclease K, and may reduce or decrease nuclease enzyme activity.

[0106] In some embodiments, variant CRISPR nuclease polypeptides derived from nuclease K provided herein may include both arginine / lysine substitutions (e.g., arginine substitutions) at one or more positions provided herein, and nickase mutations, for example, nickase mutations at one or more positions also provided herein. In some cases, variant CRISPR nuclease polypeptides of nuclease K may include arginine / lysine substitutions (e.g., arginine substitutions) listed in Tables 18 and 20. Such variant CRISPR nuclease polypeptides may further include one or more mutations that reduce or inactivate the RuvC and / or HNH nuclease domain (e.g., one or more mutations within the RuvC and / or HNH nuclease domain). For example, the variant CRISPR nuclease polypeptide derived from nuclease K may contain an arginine / lysine substitution (e.g., arginine substitution) at positions D582, I53, and T734 in SEQ ID NO: 65, and an amino acid substitution at H375 (e.g., H375A).

[0107] (C) Nuclease M and its manipulated variants In some embodiments, the nuclease polypeptide is derived from nuclease M, where nuclease M includes both wild-type nuclease M (SEQ ID NO: 80), its variants, e.g., those disclosed herein, or a fusion polypeptide comprising the same. The variant nuclease polypeptide of nuclease M may be produced by introducing one or more mutations into nuclease M to modulate (e.g., enhance or reduce) one or more activities of the nuclease.

[0108] Reference nuclease M of SEQ ID NO: 80 (see Table 21 below) is an RNA-guided nuclease containing both a RuvC nuclease domain (located at residues 52-84, 160-192, and 289-367 of SEQ ID NO: 80) and an HNH domain (located at residues 193-288 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate DNA strand breaks adjacent to a 5'-WTAAH-3' PAM motif, where W is A or T and H is A, C, or T. In one example, the PAM motif is 5'-TTAAA-3'.

[0109] Positions D58, E189, and D341 are considered to be active sites in the RuvC domain, and positions H243, H244, and H267 are considered to be active sites in the HNH domain. R329 and H338 may also be important for the nuclease activity of the RuvC domain. In addition to the nuclease domain, nuclease M of SEQ ID NO: 80 also contains a PLMP domain (residues 1-51 of SEQ ID NO: 1), a BH domain (residues 85-117 of SEQ ID NO: 1), a REC domain (residues 118-159 of SEQ ID NO: 1), a PLL domain (residues 368-380), a WED domain (residues 381-427), and a PID domain (residues 428-497).

[0110] Compared to the CRISPR-Cas9 nuclease from Streptococcus pyogenes, the reference nuclease M of SEQ ID NO: 80 disclosed herein is smaller. The scaffolds utilized by the reference CRISPR nuclease of SEQ ID NO: 80 and its variants can be miniaturized. This scaffold contains a characteristic structure compared to the SpCas9 nuclease scaffold. The characteristic scaffold allows for a reduced size of several domains (e.g., the REC domain), and is therefore expected to contribute to a smaller nuclease size. These features are beneficial for delivery. Arginine substitutions can be introduced into nuclease M SEQ ID NO: 80 to increase indel activity. Compared to SpCas9, nuclease M recognizes a different PAM motif in genomic targets, which allows for the editing of additional gene targets.

[0111] Furthermore, the RuvC domain of nuclease M cleaves the non-target strand within 3 to 12 nucleotides upstream of the PAM, and the HNH domain cleaves the target strand within 3 to 4 nucleotides upstream of the PAM, each resulting in a cleavage site with an overhang of 0 to 9 nucleotides (most commonly, a 5-nucleotide overhang). This cleavage pattern differs from that of SpCas9 described above. Also, unlike SpCas9, the RNA-guided nuclease of sequence number 80 contains a PLMP domain. The PLMP domain is expected to bind to a helix at the 3' end of the scaffold, providing the RNA-guided nuclease's increased binding affinity to its congener scaffold.

[0112] Additionally, since nuclease M of sequence number 80 contains a RuvC domain and an HNH domain, nickasase variants can be manipulated, for example, by disrupting the nuclease activity of either the RuvC domain or the HNH domain.

[0113] A variant nuclease polypeptide of nuclease M contains one or more mutations in nuclease M (SEQ ID NO: 80) and can modulate (e.g., enhance or reduce) one or more activities of the nuclease, such as one or more amino acid substitutions, one or more deletions, one or more insertions, one or more fusions, or a combination thereof. In some cases, the mutation may be introduced within the BH domain, REC domain, PLL domain, WED domain, PID domain, or a combination thereof. In some cases, the mutation is not introduced within the RuvC and / or HNH nuclease domain, or within the active site and / or activity-involved site of these domains provided herein. Alternatively, a conservative amino acid substitution may be introduced within SEQ ID NO: 80, including within the RuvC and / or HNH nuclease domain.

[0114] In some embodiments, the variant nuclease polypeptides of nuclease M provided herein may contain one or more arginine substitutions, one or more lysine substitutions, or combinations thereof with respect to SEQ ID NO: 80. In some examples, the variant nuclease polypeptide derived from nuclease M may contain up to 20 arginine substitutions, up to 20 lysine substitutions, or combinations thereof, for example, up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine substitutions, lysine substitutions, or combinations thereof. In some specific examples, the variant nuclease polypeptide derived from nuclease M may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine substitutions, lysine substitutions, or combinations thereof. In some specific examples, the variant nuclease polypeptide derived from nuclease M contains arginine substitutions with respect to SEQ ID NO: 80.

[0115] In some cases, the variant nuclease polypeptide of nuclease M may contain one or more arginine and / or lysine substitutions (e.g., arginine substitutions) at positions E88, S95, L92, E401, E83, N371, P481, and / or A373 in SEQ ID NO: 80. In some cases, the variant nuclease polypeptide derived from nuclease M may contain at least two arginine and / or lysine substitutions, e.g., at least two arginine substitutions, at these positions.

[0116] In some cases, arginine and / or lysine substitutions (e.g., arginine substitutions) may be located in any of the BH domain, REC domain, PLL domain, WED domain, PID domain, or a combination thereof in nuclease M. Alternatively or in addition, arginine and / or lysine substitutions (e.g., arginine substitutions) may be located within the RuvC and / or HNH nuclease domain of nuclease M. In some cases, arginine and / or lysine substitutions (e.g., arginine substitutions) may be located within the RuvC nuclease domain of nuclease M, reducing or inactivating the RuvC domain of nuclease M (e.g., at positions D58, E189, R329, H338, and / or D341 in the RuvC domain of SEQ ID NO: 80). In other examples, arginine and / or lysine substitutions (e.g., arginine substitutions) may be located within the HNH domain of nuclease M (SEQ ID NO: 80), for example, at positions H243, H244, and / or H267 within the HNH domain. In other examples, arginine and / or lysine substitutions (e.g., arginine substitutions) may be located within both the RuvC domain and the HNH domain of nuclease M, and may reduce or decrease nuclease enzyme activity.

[0117] Alternatively, arginine and / or lysine substitutions (e.g., arginine substitutions) may not be located in the active site and / or activity-involved site of the RuvC and / or HNH nuclease domain of nuclease M (e.g., not at positions D58, E189, R329, H338, and / or D341 in the RuvC domain of SEQ ID NO: 80, and / or not at positions H243, H244, and / or H267 in the HNH domain). In some examples, arginine and / or lysine substitutions (e.g., arginine substitutions) may not be located in the RuvC and / or HNH domain of nuclease M.

[0118] Exemplary arginine-substituted variants of nuclease M are listed in Table 24 below, each of which is within the scope of this disclosure.

[0119] Alternatively, or in addition, variant nuclease polypeptides of nuclease M provided herein may contain one or more mutations (i.e., nickase mutations) within either the RuvC or HNH nuclease domain to reduce or eliminate the nuclease activity of a target domain, thereby producing a variant with nickase activity. Such mutations may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation within either the RuvC or HNH nuclease domain is an amino acid substitution, and the substituted amino acid residue of the amino acid substitution is not a conserved substitution of the native amino acid residue at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue except K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue except R. Bases for conserved amino acid residue substitutions are provided herein.

[0120] In some cases, one or more mutations may be located within the RuvC nuclease domain and may reduce or inactivate the RuvC domain of nuclease M (e.g., at positions D58, E189, R329, H338, and / or D341 in the RuvC domain of SEQ ID NO: 80). Alternatively, one or more nickase mutations may be located within the HNH domain of nuclease M (e.g., at positions H243, H244, and / or H267 in the HNH domain of SEQ ID NO: 80). In specific examples, variant nuclease polypeptides of nuclease M may contain one or more nickase mutations at one or more of the positions D58, E189, D341, H243, H244, H267, R329, and / or H338 in SEQ ID NO: 80. In some cases, one or more of the original residues in either the RuvC domain or the HNH domain of nuclease M may be replaced with alanine (A). Alternatively, one or more of the original residues in either the RuvC domain or the HNH domain may be replaced with amino acid residues similar to A, such as G, S, or L.

[0121] In other examples, one or more mutations may be located within both the RuvC domain and the HNH domain of nuclease M, and may reduce or decrease nuclease enzyme activity.

[0122] In some embodiments, the variant nuclease polypeptide derived from nuclease M provided herein may include both arginine / lysine substitutions (e.g., arginine substitutions), e.g., arginine / lysine substitutions at one or more positions provided herein, and nickase mutations, e.g., nickase mutations at one or more positions also provided herein.

[0123] In some cases, variant nuclease polypeptides derived from nuclease M may share at least 90% (e.g., 95%, 97%, 98%, 99%, 99.5%, or more) sequence identity with SEQ ID NO: 80.

[0124] In some embodiments, the CRISPR nuclease polypeptides derived from nuclease A, nuclease K, or nuclease M provided herein may be fusion polypeptides comprising the CRISPR nuclease moiety disclosed herein and one or more additional functional elements. In some cases, the one or more additional functional elements may be heterogeneous to the CRISPR nuclease moiety.

[0125] As used herein, the terms “fusion” and “fused” refer to the linkage of at least two nucleotides or protein molecules. For example, “fusion” and “fused” may refer to the linkage of at least two polypeptide domains that are naturally encoded by distinct genes. The fusion may be an N-terminal fusion, a C-terminal fusion, or an intramolecular fusion. In some embodiments, the domains are transcribed and translated to produce a single polypeptide.

[0126] In some cases, the CRISPR nuclease moiety in the fusion polypeptide may be reference CRISPR nuclease A (SEQ ID NO: 1), nuclease K (SEQ ID NO: 65), or nuclease M (SEQ ID NO: 80). Alternatively, the CRISPR nuclease moiety in the fusion polypeptide may be a variant of any of the reference nucleases disclosed herein. Exemplary additional functional moieties that may be included within the fusion polypeptide include peptide tags, fluorescent proteins, base editing domains, DNA methylation domains, histone residue modification domains, localization factors, transcription modifiers, photo-gate regulators, chemoinducible factors, chromatin visualization factors, or combinations thereof.

[0127] In some embodiments, additional functional components may include a nuclear localization signal (NLS), a nuclear export signal (NES), or a combination thereof. In some examples, the fusion polypeptide may contain an NLS, which may be located at either the N-terminus or the C-terminus. In specific examples, the fusion polypeptide may contain a first NLS located at the N-terminus and a second NLS located at the C-terminus. The first and second NLS fragments may be identical. Alternatively, the two NLS fragments may be different. In some embodiments, the fusion polypeptide may contain the NLS near the N-terminus and / or C-terminus (e.g., within about one, two, three, four, or five amino acids from the first or last amino acid of the CRISPR nuclease). In some embodiments, the fusion polypeptide may contain the NLS within the mobile loop of the CRISPR nuclease. Exemplary fusion CRISPR nuclease polypeptides containing one or more NLS signals are provided in Tables 1, 15, and 21.

[0128] In some embodiments, additional functional components may be mobile peptide linkers, such as XTEN peptide linkers or G / S-rich peptide linkers. Examples of such peptide linkers are provided in Example 1 below, and these may be applicable to any of the CRISPR nuclease polypeptides disclosed herein.

[0129] Preparation of B.CRISPR nuclease polypeptide The CRISPR nuclease polypeptides disclosed herein may be prepared by conventional methods or by methods disclosed herein. For example, CRISPR nuclease polypeptides may be prepared by culturing host cells capable of producing nuclease polypeptides, such as bacterial or mammalian cells, isolating the nuclease polypeptides thus produced, and optionally purifying the nuclease polypeptides. The CRISPR nuclease polypeptides thus prepared may form complexes with gRNA.

[0130] CRISPR nuclease polypeptides can also be prepared by an in vitro coupled transcription-translation system and, optionally, form complexes with gRNA. Bacteria that can be used for the preparation of CRISPR nuclease polypeptides are not particularly limited, as long as they can produce CRISPR nuclease polypeptides. Some non-limiting examples of bacteria include E. coli cells described herein.

[0131] Unless otherwise noted, all compositions, complexes, and polypeptides provided herein are prepared with respect to the activity level of the composition, complex, or polypeptide, excluding impurities that may be present in commercially available sources, such as residual solvents or by-products. The enzymatic component weight is based on the total active protein. Unless otherwise indicated, all percentages and ratios are calculated by weight. Unless otherwise indicated, all percentages and ratios are calculated based on the total composition. In exemplary compositions, the enzyme level is expressed by the weight of pure enzyme in the total composition, and unless otherwise specified, components are expressed by the weight of the total composition.

[0132] (i) Vector This disclosure provides vectors for expressing CRISPR nuclease polypeptides. In some embodiments, the vectors disclosed herein include a nucleotide sequence encoding the CRISPR nuclease polypeptide provided herein. In some embodiments, the vectors include a Pol II promoter or a Pol III promoter.

[0133] The expression of native or synthetic polynucleotides is typically achieved by operably ligating a polynucleotide encoding a CRISPR nuclease polypeptide to a promoter and incorporating the construct into an expression vector. The expression vector is not particularly limited, as long as it contains a polynucleotide encoding a CRISPR nuclease polypeptide and is suitable for replication and incorporation in eukaryotic cells.

[0134] A typical expression vector contains transcription and translation terminators, start sequences, and promoters useful for the expression of a desired polynucleotide. For example, plasmid vectors containing recognition sequences for RNA polymerase (e.g., pSP64, pBluescript) can be used. Vectors containing retroviruses, such as those derived from lentiviruses, are suitable tools for achieving long-term gene transfer, as they allow for the long-term stable integration of the transgene and its proliferation in daughter cells. Examples of vectors include expression vectors, replication vectors, probe-generating vectors, and sequencing vectors. Expression vectors can be supplied to cells in the form of viral vectors.

[0135] Viral vector technology is well-known in the art and is described in various virology and molecular biology manuals. Useful viruses as vectors include, but are not limited to, phage viruses, retroviruses, adenoviruses, adeno-associated viruses, herpesviruses, and lentiviruses. Generally, a suitable vector contains a functional replication origin, promoter sequence, convenient restriction endonuclease site, and one or more selectable markers in at least one organism.

[0136] The type of vector is not particularly limited, and any vector that can be expressed in a host cell can be appropriately selected. More specifically, depending on the type of host cell, a promoter sequence is appropriately selected to ensure the expression of a polypeptide(s) from a polynucleotide, and this promoter sequence and polynucleotide are inserted into one of various plasmids or other materials for the preparation of the expression vector.

[0137] Additional promoter elements, such as enhancing sequences, regulate the frequency of transcription initiation. Typically, these are located 30–110 bp upstream of the start site, although some promoters have recently been shown to also contain functional elements downstream of the start site. Depending on the promoter, individual elements appear to be able to activate transcription either cooperatively or independently.

[0138] Furthermore, this disclosure should not be limited to the use of constitutive promoters. Inducible promoters are also intended as part of this disclosure. The use of an inducible promoter provides a molecular switch that can turn on the expression of the polynucleotide sequence to which it is operably bound when such expression is desired, or turn off the expression when expression is not desired. Examples of inducible promoters include, but are not limited to, metallothione promoters, glucocorticoid promoters, progesterone promoters, and tetracycline promoters.

[0139] The introduced expression vector may also contain either or both a selectable marker gene or a reporter gene to facilitate the identification and selection of expression cells from a population of cells intended to be transfected or infected via the viral vector. In other embodiments, the selectable marker may be vested in a separate DNA fragment or used in a co-transfection procedure. Both the selectable marker and reporter gene may be flanked by appropriate transcriptional regulatory sequences to enable expression in host cells. Examples of such markers include dihydrofolate reductase genes and neomycin resistance genes for eukaryotic cell culture; as well as tetracycline resistance genes and ampirin resistance genes for E. coli and other bacterial cultures. The use of such selectable markers may allow for confirmation that the polynucleotide encoding the polypeptide(s) of the present invention has been transferred into host cells and subsequently expressed without failure.

[0140] The preparation methods using recombinant expression vectors are not particularly limited and include methods using plasmids, phages, or cosmids.

[0141] (ii) Method of expression This disclosure includes a method for protein expression, comprising translating a CRISPR nuclease polypeptide as described herein.

[0142] In some embodiments, the host cells described herein are used to express CRISPR nuclease polypeptides. The host cells are not particularly limited, and a variety of known cells may be preferably used. Specific examples of host cells include bacteria, e.g., E. coli, yeast (budding yeast Saccharomyces cerevisiae and fission yeast Schizosaccharomyces pombe), nematodes (Caenorhabditis elegans), Xenopus laevis oocytes, and animal cells (e.g., CHO cells, COS cells, and HEK293 cells). The method for transferring the expression vector into the host cells, i.e., the transformation method, is not particularly limited, and known methods, e.g., electroporation, calcium phosphate methods, liposome methods, and DEAE dextran methods, may be used.

[0143] After the host has been transformed with the expression vector, the host cells may be cultured, cultivated, or propagated for the production of CRISPR nuclease polypeptides. After expression, the host cells may be collected, and the CRISPR nuclease polypeptides may be purified from the culture or other source according to conventional methods (e.g., filtration, centrifugation, cell disruption, gel filtration chromatography, ion exchange chromatography, etc.).

[0144] Various methods can be used to determine the level of production of mature CRISPR nuclease polypeptides in host cells. Such methods include, but are not limited to, methods using protein-specific polyclonal or monoclonal antibodies or labeling tags as described elsewhere herein. Exemplary methods include, but are not limited to, enzyme-linked immunosorbent assays (ELISA), radioimmunoassays (MA), fluorescence immunoassays (FIA), and fluorescence-activated cell sorting (FACS). These and other assays are well known in the art (see, for example, Maddox et al., J.Exp.Med.158:1211

[1983] ).

[0145] This disclosure provides a method for in vivo expression of a CRISPR nuclease polypeptide (and optionally, gRNA in a gene editing system disclosed herein). Such a method may include providing a polyribonucleotide encoding a CRISPR nuclease polypeptide to a host cell in a subject (e.g., a human subject), the polyribonucleotide encoding the CRISPR nuclease polypeptide that causes the cell to express the CRISPR nuclease polypeptide.

[0146] II. Gene Editing Systems In some embodiments, the disclosure provides a gene editing system having enhanced gene editing efficiency. The gene editing system comprises one of the CRISPR nuclease polypeptides derived from nuclease A, nuclease K, or nuclease M disclosed herein, or a nucleus encoding a CRISPR nuclease, and one or more guide RNAs (gRNAs) or nucleic acids (may be multiple) encoding gRNAs.

[0147] CRISPR nuclease In some embodiments, the gene editing systems disclosed herein comprise a CRISPR nuclease polypeptide provided herein, e.g., any of reference CRISPR nuclease A, nuclease K, or nuclease M, or a variant thereof, the CRISPR nuclease polypeptide comprising, for example, one or more arginine and / or lysine substitutions (e.g., arginine substitutions), one or more nickase mutations, additional mutations, e.g., N-terminal cleavage where applicable, or a combination thereof. See above disclosure. Such protein components may form complexes with gRNAs in the same gene editing system. Alternatively, the gene editing system comprises a nucleic acid encoding a CRISPR nuclease polypeptide. In some cases, the nucleic acid may be an expression vector (e.g., a viral vector) for producing the encoded nuclease polypeptide in a host cell. In some cases, the expression vector may further comprise a coding sequence for producing one or more gRNAs of the gene editing system.

[0148] Guide RNA The gene editing systems disclosed herein further comprise one or more gRNAs or encoding nucleic acids. As used herein, the terms “RNA guide,” “RNA guide sequence,” or “guide RNA (gRNA)” refer to an RNA molecule or modified RNA molecule that facilitates the CRISPR nuclease described herein to target a genomic site of interest. For example, an RNA guide may be a molecule comprising a spacer sequence and a scaffold sequence. The spacer sequence recognizes (e.g., binds to) a site on a non-PAM strand that is complementary to the target sequence on the PAM strand, for example, a site on a non-PAM strand designed to be complementary to a specific nucleic acid sequence. The scaffold sequence contains a nuclease-binding sequence for binding to the CRISPR nuclease. In some embodiments, the scaffold is an RNA sequence.

[0149] In some cases, the gRNAs disclosed herein may further include a linker sequence, a 5' end and / or a 3' end protection fragment, or a combination thereof.

[0150] (i) Spacer array As used herein, the terms “spacer” and “spacer sequence” (also known as DNA-binding sequence) refer to a portion (DNA sequence) of an RNA guide that is the RNA equivalent of the target sequence. A spacer contains a sequence that can bind to the non-PAM strand via base pairing at a site complementary to the target sequence (which is located on the PAM strand). Such spacers are also known to be specific to the target sequence. In some cases, a spacer may be at least 75% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%) identical to the target sequence, except for the RNA-DNA sequence difference. In some cases, a spacer may be 100% identical to the target sequence, except for the RNA-DNA sequence difference.

[0151] The gene editing systems disclosed herein comprise one or more gRNAs, each comprising a spacer sequence specific to a target sequence at a genomic site of interest, and a scaffold sequence recognizable by a CRISPR nuclease polypeptide contained within the gene editing system.

[0152] In relation to nuclease A or its variants disclosed herein, the target sequence may be adjacent to (e.g., upstream of or at 5') the 5'-NGG-3' PAM, where N represents any of the nucleotides.

[0153] In relation to nuclease K or its variants disclosed herein, the target sequence may be adjacent to (e.g., upstream of or at 5') the 5'-NGG-3' PAM, where N represents any of the nucleotides.

[0154] In relation to nuclease M, the target sequence can be adjacent to (e.g., upstream of or at 5') the PAM of 5'-WTAAH-3', where W is A or T, and H is A, C, or T. In some examples, the PAM can be 5'-TTAAA-3'.

[0155] As used herein, the terms “protospacer-adjacent motif” or “PAM sequence” refer to a DNA sequence adjacent to a target sequence. In some embodiments, the PAM sequence is required for CRISPR nuclease binding and / or indel activity. In a double-stranded DNA molecule, the strand containing the PAM motif is referred to as the “PAM strand,” and the complementary strand is referred to as the “non-PAM strand.” The gRNA binds to a site on the non-PAM strand that is complementary to the target sequence disclosed herein, and the PAM sequence described herein is present on the PAM strand. The PAM motif may be located upstream of the target sequence.

[0156] As used herein, the term “adjacent to” means that a nucleotide or amino acid sequence is in close proximity to another nucleotide or amino acid sequence. In some embodiments, if there are no nucleotides separating the two sequences, the nucleotide sequence is adjacent to (i.e., directly adjacent to) another nucleotide sequence. In some embodiments, if a small number of nucleotides (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides) separate the two sequences, the nucleotide sequence is adjacent to another nucleotide sequence. In some embodiments, if the two sequences are separated by about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides, the first sequence is adjacent to the second sequence. In some embodiments, if the two sequences are separated by up to 2 nucleotides, up to 5 nucleotides, up to 8 nucleotides, up to 10 nucleotides, up to 12 nucleotides, or up to 15 nucleotides, the first sequence is adjacent to the second sequence. In some embodiments, if the two sequences are separated by 2 to 5 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 10 to 12 nucleotides, 10 to 15 nucleotides, or 12 to 15 nucleotides, the first sequence is adjacent to the second sequence.

[0157] In a specific example, the spacer is specific to a target sequence at the desired genomic site, and the target sequence is immediately adjacent to the PAM motif. In another specific example, the target sequence and the PAM motif may have a small gap of fewer than five nucleotides (e.g., one, two, three, four, or five).

[0158] Spacer sequences disclosed herein may have a length of about 15 to about 30 nucleotides. For example, a spacer may have a length of about 15 to about 20 nucleotides, about 15 to about 25 nucleotides, about 20 to about 25 nucleotides, or about 20 to about 30 nucleotides. In some embodiments, spacers in gRNA may be designed to generally have a length of 15 to 25 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, and 25) and to be complementary to a specific target sequence. In some embodiments, spacer sequences may be designed to have a length of 18 to 22 nucleotides (e.g., 20 nucleotides).

[0159] In some embodiments, the spacer sequence may have at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.5% sequence identity with the target sequence described herein, and may bind to a complementary region of the target sequence via base pairing.

[0160] In some embodiments, the spacer sequence contains only RNA bases. In some embodiments, the spacer sequence contains DNA bases (e.g., the spacer contains at least one thymine). In some embodiments, the spacer sequence contains both RNA bases and DNA bases (e.g., the DNA-binding sequence contains at least one thymine and at least one uracil).

[0161] (ii) Scaffolding arrangement Scaffold sequences in gRNAs can similarly be recognized by CRISPR nuclease polypeptides in gene editing systems.

[0162] In some cases, the scaffolding array is recognizable by nuclease A (SEQ ID NO: 1) or any modified variant thereof as disclosed herein. Such a scaffolding array may include SEQ ID NO: 2 below. GUUACAGUUAAGGCUCUUUGGAAACAAAGAAGCCUUAAUUGUAAAACGCCUAUAUGGUAAAGUGAUGUACGUUUGGGUAUAUAUCGCCAGCCUGAACCUCUACGCCAGAAAUGGCAGCUUUAUCAUGGGUUAGGACGAUAUUUAAAAAACUUUCUGCGCUUGCUUACUUUAGUAAGCUUUGUGGCUGAGGCAGAAUUCCUU (Sequence number 2)

[0163] The predicted secondary structure of the reference scaffold array of sequence number 2 is provided in Figure 3A.

[0164] In other instances, a scaffold sequence recognizable by nuclease A (SEQ ID NO: 1) or any of its manipulated variants may be a variant derived from SEQ ID NO: 2. Such a variant scaffold sequence may contain nucleotide sequences identical to SEQ ID NO: 2 by at least 80% (e.g., at least 85%, 90%, 95%, 98%, or more). Alternatively or in addition, the variant scaffold sequence may include deletions, nucleotide substitutions, or combinations thereof. A variant CRISPR nuclease polypeptide may have increased binding to the variant scaffold sequence compared to the scaffold of SEQ ID NO: 2. In some examples, the variant scaffold may be a fragment of SEQ ID NO: 2 disclosed herein or a variant thereof. For example, the variant scaffold for use in gRNA provided herein may have a length in the range of 100–150.

[0165] In some embodiments, the scaffold sequence recognizable by nuclease A (SEQ ID NO: 1) or an engineered variant thereof may be a cleavage variant of SEQ ID NO: 2. Any of the cleavage variants provided herein may be about 110 to 140 nucleotides in length.

[0166] In some cases, a cleavage-type variant scaffold sequence recognizable by nuclease A (sequence number 1) or any of its manipulated variants may have a 3' cleavage relative to sequence number 2. Such a 3' cleavage may remove one or more of the stem structures P6a, P6b, and P6c shown in Figure 3A. In specific examples, the 3' cleavage may include a deletion within residues 143-202 of sequence number 2, e.g., a deletion of the 143-202 fragment of sequence number 2. See Figure 3B.

[0167] Alternatively, or in addition, the variant scaffold sequence may be a cleavage variant of SEQ ID NO: 2 having an internal cleavage relative to SEQ ID NO: 2. For example, such a cleavage variant may include a deletion of all or part of the stem structure P1 shown in Figure 3A. In some cases, the cleavage variant may include deletions within residues 14-20 of SEQ ID NO: 2, within residues 25-32 of SEQ ID NO: 2, or a combination thereof. In specific examples, the cleavage variant may include deletions of the 14-20 and 25-32 fragments of SEQ ID NO: 2. See, for example, Figures 3B and 3C. In some cases, the cleavage variant may include a deletion in the loop of residues 77-86 of SEQ ID NO: 2 (see Figure 3C), which may result in a shortening of the stem structure P5 shown in Figure 3A. In some cases, the cleavage variant may include deletions within residues 81-85 of SEQ ID NO: 2 (e.g., a deletion of the 81-85 fragment).

[0168] In some examples, the variant scaffold sequence recognizable by nuclease A (SEQ ID NO: 1) or any of its manipulated variants may include a combination of a 3' cleavage and one or more of the internal cleavages disclosed herein. Alternatively or in addition, the variant scaffold sequence may include one or more nucleotide variations for the corresponding residue in SEQ ID NO: 2.

[0169] In one specific example, the scaffold sequence recognizable by nuclease A (sequence number 1) or any of its manipulated variants contains (e.g., consists of) the nucleotide sequence of sequence number 27. In another specific example, the scaffold sequence contains (e.g., consists of) sequence number 28. Such scaffold sequences may be approximately 115–130 nt in length.

[0170] In some cases, the scaffolding sequence is recognizable by nuclease K (SEQ ID NO: 65) or any modified variant thereof disclosed herein. Such a scaffolding sequence may include SEQ ID NO: 79. GUUACAGUUAAGGCUCUGAAAAGAGCCUUAAUUGUAAAACGCCUAUACAGUGAAGGGAUAUACGCUUGGGUUUGUCCAGCCUGAGCCUCUAUGCCAGAAAUGGCGCCUUCAUCGUGGGUUAGGACAUUUAAAUUUAAAAAACUAUUCAGCACUGUUUGCUCUUGUCAGCUUGGUGGCAGA (Sequence number 79)

[0171] In other instances, a scaffold sequence recognizable by nuclease K (SEQ ID NO: 65) or any manipulated variant thereof may be a variant derived from SEQ ID NO: 79. Such a variant scaffold sequence may contain nucleotide sequences identical to SEQ ID NO: 79 by at least 80% (e.g., at least 85%, 90%, 95%, 98%, or more). Alternatively or in addition, the variant scaffold sequence may include deletions, nucleotide substitutions, or combinations thereof. A variant CRISPR nuclease polypeptide may have increased binding to the variant scaffold sequence compared to the scaffold of SEQ ID NO: 79. In some examples, the variant scaffold may be a fragment of SEQ ID NO: 79 disclosed herein or a variant thereof. For example, the variant scaffold sequences for use in gRNA provided herein may have lengths in the range of 100–150 nt.

[0172] In some cases, the scaffolding sequence is recognizable by nuclease M (SEQ ID NO: 80) or any modified variant thereof disclosed herein. Such a scaffolding sequence may include SEQ ID NO: 94. GUCAACUACCCCCGUCUAAAGACGGAGGCAUGAGGUUUCGUAACCAAGUGUUGUACCUGCGGGUACAGUAGUUGAACAGGCGGCGAUGCGGCUGGGCACUCCAGGAUGCCACUCCCAGUCCCGGACACUGCCGACGAGCCGCAUCAAGCCGGGGGAGACCAACCGGCUAACGAUAGCCGAGCAAUUACCUAAAAAGAGGUGCAAAGGAAAUGGUAU (SEQ ID NO: 94)

[0173] In other instances, a scaffold sequence recognizable by nuclease M (SEQ ID NO: 80) or any manipulated variant thereof may be a variant derived from SEQ ID NO: 94. Such a variant scaffold sequence may contain nucleotide sequences identical to SEQ ID NO: 94 by at least 80% (e.g., at least 85%, 90%, 95%, 98%, or more). Alternatively or in addition, the variant scaffold sequence may include deletions, nucleotide substitutions, or combinations thereof. A variant CRISPR nuclease polypeptide may have increased binding to the variant scaffold sequence compared to the scaffold of SEQ ID NO: 94. In some examples, the variant scaffold may be a fragment of SEQ ID NO: 94 disclosed herein or a variant thereof. For example, the variant scaffold for use in gRNA provided herein may have a length in the range of 150–200 nt.

[0174] In any of the gRNAs disclosed herein, the scaffold sequence may be located at the 3' end of the spacer sequence. In some cases, the scaffold sequence and the spacer sequence are directly linked. In other cases, the scaffold sequence and the spacer sequence may be linked via a nucleotide linker.

[0175] Nucleic acid modification Either the guide RNA or coding nucleic acid (e.g., mRNA encoding a CRISPR nuclease polypeptide) in the gene editing systems disclosed herein may include one or more modifications.

[0176] Exemplary modifications may include modifications to sugars, nucleic acid bases, nucleoside bonds (e.g., binding phosphate / phosphodiester bond / phosphodiester skeleton), and any combination thereof. Some of the exemplary modifications provided herein are described in detail below.

[0177] Any of the gRNAs or nucleic acid sequences encoding components of the composition may include any useful modification, e.g., modifications to sugars, nucleic acid bases, or nucleoside bonds (e.g., binding phosphate / phosphodiester bond / phosphodiester skeleton). One or more atoms of pyrimidine nucleic acid bases may be replaced or substituted with optionally substituted aminos, optionally substituted thiols, optionally substituted alkyls (e.g., methyl or ethyl), or halos (e.g., chloro or fluoro). One or more atoms of purine nucleic acid bases may be replaced or substituted with optionally substituted aminos, optionally substituted thiols, optionally substituted alkyls (e.g., methyl or ethyl), or halos (e.g., chloro or fluoro). In some embodiments, modifications (e.g., one or more modifications) are present in the sugars and nucleoside bonds, respectively. In some embodiments, any of the gRNAs or nucleic acid sequences encoding components of the composition may include debase sites (i.e., positions without purines or pyrimidines). Modifications may include modifications from ribonucleic acid (RNA) to deoxyribonucleic acid (DNA), threose nucleic acid (TNA), glycol nucleic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), or hybrids thereof. Additional modifications are described herein.

[0178] In some embodiments, modifications may include chemical or cell-inducible modifications. For example, some non-limiting examples of intracellular RNA modifications are described by Lewis and Pan in “RNA modifications and structures cooperate to guide RNA-protein interactions” from Nat Reviews Mol Cell Biol, 2017, 18:202-210.

[0179] Different sugar modifications, nucleotide modifications, and / or nucleoside bonds (e.g., skeletal structures) can be present at various positions in the sequence. It will be understood by those skilled in the art that nucleotide analogs or other modifications may be located at any position in the sequence such that the function of the sequence is not substantially diminished. The sequence may contain approximately 1% to 100% modified nucleotides (relative to the total nucleotide content, or to one or more types of nucleotides, i.e., A, G, U, or C), or any percentage of modification (e.g., 1% to 20%, 1% to 25%, 1% to 50%, 1% to 60%, 1% to 70%, 1% to 80%, 1% to 90%, 1% to 95%, 10% to 20%, 10% to 25%, 10% to 50%, 10% to 60%, 10% to 70%, 10% to 80%, 10% to 90%, 10% to 90%). It may contain modified nucleotides in the following percentages: 5%, 10%-100%, 20%-25%, 20%-50%, 20%-60%, 20%-70%, 20%-80%, 20%-90%, 20%-95%, 20%-100%, 50%-60%, 50%-70%, 50%-80%, 50%-90%, 50%-95%, 50%-100%, 70%-80%, 70%-90%, 70%-95%, 70%-100%, 80%-90%, 80%-95%, 80%-100%, 90%-95%, 90%-100%, and 95%-100%.

[0180] In some embodiments, sugar modifications (e.g., at the 2' or 4' position), or sugar substitutions in one or more ribonucleotides of a sequence, and skeletal modifications may include modifications or substitutions of phosphodiester bonds. Specific examples of sequences include, but are not limited to, sequences containing a modified skeleton, or sequences containing nucleoside modifications including unnatural nucleoside bonds, such as modifications or substitutions of phosphodiester bonds. Sequences having a modified skeleton include, among other things, those that do not have a phosphorus atom in the skeleton. For the purposes of this application and as is often referred to in the art, modified RNA that does not have a phosphorus atom in the nucleoside skeleton may also be considered an oligonucleoside. In certain embodiments, the sequence contains a ribonucleotide, and the ribonucleotide has a phosphorus atom in its nucleoside skeleton.

[0181] The modified sequence skeleton may include, for example, phosphorothioates; chiral phosphorothioates; phosphorodithioates; phosphotriesters; aminoalkyl phosphotriesters; methyl and other alkylphosphonates, e.g., 3'-alkylene phosphonates and chiral phosphonates; phosphinates; phosphoramides, e.g., 3'-aminophosphoramides and aminoalkylphosphoramides; thionophosphoramides; thionoalkyl phosphonates; thionoalkyl phosphotriesters; and boranophosphates having linear 3'-5' linkages, their 2'-5' linked analogues, and those having opposite polarity with adjacent pairs of nucleoside units linked from 3'-5' to 5'-3' or 2'-5' to 5'-2'. Also included are various salts, mixed salts, and free acid forms. In some embodiments, the sequences may have a negative or positive charge.

[0182] Modified nucleotides that can be incorporated into a sequence may be modified on nucleoside bonds (e.g., phosphate backbone). In this specification, the terms “phosphate” and “phosphodiester” are used interchangeably in the context of polynucleotide backbones. A phosphate group on a backbone may be modified by substituting one or more oxygen atoms with different substituents. Furthermore, modified nucleosides and nucleotides may include extensive substitutions of unmodified phosphate moieties at other nucleoside bonds as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, phosphorodiamidates, alkyl or arylphosphonates, and phosphotryesters. Phosphorodithioates have both sulfur-substituted and unbound oxygen atoms. Phosphate linkers can also be modified by substituting bound oxygen with nitrogen (bridged phosphoramide), sulfur (bridged phosphorothioate), and carbon (bridged methylene phosphonate).

[0183] The α-thio-substituted phosphate moiety is provided to confer stability to RNA and DNA polymers via non-natural phosphorothioate backbone binding. Phosphothioate DNA and RNA exhibit increased nuclease resistance and subsequently have a longer half-life in the cellular environment.

[0184] In specific embodiments, the modified nucleoside includes alpha-thio-nucleosides (e.g., 5'-O-(1-thiophosphate)-adenosine, 5'-O-(1-thiophosphate)-cytidine (α-thiocytidine), 5'-O-(1-thiophosphate)-guanosine, 5'-O-(1-thiophosphate)-uridine, or 5'-O-(1-thiophosphate)-psoidouridine).

[0185] Other nucleoside bonds that may be used in the present invention, including nucleoside bonds that do not contain phosphorus atoms, are described herein.

[0186] In some embodiments, the sequence may contain one or more cytotoxic nucleosides. For example, cytotoxic nucleosides may be incorporated into the sequence, such as through bifunctional modification. Cytotoxic nucleosides may include, but are not limited to, adenosine arabinoside, 5-azacitidine, 4'-thio-aracitidine, cyclopentenylcytosine, cladribine, clofarabine, cytarabine, cytosine arabinoside, 1-(2-C-cyano-2-deoxy-beta-D-arabino-pentofuranosyl)-cytosine, decitabine, 5-fluorouracil, fludarabine, phloxuridine, gemcitabine, combinations of tegafur and uracil, tegafur ((RS)-5-fluoro-1-(tetrahydrofuran-2-yl)pyrimidine-2,4(1H,3H)-dione), troxacitabine, tezacitabine, 2'-deoxy-2'-methylidencytidine (DMDC), and 6-mercaptopurine. Additional examples include fludarabine phosphate, N4-behenoyl-1-beta-D-arabinofuranosylcytosine, N4-octadecyl-1-beta-D-arabinofuranosylcytosine, N4-palmitoyl-1-(2-C-cyano-2-deoxy-beta-D-arabino-pentofuranosyl)cytosine, and P-4055 (cytarabine 5'-elaidic acid ester).

[0187] In some embodiments, the sequence includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, poly-A sequence, methylation, acylation, phosphorylation, methylation and acetylation of lysine and arginine residues, and nitrosylation of thiol and tyrosine residues). One or more post-transcriptional modifications can be any of the more than 100 different nucleoside modifications identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucleoside Acids Res 27: 196-197). In some embodiments, the first isolated nucleic acid includes messenger RNA (mRNA). In some embodiments, mRNA is pyridine-4-onribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-psoiduridine, 2-thio-psoiduridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyluridine, 1-carboxymethyl-psoiduridine, 5-propynyluridine, 1-propynyl-psoiduridine, 5-taurinomethyluridine, 1-taurinomethyl-psoiduridine, 5-taurinomethyl-2-thiouridine, 1-taurinomethyl-4-thiouridine, 5 It comprises at least one nucleoside selected from the group consisting of -methyluridine, 1-methylpsoiduridine, 4-thio-1-methylpsoiduridine, 2-thio-1-methylpsoiduridine, 1-methyl-1-deaz-psoiduridine, 2-thio-1-methyl-1-deaz-psoiduridine, dihydrouridine, dihydropsoiduridine, 2-thio-dihydrouridine, 2-thio-dihydropsoiduridine, 2-methoxyuridine, 2-methoxy-4-thiouridine, 4-methoxypsoiduridine, and 4-methoxy-2-thiopsoiduridine.In some embodiments, mRNA is 5-aza-cytidine, pseudoisocytidine, 3-methylcytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-psoidisocytidine, pyrrolo-cytidine, pyrrolo-psoidisocytidine, 2-thiocytidine, 2-thio-5-methylcytidine, 4-thio-psoidisocytidine, 4-thio-1-methyl-psoidisocytidine, 4- It comprises at least one nucleoside selected from the group consisting of thio-1-methyl-1-deazal-psoidisocytidine, 1-methyl-1-deazal-psoidisocytidine, zebralin, 5-aza-zebralin, 5-methyl-zebralin, 5-aza-2-thio-zebralin, 2-thio-zebralin, 2-methoxy-cytidine, 2-methoxy-5-methylcytidine, 4-methoxy-psoidisocytidine, and 4-methoxy-1-methyl-psoidisocytidine. In some embodiments, mRNA is 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyiso It comprises at least one nucleoside selected from the group consisting of pentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine.In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of inosine, 1-methylinosine, waiosin, waibutosin, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.

[0188] The sequence may or may not be uniformly modified along the entire length of the molecule. For example, one or more or all types of nucleotides (e.g., spontaneously occurring nucleotides, purines, or pyrimidines, or one or more or all of A, G, U, C, I, pU) may or may not be uniformly modified in or within a given given sequence region of the sequence. In some embodiments, the sequence contains pseudouridine. In some embodiments, the sequence contains inosine, which may assist the immune system in characterizing the sequence as endogenous antiviral RNA. Inosine incorporation may also mediate improved RNA stability / reduced degradation. See, for example, Yu, Z. et al. (2015) RNA editing by ADAR1 marks dsRNA as “self”. Cell Res. 25, 1283-1284, which is incorporated in its entirety by reference.

[0189] In some embodiments, any RNA sequence described herein may include terminal modifications (e.g., 5'-terminal modifications or 3'-terminal modifications). In some embodiments, the terminal modifications are chemical modifications. In some embodiments, the terminal modifications are structural modifications. See the disclosures herein for further details.

[0190] When the gene editing systems disclosed herein include nucleic acids encoding CRISPR nucleases, such nucleic acid molecules may, where applicable, contain any of the modifications disclosed herein.

[0191] III. Gene Editing Methods Any of the gene editing systems may be used to genetically modify (edit) a target nucleic acid, which can be a gene site of interest, such as a gene site where gene editing is needed, for example, to repair a gene mutation, to introduce a protective mutation, or to introduce a modification to regulate gene expression.

[0192] The gene editing systems and compositions disclosed herein are applicable to editing and the introduction of editing into various target sequences. In some embodiments, the target sequence is a DNA molecule, for example, a DNA locus (referred herein to as a target sequence or on-target sequence).

[0193] For gene editing systems containing nuclease A (SEQ ID NO: 1) or its manipulated variants, the target sequence is adjacent to a 5'-NGG-3' PAM motif, where N refers to any nucleotide. In some cases, the PAM motif is 3' (downstream) of the target sequence.

[0194] For gene editing systems containing nuclease K (SEQ ID NO: 65) or its manipulated variants, the target sequence is adjacent to a 5'-NGG-3' PAM motif, where N refers to any nucleotide. In some cases, the PAM motif is 3' (downstream) of the target sequence.

[0195] For gene editing systems containing nuclease M (SEQ ID NO: 80) or its manipulated variants, the target sequence is adjacent to the 5'-WTAAH-3' PAM motif, where W is A or T and H is A, C, or T. In one example, the PAM is 5'-TTAAA-3'. In several examples, the PAM motif is 3' (downstream) relative to the target sequence.

[0196] In some embodiments, the target nucleic acid is a genomic region in a cell. In some cases, the target nucleic acid on which gene editing occurs may be located within a protein-coding region. Alternatively, the target nucleic acid may be located within a regulatory region, such as a promoter, enhancer, or 5' or 3' untranslated region. In other cases, the target nucleic acid may be located in a non-coding gene, such as a transposon, miRNA, tRNA, ribosomal RNA, ribozyme, or lincRNA.

[0197] A. Gene editing Any of the gene editing systems disclosed herein may be used to edit a target gene of interest, for example, a gene involved in a disease (e.g., a genetic disorder). In some embodiments, the target gene may be one involved in an immune response in the subject. For example, the target gene may be an immune checkpoint gene or a member of the tumor necrosis factor receptor superfamily. Gene editing may occur in exons (e.g., in coding regions). Alternatively, gene editing may occur in introns or regulatory elements (e.g., promoters, enhancers, inhibitory elements, etc.). In some cases, gene editing may result in a reduction or elimination of the expression of the target gene. In other cases, gene editing may result in an enhancement of the expression of the target gene (e.g., destruction of an inhibitor).

[0198] In some embodiments, methods are provided herein for introducing at least one edit into a target nucleic acid (e.g., a genomic site of interest, e.g., a genomic site in any of the target genes disclosed herein) using a gene editing system described herein.

[0199] As used herein, the term “editing” refers to the introduction of one or more modifications within a nucleotide sequence in a target nucleic acid, for example, within a nucleotide sequence at a genomic site of interest. Editing may occur within a target sequence as defined herein. Alternatively, editing may occur outside a target sequence (for example, adjacent to a target sequence). Editing may consist of one or more substitutions, one or more insertions, one or more deletions, or a combination thereof.

[0200] A deletion refers to the loss of one or more nucleotides in a nucleic acid sequence relative to a reference sequence. No specific process for constructing a sequence containing a deletion is suggested. For example, a sequence containing a deletion may be synthesized directly from individual nucleotides. In other embodiments, the deletion is made by providing a reference sequence and then modifying it. Nucleic acid sequences may be located within the genome of an organism. Nucleic acid sequences may be located within cells. Nucleic acid sequences may be DNA sequences. Deletions may be frameshift mutations or non-frameshift mutations. As used herein, a deletion refers to an insertion of up to several kilobases.

[0201] An insertion refers to the acquisition of one or more nucleotides in a nucleic acid sequence relative to a reference sequence. No specific process for how to construct a sequence containing an insertion is implied. For example, a sequence containing an insertion may be synthesized directly from individual nucleotides. In other embodiments, the insertion is performed by providing a reference sequence and then modifying it. Nucleic acid sequences may be located within the genome of an organism. Nucleic acid sequences may be located within cells. Nucleic acid sequences may be DNA sequences. Insertions may be frameshift mutations or non-frameshift mutations. As described herein, an insertion refers to an insertion of up to several kilobases.

[0202] In some embodiments, the gene editing methods disclosed herein may introduce editing, including substitution, insertion, deletion, or a combination thereof, into a target nucleic acid.

[0203] In some examples, the edit may include at least one substitution, at least one insertion, and / or at least one deletion. In some embodiments, the edit includes at least one substitution, insertion, or deletion. In some embodiments, the substitution, insertion, or deletion is at least 1 to 500 nucleotides (e.g., 1 to 10 nucleotides, 10 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 100 to 200 nucleotides, 200 to 300 nucleotides, 300 to 400 nucleotides, or 400 to 500 nucleotides).

[0204] In some examples, editing may occur within approximately 500 nucleotides of any of the PAM sequences disclosed herein (e.g., PAM sequences associated with nuclease A, nuclease K, or nuclease M disclosed herein). In some embodiments, editing occurs adjacent to the PAM sequence, for example, within approximately 1 to 500 nucleotides upstream or downstream of the PAM sequence. In some embodiments, editing may occur within approximately 1 to 10 nucleotides, 10 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 100 to 200 nucleotides, 200 to 300 nucleotides, 300 to 400 nucleotides, or 400 to 500 nucleotides upstream of the PAM sequence. Alternatively, or in addition, editing may occur within approximately 1 to 10 nucleotides, 10 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 100 to 200 nucleotides, 200 to 300 nucleotides, 300 to 400 nucleotides, or 400 to 500 nucleotides downstream of the PAM sequence.

[0205] In some embodiments, editing begins at the PAM sequence. In some embodiments, editing may begin within approximately 1 to 30 nucleotides downstream of the PAM. Alternatively, editing may begin within approximately 1 to 30 nucleotides upstream of the PAM.

[0206] In some embodiments, the editing may terminate within approximately 1 to 300 nucleotides upstream of the PAM sequence, for example, within approximately 1 to 10 nucleotides, 10 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 100 to 200 nucleotides, or 200 to 300 nucleotides upstream of the PAM sequence. Alternatively, the editing may terminate within approximately 1 to 300 nucleotides downstream of the PAM sequence, for example, within approximately 1 to 10 nucleotides, 10 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 100 to 200 nucleotides, or 200 to 300 nucleotides downstream of the PAM sequence.

[0207] In some embodiments, the editing may terminate in the PAM sequence. In some embodiments, the editing terminates within approximately 1 to 30 nucleotides downstream of the PAM. In other embodiments, the editing may terminate within approximately 1 to 30 nucleotides upstream of the PAM.

[0208] B. Gene editing in cells In some embodiments, methods are provided herein for editing a target genomic site in a cell (e.g., a target gene disclosed herein) using one of the gene editing systems disclosed herein. To carry out this method, the gene editing system may be delivered to a population of cells or introduced into a population of cells. In some cases, cells containing the desired gene edit may be collected, optionally cultured in vitro, and grown.

[0209] The cells described herein can be a variety of cells. In some embodiments, the cells are isolated cells. In some embodiments, the cells are in cell culture or co-culture of two or more cell types. In some embodiments, the cells are ex vivo. In some embodiments, the cells are obtained from living organisms and maintained in cell culture. In some embodiments, the cells are single-celled organisms.

[0210] In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are bacterial cells or derived from bacterial cells. In some embodiments, the cells are archaeal cells or derived from archaeal cells.

[0211] In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are plant cells or derived from plant cells. In some embodiments, the cells are fungal cells or derived from fungal cells. In some embodiments, the cells are animal cells or derived from animal cells. In some embodiments, the cells are invertebrate cells or derived from invertebrate cells. In some embodiments, the cells are vertebrate cells or derived from vertebrate cells. In some embodiments, the cells are mammalian cells or derived from mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the cells are zebrafish cells. In some embodiments, the cells are primate cells. In some embodiments, the cells are rodent cells. In some embodiments, the cells are synthetically produced and are often referred to as artificial cells.

[0212] In some embodiments, the cells are derived from cell lines. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, HEK293T, MF7, K562, HeLa, CHO, and their transgenic variants. Cell lines are available from various sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC) (Manassas, Va.)). In some embodiments, the cells are immortal or immortalized cells. In some embodiments, the cells are stem cells, e.g., totipotent stem cells (e.g., universally pluripotent), pluripotent stem cells, duplexitic stem cells, oligopotent stem cells, or unipotent stem cells. In some embodiments, the cells are induced pluripotent stem cells (iPSCs) or derived from iPSCs. In some embodiments, the cells are mesenchymal stem cells. In some embodiments, the cells are embryonic stem cells. In some embodiments, the cells are hematopoietic stem cells. In some embodiments, the cells are differentiated cells. For example, in some embodiments, differentiated cells are muscle cells (e.g., myocytes), fat cells (e.g., adipocytes), bone cells (e.g., osteoblasts, osteocytes, osteoclasts), blood cells (e.g., monocytes, lymphocytes, neutrophils, eosinophils, basophils, macrophages, erythrocytes, or platelets), nerve cells (e.g., neurons), epithelial cells, immune cells (e.g., lymphocytes, neutrophils, monocytes, or macrophages), liver cells (e.g., hepatocytes), fibroblasts, or sex cells. In some embodiments, the cells are terminally differentiated cells. For example, in some embodiments, terminally differentiated cells are neuronal cells, adipocytes, cardiomyocytes, skeletal muscle cells, epidermal cells, or intestinal cells. In some embodiments, the cells are glial cells. In some embodiments, the cells are islet cells, which include alpha cells, beta cells, delta cells, or enterochromaffin cells. In some embodiments, the cells are immune cells. In some embodiments, the immune cells are T cells.In some embodiments, the immune cells are B cells. In some embodiments, the immune cells are natural killer (NK) cells. In some embodiments, the immune cells are tumor-infiltrating lymphocytes (TILs). In some embodiments, the cells are mammalian cells, such as human cells or primate cells or mouse cells. In some embodiments, the mouse cells are derived from wild-type mice, immunosuppressed mice, or disease-specific mouse models. In some embodiments, the cells are living tissue, organs, or cells within an organism.

[0213] In some embodiments, the cells are primary cells. For example, a culture of primary cells may be passaged 0, 1, 2, 4, 5, 10, or 15 or more times. In some embodiments, primary cells are collected from an individual by any known method. For example, leukocytes may be collected by apheresis, leukocyte apheresis, density gradient separation, etc. Cells from tissues, such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, etc., may be collected by biopsy. A suitable solution may be used to disperse or suspend the collected cells. Such a solution may generally be an equilibrium salt solution (e.g., ordinary saline, phosphate-buffered saline (PBS), Hanks equilibrium salt solution, etc.) conveniently complemented with fetal bovine serum or other spontaneously occurring factors along with an acceptable buffer at a low concentration. The buffer may include HEPES, phosphate buffer, lactate buffer, etc. The cells may be used immediately or may be stored (e.g., by freezing). Frozen cells may be thawed and reused. Cells can be frozen in DMSO, serum, culture buffer (e.g., 10% DMSO, 50% serum, 40% buffered medium), and / or other common solutions used to store cells at freezing temperatures.

[0214] In embodiments in which the gene editing system disclosed herein is introduced into a plurality of cells, at least about 0.5% of the cells contain the desired editing. In some embodiments, at least about 1% of the cells contain the desired editing. In some embodiments, at least about 2% of the cells contain the desired editing. In some embodiments, at least about 3% of the cells contain the desired editing. In some embodiments, at least about 4% of the cells contain the desired editing. In some embodiments, at least about 5% of the cells contain the desired editing. In some embodiments, at least about 10% of the cells contain the desired editing. In some embodiments, at least about 20% of the cells contain the desired editing. In some embodiments, at least about 30% of the cells contain the desired editing. In some embodiments, at least about 40% of the cells contain the desired editing. In some embodiments, at least about 50% of the cells contain the desired editing.

[0215] Cells possessing desired gene editing, for example, cells produced using any of the gene editing systems also disclosed herein by the methods disclosed herein, are also within the scope of this disclosure. In some cases, cells modified with the CRISPR nuclease polypeptides disclosed herein may be useful as expression systems for producing biomolecules. For example, modified cells may be useful for producing biomolecules, such as proteins (e.g., cytokines, antibodies, antibody-based molecules), peptides, lipids, carbohydrates, nucleic acids, amino acids, and vitamins. In other embodiments, modified cells may be useful in the production of viral vectors, such as lentiviruses, adenoviruses, adeno-associated viruses, and oncolytic viral vectors. In some embodiments, modified cells may be useful in cytotoxicity studies. In some embodiments, modified cells may be useful as disease models. In some embodiments, modified cells may be useful in vaccine production. In some embodiments, modified cells may be useful in therapeutics. For example, in some embodiments, modified cells may be useful in cell therapies, such as infusions and transplants.

[0216] In some embodiments, cells modified by the CRISPR nuclease polypeptides disclosed herein may be useful for establishing new cell lines containing modified genomic sequences. In some embodiments, the modified cells of the Disclosure are modified stem cells (e.g., modified totipotent / pluripotent stem cells, modified pluripotent stem cells, modified duplexipotent stem cells, modified oligopotent stem cells, or modified unipotent stem cells) which differentiate into one or more cell lineages containing modified stem cell deletions. The Disclosure further provides organisms (e.g., animals, plants, or fungi) containing or produced from the modified cells of the Disclosure.

[0217] Delivery of gene editing systems to C. cells In some embodiments, the gene editing systems disclosed herein or any of their components may be formulated, for example, comprising a carrier, e.g., a carrier and / or polymer carrier, e.g., liposomes or lipid nanoparticles, and delivered to cells (e.g., prokaryotes, eukaryotes, plants, mammals, etc.) by known methods. Such methods include, but are not limited to, transfection (e.g., lipid-mediated, cationic polymers, calcium phosphate, dendrimers), electroporation or other methods of membrane disruption (e.g., nucleofection), viral delivery (e.g., lentiviruses, retroviruses, adenoviruses, AAVs), microinjection, microparticle guns ("gene guns"), fugene, direct sonic loading, cell squeezing, phototransfection, protoplast fusion, imparefection, magnetofection, exosome-mediated transfer, lipid nanoparticle-mediated transfer, and any combination thereof.

[0218] In some embodiments, the method comprises delivering one or more nucleic acids (e.g., CRISPR nuclease polypeptides and / or gRNAs and / or nucleic acids encoding pre-formed ribonucleoproteins) to cells. Exemplary intracellular delivery methods include viruses or viral drugs; chemical-based transfection methods, e.g., using calcium phosphate, dendrimers, liposomes, or cationic polymers (e.g., DEAE-dextran or polyethyleneimine); and non-chemical methods, e.g., microinjection, electroporation, cell squeezing, sonoporation, phototransfection, imparefection, protoplast fusion, bacterial conjugation, plasmids, or transfections. Delivery of posons; particle-based methods, e.g., gene guns, magnetofection or magnetic-assisted transfection, use of particle guns; and hybrid methods, e.g., nucleofection. In some embodiments, the present application further provides cells produced by such methods, and organisms (e.g., animals, plants, or fungi) containing or produced from such cells. In some embodiments, the compositions of the present invention are further delivered together with agents (e.g., compounds, molecules, or biomolecules) that affect DNA repair or DNA repair mechanisms. In some embodiments, the compositions of the present invention are further delivered together with agents (e.g., compounds, molecules, or biomolecules) that affect the cell cycle.

[0219] In some embodiments, a first composition containing a CRISPR nuclease polypeptide is delivered to cells. In some embodiments, a second composition containing gRNA is delivered to cells. In some embodiments, the first composition contacts cells before the second composition contacts the cells. In some embodiments, the first composition contacts cells at the same time as the second composition contacts the cells. In some embodiments, the first composition contacts cells after the second composition has contacted the cells. In some embodiments, the first composition is delivered by a first delivery method, and the second composition is delivered by a second delivery method. In some embodiments, the first delivery method is the same as the second delivery method. For example, in some embodiments, the first and second compositions are delivered via viral delivery. In some embodiments, the first delivery method is different from the second delivery method. For example, in some embodiments, the first composition is delivered by viral delivery and the second composition is delivered by lipid nanoparticle-mediated transfer, and the second composition is delivered by viral delivery, or the first composition is delivered by lipid nanoparticle-mediated transfer and the second composition is delivered by viral delivery.

[0220] In some cases, any of the gene editing systems provided herein comprises one or more lipid excipients, which associate with components of the gene editing system to facilitate delivery of the gene editing system to host cells. In some cases, one or more lipid excipients may form lipid nanoparticles (LNPs), which associate with or encapsulate components of the gene editing system (e.g., CRISPR nuclease polypeptides and mRNA molecules encoding gRNA). Such LNP-containing gene editing systems can come into contact with host cells (e.g., be administered to subjects requiring gene editing), and the LNPs can facilitate delivery of the gene editing system to host cells.

[0221] In other instances, the gene editing systems provided herein may be delivered to host cells via a viral vector-mediated method. Such gene editing systems may include a viral vector, which may contain a transgene encoding a CRISPR nuclease polypeptide and, optionally, a transgene encoding a gRNA. In some cases, the viral vector may be an AAV vector. The viral vector may facilitate the delivery of the transgene(s) into the host cell, where the transgene(s) produce a CRISPR nuclease polypeptide and, optionally, a gRNA to perform gene editing on the target gene.

[0222] IV. Therapeutic uses Any gene editing system disclosed herein, or any modified cells produced using such a gene editing system, may be used to treat diseases that may be benefits from gene editing introduced by the gene editing system or possessed by the modified cells. For example, the disease may be a genetic disease, and gene editing repairs gene mutations associated with the genetic disease. Alternatively, the disease may be associated with abnormal gene expression, and gene editing rescues such abnormal expression.

[0223] In some embodiments, methods for treating a disease are provided herein, comprising administering one of the gene editing systems disclosed herein to a subject in need of treatment (e.g., a human patient). The gene editing system may be delivered to a specific tissue or a specific cell type in which gene editing is required. The gene editing system may include LNPs, which comprise one or more components, one or more vectors (e.g., viral vectors) encoding one or more components, or a combination thereof. The components of the gene editing system may be formulated to form a pharmaceutical composition, which may further comprise one or more pharmaceutically acceptable carriers.

[0224] In some embodiments, modified cells produced using any of the gene editing systems disclosed herein may be administered to subjects in need of treatment (e.g., human patients). Modified cells may include substitutions, insertions, and / or deletions described herein. In some examples, modified cells may include cell lines modified with CRISPR nuclease polypeptides and gRNAs disclosed herein. In some cases, modified cells may be a heterogeneous population, which may include cells with different types of gene editing. Alternatively, modified cells may include a substantially homogeneous cell population (e.g., at least 80% of the cells in the whole population), which may include one specific gene editing. In some examples, cells may be suspended in a suitable culture medium.

[0225] In some embodiments, gene editing systems or components thereof, or compositions comprising modified cells are provided herein. Such compositions may be pharmaceutical compositions. Useful pharmaceutical compositions may be prepared, packaged, or marketed in formulations suitable for oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, lesional, buccal, ocular, intravenous, intra-organal, or other routes of administration. Pharmaceutical compositions of this disclosure may be prepared, packaged, and / or marketed in bulk as single unit doses or as multiple single unit doses. As used herein, “unit dose” is a distinct amount of a pharmaceutical composition containing a predetermined number of cells. The number of cells is generally equal to the dose of cells administered to a subject, or a convenient fraction of such a dose, for example, half or one-third of such a dose.

[0226] Formulations of pharmaceutical compositions suitable for parenteral administration may include an activator (e.g., a gene editing system or its components, or modified cells) combined with a pharmaceutically acceptable carrier, such as sterile water or sterile isotonic saline. Such formulations may be prepared, packaged, or sold in forms suitable for bolus or serial administration. Some injectable formulations may be prepared, packaged, or sold in unit dosage forms, for example, in ampoules or multi-dose containers containing preservatives. Some formulations for parenteral administration include, but are not limited to, suspensions, solutions, emulsions in oily or aqueous vehicles, pastes, and embedded sustained-release or biodegradable formulations. Some formulations may further include, but are not limited to, one or more additional components, including suspending agents, stabilizers, or dispersants.

[0227] Pharmaceutical compositions may be in the form of sterile, injectable aqueous or oily suspensions or solutions. These suspensions or solutions may be formulated by known techniques and may contain, in addition to cells, additional components, such as dispersants, wetting agents, or suspending agents as described herein. Such sterile, injectable formulations may be prepared using non-toxic, parenterally acceptable diluents or solvents, such as water or saline. Other acceptable diluents and solvents include, but are not limited to, Ringer's solution, isotonic sodium chloride solution, and fixative oils such as synthetic mono or diglycerides. Other useful parental administerable formulations include those that may contain cells in packaged form, within liposome preparations, or as components of biodegradable polymer systems. Some compositions for sustained release or embedding may contain pharmaceutically acceptable polymers or hydrophobic materials, such as emulsions, ion exchange resins, sparingly soluble polymers, or sparingly soluble salts.

[0228] V. Kits and their use This disclosure also provides kits or systems that can be used, for example, to carry out the methods described herein. In some embodiments, the kit or system comprises a CRISPR nuclease polypeptide and, optionally, a gRNA. In some embodiments, the kit or system comprises a polynucleotide, the polynucleotide encoding the CRISPR nuclease polypeptide and, optionally, encoding the gRNA. The gRNA in the kit may be designed to target a sequence of interest. The CRISPR nuclease polypeptide and gRNA may be packaged in the same vial or other container within the kit or system, or in separate vials or other containers, and their contents may be mixed before use. In addition, the kit or system may optionally include a buffer, and / or instructions on how to use the CRISPR nuclease polypeptide and gRNA.

[0229] In some embodiments, the kit comprises a first composition comprising a CRISPR nuclease polypeptide disclosed herein. In some embodiments, the kit comprises a second composition comprising a gRNA also disclosed herein. In some embodiments, the first and second compositions are packaged in the same vial. In some embodiments, the first and second compositions are packaged in different vials.

[0230] In some embodiments, the kit may be useful for research purposes. For example, in some embodiments, the kit may be useful for studying gene function.

[0231] General technology Unless otherwise specified, the implementation of this invention will utilize conventional techniques within the scope of the art, including molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology. Such techniques are described, for example, in Molecular Cloning: A Laboratory Manual, second edition (Sambrook, et al., 1989) Cold Spring Harbor Press, Oligonucleotide Synthesis (MJ Gait, ed. 1984), Methods in Molecular Biology, Humana Press, Cell Biology: A Laboratory Notebook (JECellis, ed., 1989) Academic Press, Animal Cell Culture (RIFreshney, ed. 1987), Introduction to Cell and Tissue Culture (JP Mather and PERoberts, 1998) Plenum Press, Cell and Tissue Culture: Laboratory Procedures (A. Doyle, JBGriffiths, and DG Newell, eds. 1993-8) J. Wiley and Sons, Methods in Enzymology (Academic Press, Inc.), Handbook of Experimental Immunology (DM Weir and CC Blackwell, eds.): Gene Transfer Vectors for Mammalian Cells (JMMiller and MPCalos, eds., 1987), Current Protocols in Molecular Biology (FMAusubel, et al. eds. 1987), PCR: The Polymerase Chain Reaction, (Mullis, et al., eds. 1994), Current Protocols in Immunology (JEColigan et al., eds., (1991), Short Protocols in Molecular Biology (Wiley and Sons, 1999), Immunobiology (C.A. Janeway and P. Travers, 1997), Antibodies (P. Finch, 1997), Antibodies: a practice approach (D. Catty., ed., IRL Press, 1988 - 1989), Monoclonal antibodies: a practical approach (P. Shepherd and C. Dean, eds., Oxford University Press, 2000), Using antibodies: a laboratory manual (E. Harlow and D. Lane (Cold Spring Harbor Laboratory Press, 1999), The Antibodies (M. Zanetti and J.D. Capra, eds. Harwood Academic Publishers, 1995), DNA Cloning: A practical Approach, Volumes I and II (D.N. Glover ed. 1985), Nucleic Acid Hybridization (B.D. Hames & S.J. Higgins eds. (1985>>, Transcription and Translation (B.D. Hames & S.J. Higgins, eds. (1984>>, Animal Cell Culture (R.I. Freshney, ed. (1986>>, Immobilized Cells and Enzymes (IRL Press, (1986>>, and are fully described in the literature such as B. Perbal, A practical Guide To Molecular Cloning (1984), F.M. Ausubel et al. (eds.).

[0232] Without further detail, those skilled in the art will likely be able to make the most of the use of the present invention based on the above description. Accordingly, the following specific embodiments should be construed as merely illustrative and in no way limit the remainder of this disclosure. All publications referenced herein are incorporated by reference for the purposes or subjects referenced herein. [Examples]

[0233] Example 1: Editing of human target genes in HEK293T cells using CRISPR nuclease polypeptide derived from nuclease A. This example describes genome editing of the AAVS1, EMX1, and VEGFA genes by introducing nuclease A (SEQ ID NO: 1) or a variant thereof into the HEK293T cell line via lipid-based transient transfection.

[0234] Nuclease A was tagged with the N-terminal SV40 nuclear localization signal (NLS) and the C-terminal nucleoplasmin nuclear localization signal (NLS). Its coding sequence was converted to a human codon-optimized DNA sequence, synthesized, and cloned into a pcDNA3.1 vector (Invitrogen) containing a CMV promoter for expression. The reference and NLS-tagged sequences used are shown in Table 1. Plasmids were purified using a midiprep kit.

[0235] RNA guides were designed, cloned into the pUC19 plasmid, and subsequently terminated with a 6x poly-T sequence. The RNA guides were designed to be specific to target sequences within the coding exons of AAVS1, EMX1, and VEGFA, which possess a 5'-NGG-3'PAM sequence (the PAM sequence being at the 3' end of the target sequence). See Table 2 for all RNA guide sequences. For more efficient transcription, the U6 PolIII promoter used +1G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences listed in Table 2. Plasmids were purified using a midiprep kit. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 2]

[0236] Approximately 16 hours prior to transfection, 25,000 HEK293T cells were seeded in DMEM / 10% FBS+Pen / Strep (D10 medium) into each well of a 96-well plate. On the day of transfection, the cells were 50–90% confluent. For each well to be transfected, a mixture of Lipofectamine 2000® (ThermoFisher Scientific) and Opti-MEM® (ThermoFisher Scientific) was prepared and incubated at room temperature for 5 minutes (Solution 1). After incubation, the Lipofectamine 2000®:Opti-MEM® mixture was added to a separate mixture containing a CRISPR nuclease plasmid (NLS tagged), an RNA guide plasmid, and Opti-MEM® (Solution 2). In the case of a negative control, the CRISPR nuclease plasmid was excluded. Solutions 1 and 2 were mixed by pipetting up and down, and then incubated at room temperature for 25 minutes. After incubation, the mixture of solutions 1 and 2 was added dropwise to each well of a 96-well plate containing cells. Approximately 72 hours after transfection, the cells were trypsinized by adding TrypLE® (ThermoFisher Scientific) to the center of each well and incubating at 37°C for approximately 5 minutes. D10 medium was then added to each well and mixed to resuspend the cells. The resuspended cells were centrifuged for 10 minutes to obtain a pellet, and the supernatant was discarded. The cell pellet was then resuspended in QuickExtract® buffer (Lucigen®), and the cells were incubated at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes.

[0237] Next-generation sequencing (NGS) samples were prepared by two rounds of PCR. Three technical replicates were analyzed for each target, reference, and variant. The first round (PCR1) was used to amplify specific genomic regions in a target-dependent manner. A second round of PCR (PCR2) was performed to add the Illumina adapter and index. The reaction products were then pooled and purified by column purification. Sequencing runs were performed using kits, e.g., 150-cycle NextSeq 500 / 550 medium or high-power v2.5 kits.

[0238] For NGS analysis, the indel mapping function used the sample's fastq file, amplicon reference sequence, and forward primer sequence. For each read, the KMER scan algorithm was used to calculate the editing behavior (match, mismatch, insertion, deletion) between the read and the reference sequence. To remove small amounts of primer dimers present in some samples, the first 30 nt of each read was required to match the reference, and reads with more than half of the mapping nucleotides being mismatched were filtered out. Up to 50,000 reads that passed these filters were used for analysis, and if a read contained an insertion or deletion, it was counted as an indel read. % indel was calculated by dividing the number of indel-containing reads by the number of reads analyzed (the maximum number of reads that passed the filter was 50,000). The QC criterion for the minimum number of reads that passed the filter was 10,000.

[0239] For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for each sample and its related protein-free control. The higher percentage of indels in targets when CRISPR nucleases were included in transfection demonstrated the DNA editing results in cells.

[0240] As shown in Figure 1, four of the six targets tested demonstrated that larger levels of indels were observed when the nuclease A (SEQ ID NO: 1) plasmid was present.

[0241] Example 2: Efficacy of CRISPR nuclease A variant for targeting mammalian genes This example describes indel evaluation for a mammalian target using a variant of CRISPR nuclease A (SEQ ID NO: 1) transfected into HEK293T cells.

[0242] Arginine phylogenetic mutagenesis was performed to individually substitute each non-arginine residue of nuclease A (SEQ ID NO: 1) with arginine. SEQ ID NO: 1 is referred herein to as the nuclease A reference sequence. This resulted in 710 single-arginine substitution variants. The nucleic acids encoding nuclease A and each CRISPR nuclease variant were then individually cloned into a pcda3.1 backbone (Invitrogen), and the plasmids were maxi-prepped and diluted. The plasmids contained a CMV promoter, a first NLS upstream of the coding sequence (KRTADGSEFESPKKKRKV; SEQ ID NO: 3), an XTEN linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS; SEQ ID NO: 95), followed by a second NLS downstream of the coding sequence (KRPAATKKAGQAKKKK; SEQ ID NO: 4). See also Example 1 above.

[0243] The RNA guide and target sequences are shown in Table 3. The RNA guide was cloned into the pUC19 backbone (New England Biolabs®), followed by the U6 PolIII promoter, and terminated with a 6x polyT sequence. For more efficient transcription, the U6 PolIII promoter was configured to use +1G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences listed in Table 3. The plasmids were then maxi-prepped and diluted. [Table 3]

[0244] HEK293T cells were transfected with plasmids expressing the CRISPR nuclease polypeptide and gRNA provided herein, according to the method provided in Example 1 above. Gene editing efficiency was determined by next-generation sequencing, according to the method also provided in Example 1.

[0245] For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for both the reference and each variant. The indel ratio used for calculating the multiplier change was the average of the two technical replicates. Then, to calculate the multiplier change in the indel ratio for a particular target, the indel ratio for each variant was divided by the indel ratio for the reference. Table 4 shows the multiplier change in the indel ratio for each target and as an average for both targets. The numbering is relative to the reference nuclease of Sequence ID No. 1 (i.e., without NLS). For example, in the AAVS1 target, the indel ratio for the I67R variant was 7.07 times that of the reference indel ratio, and in the VEGFA target, the indel ratio for the I67R variant was 4.43 times that of the reference indel ratio. The average multiplier change in the indel ratio for the I67R variant in both targets was 5.75 times that of the reference indel ratio in both targets.

[0246] As shown in Table 4, the 21 variants with a single arginine substitution (left column) were characterized as resulting in an increase in the indel ratio of at least 1.5 times compared to the reference indel ratio when averaged across both targets (right column). [Table 4]

[0247] The following 155 variants were analyzed as having indel ratios between 1 and 1.5 times the reference indel ratio when averaged across both targets. E206R, K74R, S280R, A319R, I96R, A354R, D323R, F265R, E120R, H442R, V284R, E364R, C653R, A365R, N586R, E294R, E490R, N459R, E460R, E 144R, S95R, N571R, S587R, P708R, K163R, G485R, V387R, I711R, W148R, K100R, G722R, A132R, Q328R, T201R, K322R, Q277R, N87R, K378R, N75 8R, I563R, G330R, D207R, K86R, G359R, K360R, K49R, T741R, I519R, E248R, E299R, K333R, Q41R, V12R, E152R, V210R, D53R, Q380R, M105R, K 145R, S548R, S761R, K28R, V390R, K17R, D537R, K585R, Q772R, K351R, T255R, K287R, K146R, G66R, A441R, E491R, K302R, N340R, T98R, T83R, S752R, E149R, A55R, S650R, K355R, N489R, N457R, I753R, E94R, V774R, K572R, E371R, K566R, N462R, D561R, K366R, L369R, K546R, Y693R, S 368R, K773R, Q16R, P771R, G385R, S180R, E11R, K614R, V10R, K422R, E129R, H243R, G209R, Q436R, K770R, K737R, K767R, K579R, K764R, G97R , K224R, P709R, L8R, K298R, K750R, Q177R, D533R, D356R, S170R, T395R, K449R, K252R, Y271R, E251R, W141R, K714R, T726R, K167R, I370R, K184R, K91R, E418R, D173R, Q295R, E131R, K350R, N332R, Y9R, A242R, K288R, E142R, N154R, F775R, H600R, L266R, E438R, V749R, and S134R. The remaining variants (534 variants) are as follows:This resulted in a decreased indel ratio relative to the reference indel ratio (a magnification change in indel ratios less than 1.0). G488R, E300R, K348R, K424R, P437R, K557R, S729R, K702R, G713R, E128R , V654R, S547R, T233R, A135R, T542R, T766R, K386R, T13R, Q64R, E453R, H293R, L543R, D536R, K75R, K363R, P560R, Q140R, K143R, K467R, P192R, M15R, K123R, L245R, K5R, T297R, K278R, K76R, S235R, H621R, E367R, K689 R, D246R, Y534R, I756R, V768R, K268R, L434R, T391R, K139R, S465R, E23 6R, K6R, K82R, K665R, D7R, K138R, Y29R, V119R, E137R, N39R, E556R, N379 R, C4R, K247R, T151R, A238R, S742R, T126R, D446R, V372R, Y471R, E357R , N549R, T392R, Q234R, I130R, L701R, P698R, G538R, D272R, I678R, E710R , N517R, M133R, K613R, D159R, E292R, K626R, E552R, E430R, T562R, G516 R, K311R, M594R, D164R, A274R, T535R, P122R, T70R, F275R, N677R, K687R , N263R, A747R, L559R, S570R, S155R, G550R, K237R, D52R, K620R, G751R , K254R, T763R, T440R, P575R, N196R, K109R, E495R, K188R, K216R, F304R , S769R, V205R, Q269R, L291R, G249R, E633R, K628R, E239R, K691R, K644 R, K712R, K715R, A760R, N667R, P493R, I226R, E176R, G754R, L684R, H260 R, D554R, H727R, N276R, S625R, Q637R, N315R, E160R, A403R, N3R, V433R , K696R, K202R, K402R, D609R, L200R, F229R, V744R, L429R, E18R, T718R,I607R、T73R、H296R、L374R、L78R、D481R、V748R、Q189R、E147R、S583R、H181R、E728R、E150R、E171R、A174R、D220R、K690R、K532R、Y724R、E124R、S759R、Y717R、H703R、P14R、K707R、A681R、L43R、L716R、C629R、L19R、K204R、G740R、T106R、K125R、T175R、E421R、M121R、K113R、K602R、D136R、G361R、I487R、K193R、E617R、G349R、L326R、G623R、K616R、E610R、I695R、E686R、Q497R、E746R、E597R、H558R、N445R、I168R、K313R、E241R、G688R、Q642R、D518R、V733R、K544R、D93R、H362R、L279R、N187R、S732R、K569R、L214R、D225R、V244R、P308R、N464R、G646R、P231R、L412R、I57R、F565R、W508R、Y504R、T162R、Y190R、E730R、D731R、G573R、H478R、G404R、N668R、Q178R、W622R、D540R、S431R、L221R、I199R、T590R、L501R、M541R、G627R、D685R、G601R、I439R、H112R、N407R、V705R、D454R、A344R、N647R、W476R、T631R、N217R、D632R、E80R、N101R、V539R、L486R、E324R、I335R、N211R、N645R、L110R、L492R、V253R、S498R、F409R、G639R、L447R、F195R、V20R、P555R、K738R、V161R、I223R、T638R、T283R、S608R、Y92R、E514R、P428R、T624R、L58R、V218R、K448R、I179R、M458R、I88R、I183R、A692R、Y577R、I325R、L222R、A45R、E102R、A670R、G89R、F329R、T427R、T574R、E634R、D117R、G432R、L232R、V303R、C381R、W660R、D36R、G281R、I584R、A307R、D604R、G679R、G510R、P327R、N316R、L318R、I262R、V227R、F256R、Y8 4R、N580R、L443R、V320R、F257R、D723R、V54R、F553R、D342R、L353R、T212R、M 649R、Y116R、G425R、F104R、H157R、Y25R、A503R、I666R、I314R、I589R、G405 R、L230R、T312R、I500R、I290R、N40R、Y44R、P435R、P103R、L582R、Y468R、V16 5R、P213R、C419R、G90R、F99R、F704R、F505R、F648R、L306R、A618R、L463R、I 388R、N197R、A522R、P545R、H676R、F285R、Y118R、N669R、Y734R、N697R、V680 R, H520R, L373R, I273R, K513R, L240R, F671R, C203R, E376R, L461R, N394R, L719R, I185R, M479R, V682R, A345R, E603R, T47R, F619R, A762R, L576R, D578 R、A410R、I451R、I494R、L81R、I331R、V408R、C208R、S530R、Y383R、K423R、M 469R、F511R、Y108R、V321R、I270R、V382R、S309R、L651R、T377R、N674R、F640 R、L658R、G31R、V107R、A455R、K672R、W509R、I23R、P743R、V42R、T596R、L45 2R、M169R、H289R、M336R、L630R、T30R、V34R、V35R、G22R、G659R、L85R、C111R 、G115R、V317R、V725R、Y683R、V595R、M50R、L615R、I736R、V526R、A529R、I6 06R、Q662R、A414R、C384R、I406R、V48R、L700R、Y611R、I466R、H397R、A635R、 L523R, T347R, F399R, I527R, G657R, C415R, A525R, V664R, L310R, V182R, D524R, F186R, E337R, V675R, L765R, L483R, L32R, K450R, S473R, G499R, P581RL21R, H521R, D396R, C416R, I398R, Y641R, T673R, L739R, G27R, V413R, M663R, L745R, P400R, C286R, T389R, N339R, A393R, I474R, T502R, N411R, A33R, F341R, S338R, G26R, L528R, G475R, N420R, D24R, Y358R, V334R, A472R, I343R, and S470R.

[0248] Based on this experiment, the following substitutions, I67R, E63R, D56R, T605R, D568R, E694R, G71R, E706R, and E655R, were selected for further manipulation of nuclease A of SEQ ID NO: 1.

[0249] Example 3: Efficacy of CRISPR nuclease A combination variant for targeting mammalian genes This example describes indel evaluation for mammalian targets using nuclease A variants containing two or more substitutions identified in Example 2 as increasing indel activity. 106 combination nuclease A variants were tested.

[0250] Each nuclease A variant and RNA guide (see Table 5) was cloned as described in Example 2. HEK293T cells were further transfected as described in Example 2, followed by NGS analysis. For each target, the indel ratio, which represents the percentage of NGS reads containing indels, was calculated for nuclease A (SEQ ID NO: 1) and each variant of nuclease A. The indel ratios shown in Table 6 were calculated as the average of two biological replicas, each containing two technical replicas. [Table 5] [Table 6-1] [Table 6-2] [Table 6-3]

[0251] As shown in Table 6, all but one of the nuclease A variants with amino acid substitution combinations showed higher indel activity than nuclease A (SEQ ID NO: 1). The eight CRISPR nuclease variants yielded an indel ratio of over 0.3 when averaged across three targets, indicating that more than 30% of the NGS reads contained indels. These eight nuclease variants of nuclease A included the following substitution combinations: a) I67R, D568R, and E706R; b) I67R, D56R, and D568R; c) I67R, T605R, and D568R; d) I67R, D56R, and E706R; e) I67R, T605R, and E706R; f) I67R, D568R, and E694R; g) I67R, D56R, E694R, and E706R; and h) I67R, D56R, T605R, and E706R. The 64 variants of nuclease A yielded an indel ratio greater than 0.2 when averaged across three targets, indicating that more than 20% of the NGS reads contained indels. The 29 variants of nuclease A yielded an indel ratio greater than 0.1 when averaged across three targets, indicating that more than 10% of the NGS reads contained indels.

[0252] Based on this experiment, nuclease A variants containing the following substitutions, I67R, D568R, and E706R, were selected for further testing. These nuclease A variants showed more than an 8-fold increase in indel activity compared to the reference nuclease A (SEQ ID NO: 1).

[0253] Editing of Human Target Genes in HEK293T Cells by RNA-Guided Nuclease A Variants, and RNA Scaffold Sequences This example describes the genome editing of exemplary target genes by nuclease A of SEQ ID NO: 1 and I67R, D568R, E706R nuclease A variants of SEQ ID NO: 6 when combined with RNA guides containing alternative scaffold sequences.

[0254] Unless otherwise stated, RNA guides were designed using the RNA scaffold sequences in Table 7, cloned into the pUC19 plasmid, followed by cloning into the U6 PolIII promoter and terminated with a 6x polyT sequence. The RNA guides were designed to be specific to the AAVS1-T3, EMX1-T7, and / or AAVS1-T3b target sequences (see Example 1 and Table 7). See all RNA guide sequences in Table 8. U6 PolIII uses +1G at the start of the transcript (i.e., the 5' end of the RNA) for more efficient transcription, which is excluded from the sequences described in Table 8. Industrial-grade plasmids were received from GenScript.

Table 7

Table 8

[0255] The RNA guides in Table 8 were combined with either nuclease A or the I67R / D568R / E706R variant of nuclease A and introduced into HEK293T cells by lipid-based transient transfection as described in Example 1. The AAVS1-T3 and EMX1-T7 RNA guides from Example 1 were further used as controls. Genomic DNA was harvested approximately 72 hours after transfection, the samples were prepared for NGS, and analyzed as described in Example 1. The percentage of NGS reads containing the indels shown in Table 9 was calculated as the mean of two biological replicates, each of which included two technical replicates. Unless otherwise indicated, the percentage of NGS reads containing the indels shown in Table 10 was calculated as the mean of three technical replicates.

Table 9

Table 10

[0256] As shown in Table 9, RNA guides containing the cleavage-type reference scaffold showed higher indel activity when paired with reference nuclease A than RNA guides containing the reference scaffold. As shown in Table 10, RNA guides containing the reference scaffold or scaffold 2 each showed high indel activity with the I67R, D568R, E706R nuclease A variant. In all cases, the observed mean percent indels were significantly higher than the background controls containing either reference nuclease A or the I67R, D568R, E706R nuclease A variant in the absence of the RNA guide.

[0257] Therefore, this example shows that the cleavage-type reference scaffold and scaffold 2 can be recognized by nuclease A (SEQ ID NO: 1) and the nuclease A variant (SEQ ID NO: 6).

[0258] Example 5: Manipulation and efficacy of CRISPR nuclease A nickerse variant for targeting mammalian genes This example describes the production of a functional nickase by introducing a mutation that disrupts either the HNH or RuvC domain into nuclease A of SEQ ID NO: 1. D396, H397, and N420 were identified as putative catalytic residues of the HNH domain. D24, E337, and D524 were identified as putative catalytic residues of the RuvC domain. These positions were identified by analyzing structural regions similar to known HNH and RuvC active sites using a model generated with AlphaFold2 (Jumper et al., Nature 596:583-9 (2021)), and / or by performing sequence alignment with other nucleases whose candidate sites had been previously identified. Examples of reference structures used to identify the HNH and RuvC active sites are represented by the following PDB IDs: 5h0m, 7eu9, 6ltu, 7odf, 7lys, 8dc2, 4cmp, 4oo8, 7z4j, 5axw, 5b2o, 6kc8, 7utn, 8csz, 8ctl, and 8dmb.

[0259] The coding sequence for nuclease A was converted to an E. coli codon-optimized DNA sequence, synthesized, and cloned into a pET-28a(+) vector (Novagen) containing lac and T7 RNA polymerase promoters for gene expression. To test nickase activity, individual alanine variants were cloned for each of the positions identified as putative active site residues in the HNH and RuvC domains. A leucine variant was also cloned for position H397. Research-grade plasmids were received from GenScript. The manipulated nickase sequences are shown in Table 11. Codons encoding substituted residues are shown in underlined bold capital letters in the nucleotide sequence, and substituted residues are shown in underlined bold letters in the amino acid sequence. The putative HNH knockout nickase was expected to cleave the non-target strand but not the target strand. The putative RuvC knockout nickase was expected to cleave the target strand but not the non-target strand. [Table 11-1] [Table 11-2] [Table 11-3] [Table 11-4] [Table 11-5] [Table 11-6] [Table 11-7] [Table 11-8]

Table 11-9

[0260] A linear DNA template sequence encoding the RNA guide was designed and ordered (IDT), and the linear DNA template sequence had a T7 promoter upstream of the guide and a T7 Te terminator sequence downstream of the guide. The RNA guide was designed to be specific for a previously tested target sequence described in Example 1 within the coding exon of AAVS1 that has a 5’-NGG-3’ PAM sequence (PAM is the 3’ end of the target sequence). The sequences of the encoded RNA guide and its individual components are shown in Table 12. The T7 promoter uses +1G at the start of the transcript (i.e., the 5’ end of the RNA) for more efficient transcription. This +1G is shown for SEQ ID NO: 49.

Table 12

[0261] The DNA target was designed and ordered as a synthetic linear DNA fragment. The target sequences from AAVS1 and 10 bases upstream and downstream within the exon were flanked by 100 bases upstream of a non-related sequence and 200 bases downstream of a non-related sequence. Extra sequences were added, thereby enabling good separation of the cleaved and uncleaved products on a gel. The target and non-target strands were labeled with 5’ IR800 and 5’ IR700 labels, respectively, via PCR amplification using labeled primers. The sequences of the DNA target, the individual components of the DNA target, and the labeled PCR primers are shown in Table 13.

Table 13

[0262] The cleavage activity of nuclease A (SEQ ID NO: 1) and putative nickase was evaluated using an in vitro cleavage assay. Each polypeptide was individually co-expressed in vitro with an RNA guide by incubating a linear DNA template for the plasmid encoding the target protein from Table 11 and the T7 transcription AAVS1-T3 sgRNA from Table 12 in PURExpress® solution (NEB) containing the SUPERase·In® RNase inhibitor (Invitrogen) for 2 hours at 37°C. The unpurified polypeptide / RNA solution was then diluted in a 1x solution of NEB Buffer 2 (NEB) containing approximately 1 ng / μl of the labeled DNA target amplification product. The solution was then incubated at 37°C for 1 hour. The reaction was stopped by incubation with RNase Cocktail (Invitrogen; final concentration approximately 1 U / μl) at 37°C for 15 minutes, followed by incubation with proteinase K (NEB; final concentration approximately 0.04 U / μl) at 55°C for 30 minutes. The DNA was then purified using CleanNGS DNA & RNA Clean-Up Magnetic Beads (Bulldog Bio).

[0263] The cleaved and uncleaved products of the target and non-target strands were separated by running the sample on a 10% TBE-urea PAGE gel. The gel was imaged using a LI-COR Odyssey M imaging system with 800 nm and 700 nm channels to visualize 5'IR800 and 5'IR700 labeling in the target and non-target strands of the target DNA substrate. Band intensities were quantified using ImageJ software.

[0264] Gel images are shown in Figures 2A–2C, and the quantification of the percentage of cleaved target and non-target chains is shown in Figure 2D. Uncleaved chain, HNH-cleaved chain, and RuvC-cleaved chain are shown. Figure 2A is a gel image captured using an 800 nm channel, showing cleavage of the target chain. Figure 2B is a gel image captured using a 700 nm channel, showing cleavage of the non-target chain. Figure 2C is an overlay of gel images from Figures 2A and 2B. As shown in Figures 2A–2D, nuclease A (SEQ ID NO: 1) cleaved both the target and non-target chains, as expected. Each of the four HNH knockout nickase constructs (D396A, H397A, H397L, and N420A) showed significantly reduced activity in the target chain while retaining activity in the non-target chain. Two of the three RuvC knockout nickase constructs (D24A and E337A) exhibited significantly reduced activity in the non-target chain while retaining activity in the target chain (Figures 2A-2D).

[0265] Therefore, this example demonstrates the successful manipulation of HNH knockout nickases and RuvC knockout nickases.

[0266] Additional variant nuclease polypeptides derived from nuclease A are provided in Table 14 below, each of which is also within the scope of this disclosure. [Table 14-1] [Table 14-2] [Table 14-3] [Table 14-4] [Table 14-5] [Table 14-6]

[0267] Example 6. Editing of human genes in HEK293T cells using CRISPR nuclease polypeptide derived from nuclease A. This example demonstrates a genetic modification of a human gene utilizing nuclease polypeptides derived from nuclease A, listed in Tables 1 and 14, constructed according to the descriptions provided in Examples 1 and 3. Specifically, the polypeptides are used for indels into human target genes when combined with an RNA guide.

[0268] RNA guides including a reference scaffold, a cleaved reference scaffold, or EMX1-T3 with scaffold 2, EMX1-T7 with scaffold 2, and AAVS-T3 with scaffold 2 are shown in Table 2 of Example 1 or Table 8 of Example 4. Samples are prepared for NGS and analyzed as described in Example 1.

[0269] The nuclease polypeptide derived from nuclease A described in this example is expected to be capable of generating indels within human genes at target sites encoded by an RNA guide, such as the one described in this example.

[0270] Example 7: Editing of human target genes in HEK293T cells using CRISPR nuclease K or its variants This example describes genome editing of the AAVS1, EMX1, and VEGFA genes by introducing nuclease K of SEQ ID NO: 65 into HEK293T cell lines via lipid-based transient transfection.

[0271] Nuclease K was tagged with the N-terminal SV40 nuclear localization sequence (NLS) and the C-terminal nucleoplasmin nuclear localization sequence (NLS), its coding sequence was converted to a human codon-optimized DNA sequence, synthesized, and cloned into a pcDNA3.1 vector (Invitrogen) containing a CMV promoter for expression. The reference and NLS-tagged sequences used are shown in Table 15. For more efficient transcription, the U6 PolIII promoter was used with +1G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences listed in Table 16. Plasmids were purified using a midiprep kit.

[0272] RNA guides were designed, cloned into the pUC19 plasmid, and subsequently terminated with a 6x polyT sequence. The RNA guides were designed to be specific to target sequences within the coding exons of AAVS1, EMX1, and VEGFA, which possess a 5'-NGG-3'PAM sequence (the PAM sequence is located at the 3' end of the target sequence). See Table 16 for all RNA guide sequences. Plasmids were purified using a midiprep kit. [Table 15] [Table 16]

[0273] HEK293T cells were transfected with plasmids expressing nuclease polypeptides and gRNAs derived from nuclease K provided herein, according to the method provided in Example 1 above. Gene editing efficiency was determined by NGS, according to the method also provided in Example 1. The NGS results were also analyzed according to the method provided in Example 1 above.

[0274] For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for each sample and its related protein-free control. The higher percentage of indels in targets when nuclease K was included in transfection demonstrated the DNA editing results in cells.

[0275] As shown in Figure 4, four of the six targets tested demonstrated that higher levels of indels were observed when the nuclease K (SEQ ID NO: 65) plasmid was present.

[0276] Example 8: Efficacy of CRISPR nuclease K variant for targeting mammalian genes This example describes indel evaluation for a mammalian target using a nuclease K variant transfected into HEK293T cells.

[0277] Arginine phylogenetic mutagenesis was performed to individually substitute each non-arginine residue of reference nuclease K (SEQ ID NO: 65) with arginine. This resulted in 682 single-arginine substitution variants. Nuclease K and the nucleic acids encoding each variant were then individually cloned into a pcda3.1 backbone (Invitrogen), and the plasmids were maxi-prepped and diluted. The plasmids contained a CMV promoter, a first NLS upstream of the coding sequence (KRTADGSEFESPKKKRKV; SEQ ID NO: 3), an XTEN linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS; SEQ ID NO: 95), followed by a second NLS downstream of the coding sequence (KRPAATKKAGQAKKKK; SEQ ID NO: 4). See also Example 7 above.

[0278] The RNA guide and target sequences are shown in Table 17. The RNA guide was cloned into the pUC19 backbone (New England Biolabs®), followed by the U6 PolIII promoter, and terminated with a 6x polyT sequence. For more efficient transcription, U6 PolIII used +1G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences listed in Table 17. The plasmids were then maxi-prepped and diluted. [Table 17]

[0279] HEK293T cells were transfected with plasmids expressing nuclease polypeptides and gRNAs derived from nuclease K provided herein, according to the method provided in Example 1 above. Gene editing efficiency was determined by NGS, according to the method also provided in Example 1. The NGS results were also analyzed according to the method provided in Example 1 above.

[0280] For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for both the reference and each variant. The indel ratio used for calculating the multiplier change was the average of the two technical replicates. Then, to calculate the multiplier change in the indel ratio for a particular target, the indel ratio for each variant was divided by the indel ratio for the reference. Table 4 shows the multiplier change in the indel ratio for each target and as an average for both targets. The numbering is relative to the reference nuclease of Sequence ID No. 1 (i.e., without NLS). For example, in the AAVS1 target, the indel ratio for the D582R variant was 3.68 times that of the reference indel ratio, and in the EMX1 target, the indel ratio for the D582R variant was 31.80 times that of the reference indel ratio. The average multiplier change in the indel ratio for the D582R variant in both targets was 17.74 times that of the reference indel ratio in both targets.

[0281] As shown in Table 18, the 29 variants with a single arginine substitution (left column) were characterized as resulting in an increase in the indel ratio of at least 1.5 times compared to the reference indel ratio when averaged across both targets (right column). [Table 18]

[0282] The following 177 variants were analyzed as having indel ratios between 1 and 1.5 times the reference indel ratio when averaged across both targets.S548R, S365R, S134R, D729R, E747R, A724R, D581R, D301R, E623R, W127R, T1 80R, V114R, A395R, E123R, T703R, K68R, S258R, K60R, K324R, A297R, A381R, K 402R, D217R, D356R, A98R, E266R, T85R, N218R, K311R, V396R, N400R, K131R , K633R, V692R, I565R, E671R, I109R, E110R, A709R, Q27R, K300R, Q125R, H25 R, S214R, T438R, V726R, K605R, D186R, K563R, L495R, P637R, E121R, T741R, K306R, E107R, N689R, K521R, L269R, V262R, Q58R, K333R, S79R, K414R, E212R , D367R, K380R, K549R, D512R, K629R, K276R, P101R, E334R, G730R, M112R, K 14R, F420R, L407R, D708R, P685R, I149R, K238R, E594R, E203R, K428R, K603R , W120R, K697R, E108R, E102R, S745R, G492R, N437R, H244R, E115R, K45R, E5 38R, K364R, K427R, S88R, K159R, T308R, V189R, K35R, K328R, D113R, N744R, K 666R, K602R, Q50R, K49R, Q326R, K72R, E116R, G343R, K543R, E590R, T105R, K124R, E103R, K142R, K731R, G464R, A322R, K329R, M701R, P686R, S316R, K53 4R, E230R, D242R, M571R, K610R, A111R, K664R, K62R, Q687R, E349R, P684R, K344R, E342R, E431R, V184R, K679R, G337R, G727R, E150R, L193R, K146R, K26 5R, A419R, C606R, S275R, E705R, S511R, L624R, K338R, K240R, D221R, T519R, P469R, K104R, N133R, H239R, S678R, P210R, H145R, K466R, I584R, and K422R.

[0283] The remaining variants (483 variants) below resulted in a reduced indel ratio relative to the reference indel ratio (a multiplier change in indel ratios less than 1.0). G357R, K216R, N644R, E586R, M100R, D424R, D199R, D545R, K443R, G527R, S106R, G529R, K739R, K562R, K620R, H274R, L200R, E155R, T272R, S 73R, S627R, N69R, E345R, K597R, K667R, K122R, K233R, P305R, Q119R, P415R, T346R, S369R, T130R, K488R, K579R, K613R, N525R, S736R, E2R, E185R, K523R, K673R, E614R, Q213R, S412R, K220R, K61R, K668R, K280R, K591R, P171R, N175R, K714R, K118R, K138R, K289R, T587R, N370R, F2 08R, K593R, T732R, L661R, I540R, S728R, I268R, E277R, G528R, K467R, K226R, Y95R, N318R, P688R, E533R, K526R, I710R, E215R, E126R, G604R , D513R, K181R, A743R, D254R, I619R, G188R, E723R, E278R, K742R, E399R, K195R, L224R, P340R, K247R, N310R, S693R, G564R, N77R, N227R, I 158R, Y250R, C577R, N435R, N622R, T154R, K546R, D225R, G509R, A147R , E156R, T695R, N255R, D531R, K256R, Y670R, K92R, E471R, L390R, Y660 R, T738R, G52R, D251R, K621R, I411R, G363R, M518R, H535R, A253R, V735R, P539R, E707R, L179R, T567R, V725R, P537R, E408R, N654R, S532R, K167R, E490R, E78R, Q248R, Q473R, N204R, K172R, E4R, I245R, T241R, G578R, N645R, G547R, T608R, V625R, F282R, E663R, F694R, E129R, I43R,L257R、C618R、E416R、S560R、E461R、E335R、Y524R、E302R、E162R、I476R、A350R、G656R、E139R、Q168R、N423R、Y510R、E574R、G550R、V197R、G616R、D143R、K183R、A222R、N465R、G514R、T601R、K348R、D440R、L384R、E231R、H271R、F599R、F174R、A41R、G665R、G690R、I178R、G706R、D609R、Y480R、D517R、V704R、G75R、T59R、G486R、I418R、K508R、K715R、A615R、D457R、Y76R、Y336R、L5R、T56R、L304R、S462R、H680R、P675R、E293R、I211R、L536R、N385R、I417R、F209R、G151R、F542R、F530R、L29R、E611R、S409R、T721R、H372R、V655R、H91R、G327R、I436R、W484R、L89R、V232R、I205R、D662R、G382R、T494R、D658R、V141R、G339R、T80R、V652R、H136R、G717R、Y561R、A463R、G403R、A433R、L64R、N651R、N166R、Y588R、A371R、Q157R、D432R、H653R、T405R、L520R、N196R、L696R、Y169R、K401R、L425R、P192R、P406R、S585R、L352R、N442R、A285R、V313R、F387R、V281R、I74R、A388R、Y454R、V140R、G410R、I682R、N190R、F263R、Y477R、I566R、D96R、I643R、M626R、V647R、N493R、V516R、L44R、S551R、D22R、T290R、N557R、T635R、L677R、I366R、L351R、A669R、F377R、E354R、A31R、F681R、N294R、K291R、G383R、E580R、C264R、G634R、N317R、F617R、A323R、V249R、L331R、F307R、D39R、K426R、V659R、F235R、E66R、P522R、G259R、N646R、D700R、H267R、S474R, L201R, C394R, V161R, G636R, N674R, F648R, L288R, Y15R, L296R, P286R, N32R, T261R, H447R, P413R, A612R, V6R, I386R, Q639R, T82 R, T325R, F236R, I292R, C359R, D38R, K649R, I575R, M640R, V295R, I164R, V40R, S446R, N176R, Y711R, Y97R, T191R, D320R, V583R, V298R, T650R, N389R, I252R, I641R, V657R, I360R, A153R, I722R, L284R, I321R, V206R, I376R, V28R, M148R, C187R, Y30R, A737R, V391R, N398R, F 83R, E315R, I202R, L628R, V309R, L219R, V144R, M421R, F481R, V21R, L553R, M455R, A505R, P378R, I303R, P720R, I713R, A498R, Y70R, G12R , C182R, V312R, V86R, L499R, S287R, A19R, M314R, T573R, L740R, L7R, T355R, I299R, L18R, I503R, Y361R, V20R, Y87R, Y11R, G94R, F441R, N 26R, F487R, M445R, L716R, F319R, L468R, H375R, V223R, G8R, A448R, D374R, V439R, A501R, V502R, M36R, L559R, F165R, G17R, L607R, V702R, A595R, L34R, K489R, L592R, S506R, W452R, F596R, C362R, L504R, T16R, L430R, S449R, D555R, H496R, D10R, H554R, G13R, G451R, A479R, L67R, T478R, P558R, D500R, I429R, I470R, Y444R, A392R, V572R, I450R, H497R, C393R, C397R, G475R, I9R, W485R, L71R, A90R, L459R, and A81R.

[0284] Based on this experiment, the following variants, D582R, D46R, I631R, F570R, I53R, G42R, E576R, T734R, C630R, E683R, S719R, and E128R, were selected for further manipulation of nuclease K of sequence number 65.

[0285] Example 9: Efficacy of combined CRISPR nuclease K variants for targeting mammalian genes This example describes indel evaluation for mammalian targets using nuclease variants of nuclease K containing two or more substitutions identified in Example 8 as increasing indel activity. 71 combination nuclease K variants were tested.

[0286] Each nuclease K variant and RNA guide (see Table 19) was cloned as described in Example 8. HEK293T cells were further transfected as described in Example 8, followed by NGS analysis. For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for the reference nuclease K (SEQ ID NO: 65) and each nuclease K variant. The indel ratios shown in Table 20 were calculated as the average of two biological replicas, each containing two technical replicas. [Table 19] [Table 20-1] [Table 20-2]

[0287] As shown in Table 20, all but three of the combined nuclease K variants showed higher indel activity than the reference nuclease K (SEQ ID NO: 65). Two nuclease K variants yielded an indel ratio greater than 0.2 when averaged across three targets, indicating that more than 20% of the NGS reads contained indels. These two nuclease K variants included the following substitutions: a) D582R, I53R, and T734R, and b) D582R and G42R. Forty-two nuclease K variants yielded an indel ratio greater than 0.1 when averaged across three targets, indicating that more than 10% of the NGS reads contained indels.

[0288] Based on this experiment, nuclease K variants containing the following substitutions, D582R and G42R, were selected for further testing. These nuclease K variants showed more than an 8-fold increase in indel activity compared to the reference nuclease K (SEQ ID NO: 65).

[0289] Example 10: Editing of human target sequences in HEK293T cells mediated by nuclease M. This example describes genome editing of the AAVS1, EMX1, and VEGFA genes by introducing nuclease M (SEQ ID NO: 80) into HEK293T cell lines via lipid-based transient transfection.

[0290] Nuclease M was tagged with the N-terminal SV40 nuclear localization sequence and the C-terminal nucleoplasmin nuclear localization sequence, converted to a human codon-optimized DNA sequence, synthesized, and cloned into a pcDNA3.1 vector (Invitrogen) containing a CMV promoter for expression. The reference and NLS-tagged sequences used are shown in Table 21. For more efficient transcription, the U6 PolIII promoter was used with +1G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences listed in Table 22. Plasmids were purified using a midiprep kit.

[0291] RNA guides were designed, cloned into the pUC19 plasmid, and subsequently terminated with a 6x polyT sequence. The RNA guides were designed to be specific to target sequences within the coding exons of AAVS1, EMX1, and VEGFA, which possess a 5'-WTAAH-3'PAM sequence (the PAM sequence is located at the 3' end of the target sequence). See Table 22 for all RNA guide sequences. Plasmids were purified using a midiprep kit. [Table 21] [Table 22]

[0292] HEK293T cells were transfected with plasmids expressing CRISPR nuclease polypeptides and gRNAs provided herein, according to the method provided in Example 1 above. Gene editing efficiency was determined by next-generation sequencing, according to the method also provided in Example 1. The NGS results were also analyzed according to the method provided in Example 1 above.

[0293] For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for each sample and its related protein-free control. The higher percentage of indels in the target when nuclease M was included in the transfection demonstrated the DNA editing effect in cells.

[0294] As shown in Figure 5, three of the six targets tested demonstrated that larger levels of indels were observed when the nuclease M (SEQ ID NO: 80) plasmid was present.

[0295] Example 11: Efficacy of variant CRISPR nuclease for targeting mammalian genes This example describes indel evaluation for a mammalian target using a CRISPR nuclease variant derived from nuclease M (SEQ ID NO: 80) transfected into HEK293T cells.

[0296] A systematic arginine mutagenesis was performed to individually substitute selected non-arginine residues of nuclease M (SEQ ID NO: 80) with arginine. This resulted in 450 single-arginine substitution variants. The nucleic acids encoding the reference CRISPR nuclease and each CRISPR nuclease variant were then individually cloned into a pcDNA3.1 backbone (Invitrogen®), and the plasmids were maxi-prep and diluted. The plasmids contained a CMV promoter, a first NLS upstream of the coding sequence (KRTADGSEFESPKKKRKV; SEQ ID NO: 3), an XTEN linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS; SEQ ID NO: 95), followed by a second NLS downstream of the coding sequence (KRPAATKKAGQAKKKK; SEQ ID NO: 4). See also Example 10 above.

[0297] The RNA guide and target sequences are shown in Table 23. The RNA guide was cloned into the pUC19 backbone (New England Biolabs®), followed by the U6 PolIII promoter, and terminated with a 6x polyT sequence. For more efficient transcription, U6 PolIII used +1G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences listed in Table 23. The plasmids were then maxi-prepped and diluted. [Table 23]

[0298] HEK293T cells were transfected with plasmids expressing the CRISPR nuclease polypeptide and gRNA provided herein, according to the method provided in Example 1 above. Gene editing efficiency was determined by NGS according to the method also provided in Example 10. The NGS results were also analyzed according to the method provided in Example 10 above.

[0299] The indel ratio, which represents the percentage of NGS reads containing indels, was calculated for the reference and each variant. The indel ratio used for calculating the multiplier change was the average of the two technical replicates. Then, to calculate the multiplier change in the indel ratio, the indel ratio for each variant was divided by the indel ratio for the reference. Table 4 shows the multiplier change in the indel ratio. The numbering is relative to reference nuclease M (i.e., without NLS) of sequence number 80.

[0300] As shown in Table 24, eight variants of nuclease M with a single arginine substitution (left column) were characterized as resulting in an increase in the indel ratio of at least 1.5 times compared to the reference indel ratio. [Table 24]

[0301] The following 173 variants of nuclease M were analyzed as having indel ratios 1 to 1.5 times that of the reference indel ratio. G211R, D52R, T54R, K237R, E207R, A327R, K283R, A133R, K212R, E282R, K 126R, L412R, K365R, K236R, S46R, F57R, A328R, V112R, L312R, N106R, Y1 41R, K143R, K277R, T414R, V116R, K10R, I102R, V81R, E27R, Y256R, E75R , G132R, G215R, V245R, A353R, Y454R, G333R, V270R, E315R, H18R, T73R, L171R, S310R, T138R, K7R, V33R, V128R, L232R, K26R, I180R, V430R, Y49 R, G351R, T294R, I356R, L219R, A225R, A335R, T398R, P11R, P182R, S79R , K140R, S253R, V4R, Q378R, Y367R, K436R, K257R, L487R, D386R, V32R, W 130R, A318R, L167R, T50R, I137R, S470R, K28R, S249R, K127R, K29R, K11 0R, K309R, P37R, P447R, S6R, L24R, D443R, Y343R, G9R, D194R, N218R, L2 5R, Y362R, L20R, K119R, N340R, V177R, Y321R, H370R, N172R, T314R, I48 8R, K111R, L347R, Y408R, E146R, I359R, L242R, P131R, I195R, A492R, M3 01R, S337R, F174R, P136R, G402R, K145R, A451R, I5R, E474R, Y222R, S33 1R, L198R, T39R, N217R, T489R, K330R, E324R, F446R, L280R, N118R, S43 8R, D45R, V69R, W158R, Q400R, Y3R, Q227R, S339R, I296R, F122R, L12R, Y 364R, G417R, S295R, I183R, C485R, T160R, I239R, E407R, S426R, I472R, Y121R, Y78R, A163R, M13R, P361R, T15R, Y44R, V434R, F38R, V2R, T169R,F149R, P125R, P144R, C375R, I317R, F374R, N151R, P14R, L186R, Q64R, C263R, C266R, N86R, M93R, Y431R, L334R, L308R, L297R, Y479R, Q366R, G252R, G60R, T162R, I389R, H244R, H271R, D376R, P161R, A377R, V444R, K384R, V89R, A30R, A115R, A55R, L166R, T120R, C450R, E189R, T 67R, A147R, V460R, V21R, H267R, A68R, D58R, I66R, C234R, L181R, I300R, I40R, Y323R, T84R, N258R, S65R, G191R, H165R, T59R, I142R, A139R, E307R, S463R, H243R, S80R, L471R, L188R, F465R, G56R, H170R, L428R, T439R, A421R, N466R, S404R, I53R, I425R, Y173R, L82R, F464R, D 380R, G261R, L304R, S342R, K336R, L262R, G322R, L382R, L405R, A433R, A391R, A346R, G449R, V452R, V461R, G468R, N393R, C231R, F193R, D341R, T325R, Y383R, M345R, H338R, and A344R.

[0302] The remaining variants of nuclease M (269 variants) resulted in a reduced indel ratio relative to the nuclease M indel ratio (a multiplier change in indel ratios less than 1.0).

[0303] Based on this experiment, variants of nuclease M (SEQ ID NO: 80) having one or more of the following substitutions, E88R, S95R, L92R, E401R, E83R, N371R, P481R, and A373R, showed enhanced activity.

[0304] Other Embodiments All features disclosed herein may be combined in any combination. Each feature disclosed herein may be replaced by an alternative feature that serves the same, equivalent, or similar purpose. Thus, unless otherwise expressly stated, each feature disclosed is merely an example of a general set of equivalent or similar features.

[0305] From the above description, those skilled in the art will readily be able to identify the essential features of the present invention and adapt it to various uses and conditions in various modified and altered forms without departing from its spirit and scope. Therefore, other embodiments are also within the scope of the claims.

[0306] Equal portions While several embodiments of the present invention are described and illustrated herein, various other means and / or structures for carrying out the functions described herein and / or obtaining one or more of the results and / or benefits will be readily conceivable to those skilled in the art, and each of such variations and / or modifications will be considered to fall within the scope of the embodiments of the present invention described herein. Furthermore, generally speaking, all parameters, dimensions, materials and configurations described herein are intended to be illustrative, and it will be readily understood to those skilled in the art that the actual parameters, dimensions, materials and / or configurations will depend on one or more specific applications in which the teachings of the present invention are used. Those skilled in the art can recognize or verify many equivalents of the specific embodiments of the present invention described herein by means of conventional experimentation alone. Thus, it will be understood that the above embodiments are presented only as examples, and within the scope of the appended claims and their equivalents, embodiments of the present invention may be practiced in a manner different from those specifically described and claimed. The embodiments of the present invention in this disclosure cover each individual feature, system, article, material, kit and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present invention as long as such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

[0307] It should be understood that all definitions defined and used herein take precedence over dictionary definitions, definitions in documents incorporated by reference, and / or the ordinary meanings of the terms being defined.

[0308] All references, patents, and patent applications disclosed herein are incorporated by reference with respect to the subject matter they refer to, and in some cases the entire document may be included.

[0309] As used herein and in the claims, the indefinite article "a" or "an" should be understood to mean "at least one" unless otherwise clearly indicated. As used herein and in the claims, the phrase “and / or” should be understood to mean “either or both” of the elements thus connected, that is, that the elements exist connectively in some cases and disjunctively in others. Multiple elements enumerated by “and / or” should be interpreted in the same manner, that is, “one or more” of the elements thus connected. Other elements other than those specifically identified by the “and / or” clause may exist at their discretion, whether related to or unrelated to those specifically identified elements. Thus, as a non-restrictive example, when used with open-ended terms, such as “including,” a reference to “A and / or B” may, in one embodiment, refer to A only (optionally including elements other than B), in another embodiment, refer to B only (optionally including elements other than A), and in yet another embodiment, refer to both A and B (optionally including other elements), and so on.

[0310] Where used herein and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in an enumeration, “or” or “and / or” is interpreted as inclusive, that is, including at least one of several elements or enumerations of elements, and including two or more of them, and optionally including additional unenumerated items. Only terms that are clearly shown to be contrary, such as “only one of” or “exactly one of” or, when used in the claims, “consisting of,” means including exactly one element of several elements or enumerations of elements. In general, where used herein, the term “or” is interpreted only as indicating an exclusive substitution (e.g., “one or the other, but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of” or “exactly one of.” Where used in the claims, “essentially consisting of” has its usual meaning as used in the field of patent law.

[0311] As used herein and in the claims, the phrase “at least one” means, when referring to an enumeration of one or more elements, at least one element selected from any one or more of the elements in the enumeration of elements, but not necessarily including at least one of every element specifically enumerated in the enumeration of elements, and not excluding any combination of elements in the enumeration of elements. This definition also allows for the existence of elements other than those specifically identified in the enumeration of elements and referred to by the phrase “at least one,” whether related to or unrelated to these specifically identified elements, at the discretion of the party. Therefore, as a non-restrictive example, "at least one of A and B" (or equally, "at least one of A or B", or equally, "at least one of A and / or B") may mean, in one embodiment, at least one A, optionally including two or more A's and no B (optionally including elements other than B); in another embodiment, at least one B, optionally including two or more B's and no A (optionally including elements other than A); and in yet another embodiment, at least one A, optionally including two or more A's and at least one B, optionally including two or more B's (optionally including other elements), and so on.

[0312] Furthermore, unless otherwise clearly indicated, in any method claimed herein that includes two or more steps or actions, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are enumerated.

Claims

1. A CRISPR nuclease polypeptide comprising a RuvC nuclease domain and an HNH nuclease domain, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence identical to at least 90% of SEQ ID NO: 1, and optionally, the CRISPR nuclease polypeptide is a variant comprising at least one mutation relative to SEQ ID NO:

1.

2. The CRISPR nuclease polypeptide is the variant comprising the at least one mutation, and the at least one mutation is (a) For SEQ ID NO: 1, one or more arginine and / or lysine substitutions, and optionally one or more arginine substitutions, (b) One or more nicasse mutations in the HNH nuclease domain or the RuvC nuclease domain of Sequence ID No. 1, (c) N-terminal cleavage, (d) C-terminal cleavage, or (e) A CRISPR nuclease polypeptide according to claim 1, comprising a combination of (a), (b), (c), and / or (d).

3. The CRISPR nuclease polypeptide according to claim 1 or 2, wherein the at least one mutation comprises (a), wherein (a) is located in a bridge helix (BH) domain, a phosphate-locked loop (PLL) domain, a wedge-type (WED) domain, a PAM interaction (PID) domain, or a combination thereof.

4. The CRISPR nuclease polypeptide according to claim 2, wherein at least one mutation comprises (a), and the one or more arginine and / or lysine substitutions, optionally, one or more arginine substitutions are located at one or more of the positions D56, E59, G60, E63, I67, G71, S564, D568, T605, E655, E694, and E706 in SEQ ID NO:

1.

5. The CRISPR nuclease polypeptide has arginine and / or lysine substitutions at the following positions relative to SEQ ID NO:

1. a) I67, D568, and E706, b) I67, D56, and D568, c) I67, T605, and D568, d) I67, D56, and E706, e) I67, T605, and E706, f) I67, D568, and E694, g) I67, D56, E694, and E706, h) I67, D56, T605, and E706, or i) The CRISPR nuclease polypeptide according to claim 4, comprising I67, D568, W593, and E706.

6. The CRISPR nuclease polypeptide according to any one of claims 1 to 5, wherein the CRISPR nuclease polypeptide has an N-terminal cleavage with respect to SEQ ID NO: 1, and optionally the N-terminal cleavage is a deletion within residues 1 to 15 of SEQ ID NO: 1, preferably the N-terminal cleavage is a deletion within residues 1 to 14 or 1 to 15 of SEQ ID NO:

1.

7. The CRISPR nuclease according to any one of claims 2 to 6, wherein the CRISPR nuclease polypeptide contains up to 20 arginine and / or lysine substitutions, optionally up to 20 arginine substitutions, and optionally the CRISPR nuclease polypeptide contains up to 15 arginine and / or lysine substitutions, optionally up to 15 arginine substitutions, with respect to SEQ ID NO:

1.

8. The CRISPR nuclease polypeptide is the following combination of arginine substitutions: a) I67R, D568R, and E706R, b) I67R, D56R, and D568R, c) I67R, T605R, and D568R, d) I67R, D56R, and E706R, e) I67R, T605R, and E706R, f) I67R, D568R, and E694R, g) I67R, D56R, E694R, and E706R, h) I67R, D56R, T605R, and E706R, or i) The CRISPR nuclease according to claim 7, comprising I67R, D568R, W593R, and E706R.

9. The CRISPR nuclease according to claim 8, wherein the CRISPR nuclease polypeptide comprises the arginine substitutions of I67R, D568R, and E706R.

10. The CRISPR nuclease according to any one of claims 2 to 9, wherein the at least one mutation in the CRISPR nuclease polypeptide comprises (b), wherein (b) comprises one or more mutations at positions H397, D396, N420, D24, E337, and / or D524 of SEQ ID NO:

1.

11. The CRISPR nuclease according to claim 10, wherein the at least one mutation comprises a mutation at position H397, and optionally the mutation is an amino acid substitution of H397A or H397L.

12. The aforementioned variant is (a) The one or more nickase mutations in the HNH nuclease domain, wherein the one or more nickase mutations are optionally located at positions D396, H397, and / or N420 of SEQ ID NO: 1, and optionally located at position H397, (b) The CRISPR nuclease polypeptide according to claim 2, comprising one or more arginine and / or lysine substitutions, optionally located at positions I67, D568, and E706 of SEQ ID NO:

1.

13. The aforementioned variant is (a) The one or more nickase mutations in the HNH nuclease domain, which are optionally located at positions D396, H397, and / or N420 of SEQ ID NO: 1, (b) The one or more arginine and / or lysine substitutions, which optionally are located at positions I67, D568, and E706 of SEQ ID NO: 1, (c) The CRISPR nuclease polypeptide according to claim 2, comprising: (c) the N-terminal cleavage within residues 1 to 15 of SEQ ID NO: 1, wherein the N-terminal cleavage is optionally a deletion of residues 1 to 14 or 1 to 15 of SEQ ID NO:

1.

14. The aforementioned variant is (a) The aforementioned nicasse mutation in H397A, (b) The arginine substitutions of I67R, D568R, and E706R, (c) The CRISPR nuclease polypeptide according to claim 13, comprising an N-terminal cleavage including the deletion of residues 1 to 14 or 1 to 15 of SEQ ID NO:

1.

15. The CRISPR nuclease polypeptide according to claim 1, wherein the CRISPR nuclease polypeptide is one of those listed in Table 1, Table 11, or Table 14.

16. The d CRISPR nuclease polypeptide according to any one of claims 1 to 15, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence identical to that of SEQ ID NO: 1 by at least 95%.

17. The CRISPR nuclease polypeptide according to claim 16, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence identical to that of SEQ ID NO: 1 by at least 98%.

18. The CRISPR nuclease polypeptide according to any one of claims 1 to 17, which is a fusion polypeptide further comprising one or more functional fragments.

19. The CRISPR nuclease polypeptide according to claim 18, wherein the one or more functional fragments comprise one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof.

20. A nucleic acid comprising a nucleotide sequence encoding the CRISPR nuclease polypeptide described in any one of claims 1 to 19.

21. The nucleic acid according to claim 20, wherein the nucleic acid is messenger RNA (mRNA).

22. The nucleic acid according to claim 20, wherein the nucleic acid is an expression vector, the nucleotide sequence encoding the CRISPR nuclease polypeptide is operably bound to a promoter, and optionally the expression vector is a viral vector.

23. A host cell comprising the nucleic acid described in claim 21 or 22.

24. (a) A CRISPR nuclease polypeptide or a first nucleic acid encoding the CRISPR nuclease, wherein the CRISPR nuclease polypeptide is a CRISPR nuclease polypeptide or a first nucleic acid encoding the CRISPR nuclease as described in any one of claims 1 to 19, (b) A gene editing system comprising a guide RNA (gRNA) or a second nucleic acid encoding the gRNA, wherein the gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease and a spacer specific to a target sequence in a genomic region of interest, and the target sequence is adjacent to a protospacer-adjacent motif (PAM), the gRNA or the second nucleic acid encoding the gRNA.

25. The scaffold sequence includes a nucleotide sequence that is at least 85% identical to sequence number 2, The gene editing system according to claim 24, wherein the scaffold sequence optionally includes one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to sequence number 2.

26. The gene editing system according to claim 25, wherein the scaffold sequence is a cleavage variant of SEQ ID NO: 2, the cleavage variant has a 3' cleavage relative to SEQ ID NO: 2, the cleavage variant is about 110 to 140 nt in length, and optionally the 3' cleavage includes a deletion within residues 143 to 202 of SEQ ID NO:

2.

27. The gene editing system according to claim 26, wherein the cleavage variant further comprises one or more deletions within residues 10 to 40 of SEQ ID NO: 2, and optionally the deletion comprises residues 14 to 20 and / or 25 to 32 of SEQ ID NO:

2.

28. The gene editing system according to claim 26 or 27, wherein the cleavage variant further comprises one or more mutations relative to SEQ ID NO: 2, and optionally, the one or more mutations comprise a deletion within residues 82-85 of SEQ ID NO:

2.

29. The gene editing system according to claim 26, wherein the scaffold sequence comprises the nucleotide sequence of sequence number 27 or sequence number 28, and the scaffold sequence has a length of about 115 to 130 nt.

30. In the aforementioned gene editing system, (a) The CRISPR nuclease polypeptide contains the amino acid sequence of SEQ ID NO: 1, and the scaffold contains the nucleotide sequence of SEQ ID NO: 27, (b) The CRISPR nuclease polypeptide is a variant of SEQ ID NO: 1, wherein the variant includes mutations at positions I67, D568, and E706 of SEQ ID NO: 1, wherein the mutations are optionally arginine substitutions I67R, D568R, and E706R, and the scaffold includes the nucleotide sequence of SEQ ID NO: 28, or (c) The gene editing system according to claim 29, wherein the CRISPR nuclease polypeptide is a variant of SEQ ID NO: 1, the variant comprises a deletion in residues 1 to 15 of SEQ ID NO: 1, or optionally a deletion in residues 1 to 14 or 1 to 15 of SEQ ID NO: 1, and the scaffold comprises the nucleotide sequence of SEQ ID NO:

28.

31. The gene editing system according to any one of claims 24 to 30, wherein the target sequence is adjacent to the 5'-NGG-3' PAM, and in the sequence, N represents any of the nucleotides.

32. The gene editing system according to any one of claims 24 to 31, further comprising one or more lipid excipients associated with element (a) and / or element (b) of the gene editing system, wherein optionally, the one or more lipid excipients form lipid nanoparticles, and the lipid nanoparticles are associated with or encapsulated with element (a) and / or element (b) of the gene editing system.

33. A gene editing method comprising delivering the gene editing system according to any one of claims 24 to 32 to a host cell to edit a genomic site targeted by the gRNA of the gene editing system.

34. A guide RNA comprising a spacer sequence and a scaffold sequence, wherein the scaffold sequence is a variant of SEQ ID NO: 2, the variant comprises one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to SEQ ID NO: 2, and the scaffold sequence is recognizable by a CRISPR nuclease polypeptide as described in any one of claims 1 to 19.

35. The guide RNA according to claim 34, wherein the scaffold sequence is a cleavage variant of SEQ ID NO: 2, the cleavage variant has a 3' cleavage relative to SEQ ID NO: 2, and the cleavage variant is approximately 110 to 140 nt in length.

36. The guide RNA according to claim 35, wherein the 3' cleavage includes a deletion within residues 143-202 of SEQ ID NO:

2.

37. The guide RNA according to claim 35 or 36, wherein the cleavage variant further comprises one or more deletions within residues 10 to 40 of SEQ ID NO:

2.

38. The guide RNA according to claim 37, wherein the deletion comprises residues 14-20 and / or 25-32 of SEQ ID NO:

2.

39. The guide RNA according to claim 37 or 38, wherein the cleavage variant further comprises one or more mutations relative to SEQ ID NO: 2, and optionally, the one or more mutations comprise deletions within residues 81-85 of SEQ ID NO:

2.

40. The guide RNA according to claim 35, wherein the scaffold sequence comprises the nucleotide sequence of SEQ ID NO: 27 or SEQ ID NO: 28, and the scaffold has a length of about 115 to 130 nt.

41. A CRISPR nuclease polypeptide comprising a RuvC nuclease domain and an HNH nuclease domain, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence identical to at least 90% of SEQ ID NO: 65, and optionally, the CRISPR nuclease polypeptide is a variant comprising at least one mutation relative to SEQ ID NO:

65.

42. The CRISPR nuclease polypeptide is a variant containing the at least one mutation, and the at least one mutation is (a) For SEQ ID NO: 65, one or more arginine and / or lysine substitutions, and optionally one or more arginine substitutions, (b) One or more nicasse mutations in the HNH nuclease domain or the RuvC nuclease domain of Sequence ID No. 65, or (c) The CRISPR nuclease polypeptide according to claim 41, comprising a combination of (a) and (b).

43. The CRISPR nuclease polypeptide according to claim 42, wherein the CRISPR nuclease polypeptide comprises a bridge helix (BH) domain, a wedge-type (WED) domain, and a PAM interaction (PID) domain, and the CRISPR nuclease polypeptide is a variant comprising one or more arginine and / or lysine substitutions of (a), wherein one or more arginine and / or lysine substitutions of (a) are located in the BH domain, the WED domain, the PID domain, or a combination thereof.

44. The CRISPR nuclease polypeptide according to claim 43, wherein the one or more arginine and / or lysine substitutions are located at one or more of the positions G42, D46, I53, F83, E128, E541, F570, E576, D582, C630, I631, E683, S719, T734 in SEQ ID NO:

65.

45. The CRISPR nuclease polypeptide has arginine and / or lysine substitutions at the following positions relative to SEQ ID NO:

65. (i) D582 and G42, or (ii) The CRISPR nuclease polypeptide according to claim 44, comprising D582, I53, and T734R.

46. The CRISPR nuclease polypeptide is subject to the following arginine substitutions: a) D582R and G42R, or b) The CRISPR nuclease polypeptide according to claim 45, comprising D582R, I53R, and T734R.

47. The CRISPR nuclease polypeptide according to any one of claims 42 to 46, wherein the CRISPR nuclease polypeptide contains up to 20 arginine and / or lysine substitutions relative to SEQ ID NO: 65, optionally up to 20 arginine substitutions, and optionally the CRISPR nuclease polypeptide contains up to 15 arginine and / or lysine substitutions relative to SEQ ID NO: 65, optionally up to 15 arginine substitutions.

48. The CRISPR nuclease polypeptide according to any one of claims 42 to 47, wherein the CRISPR nuclease polypeptide comprises one or more nickase mutations of (b), the one or more nickase mutations of (b) are located at positions D374, H375, and / or N398 in SEQ ID NO:

65.

49. The CRISPR nuclease polypeptide according to claim 48, wherein the nickase mutation is located at position H375, and optionally the mutation is an amino acid substitution at H375A.

50. The CRISPR nuclease polypeptide according to any one of claims 41 to 49, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 95% identical to that of SEQ ID NO:

65.

51. The CRISPR nuclease polypeptide according to claim 50, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 98% identical to that of SEQ ID NO:

65.

52. A CRISPR nuclease polypeptide according to any one of claims 41 to 51, which is a fusion polypeptide comprising one or more additional functional elements.

53. The CRISPR nuclease polypeptide according to claim 52, wherein the one or more additional functional elements include one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof.

54. A nucleic acid comprising a nucleotide sequence encoding a CRISPR nuclease polypeptide as described in any one of claims 41 to 53.

55. The nucleic acid according to claim 54, wherein the nucleic acid is an expression vector, the nucleotide sequence encoding the CRISPR nuclease polypeptide is operably bound to a promoter, and optionally the expression vector is a viral vector.

56. The nucleic acid according to claim 54, wherein the nucleic acid is messenger RNA (mRNA).

57. A host cell comprising the nucleic acid described in any one of claims 54 to 56.

58. (a) A CRISPR nuclease polypeptide according to any one of claims 41 to 53 or a first nucleic acid encoding the CRISPR nuclease polypeptide, (b) A gene editing system comprising a guide RNA (gRNA) or a second nucleic acid encoding the gRNA, wherein the gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease polypeptide and a spacer sequence specific to a target sequence in a genomic region of interest, and the target sequence is adjacent to a protospacer adjacency motif (PAM), the gRNA or the second nucleic acid encoding the gRNA.

59. The gene editing system according to claim 58, wherein the scaffold sequence includes a nucleotide sequence that is at least 85% identical to sequence number 79.

60. The gene editing system according to claim 59, wherein the scaffold sequence is a fragment of sequence number 79.

61. The gene editing system according to any one of claims 58 to 60, wherein the scaffold sequence comprises one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to SEQ ID NO:

79.

62. The gene editing system according to any one of claims 58 to 61, wherein the target sequence is upstream of the 5'-NGG-3' PAM, and in the sequence, N represents any of the nucleotides.

63. The gene editing system according to any one of claims 58 to 62, further comprising one or more lipid excipients associated with element (a) and / or element (b) of the gene editing system, wherein optionally, the one or more lipid excipients form lipid nanoparticles, and the lipid nanoparticles are associated with or encapsulated with element (a) and / or element (b) of the gene editing system.

64. A gene editing method comprising delivering the gene editing system according to any one of claims 58 to 63 to a host cell to edit a genomic site targeted by the gRNA of the gene editing system.

65. A nuclease polypeptide comprising a RuvC nuclease domain and an HNH nuclease domain, wherein the nuclease polypeptide comprises an amino acid sequence identical to at least 90% of SEQ ID NO: 80, and optionally, the nuclease polypeptide is a variant of SEQ ID NO: 80 comprising at least one mutation relative to SEQ ID NO:

80.

66. The nuclease polypeptide is the variant comprising the at least one mutation, and the at least one mutation is (a) For SEQ ID NO: 80, one or more arginine and / or lysine substitutions, and optionally one or more arginine substitutions, (b) One or more nicasse mutations in the HNH nuclease domain or the RuvC nuclease domain of Sequence ID No. 80, or (c) The nuclease polypeptide according to claim 65, comprising a combination of (a) and (b).

67. The nuclease polypeptide according to claim 66, wherein the at least one mutation comprises (a), wherein (a) is located in a bridge helix (BH) domain, a nucleic acid recognition (REC) domain, a phosphate-locked loop (PLL) domain, a wedge-type (WED) domain, a PAM interaction (PID) domain, a nuclease domain, or a combination thereof.

68. The nuclease polypeptide according to claim 66 or 67, wherein the nuclease polypeptide contains up to 20 arginine and / or lysine substitutions, optionally up to 20 arginine substitutions, relative to SEQ ID NO:

80.

69. The nuclease polypeptide according to claim 68, wherein the nuclease polypeptide contains up to 15 arginine and / or lysine substitutions, optionally up to 15 arginine substitutions, relative to SEQ ID NO:

80.

70. The nuclease polypeptide according to any one of claims 67 to 69, wherein the one or more arginine and / or lysine substitutions, optionally, the arginine substitution is at position E88, S95, L92, E401, E83, N371, P481, and / or A373 in SEQ ID NO:

80.

71. The nuclease polypeptide according to any one of claims 66 to 70, wherein the at least one mutation comprises (b), and the one or more nickase mutations are located at one or more of the positions D58, E189, D341, H243, H244, H267, R329, and / or H338 in SEQ ID NO:

80.

72. The nuclease polypeptide according to any one of claims 65 to 71, wherein the nuclease polypeptide comprises an amino acid sequence that is at least 95% identical to that of SEQ ID NO:

80.

73. The nuclease polypeptide according to claim 72, wherein the nuclease polypeptide comprises an amino acid sequence that is at least 98% identical to that of SEQ ID NO:

80.

74. The nuclease polypeptide according to any one of claims 65 to 73, wherein the nuclease polypeptide is a fusion polypeptide, and the fusion polypeptide further comprises one or more functional fragments, the one or more functional fragments being heterogeneous to the nuclease moiety of the fusion polypeptide.

75. The nuclease polypeptide according to claim 74, wherein the one or more functional fragments comprise one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof.

76. A nucleic acid comprising a nucleotide sequence encoding a nuclease polypeptide according to any one of claims 65 to 75.

77. The nucleic acid according to claim 76, wherein the nucleic acid is an expression vector, the nucleotide sequence encoding the nuclease is operably bound to a promoter, and optionally the expression vector is a viral vector.

78. The nucleic acid according to claim 77, wherein the nucleic acid is messenger RNA (mRNA).

79. A host cell comprising the nucleic acid according to any one of claims 76 to 78.

80. (a) a nuclease polypeptide or a first nucleic acid encoding the nuclease, wherein the nuclease polypeptide is a nuclease polypeptide or a first nucleic acid encoding the nuclease as described in any one of claims 65 to 75, (b) A gene editing system comprising a guide RNA (gRNA) or a second nucleic acid encoding the gRNA, wherein the gRNA comprises a scaffold sequence recognizable by the nuclease polypeptide and a spacer sequence specific to a target sequence in a genomic region of interest, and the target sequence is adjacent to a protospacer adjacency motif (PAM), the gRNA or the second nucleic acid encoding the gRNA.

81. The gene editing system according to claim 80, wherein the scaffold sequence includes a nucleotide sequence that is at least 85% identical to sequence number 94.

82. The gene editing system according to claim 81, wherein the scaffold sequence is a fragment of sequence number 94.

83. The gene editing system according to claim 82, wherein the scaffold sequence comprises one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to sequence number 94.

84. The gene editing system according to any one of claims 80 to 83, wherein the PAM is 5'-WTAAH-3', in which W is A or T, H is A, C or T, and optionally the PAM is 5'-TTAAA-3'.

85. The gene editing system according to any one of claims 80 to 84, further comprising one or more lipid excipients associated with element (a) and / or element (b) of the gene editing system, wherein optionally, the one or more lipid excipients form lipid nanoparticles, and the lipid nanoparticles are associated with or encapsulated with element (a) and / or element (b) of the gene editing system.

86. A gene editing method comprising delivering the gene editing system according to any one of claims 80 to 85 to a host cell to edit a genomic site targeted by the gRNA of the gene editing system.