CRISPR nuclease polypeptides and gene editing systems comprising such CRISPR nuclease polypeptides
By modifying CRISPR nucleases through arginine and lysine substitution and nicking enzyme mutation, the enzymatic activity and gene editing efficiency of CRISPR nucleases have been enhanced, overcoming the shortcomings of existing CRISPR-Cas systems in gene editing efficiency and enzymatic activity, and achieving more efficient gene editing results.
Patent Information
- Application Number
- CN202480034731.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2024-03-29
- Publication Date
- 2025-12-30
AI Technical Summary
Existing CRISPR-Cas systems have shortcomings in gene editing efficiency and enzymatic activity, making it difficult to achieve efficient gene editing results.
CRISPR nuclease peptides are provided, including reference CRISPR nucleases A, K, and M and their variants, which enhance enzymatic activity and binding ability to homologous guide RNA by introducing arginine and lysine substitutions at specific positions, nicking enzyme mutations, and N-terminal truncation.
This improved the nicking enzyme activity and gene editing efficiency of CRISPR nucleases, resulting in more efficient gene editing effects.
Smart Images

Figure CN121241133A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 493,360, filed March 31, 2023; U.S. Provisional Application No. 63 / 516,246, filed July 28, 2023; U.S. Provisional Application No. 63 / 566,661, filed March 18, 2024; U.S. Provisional Application No. 63 / 493,355, filed March 31, 2023; U.S. Provisional Application No. 63 / 562,132, filed March 6, 2024; and U.S. Provisional Application No. 63 / 493,363, filed March 31, 2023. Each priority application in the priority applications is incorporated herein by reference in its entirety.
[0003] sequence list
[0004] This application contains a sequence list that has been electronically submitted in XML format and is hereby incorporated in its entirety by reference. The XML copy created on March 27, 2024, is named 063586-514001WO_SeqList_ST26.xml and has a size of 137,220 bytes. Background Technology
[0005] Clusters of regularly spaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes are collectively referred to as the CRISPR-Cas or CRISPR / Cas system. This system is an adaptive immune system in archaea and bacteria that protects specific species from the influence of foreign genetic factors.
[0006] CRISPR-Cas systems typically contain a CRISPR nuclease and one or more RNA components that guide the CRISPR nuclease to a target genomic site for gene editing. Developing efficient CRISPR nucleases to improve gene editing efficiency is of great interest. Summary of the Invention
[0007] This disclosure provides CRISPR nucleases that exhibit favorable enzymatic activities (e.g., nicking enzyme activity, high indel activity, high binding activity to homologous guide RNA scaffolds, and / or high DNA cleavage activity). Therefore, the CRISPR nucleases provided herein are expected to demonstrate superior efficiency when used for gene editing.
[0008] Therefore, this document provides CRISPR nuclease peptides derived from reference CRISPR nucleases, including nuclease A (SEQ ID NO: 1), nuclease K (SEQ ID NO: 65), or nuclease M (SEQ ID NO: 80). Such CRISPR nucleases contain a RuvC nuclease domain and an HNH nuclease domain. In some cases, the CRISPR nuclease peptides provided herein share at least 85% (e.g., at least 90%) of the same amino acid sequence as the reference CRISPR nuclease provided herein. In some cases, the CRISPR nuclease peptides disclosed herein comprise variant CRISPR nucleases (also known as engineered CRISPR nucleases) containing at least one mutation relative to the reference CRISPR nuclease.
[0009] In some aspects, this disclosure provides a CRISPR nuclease polypeptide derived from a reference nuclease A, a gene editing system comprising the CRISPR nuclease polypeptide derived from said nuclease A and a guide RNA targeting a genomic site of interest, and the use of said gene editing system for modifying a genomic site of interest in a host cell.
[0010] In some embodiments, the CRISPR nuclease polypeptide is an engineered variant of a reference nuclease A (SEQ ID NO: 1), the engineered variant comprising at least one mutation relative to nuclease A. In some cases, the at least one mutation may (a) comprise one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitution) relative to nuclease A, (b) comprise one or more nickase mutations in the HNH nuclease domain or the RuvC nuclease domain of nuclease A, (c) comprise an N-terminal truncation relative to nuclease A, or any combination of (a), (b), and (c).
[0011] Similar to nuclease A, the CRISPR nuclease polypeptide derived from nuclease A may comprise a bridged helix (BH) domain, a phosphate-locked ring (PLL) domain, a wedge (WED) domain, and a PAM interaction (PID) domain. In some instances, the nuclease A variants provided herein may comprise one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitution). In some cases, the one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitution) are located in the BH domain, the PLL domain, the WED domain, the PID domain, or a combination thereof. In some cases, the one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitution) may be located at one or more of positions D56, E59, G60, E63, I67, G71, S564, D568, W593, T605, E655, E694, and E706 in SEQ ID NO: 1. In some specific instances, the engineered variants of nuclease A comprise arginine substitutions and / or lysine substitutions (e.g., arginine substitutions) at the following positions relative to SEQ ID NO: 1: a) I67, D568, and E706; b) I67, D56, and D568; c) I67, T605, and D568; d) I67, D56, and E706; e) I67, T605, and E706; f) I67, D568, and E694; g) I67, D56, E694, and E706; h) I67, D56, T605, and E706; or i) I67, D568, W593, and E706. In one specific instance, the engineered variants of nuclease A comprise arginine substitutions and / or lysine substitutions at I67, D568, and E706.
[0012] In some instances, the engineered variant nuclease polypeptide of nuclease A contains (a) arginine substitutions of I67R, D568R, and E706R. In some instances, the engineered CRISPR nuclease polypeptide contains (b) arginine substitutions of I67R, D56R, and D568R. In some instances, the engineered CRISPR nuclease polypeptide contains (c) arginine substitutions of I67R, T605R, and D568R. In some instances, the engineered CRISPR nuclease polypeptide contains (d) arginine substitutions of I67R, D56R, and E706R. In some instances, the engineered CRISPR nuclease polypeptide contains (e) arginine substitutions of I67R, T605R, and E706R. In some instances, the engineered CRISPR nuclease polypeptide contains (f) arginine substitutions of I67R, D568R, and E694R. In some instances, the engineered CRISPR nuclease peptide contains arginine substitutions of (g) I67R, D56R, E694R, and E706R. In some instances, the engineered CRISPR nuclease peptide contains arginine substitutions of (h) I67R, D56R, T605R, and E706R. In some instances, the engineered CRISPR nuclease peptide contains arginine substitutions of (i) I67R, D568R, W593R, and E706R. In some specific instances, the engineered CRISPR nuclease peptide contains arginine substitutions of I67R, D568R, and E706R.
[0013] Any engineered variant of nuclease A disclosed herein may contain up to 20 arginine substitutions and / or lysine substitutions (e.g., arginine substitutions), such as up to 15 arginine substitutions and / or lysine substitutions (e.g., arginine substitutions).
[0014] Alternatively or additionally, engineered variants of nuclease A may have an N-terminal truncation relative to nuclease A (SEQ ID NO: 1). In some instances, the N-terminal truncation is a deletion of residues 1-15 of SEQ ID NO: 1. In some specific instances, the N-terminal truncation is a deletion of residues 1-14. In another specific instance, the N-terminal truncation is a deletion of residues 1-15 of SEQ ID NO: 1.
[0015] Furthermore, the CRISPR nuclease polypeptides disclosed herein can be nicking enzyme variants of nuclease A exhibiting nicking enzyme activity. Such nicking enzyme variants may contain one or more mutations located at positions H397, D396, N420, D24, E337, and / or D524 of SEQ ID NO: 1. In some instances, the nicking enzyme variant may contain a mutation located at position H397. In some specific instances, the mutation is an amino acid substitution of H397A or H397L.
[0016] In some embodiments, engineered variants of nuclease A as disclosed herein may comprise any combination of mutations disclosed herein, such as one or more arginine / lysine substitutions, one or more nickase mutations, and / or N-terminal truncation. For example, the variant may comprise: (a) one or more nickase mutations in the HNH nuclease domain (e.g., located at positions D396, H397, and / or N420 relative to SEQ ID NO: 1); and the one or more arginine and / or lysine substitutions. In a specific example, the nickase mutation is located at position H397 of SEQ ID NO: 1, and the arginine / lysine substitution is located at positions I67, D568, and E706.
[0017] In another example, the engineered variant of nuclease A may comprise: (a) one or more nickase mutations in the HNH nuclease domain (e.g., located at positions D396, H397, and / or N420 relative to SEQ ID NO: 1); (b) one or more arginine substitutions and / or lysine substitutions (e.g., located at positions I67, D568, and E706); and (c) the N-terminal truncation within residues 1-15 of SEQ ID NO: 1 (e.g., deletion of residues 1-14 or 1-15 of SEQ ID NO: 1). Alternatively or additionally, the engineered variant of nuclease A may comprise a truncation relative to the C-terminus of nuclease A (e.g., within residues 1-5 of the C-terminus of SEQ ID NO: 1).
[0018] In yet another example, the variant of nuclease A comprises: (a) the nickase mutation of H397A; (b) the arginine substitution of I67R, D568R and E706R; and (c) the N-terminal truncation comprising the deletion of residues 1-14 or 1-15 of SEQ ID NO: 1.
[0019] Any engineered variant of nuclease A disclosed herein may contain at least 95% of the same amino acid sequence as SEQ ID NO: 1. In some instances, the engineered variant may contain at least 98% of the same amino acid sequence as SEQ ID NO: 1.
[0020] Exemplary CRISPR nuclease peptides derived from nuclease A, as provided herein, are listed in Tables 1, 11, or 14, each of which is within the scope of this disclosure.
[0021] Any CRISPR nuclease polypeptide derived from nuclease A disclosed herein may be a fusion polypeptide, which may further comprise one or more functional fragments. In some embodiments, the one or more functional fragments may comprise one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof. The one or more NLS may be located at an N-terminus, a C-terminus, or both.
[0022] This document also provides a nucleic acid comprising a nucleotide sequence encoding any CRISPR nuclease polypeptide derived from nuclease A as disclosed herein. In some cases, the nucleic acid is an expression vector wherein the nucleotide sequence encoding the CRISPR nuclease polypeptide is operatively linked to a promoter. Further, this disclosure provides a host cell comprising a nucleic acid encoding a CRISPR nuclease polypeptide as disclosed herein.
[0023] In some embodiments, this disclosure features a gene editing system comprising: (a) any CRISPR nuclease polypeptide derived from nuclease A as disclosed herein, or a first nucleic acid encoding said CRISPR nuclease polypeptide; and (b) a guide RNA (gRNA) or a second nucleic acid encoding said gRNA. The gRNA comprises a scaffold sequence recognizable by said CRISPR nuclease polypeptide and a spacer sequence specific to a target sequence at a genomic site of interest. The target sequence is adjacent to a protospacer adjacent motif (PAM). In some embodiments, the target sequence is located upstream (5'-) of a 5'-NGG-3' protospacer adjacent motif (PAM), where N represents any nucleotide.
[0024] In some cases, the scaffold sequence that can be utilized by a CRISPR nuclease polypeptide derived from nuclease A as disclosed herein may contain at least 85% identical nucleotide sequences to SEQ ID NO: 2. In some cases, the scaffold sequence contains one or more deletions, one or more nucleotide substitutions, or combinations thereof compared to SEQ ID NO: 2.
[0025] In some cases, the scaffold sequence is a truncated variant of SEQ ID NO: 2. Such truncated variants can be approximately 110-140 nt in length. In some instances, the truncated variant may have a 3' truncation relative to SEQ ID NO: 2. For example, the 3' truncation may include deletions within residues 143-202 of SEQ ID NO: 2. In other instances, the truncated variant of SEQ ID NO: 2 may have an internal truncation relative to SEQ ID NO: 2. For example, the variant may include one or more deletions within residues 10-40 of SEQ ID NO: 2. In specific instances, the deletion may include residues 14-20 and / or 25-32 of SEQ ID NO: 2. In other instances, the truncated variant of SEQ ID NO: 2 may have a 3' truncation relative to SEQ ID NO: 2 (e.g., the 3' truncation disclosed herein) and an internal truncation (e.g., the internal truncation disclosed herein). Alternatively or additionally, the truncated variant may further comprise one or more mutations relative to SEQ ID NO: 2, such as deletions and / or substitutions within residues 81-85 of SEQ ID NO: 2. In one specific example, the scaffold sequence comprises (e.g., constitutes) the nucleotide sequence of SEQ ID NO: 27. In another specific example, the scaffold sequence comprises (e.g., constitutes) SEQ ID NO: 28. The length of such scaffold sequences can be approximately 115-130 nt.
[0026] Guide RNAs containing spacer subsequences as defined herein and any variant scaffold sequence relative to SEQ ID NO: 2 are also within the scope of this disclosure.
[0027] In some cases, the CRISPR nuclease polypeptide included in the gene editing system disclosed herein is the reference CRISPR nuclease of SEQ ID NO: 1, and the scaffold sequence comprises the nucleotide sequence of SEQ ID NO: 27. Alternatively, the CRISPR nuclease polypeptide in the gene editing system may be a variant of SEQ ID NO: 1, for example, containing mutations at positions I67, D568, and E706 (e.g., arginine substitutions for I67R, D568R, and E706R), and the scaffold sequence comprises the nucleotide sequence of SEQ ID NO: 28.
[0028] In other respects, this document provides a CRISPR nuclease polypeptide derived from CRISPR nuclease K (SEQ ID NO: 65), the CRISPR nuclease polypeptide comprising a RuvC nuclease domain and an HNH nuclease domain. The CRISPR nuclease polypeptide derived from nuclease K may contain at least 90% of the same amino acid sequence as SEQ ID NO: 65. In addition to the RuvC nuclease domain and the HNH nuclease domain, the CRISPR nuclease polypeptide may also contain a bridged helix (BH) domain, a wedge (WED) domain, and a PAM interaction (PID) domain.
[0029] In some embodiments, the CRISPR nuclease polypeptide may be a variant of nuclease K containing at least one mutation relative to SEQ ID NO: 65. Such variant CRISPR nuclease polypeptides may have enhanced enzymatic activity compared to the reference CRISPR nuclease K. In some cases, the at least one mutation may comprise (a) one or more arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 65, optionally one or more arginine substitutions; (b) one or more nickase mutations in the HNH nuclease domain or the RuvC nuclease domain of SEQ ID NO: 65; or (c) a combination of (a) and (b).
[0030] In some instances, variants of the nuclease K may include variants of (a) one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitution). The variant CRISPR nuclease polypeptides disclosed herein may contain up to 20 arginine substitutions and / or lysine substitutions (e.g., up to 20 arginine substitutions), such as up to 15 arginine substitutions and / or lysine substitutions (e.g., up to 15 arginine substitutions). In some cases, the one or more arginine substitutions and / or lysine-arginine substitutions are located in the BH domain, the WED domain, the PID domain, or a combination thereof. In a specific instance, the one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitutions) may be located at one or more of the following positions in SEQ ID NO: 65: G42, D46, I53, F83, E128, E541, F570, E576, D582, C630, I631, E683, S719, and T734.
[0031] In some cases, variants of the nuclease K may contain a combination of arginine substitutions and / or lysine substitutions. For example, the variant may contain the arginine substitutions and / or lysine substitutions (e.g., G42R and D582R) located at positions G42 and D582 of SEQ ID NO: 65. In another example, the variant may contain the arginine substitutions and / or lysine substitutions (e.g., I53R, D582R, and T734R) located at positions I53, D582, and T734 of SEQ ID NO: 65.
[0032] In some embodiments, the engineered CRISPR nuclease polypeptide derived from nuclease K provided herein may contain one or more mutations that produce nicking enzyme activity, for example, one or more mutations located at positions D374, H375, and / or N398 in SEQ ID NO: 65. In some instances, the nicking enzyme variant may contain a mutation located at position H375; optionally, said mutation is an amino acid substitution of H375A. In specific instances, the engineered CRISPR nuclease polypeptide may contain arginine substitutions and / or lysine substitutions (e.g., arginine substitutions) located at positions I53, D582, and T734 in SEQ ID NO: 65, and further contain the amino acid substitution of H375A.
[0033] In some embodiments, the engineered CRISPR nuclease polypeptide derived from nuclease K disclosed herein may contain at least 90% of the same amino acid sequence as SEQ ID NO: 65. In some instances, the engineered CRISPR nuclease derived from nuclease K may contain at least 95% of the same amino acid sequence as SEQ ID NO: 65. In a specific instance, the engineered CRISPR nuclease derived from nuclease K may contain at least 98% of the same amino acid sequence as SEQ ID NO: 65.
[0034] In some embodiments, the engineered CRISPR nuclease polypeptide derived from nuclease K, as disclosed herein, may be a fusion polypeptide comprising a CRISPR nuclease and one or more additional fragments heterologous to the CRISPR nuclease. In some embodiments, the one or more functional fragments may comprise one or more NLSs, one or more peptide linkers, or a combination thereof. The one or more NLSs may be located at an N-terminus, a C-terminus, or both.
[0035] This document also provides nucleic acids comprising a nucleotide sequence encoding any engineered CRISPR nuclease polypeptide derived from nuclease K disclosed herein. In some cases, the nucleic acid is an expression vector in which the nucleotide sequence encoding the engineered CRISPR nuclease polypeptide is operatively linked to a promoter. Further, this disclosure provides a host cell comprising a nucleic acid encoding an engineered CRISPR nuclease polypeptide as disclosed herein.
[0036] On the other hand, this disclosure features a gene editing system comprising: (a) any engineered CRISPR nuclease polypeptide derived from nuclease K disclosed herein, or a first nucleic acid encoding a CRISPR nuclease polypeptide; and (b) a guide RNA (gRNA) comprising a scaffold sequence and a spacer recognizable by the engineered CRISPR nuclease polypeptide, or a second nucleic acid encoding the gRNA. The gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease polypeptide and a spacer sequence specific to a target sequence at a genomic site of interest. The target sequence is adjacent to a protospacer adjacent motif (PAM). In some embodiments, the target sequence is located upstream (5'-) of a 5'-NGG-3' protospacer adjacent motif (PAM), where N represents any nucleotide.
[0037] In some embodiments, a scaffold capable of being utilized by a CRISPR nuclease polypeptide derived from nuclease A as disclosed herein may comprise at least 85% of the same nucleotide sequence as SEQ ID NO: 79. Alternatively or additionally, the scaffold may be a fragment of SEQ ID NO: 79. In some cases, the scaffold comprises one or more deletions, one or more nucleotide substitutions, or combinations thereof compared to SEQ ID NO: 79.
[0038] In another aspect, this disclosure provides a nuclease polypeptide derived from a reference nuclease M (SEQ ID NO: 80), the nuclease polypeptide comprising a RuvC nuclease domain and an HNH nuclease domain. Such nuclease polypeptides may contain at least 90% of the same amino acid sequence as SEQ ID NO: 80 (referring to the reference nuclease M disclosed herein). In some embodiments, the nuclease polypeptide derived from nuclease M may be a variant of nuclease M, the variant comprising at least one mutation relative to nuclease M. Compared to nuclease M, such variants of nuclease M may have enhanced enzymatic activity.
[0039] In some cases, the at least one mutation in the variant of nuclease M may comprise (a) one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitutions) relative to nuclease M (SEQ ID NO: 80); (b) one or more nickase mutations in the HNH nuclease domain or the RuvC nuclease domain of SEQ ID NO: 80; or (c) a combination of (a) and (b).
[0040] In some instances, variants of nuclease M may contain one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitutions), which may be located in one or more of the following: a bridged helix (BH) domain, a nucleic acid recognition (REC) domain, a phosphate-locked ring (PLL) domain, a wedge (WED) domain, a PAM interaction (PID) domain, a nuclease domain, or a combination thereof. In some cases, a variant nuclease polypeptide of nuclease M disclosed herein may contain up to 20 arginine substitutions and / or lysine substitutions (e.g., arginine substitutions) relative to the reference nuclease M. For example, the variant nuclease polypeptide may contain up to 15 arginine substitutions and / or lysine substitutions (e.g., up to 12 arginine / lysine substitutions) relative to the reference nuclease M, such as up to 15 arginine substitutions or up to 12 arginine substitutions.
[0041] In a specific instance, the one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitutions) may be located at positions E88, S95, L92, E401, E83, N371, P481 and / or A373 in SEQ ID NO: 80.
[0042] Alternatively or additionally, the variant nuclease polypeptide of said nuclease M may contain one or more mutations that produce nicking enzyme activity. In some embodiments, said one or more nicking enzyme mutations are located at one or more of positions D58, E189, D341, H243, H244, H267, R329 and / or H338 in SEQ ID NO: 80.
[0043] Any variant of nuclease M provided herein may contain at least 90% of the same amino acid sequence as SEQ ID NO: 80. In some instances, the nuclease polypeptide derived from nuclease M may contain at least 95% of the same amino acid sequence as SEQ ID NO: 80. In other instances, nuclease M derived from the nuclease polypeptide may contain at least 98% of the same amino acid sequence as SEQ ID NO: 80.
[0044] In some embodiments, the nuclease polypeptide derived from nuclease M may be a fusion polypeptide, which may further comprise one or more functional fragments. In some cases, the one or more functional fragments may be heterologous to the nuclease portion of the fusion polypeptide. In some cases, the one or more functional fragments may comprise one or more NLS, one or more peptide linkers, or combinations thereof, which may be located at the N-terminus, C-terminus, or both.
[0045] This document also provides a nucleic acid comprising a nucleotide sequence encoding a nuclease polypeptide derived from any nuclease M as disclosed herein. In some cases, the nucleic acid is an expression vector wherein the nucleotide sequence encoding the nuclease polypeptide is operatively linked to a promoter. Alternatively, the nucleic acid may be a messenger RNA (mRNA) molecule. Further, this disclosure provides a host cell comprising a nucleic acid encoding a nuclease polypeptide derived from nuclease M as disclosed herein.
[0046] Additionally, this disclosure features a gene editing system comprising: (a) any nuclease polypeptide derived from nuclease M as disclosed herein, or a first nucleic acid encoding such a nuclease polypeptide derived from nuclease M; and (b) a guide RNA (gRNA) or a second nucleic acid encoding said gRNA, said gRNA comprising a scaffold sequence recognizable by said nuclease polypeptide derived from nuclease M and a spacer specific to a target sequence within a genomic site, said target sequence being adjacent to a protospacer adjacent motif (PAM). In some embodiments, said PAM is 5'-WTAAH-3', where W is A or T, and H is A, C, or T. In some instances, said PAM is 5'-TTAAA-3'.
[0047] In some embodiments, a scaffold sequence that can be utilized by a nuclease polypeptide derived from nuclease M as disclosed herein may comprise at least 85% of the same nucleotide sequence as SEQ ID NO: 94. Alternatively or additionally, the scaffold sequence may be a fragment of SEQ ID NO: 94. In some cases, the scaffold sequence comprises one or more deletions, one or more nucleotide substitutions, or combinations thereof compared to SEQ ID NO: 94.
[0048] Any gene editing system disclosed herein may further comprise one or more lipid excipients associated with element (a) and / or element (b) of the gene editing system. In some instances, the one or more lipid excipients form lipid nanoparticles that associate with or encapsulate element (a) and / or element (b) of the gene editing system.
[0049] Alternatively, any gene editing system disclosed herein may contain a viral vector, such as an adeno-associated virus (AAV) vector containing the coding sequences of a CRISPR nuclease polypeptide and the gRNA described herein.
[0050] Furthermore, this document provides a gene editing method comprising delivering any gene editing system as disclosed herein to a host cell to edit a genomic site targeted by the gRNA of said gene editing system. In some cases, the host cell is cultured in vitro. In other cases, the host cell is located in a subject requiring gene editing.
[0051] Details of one or more embodiments of the invention are set forth in the following description. Other features or advantages of the invention will become apparent from the following drawings and the following detailed description of several embodiments, and also from the appended claims. Attached Figure Description
[0052] The following drawings form part of this specification and are included to further illustrate certain aspects of this disclosure. A better understanding of certain aspects of this disclosure can be achieved by referring to the drawings in combination with a detailed description of the specific embodiments presented herein.
[0053] Figure 1 This is a graph showing the percentage of NGS reads contained in the indels at the six genetic loci, as indicated by the presence or absence of the CRISPR nuclease SEQ ID NO: 1 (nuclease A).
[0054] Figure 2A-2D These are gel images showing the in vitro cleavage of the target or non-target strands of the target DNA substrate and bar graphs showing the quantification of nuclease activity of nuclease A and its variants. Figure 2A These are gel images captured using an 800 nm channel, showing the target strand (labeled with IR800 dye at the 5' end) of a reference CRISPR nuclease, a hypothetical HNH knockout cleavage enzyme, or a hypothetical RuvC knockout cleavage enzyme cleaving the target DNA substrate in vitro. Figure 2B These are gel images captured using a 700 nm channel, showing the in vitro cleavage of the non-target strand of the target DNA substrate by a reference CRISPR nuclease, a hypothetical HNH knockout cleavage enzyme, or a hypothetical RuvC knockout cleavage enzyme (labeled with IR700 dye at the 5' end). Figure 2C Is using Figure 2A and Figure 2B Overlay images captured by the 800 nm and 700 nm channels. Figure 2DThis is a bar chart that quantitatively shows the percentage of target DNA and non-target DNA cleaved by tests using a reference CRISPR nuclease, a hypothetical HNH knockout cleavage enzyme, and a hypothetical RuvC knockout cleavage enzyme.
[0055] Figures 3A-3C It is a schematic diagram depicting the predicted secondary structures of a stent sequence. Figure 3A Reference stent (SEQ IDNO: 2). Figure 3B : The truncated reference stent (SEQ ID NO: 27). Figure 3C : Stent 2 (SEQ ID NO: 28).
[0056] Figure 4 This is a graph showing the percentage of NGS reads contained in the indels at the six genetic loci, as indicated by the presence or absence of the CRISPR nuclease SEQ ID NO: 65 (nuclease K).
[0057] Figure 5 This is a graph showing the percentage of NGS reads contained in the indels at the six genetic loci, as indicated by the presence or absence of the nuclease SEQ ID NO: 80 (Nuclease M). Detailed Implementation
[0058] This document provides CRISPR nuclease peptides such as reference nuclease A (SEQ ID NO: 1), nuclease K (SEQ ID NO: 65), or nuclease M (SEQ ID NO: 80). Such CRISPR nuclease peptides may contain a reference nuclease as disclosed herein or a variant thereof. In some cases, the CRISPR nuclease peptides provided herein may be variant nucleases of the reference nuclease containing at least one mutation relative to the reference nuclease.
[0059] In some embodiments, the variant CRISPR nuclease polypeptide may contain one or more mutations relative to a reference CRISPR nuclease (e.g., arginine substitution, lysine substitution, or a combination thereof). Alternatively or additionally, the variant CRISPR nuclease polypeptide may contain one or more mutations in the RuvC nuclease domain or the HNH nuclease domain. Such mutations (i.e., nicking enzyme mutations) can reduce or eliminate the nuclease activity of the RuvC nuclease domain or the HNH nuclease domain, thereby producing a variant exhibiting nicking enzyme activity. As used herein, the term "nicking enzyme" refers to an enzyme that cleaves one strand of double-stranded DNA at a specific recognition nucleotide sequence (e.g., the target sequence disclosed herein). A nicking enzyme can interact with one strand of a DNA duplex to produce a DNA molecule cleaved at one strand (also known as a nicked molecule). In some embodiments, the nicking enzyme is a variant of a CRISPR nuclease containing an inactivated HNH domain. In some embodiments, the nicking enzyme is a variant of a CRISPR nuclease containing an inactivated RuvC domain.
[0060] Any variant CRISPR nuclease polypeptide provided herein may share high sequence homology (e.g., at least 85% sequence identity, such as at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or higher sequence identity) relative to reference CRISPR nuclease A (SEQ ID NO: 1), nuclease K (SEQ ID NO: 65), or nuclease M (SEQ ID NO: 80).
[0061] The variant CRISPR nuclease peptides provided herein are expected to possess advantageous features compared to the reference CRISPR nuclease, such as increased binding to homologous guide RNA and higher nuclease activity. Therefore, the variant CRISPR nuclease peptides disclosed herein are expected to exhibit better activity (e.g., nicking enzyme activity) in gene editing compared to the reference CRISPR nuclease, and to demonstrate greater efficiency and accuracy in gene editing involving strand substitution.
[0062] Alternatively or additionally, the variant CRISPR nuclease polypeptide may be a fusion polypeptide comprising a CRISPR nuclease (e.g., a reference nuclease A, nuclease K, or nuclease M, or a variant of a reference nuclease as provided herein) (the nuclease portion) and one or more additional functional fragments, such as those described herein (e.g., NLS and / or peptide linkers). In addition to the advantageous features described above, the fusion polypeptide also has additional functions attributable to the fusion partner.
[0063] Therefore, this disclosure provides CRISPR nuclease peptides derived from reference CRISPR nucleases (nuclease A, nuclease K, or nuclease M), gene editing systems comprising such CRISPR nuclease peptides, and gene editing methods using such CRISPR nuclease peptides.
[0064] I. CRISPR nuclease polypeptide
[0065] As used herein, the term "CRISPR nuclease" refers to an effector that is capable of binding to nucleic acids and introducing single-strand or double-strand breaks. CRISPR nucleases typically contain multiple functional domains, such as nuclease domains (e.g., RuvC and / or HNH), PLMP domains, bridged helix (BH) domains, nucleic acid recognition (REC) domains, phosphate-locked loop (PLL) domains, wedge-shaped domains (WED), PAM interaction domains (PID), or combinations thereof. As used herein, the term "domain" refers to a distinct functional and / or structural unit of a polypeptide. In some cases, the functional domain may be linear. In others, the functional domain may be discontinuous and conformational. In some embodiments, the domain may contain a conserved amino acid sequence.
[0066] As used herein, the term “variant CRISPR nuclease polypeptide” refers to a CRISPR nuclease polypeptide that contains alterations (e.g., substitution, insertion, deletion, and / or fusion) at one or more residue sites compared to a reference CRISPR nuclease (nuclease A, nuclease K, or nuclease M).
[0067] The variant CRISPR nuclease peptides disclosed herein are expected to exhibit one or more regulated activities (e.g., enhanced or reduced) relative to a reference CRISPR nuclease. As used herein, the term "activity" refers to biological activity. In some embodiments, activity includes enzymatic activity, such as the catalytic ability of an effector. For example, activity may include nuclease activity. In some embodiments, activity includes binding activity, such as the binding of an effector (e.g., a CRISPR nuclease) to an RNA guide and / or target nucleic acid. In some instances, the variant CRISPR nuclease peptides disclosed herein have enhanced binding to homologous guide RNA (gRNA) compared to a reference CRISPR nuclease, for example, binding activity that is at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2-fold, 2-fold, 5-fold, 10-fold, or higher than the binding activity of the reference CRISPR nuclease. Homologous gRNA refers to gRNA having a scaffold that can be recognized by a CRISPR nuclease.
[0068] In some instances, the variant CRISPR nuclease peptides disclosed herein exhibit enhanced enzymatic activity relative to a reference CRISPR nuclease, for example, at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2-fold, 2-fold, 5-fold, 10-fold, or higher than the enzymatic activity of the reference CRISPR nuclease. In other instances, the variant CRISPR nuclease peptides disclosed herein exhibit reduced enzymatic activity relative to a reference CRISPR nuclease, for example, at least 20%, 30%, 40%, 50%, 60%, or 70% lower than the enzymatic activity of the reference CRISPR nuclease. In some cases, the reduced enzymatic activity is achieved by reducing or attenuating the nuclease activity of the RuvC domain. In some cases, the reduced enzymatic activity is achieved by reducing or attenuating the nuclease activity of the HNH domain.
[0069] In some cases, the variant CRISPR nuclease peptides disclosed herein exhibit enhanced indel activity relative to a reference CRISPR nuclease. As used herein, the term "indel activity" refers to the ability of a CRISPR nuclease to introduce an indel (insertion / deletion) into a sequence (e.g., a genomic target).
[0070] In some embodiments, the variant CRISPR nuclease peptides provided herein share high sequence homology with a reference CRISPR nuclease. For example, the variant CRISPR nuclease peptide may contain at least 70% (e.g., at least 80%, 85%, 90%, 95%, or higher) of the same amino acid sequence as SEQ ID NO: 1, SEQ ID NO: 65, or SEQ ID NO: 80. In some cases, the variant CRISPR nuclease peptide may contain at least 90% of the same amino acid sequence as SEQ ID NO: 1, SEQ ID NO: 65, or SEQ ID NO: 80. In some cases, the variant CRISPR nuclease peptide may contain at least 95% of the same amino acid sequence as SEQ ID NO: 1, SEQ ID NO: 65, or SEQ ID NO: 80. In other cases, the variant CRISPR nuclease peptide may contain at least 97% (e.g., 98%, 99%, 99.5%, or higher) of the same amino acid sequence as SEQ ID NO: 1, SEQ ID NO: 65, or SEQ ID NO: 80.
[0071] The "percentage of identity" (also known as sequence identity) between two nucleic acid or two amino acid sequences is determined using the algorithm described in Karlin and Altschul, Proceedings of the National Academy of Sciences (PNAS) 87:2264-68, 1990, modified as described in Karlin and Altschul, PNAS 90:5873-77, 1993. This algorithm is incorporated into the NBLAST and XBLAST programs (version 2.0) of Altschul et al., Journal of Molecular Biology 215:403-10, 1990. BLAST nucleotide searches can be performed using the NBLAST program with a score of 100 and a word length of -12 to obtain nucleotide sequences homologous to the nucleic acid molecules of this invention. BLAST protein searches can be performed using the XBLAST program with a score of 50 and a word length of 3 to obtain amino acid sequences homologous to the protein molecules of this invention. In cases where gaps exist between two sequences, gapped BLAST can be used, as described by Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997. When using BLAST and gapped BLAST programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used.
[0072] In some cases, the variant CRISPR nuclease polypeptides disclosed herein may contain one or more arginine and / or lysine substitutions relative to the reference nuclease A, nuclease K, or nuclease M provided herein. “Arginine substitution” and / or “lysine substitution” means replacing a non-arginine or non-lysine residue in SEQ ID NO: 1 (nuclease A), SEQ ID NO: 65 (nuclease K), or SEQ ID NO: 80 (nuclease M) with an arginine or lysine residue.
[0073] In some cases, one or more substituted arginine residues may be replaced by conserved amino acid residues such as lysine or histidine. In some embodiments, the variant CRISPR nuclease polypeptides provided herein may contain one or more arginine substitutions, one or more lysine substitutions, or combinations thereof.
[0074] In some cases, the variant CRISPR nuclease peptides presented herein may contain one or more conserved amino acid residues substituted, either alone or in combination with other types of mutations disclosed herein.
[0075] As used herein, “conservative amino acid substitution” refers to an amino acid substitution that does not alter the relative charge or size properties of the protein to which the amino acid substitution is made. Variants can be prepared according to methods for altering polypeptide sequences known to those skilled in the art, as found in references that compile such methods, such as *Molecular Cloning: A Laboratory Manual*, edited by J. Sambrook et al., 2nd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989, or *Current Protocols in Molecular Biology*, edited by FM Ausubel et al., John Wiley & Sons, Inc., New York. Conservative substitutions of amino acids include substitutions between amino acids in the following groups: (a) M, I, L, V; (b) F, Y, W; (c) K, R, H; (d) A, G; (e) S, T; (f) Q, N; and (g) E, D.
[0076] (A) Nuclease A and its engineered variants
[0077] In some embodiments, the CRISPR nuclease polypeptide is derived from nuclease A, which includes wild-type nuclease A (SEQ ID NO: 1), variants thereof (such as those disclosed herein), or fusion polypeptides comprising such wild-type nuclease A and its variants. Variant CRISPR nuclease polypeptides of nuclease A can be generated by introducing one or more mutations into a reference CRISPR nuclease to modulate (e.g., enhance or reduce) one or more activities of the nuclease.
[0078] The reference CRISPR nuclease for nuclease A (SEQ ID NO: 1) (see Table 1 below) is a CRISPR nuclease comprising both a RuvC nuclease domain (located at residues 15-53, 308-338, and 472-566 of SEQ ID NO: 1) and an HNH domain (located at residues 339-471 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate the cleavage of DNA strands adjacent to the 5'-NGG-3' PAM motif, where N represents any nucleotide. Positions D24, E335, and D524 are considered active sites in the RuvC domain, and positions D396, H397, and N420 are considered active sites in the HNH domain. R506 and H521 may also be important for the nuclease activity of the RuvC domain. In addition to the nuclease domain, the reference CRISPR nuclease of SEQ ID NO: 1 also includes the BH domain (residues 54-89 of SEQ ID NO: 1), the REC domain (residues 90-307 of SEQ ID NO: 1), the PLL domain (residues 567-580 of SEQ ID NO: 1), the WED domain (residues 581-672 of SEQ ID NO: 1), and the PID domain (residues 673-776 of SEQ ID NO: 1).
[0079] The reference CRISPR nuclease of SEQ ID NO: 1 disclosed herein is smaller than that of the CRISPR-Cas9 nuclease from *Streptococcus pyogenes*. The scaffold utilized by the reference CRISPR nuclease of SEQ ID NO: 1 and its variants can be miniaturized. This scaffold contains a unique structure compared to the SpCas9 nuclease scaffold. It is anticipated that this unique scaffold will allow for a reduction in the size of some domains (e.g., the REC domain), and thus contribute to a smaller CRISPR nuclease size. These features will be beneficial for delivery. Arginine and / or lysine substitutions (e.g., arginine substitution) can be introduced into the CRISPR nuclease of SEQ ID NO: 1 to increase indel activity.
[0080] Furthermore, the RuvC domain of nuclease A (SEQ ID NO: 1) cleaves the non-target strand of the target nucleic acid within 4 to 8 nucleotides upstream of PAM, and the HNH domain cleaves the target strand upstream of PAM within 3 to 4 nucleotides, each producing a cleavage site with a 0-5 nucleotide overhang (most commonly a 3-5 nucleotide overhang). In contrast, the RuvC domain of SpCas9 cleaves the non-target strand of the target nucleic acid within 3 to 5 nucleotides upstream of homologous PAM (5'-NGG-3', where N is any nucleotide), and the HNH domain cleaves the target strand upstream of PAM within 3 to 4 nucleotides, each producing a 0-3 nucleotide overhang (most commonly a flat cleavage).
[0081] Furthermore, using a gene editing system containing nuclease A (SEQ ID NO: 1) or a variant thereof allows for the introduction of indels larger than those that can be introduced by SpCas9 into the target nucleic acid. For example, insertions induced by nuclease A or a variant thereof can range from about 1 nucleotide to about 7 nucleotides (most commonly about 4 nucleotides). Deletions induced by nuclease A or a variant thereof can range from about 1 nucleotide to about 25 nucleotides or more (most commonly about 11 nucleotides). Further, since the reference CRISPR nuclease of SEQ ID NO: 1 contains a RuvC domain and an HNH domain, nickase variants can be engineered, for example, by disrupting the nuclease activity of either the RuvC domain or the HNH domain.
[0082] The CRISPR nuclease polypeptide variant of nuclease A provided herein may contain one or more alterations relative to nuclease A (SEQ ID NO: 1), such as substitution of one or more amino acid residues, deletion of one or more, insertion of one or more, fusion of one or more, or combinations thereof. In some cases, alterations may be introduced into the BH domain, PLL domain, WED domain, PID domain, or combinations thereof. In some cases, no alterations are introduced into the RuvC nuclease domain and / or the HNH nuclease domain, or at the active site and / or at sites involved in the activity of these domains as provided herein. Alternatively, conserved amino acid substitutions may be introduced into SEQ ID NO: 1, including in the RuvC nuclease domain and / or the HNH nuclease domain.
[0083] In some embodiments, relative to SEQ ID NO: 1, the variant CRISPR nuclease polypeptide of nuclease A may contain one or more arginine substitutions, one or more lysine substitutions, or combinations thereof. In some instances, the variant CRISPR nuclease polypeptide may contain up to 20 arginine substitutions and / or lysine substitutions (e.g., up to 20 arginine substitutions, up to 20 lysine substitutions, or combinations thereof), for example, up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 arginine substitutions, lysine substitutions, or combinations thereof. In specific instances, the variant CRISPR nuclease polypeptide may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 arginine substitutions, lysine substitutions, or combinations thereof. In some instances, the variant CRISPR nuclease polypeptide provided herein contains an arginine substitution.
[0084] In some cases, arginine substitution and / or lysine substitution may be located in the BH domain, PLL domain, WED domain, PID domain, or any combination thereof of nuclease A. In some instances, variant CRISPR nuclease polypeptides derived from nuclease A may contain one or more substitutions of arginine and / or lysine (e.g., arginine substitution) at one or more of the following positions in SEQ ID NO: 1: D56, E59, G60, E63, I67, G71, S564, D568, T605, E655, E694, and E706. In some instances, the variant CRISPR nuclease polypeptide derived from nuclease A contains one or more of the following arginine substitutions relative to SEQ ID NO: 1: D56R, E59R, G60R, E63R, I67R, G71R, S564R, D568R, T605R, E655R, E694R, and E706R.
[0085] In some cases, arginine substitution and / or lysine substitution may be located within the RuvC nuclease domain and / or HNH nuclease domain of nuclease A. In some instances, arginine substitution and / or lysine substitution may be located within the RuvC nuclease domain of nuclease A to reduce or inactivate the RuvC domain (e.g., at positions D24, E335, R506, H521, and / or D524 in the RuvC domain of SEQ ID NO: 1). In other instances, arginine substitution and / or lysine substitution may be located within the HNH domain of nuclease A, for example, at positions D396, H397, and / or N420 in the HNH domain of SEQ ID NO: 1. In still other instances, arginine substitution and / or lysine substitution may be located within both the RuvC domain and the HNH domain of nuclease A to reduce or weaken the enzymatic activity of the nuclease.
[0086] Alternatively, arginine substitution and / or lysine substitution may not be located at the active site and / or sites involved in activity within the RuvC nuclease domain and / or HNH nuclease domain of nuclease A (e.g., not at positions D24, E335, R506, H521 and / or D524 in the RuvC domain of SEQ ID NO: 1 and / or positions D396, H397 and / or N420 in the HNH domain). In some instances, arginine substitution and / or lysine substitution may not be present in the RuvC domain and / or HNH domain of nuclease A.
[0087] This article reports that arginine substitutions at the following positions in SEQ ID NO: 1 lead to decreased or absent nuclease activity, indicating that these positions are intolerable to mutations in nuclease activity: L21, G22, I23, D24, G26, G27, T30, G31, L32, A33, V34, V35, V42, V48, M50, L85, V107, Y108, C111, G115, M169, V182, F186, I270, C286, H289, S309, L310, V317, V3 21. M336, I343, V334, E337, S338, N339, F341, T347, Y358, T377, V382, Y383, C384 , T389, A393, D396, H397, I398, F399, P400, I406, N411, V413, A414, C415, C416, N 420, K423, K450, L452, A455, I466, M469, S470, A472, S473, I474, G475, L483, G49 9. T502, W509, F511, H521, L523, D524, A525, V526, I527, L528, A529, P581, V595, T596, I606, Y611, L615, L630, A635, F640, Y641, L651, G657, L658, G659, Q662, M663, V664, K672, T673, N674, V675, Y683, L700, V725, I736, L739, P743, L745, and L765.
[0088] In some embodiments, the CRISPR nuclease A variants provided herein exhibit nuclease activity and may not have mutations at the aforementioned positions. Alternatively, the variant CRISPR nuclease may be a dead nuclease (e.g., for base editing) with mutations at one or more of these positions.
[0089] In some cases, the variant CRISPR nuclease contains arginine substitutions and / or lysine substitutions at combinations of two or more (e.g., three or four) positions of, for example, D56, I67, D568, T605, E694, and E706 of SEQ ID NO: 1. Examples include, but are not limited to, (a) I67, D568, and E706; (b) D56, I67, and D568; (c) I67, D568, and T605; (d) D56; I67, and E706; (e) I67, T605, and E706; (f) I67, D568, and E694; (g) D56, I67, E694, and E706; (h) D56, I67, T605, and E706; and (i) I67, D568, W593, and E706. In some specific instances, engineered CRISPR nuclease peptides contain arginine substitutions at I67, D568, and E706. In some instances, variant CRISPR nucleases contain arginine substitutions at the combined positions listed herein.
[0090] Further exemplary combinations of arginine substitution and / or lysine substitution (e.g., arginine substitution) include (relative to SEQ ID NO: 1): (i) at least two positions of D56, E59, G60, E63, I67, G71, S564, D568, T605, E655, E694, and E706; (ii) at least two positions of D56, I67, T605, E694, and E706; (iii) at least two positions of D56, I67, D568, T605, E694, and E706; (iv) at least two positions of D56, E63, G71, D568, T605, E655, E694, and E706; and (v) at least two positions of D56, I67, G71, D568, T605, E655, E694, and E706.
[0091] Alternatively or additionally, the engineered variant CRISPR nuclease polypeptide of nuclease A disclosed herein may have an N-terminal truncation relative to nuclease A of SEQ ID NO: 1. Such engineered variant CRISPR nuclease polypeptides have a segment deleted at the N-terminus of nuclease A. In some cases, the deleted N-terminal segment has up to 80 amino acid residues, for example, up to 70 amino acid residues, up to 60 amino acid residues, up to 50 amino acid residues, up to 40 amino acid residues, up to 30 amino acid residues, or up to 20 amino acid residues. In some instances, the N-terminal truncation is a deletion of residues 1-15 of SEQ ID NO: 1. In one specific instance, the N-terminal truncation is a deletion of residues 1-14 of SEQ ID NO: 1. In another specific instance, the N-terminal truncation is a deletion of residues 1-15 of SEQ ID NO: 1.
[0092] Alternatively or additionally, the engineered CRISPR nuclease peptides disclosed herein may have a C-terminal truncation relative to the reference CRISPR nuclease A of SEQ ID NO: 1. Such engineered CRISPR nuclease peptides have a segment deleted at the C-terminus of the reference CRISPR nuclease. In some cases, the deleted C-terminal segment has up to 10 amino acid residues, for example, up to 5 amino acid residues, up to 4 amino acid residues, up to 3 amino acid residues, up to 2 amino acid residues, or up to 1 amino acid residue. In some instances, the C-terminal truncation is the deletion of the last residue, the last two residues, the last three residues, the last four residues, or the last five residues of SEQ ID NO: 1. In some cases, variants of the C-terminal truncation may further include one or more mutations among those disclosed herein, for example, one or more arginine substitutions / lysine substitutions, one or more mutations producing nicking enzyme activity, or combinations thereof. In some specific instances, the C-terminal truncated variants further comprise arginine substitutions relative to I67R, D568R, and E706R of SEQ ID NO: 1. In other specific instances, the C-terminal truncated variants further comprise arginine substitutions relative to I67R, D568R, W593R, and E706R of SEQ ID NO: 1. In some instances, the C-terminal truncated variants further comprise N-terminal truncation. In some instances, the C-terminal truncated variants further comprise the deletion of residues 1-14 or 1-15 of SEQ ID NO: 1.
[0093] In some cases, the truncated variants may further include one or more mutations disclosed herein, such as one or more arginine substitutions / lysine substitutions, one or more mutations producing nicking enzyme activity (nicking enzyme mutations), or combinations thereof. In some specific instances, the truncated variants may further include arginine substitutions / lysine substitutions at positions I67, D568, and E706 of SEQ ID NO: 1 (e.g., arginine substitutions I67R, D568R, and E706R). In other specific instances, the truncated variants further include arginine substitutions / lysine substitutions at positions I67, D568, W593, and E706 of SEQ ID NO: 1 (e.g., arginine substitutions at I67R, D568R, W593R, and E706R).
[0094] Alternatively or additionally, any variant CRISPR nuclease polypeptide of nuclease A provided herein may contain one or more nicking enzyme mutations located within the RuvC nuclease domain or the HNH nuclease domain to reduce or eliminate the nuclease activity of the target domain, thereby producing a variant with nicking enzyme activity. Such mutations may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation within the RuvC nuclease domain or the HNH nuclease domain is an amino acid substitution, wherein the substituted amino acid residue is not a conserved substitution of the native amino acid residue located at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue other than K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue other than R. A set of conserved amino acid residue substitutions is provided herein.
[0095] Positions D24, E337, and D524 are identified as putative catalytic residues in the RuvC domain of nuclease A, and positions D396, H397, and N420 are identified as putative catalytic residues in the HNH domain of nuclease A of SEQ ID NO: 1. In some instances, one or more mutations may be located within the RuvC nuclease domain to reduce or inactivate the RuvC domain (e.g., at positions D24, E337, and / or D524 within the RuvC domain). In some specific instances, variant CRISPR nuclease polypeptides derived from nuclease A may contain substitutions for D24A, E337A, and / or D524A. In other instances, any one of D24, E337, and D524 may be substituted with an amino acid residue similar to A (e.g., G, S, or L).
[0096] In some instances, one or more mutations may be located within the HNH domain to reduce or inactivate the HNH domain (e.g., at positions D396, H397, and / or N420 within the HNH domain). In some specific instances, the variant CRISPR nuclease polypeptide may contain substitutions for H397A, D396A, and / or N420A. Alternatively, any of H397, D396, and N420 may be replaced by an amino acid residue similar to A (e.g., G, L, or S). In one specific instance, the variant CRISPR nuclease polypeptide may contain substitutions for H397A or H397L.
[0097] In other instances, one or more mutations may be located within both the RuvC and HNH domains of nuclease A to reduce or weaken the enzymatic activity of the nuclease.
[0098] In some embodiments, the variant CRISPR nuclease peptides provided herein may comprise, for example, both arginine / lysine substitutions (e.g., arginine substitutions) and nickase mutations at one or more positions provided herein. For example, the variant CRISPR nuclease peptide may comprise arginine / lysine substitutions (e.g., arginine substitutions) at positions I67, D568, and / or E706, and further comprise an amino acid substitution at H397 (e.g., H1397A or H397L) in SEQ ID NO: 1. In some cases, such variants may further comprise N-terminal truncation as disclosed herein, for example, deletions within residues 1-15 of SEQ ID NO: 1, such as deletions of residues 1-14 or 1-15aa in SEQ ID NO: 1. In some cases, such variants may further comprise C-terminal truncation as disclosed herein, for example, deletions within the last five residues of SEQ ID NO: 1.
[0099] In some instances, the variant CRISPR nuclease peptides provided herein may share at least 90% (e.g., 95%, 97%, 98%, 99%, 99.5% or higher) sequence identity with SEQ ID NO: 1. Exemplary engineered variant CRISPR nuclease peptides of nuclease A are listed in Tables 1, 11, or 14, each of which is within the scope of this disclosure. In specific instances, the variant CRISPR nuclease peptide may comprise (e.g., consist of) the amino acid sequence of any of SEQ ID NOs: 6-13, 34, 36, 38, 40, 42, 44, 46, 58, 62, and 64 (e.g., SEQ ID NO: 6 or SEQ ID NO: 36).
[0100] (B) Nuclease K and its engineered variants
[0101] In some embodiments, the CRISPR nuclease polypeptide is derived from nuclease K, which includes wild-type nuclease K (SEQ ID NO: 65), variants thereof (such as those disclosed herein), or fusion polypeptides comprising such wild-type nuclease K and its variants. Variant CRISPR nuclease polypeptides of nuclease K can be generated by introducing one or more mutations into a reference CRISPR nuclease to modulate (e.g., enhance or reduce) one or more activities of the nuclease.
[0102] Reference CRISPR nuclease K (SEQ ID NO: 65, see Table 15 below) is a CRISPR nuclease comprising both a RuvC nuclease domain (located at residues 1-39, 286-320, and 448-543 of SEQ ID NO: 1) and an HNH domain (located at residues 321-447 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate the cleavage of DNA strands adjacent to the 5'-NGG-3' PAM motif, where N represents any nucleotide. Positions D10, E315, and D500 are considered active sites in the RuvC domain, and positions D374, H375, and N398 are considered active sites in the HNH domain. R482 and H497 may also be important for the nuclease activity of the RuvC domain. In addition to the nuclease domain, the reference CRISPR nuclease of SEQ ID NO: 1 also includes the BH domain (residues 40-77 of SEQ ID NO: 1), the REC domain (residues 78-285 of SEQ ID NO: 1), the PLL domain (residues 544-556), the WED domain (residues 557-648), and the PID domain (residues 649-747).
[0103] The reference CRISPR nuclease K of SEQ ID NO: 65 disclosed herein is smaller than that of the CRISPR-Cas9 nuclease from *Streptococcus pyogenes*. The scaffold utilized by the reference CRISPR nuclease of SEQ ID NO: 65 and its variants can be miniaturized. This scaffold contains a unique structure compared to the SpCas9 nuclease scaffold. It is anticipated that this unique scaffold will allow for a reduction in the size of some domains (e.g., the REC domain), and thus contribute to a smaller CRISPR nuclease size. These features will be beneficial for delivery. Arginine substitution and / or lysine substitution may be introduced into the CRISPR nuclease of SEQ ID NO: 65 to increase indel activity.
[0104] Furthermore, the RuvC domain of nuclease K cleaves the non-target strand upstream of PAM within 4 to 6 nucleotides, and the HNH domain cleaves the target strand upstream of PAM within 3 to 4 nucleotides, each producing a cleavage site with a 0-3 nucleotide overhang (most commonly a 3 nucleotide overhang). This cleavage pattern differs from the SpCas9 cleavage pattern described above. Additionally, using a gene editing system containing nuclease K or a variant thereof allows for the introduction of indels larger than those that can be introduced by SpCas9 into the target nucleic acid. For example, insertions induced by nuclease K or a variant thereof can range from about 1 nucleotide to about 10 nucleotides (most commonly about 4 nucleotides). Deletions induced by nuclease A or a variant thereof can range from about 1 nucleotide to about 25 nucleotides or more (most commonly about 11 nucleotides). Furthermore, because nuclease K contains both the RuvC and HNH domains, cleavage enzyme variants can be engineered, for example, by disrupting the nuclease activity of either the RuvC or HNH domains.
[0105] The variants of the nuclease K polypeptide provided herein may contain one or more mutations relative to nuclease K (SEQ ID NO: 65), such as substitution of one or more amino acid residues, deletion of one or more, insertion of one or more, fusion of one or more, or combinations thereof. In some cases, changes may be introduced into the BH domain, WED, PID, or combinations thereof of nuclease K. In some cases, no changes are introduced into the RuvC nuclease domain and / or HNH nuclease domain of nuclease K, or at the active site and / or sites in these domains involved in activity as provided herein. Alternatively, conserved amino acid substitutions may be introduced into SEQ ID NO: 65, including in the RuvC nuclease domain and / or HNH nuclease domain of nuclease K.
[0106] In some embodiments, the variant CRISPR nuclease polypeptide derived from nuclease K provided herein, relative to SEQ ID NO: 65, may contain one or more arginine substitutions, one or more lysine substitutions, or combinations thereof. In some instances, the variant CRISPR nuclease polypeptide derived from nuclease K may contain up to 20 arginine substitutions and / or lysine substitutions (e.g., up to 20 arginine substitutions, up to 20 lysine substitutions, or combinations thereof), for example, up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 arginine substitutions, lysine substitutions, or combinations thereof. In specific instances, the variant CRISPR nuclease polypeptide derived from nuclease K may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 arginine substitutions, lysine substitutions, or combinations thereof. In some instances, the CRISPR nuclease peptides derived from nuclease K presented herein contain arginine substitutions.
[0107] In some cases, arginine substitution and / or lysine substitution may be located in the BH domain, WED domain, PID domain, or any combination thereof of nuclease K. For example, arginine substitution and / or lysine substitution may be introduced into one or more of the following positions: G42, D46, I53, F83, E128, E541, F570, E576, D582, C630, I631, E683, S719, and T734 of SEQ ID NO: 65. In some cases, arginine substitution and / or lysine substitution may be located at positions G42 and D582 of SEQ ID NO: 65. In other cases, arginine substitution and / or lysine substitution may be located at positions I53, D582, and T734 of SEQ ID NO: 65.
[0108] In some instances, variant CRISPR nuclease peptides derived from nuclease K (SEQ ID NO: 65) may contain one or more arginine substitutions from the following: G42R, D46R, I53R, E128R, D582R, F570R, E576R, C630R, I631R, E683R, T734R, and S719R. In other instances, variant CRISPR nuclease peptides derived from nuclease K may contain one or more arginine substitutions from the following: G42R, D46R, I53R, F83R, E541R, F570R, D582R, E683R, and S719R. In some specific instances, variant CRISPR nuclease peptides derived from nuclease K may contain arginine substitutions of D582R and G42R. In other specific instances, variant CRISPR nuclease peptides may contain arginine substitutions for D582R, I53R, and T734R.
[0109] In some cases, arginine substitution and / or lysine substitution may be located within the RuvC nuclease domain and / or HNH nuclease domain of nuclease K. In some instances, arginine substitution and / or lysine substitution may be located within the RuvC nuclease domain of nuclease A to reduce or inactivate the RuvC domain (e.g., at positions D10, E315, R482, H497, and / or D500 in the RuvC domain of SEQ ID NO: 65). In other instances, arginine substitution and / or lysine substitution may be located within the HNH domain of nuclease K, for example, at positions D374, H375, and / or N398 in the HNH domain of SEQ ID NO: 65. In still other instances, arginine substitution and / or lysine substitution may be located within both the RuvC and HNH domains of nuclease K to reduce or weaken the enzymatic activity of the nuclease.
[0110] Alternatively, arginine substitution and / or lysine substitution may not be located at the active site and / or sites involved in activity within the RuvC and / or HNH nuclease domains of nuclease K (e.g., not at positions D10, E315, R482, H497, and / or D500 in the RuvC domain and / or positions D374, H375, and / or N398 in the HNH domain). In some instances, arginine substitution and / or lysine substitution may not be located within the RuvC and / or HNH domains of nuclease K.
[0111] Alternatively or additionally, the variant CRISPR nuclease peptides derived from nuclease K provided herein may contain one or more mutations (e.g., nickase mutations) within the RuvC or HNH nuclease domains to reduce or eliminate the nuclease activity of the target domain, thereby producing variants with nickase activity. Such mutations may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation within the RuvC or HNH nuclease domain of nuclease K is an amino acid substitution, wherein the substituted amino acid residue is not a conserved substitution of the native amino acid residue at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue other than K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue other than R. A set of conserved amino acid residue substitutions is provided herein.
[0112] Positions D10, E315, and D500 are considered active sites within the RuvC domain of nuclease K (SEQ ID NO: 65), and positions D374, H375, and N398 are considered active sites within the HNH domain. R482 and H497 may also be important for the nuclease activity of the RuvC domain of nuclease K. In some instances, one or more mutations may be located within the RuvC nuclease domain to reduce or inactivate the RuvC domain (e.g., at positions D10, E315, R482, H497, and / or D500 within the RuvC domain). In some specific instances, variant CRISPR nuclease peptides may contain substitutions for D10A, E315A, and / or D500A. In other instances, any one of D10, E315, and D500 may be substituted with an amino acid residue similar to A (e.g., G, S, or L).
[0113] In some instances, one or more mutations may be located within the HNH domain (e.g., at positions H375, D374, and / or N398 in the HNH domain of SEQ ID NO: 65). In some specific instances, variant CRISPR nuclease peptides derived from nuclease K may contain substitutions for H375A, D374A, and / or N398A. Alternatively, any of H375, D374, and N398 may be replaced by an amino acid residue similar to A (e.g., G, L, or S). In one specific instance, variant CRISPR nuclease peptides derived from nuclease K may contain substitutions for H375A or H375L.
[0114] In other instances, one or more mutations may be located within both the RuvC and HNH domains of nuclease K to reduce or weaken the enzymatic activity of the nuclease.
[0115] In some embodiments, the variant CRISPR nuclease peptides derived from nuclease K provided herein may include, for example, arginine substitution / lysine substitution (e.g., arginine substitution) at one or more positions provided herein, and both nickase mutations at one or more positions provided herein. In some cases, the variant CRISPR nuclease peptides derived from nuclease K may include arginine substitution / lysine substitution (e.g., arginine substitution) listed in Tables 18 and 20. Such variant CRISPR nuclease peptides may further include one or more mutations that reduce or inactivate the RuvC nuclease domain and / or the HNH nuclease domain (e.g., one or more mutations located within the RuvC nuclease domain and / or the HNH nuclease domain). For example, the variant CRISPR nuclease peptides derived from nuclease K may include arginine substitution / lysine substitution (e.g., arginine substitution) at positions D582, I53, and / or T734, and an amino acid substitution at H375 (e.g., H375A) in SEQ ID NO: 65.
[0116] (C) Nuclease M and its engineered variants
[0117] In some embodiments, the nuclease polypeptide is derived from nuclease M, which includes wild-type nuclease M (SEQ ID NO: 80), variants thereof (such as those disclosed herein), or fusion polypeptides comprising such wild-type nuclease M and its variants. Variant nuclease polypeptides of nuclease M can be generated by introducing one or more mutations into nuclease M to modulate (e.g., enhance or reduce) one or more activities of the nuclease.
[0118] The reference nuclease M of SEQ ID NO: 80 (see Table 21 below) is an RNA-guided nuclease comprising both a RuvC nuclease domain (located at residues 52-84, 160-192, and 289-367 of SEQ ID NO: 80) and an HNH domain (located at residues 193-288 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate the cleavage of the DNA strand adjacent to the 5'-WTAAH-3' PAM motif, where W is A or T, and H is A, C, or T. In one example, the PAM motif is 5'-TTAAA-3'.
[0119] Positions D58, E189, and D341 are considered active sites within the RuvC domain, and positions H243, H244, and H267 are considered active sites within the HNH domain. R329 and H338 may also be important for the nuclease activity of the RuvC domain. In addition to the nuclease domains, the nuclease M of SEQ ID NO: 80 also includes the PLMP domain (residues 1-51 of SEQ ID NO: 1), the BH domain (residues 85-117 of SEQ ID NO: 1), the REC domain (residues 118-159 of SEQ ID NO: 1), the PLL domain (residues 368-380), the WED domain (residues 381-427), and the PID domain (residues 428-497).
[0120] The reference nuclease M of SEQ ID NO: 80 disclosed herein is smaller than the CRISPR-Cas9 nuclease from *Streptococcus pyogenes*. The scaffold utilized by the reference CRISPR nuclease of SEQ ID NO: 80 and its variants can be miniaturized. This scaffold contains a unique structure compared to the SpCas9 nuclease scaffold. The unique scaffold is expected to allow for reduction in the size of some domains (e.g., the REC domain), and thus contribute to a smaller nuclease size. These features will be beneficial for delivery. Arginine substitutions can be introduced into nuclease M SEQ ID NO: 80 to increase indel activity. Compared to SpCas9, nuclease M recognizes different PAM motifs in genomic targets, thereby allowing for the editing of additional gene targets.
[0121] Furthermore, the RuvC domain of nuclease M cleaves the non-target strand upstream of PAM within 3 to 12 nucleotides, and the HNH domain cleaves the target strand upstream of PAM within 3 to 4 nucleotides, each producing a cleavage site with a 0-9 nucleotide overhang (most commonly a 5 nucleotide overhang). This cleavage pattern differs from the SpCas9 cleavage pattern described above. Additionally, unlike SpCas9, the RNA-guided nuclease of SEQ ID NO: 80 contains a PLMP domain. The PLMP domain is expected to bind to the helix at the 3' end of the scaffold and provide increased binding affinity between the RNA-guided nuclease and its homologous scaffold.
[0122] In addition, since nuclease M of SEQ ID NO: 80 contains a RuvC domain and an HNH domain, nickase variants can be engineered, for example, by disrupting the nuclease activity of either the RuvC domain or the HNH domain.
[0123] Variant nuclease polypeptides of nuclease M may contain one or more mutations relative to nuclease M (SEQ ID NO: 80) to regulate (e.g., enhance or reduce) one or more activities of the nuclease, such as substitution of one or more amino acid residues, deletion of one or more, insertion of one or more, fusion of one or more, or combinations thereof. In some cases, mutations may be introduced into the BH domain, REC domain, PLL domain, WED domain, PID domain, or combinations thereof. In some cases, no mutations are introduced into the RuvC nuclease domain and / or HNH nuclease domain, or at active sites and / or sites involved in the activity of these domains as provided herein. Alternatively, conserved amino acid substitutions may be introduced into SEQ ID NO: 80, including in the RuvC nuclease domain and / or HNH nuclease domain.
[0124] In some embodiments, the variant nuclease polypeptide of nuclease M provided herein, relative to SEQ ID NO: 80, may contain one or more arginine substitutions, one or more lysine substitutions, or combinations thereof. In some instances, the variant nuclease polypeptide derived from nuclease M may contain up to 20 arginine substitutions, up to 20 lysine substitutions, or combinations thereof, for example, up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 arginine substitutions, lysine substitutions, or combinations thereof. In specific instances, the variant nuclease polypeptide derived from nuclease M may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 arginine substitutions, lysine substitutions, or combinations thereof. In some specific instances, the variant nuclease polypeptide derived from nuclease M contains arginine substitutions relative to SEQ ID NO: 80.
[0125] In some instances, the variant nuclease polypeptide of nuclease M may contain one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitutions) located at positions E88, S95, L92, E401, E83, N371, P481, and / or A373 in SEQ ID NO: 80. In some cases, the variant nuclease polypeptide derived from nuclease M may contain at least two arginine substitutions and / or lysine substitutions located at these positions, such as at least two arginine substitutions.
[0126] In some cases, arginine substitution and / or lysine substitution (e.g., arginine substitution) may be located in the BH domain, REC domain, PLL domain, WED domain, PID domain, or any combination thereof of nuclease M. Alternatively or additionally, arginine substitution and / or lysine substitution (e.g., arginine substitution) may be located within the RuvC nuclease domain and / or HNH nuclease domain of nuclease M. In some instances, arginine substitution and / or lysine substitution (e.g., arginine substitution) may be located within the RuvC nuclease domain of nuclease M to reduce or inactivate the RuvC domain of nuclease M (e.g., at positions D58, E189, R329, H338, and / or D341 in the RuvC domain of SEQ ID NO: 80). In other instances, arginine substitution and / or lysine substitution (e.g., arginine substitution) may be located within the HNH domain of nuclease M (SEQ ID NO: 80), for example, at positions H243, H244, and / or H267 within the HNH domain. In still other instances, arginine substitution and / or lysine substitution (e.g., arginine substitution) may be located within both the RuvC domain and the HNH domain of nuclease M to reduce or weaken the enzymatic activity of the nuclease.
[0127] Alternatively, arginine substitution and / or lysine substitution (e.g., arginine substitution) may not be located at the active site and / or sites involved in activity of the RuvC nuclease domain and / or HNH nuclease domain of nuclease M (e.g., not at positions D58, E189, R329, H338 and / or D341 in the RuvC domain of SEQ ID NO: 80 and / or positions H243, H244 and / or H267 in the HNH domain). In some instances, arginine substitution and / or lysine substitution (e.g., arginine substitution) may not be located in the RuvC domain and / or HNH domain of nuclease M.
[0128] Exemplary arginine substitution variants of nuclease M are listed in Table 24 below, each of which is within the scope of this disclosure.
[0129] Alternatively or additionally, the variant nuclease polypeptides of nuclease M provided herein may contain one or more mutations (e.g., nickase mutations) located within the RuvC nuclease domain or the HNH nuclease domain to reduce or eliminate the nuclease activity of the target domain, thereby producing variants with nickase activity. Such mutations may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation within the RuvC nuclease domain or the HNH nuclease domain is an amino acid substitution, wherein the substituted amino acid residue is not a conserved substitution of the native amino acid residue located at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue other than K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue other than R. A set of conserved amino acid residue substitutions is provided herein.
[0130] In some instances, one or more mutations may be located within the RuvC nuclease domain to reduce or inactivate the RuvC domain of nuclease M (e.g., at positions D58, E189, R329, H338, and / or D341 in the RuvC domain of SEQ ID NO: 80). Alternatively, one or more nickase mutations may be located within the HNH domain of nuclease M (e.g., at positions H243, H244, and / or H267 in the HNH domain of SEQ ID NO: 80). In specific instances, variant nuclease polypeptides of nuclease M may contain one or more nickase mutations at one or more of positions D58, E189, D341, H243, H244, H267, R329, and / or H338 in SEQ ID NO: 80. In some cases, one or more original residues in the original RuvC or HNH domain of nuclease M may be replaced by alanine (A). Alternatively, one or more of the original residues in the RuvC domain or HNH domain may be replaced by an amino acid residue similar to A (e.g., G, S, or L).
[0131] In other instances, one or more mutations may be located within both the RuvC domain and the HNH domain of nuclease M to reduce or weaken the enzymatic activity of the nuclease.
[0132] In some embodiments, the variant nuclease polypeptide derived from nuclease M provided herein may include, for example, arginine substitution / lysine substitution (e.g., arginine substitution) at one or more positions provided herein, and both, for example, nickase mutations at one or more positions provided herein.
[0133] In some instances, variant nuclease polypeptides derived from nuclease M may share at least 90% (e.g., 95%, 97%, 98%, 99%, 99.5% or higher) sequence identity with SEQ ID NO: 80.
[0134] In some embodiments, the CRISPR nuclease polypeptide derived from nuclease A, nuclease K, or nuclease M, as provided herein, may be a fusion polypeptide comprising the CRISPR nuclease moiety as disclosed herein and one or more additional functional elements. In some cases, one or more additional functional elements may be heterologous to the CRISPR nuclease moiety.
[0135] As used herein, the term "fusion" refers to the connection of at least two nucleotides or protein molecules. For example, "fusion" can refer to the connection of at least two polypeptide domains encoded by a single gene in nature. Fusions can be N-terminal fusions, C-terminal fusions, or intramolecular fusions. In some respects, the domains are transcribed and translated to produce a single polypeptide.
[0136] In some cases, the CRISPR nuclease moiety in the fusion peptide may be a reference CRISPR nuclease A (SEQ ID NO: 1), nuclease K (SEQ ID NO: 65), or nuclease M (SEQ ID NO: 80). Alternatively, the CRISPR nuclease moiety in the fusion peptide may be a variant of any reference nuclease as disclosed herein. Exemplary additional functional moieties included in the fusion peptide include peptide tags, fluorescent proteins, base editing domains, DNA methylation domains, histone residue modification domains, localization factors, transcriptional modification factors, light-gated control factors, chemically induced factors, chromatin visualization factors, or combinations thereof.
[0137] In some embodiments, additional functional portions may include nuclear localization signals (NLS), nuclear output signals (NES), or combinations thereof. In some instances, the fusion peptide may contain an NLS, which may be located at the N-terminus or the C-terminus. In specific instances, the fusion peptide may contain a first NLS located at the N-terminus and a second NLS located at the C-terminus. The first NLS fragment and the second NLS fragment may be identical. Alternatively, the two NLS fragments may be different. In some embodiments, the fusion peptide may contain an NLS near the N-terminus and / or near the C-terminus (e.g., within about 1, 2, 3, 4, or 5 amino acids of the first or last amino acid of the CRISPR nuclease). In some embodiments, the fusion peptide may contain an NLS located within the flexible loop of the CRISPR nuclease. Exemplary fusion CRISPR nuclease peptides containing one or more NLS signals are provided in Tables 1, 15, and 21.
[0138] In some embodiments, an additional functional component may be a flexible peptide linker, such as an XTEN peptide linker or a G / S-rich peptide linker. Examples of such peptide linkers are provided in Example 1 below, which may be applicable to any CRISPR nuclease peptide disclosed herein.
[0139] B. Preparation of CRISPR nuclease peptides
[0140] The CRISPR nuclease peptides disclosed herein can be prepared using conventional methods or methods disclosed herein. For example, CRISPR nuclease peptides can be prepared by culturing host cells (such as bacterial or mammalian cells) capable of producing nuclease peptides, isolating the nuclease peptides thus produced, and optionally purifying the nuclease peptides. CRISPR nuclease peptides prepared in this way can be complexed with gRNA.
[0141] CRISPR nuclease peptides can also be prepared via an in vitro coupled transcription-translation system and optionally complexed with gRNA. The bacteria that can be used to prepare CRISPR nuclease peptides are not particularly limited, provided that the bacteria are capable of producing CRISPR nuclease peptides. Some non-limiting examples of bacteria include *E. coli* cells described herein.
[0142] Unless otherwise stated, all compositions, complexes, and peptides provided herein are prepared with reference to the activity levels of the stated composition, complex, or peptide and are free from impurities that may be present in commercially available sources, such as residual solvents or byproducts. Enzymatic component weights are based on total active protein. Unless otherwise stated, all percentages and ratios are by weight. Unless otherwise stated, all percentages and ratios are based on the total composition. In the illustrated compositions, the enzymatic level is expressed as the weight of the pure enzyme in the total composition, and unless otherwise stated, components are expressed by weight of the total composition.
[0143] (i) Carrier
[0144] This disclosure provides vectors for expressing CRISPR nuclease peptides. In some embodiments, the vectors disclosed herein comprise nucleotide sequences encoding the CRISPR nuclease peptides as provided herein. In some embodiments, the vectors comprise a Pol II promoter or a Pol III promoter.
[0145] Expression of natural or synthetic polynucleotides is typically achieved by operatively linking a polynucleotide encoding a CRISPR nuclease polypeptide to a promoter and incorporating the construct into an expression vector. The expression vector is not particularly limited, as long as it contains a polynucleotide encoding a CRISPR nuclease polypeptide and is suitable for replication and integration in eukaryotic cells.
[0146] Typical expression vectors include transcription and translation terminators, initiation sequences, and promoters for the desired polynucleotide expression. For example, plasmid vectors carrying recognition sequences for RNA polymerase (pSP64, pBluescript, etc.) can be used. Vectors derived from retroviruses, such as lentiviruses, are suitable tools for achieving long-term gene transfer because they allow for the long-term stable integration of transgenes and their propagation in daughter cells. Examples of vectors include expression vectors, replication vectors, probe-generating vectors, and sequencing vectors. Expression vectors can be delivered to cells in the form of viral vectors.
[0147] Viral vector technology is well-known in the field and is described in various virology and molecular biology handbooks. Viruses that can be used as vectors include, but are not limited to, bacteriophages, retroviruses, adenoviruses, adeno-associated viruses, herpesviruses, and lentiviruses. Generally, a suitable vector contains an origin of replication that is functional in at least one organism, a promoter sequence, a convenient restriction endonuclease site, and one or more selectable markers.
[0148] The type of vector is not particularly limited, and vectors that can be expressed in host cells can be appropriately selected. More specifically, depending on the type of host cell, a promoter sequence that ensures the expression of the polypeptide from the polynucleotide is appropriately selected, and this promoter sequence and polynucleotide are inserted into any different plasmid, etc., for the preparation of the expression vector.
[0149] Additional promoter elements (e.g., enhancer sequences) regulate the frequency of transcription initiation. These are typically located in a region 30–110 bp upstream of the start site, but many promoters have recently been shown to also contain functional elements downstream of the start site. Depending on the promoter, it appears that individual elements can act synergistically or independently to activate transcription.
[0150] Furthermore, this disclosure is not limited to the use of constitutive promoters. Inducible promoters are also contemplated as part of this disclosure. The use of inducible promoters provides a molecular switch capable of turning on the expression of a polynucleotide sequence operatively linked to it when such expression is desired, or turning off said expression when it is not desired. Examples of inducible promoters include, but are not limited to, metallothionein promoters, glucocorticoid promoters, progesterone promoters, and tetracycline promoters.
[0151] The expression vector to be introduced may also contain a selectable marker gene or a reporter gene, or both, to facilitate the identification and selection of expressing cells from a population of cells attempting to be transfected or infected by a viral vector. In other respects, the selectable marker may be carried on a separate DNA fragment and used in a co-transfection procedure. Both the selectable marker and the reporter gene may be side-linked with an appropriate transcriptional control sequence to enable expression in host cells. Examples of such markers include dihydrofolate reductase genes and neomycin resistance genes for eukaryotic cell culture, and tetracycline resistance genes and ampicillin resistance genes for culturing *E. coli* and other bacteria. By using such selectable markers, it can be confirmed whether the polynucleotide encoding the polypeptide of the present invention has been transferred to the host cell and then successfully expressed.
[0152] The methods for preparing recombinant expression vectors are not particularly limited, and examples include methods using plasmids, phages, or granules.
[0153] (ii) Methods of expression
[0154] This disclosure includes methods for protein expression, said methods comprising translating the CRISPR nuclease polypeptide described herein.
[0155] In some embodiments, the host cells described herein are used to express CRISPR nuclease peptides. The host cells are not particularly limited, and a variety of known cells may preferably be used. Specific examples of host cells include bacteria such as *Escherichia coli*, yeasts (budding yeasts, *Saccharomyces cerevisiae*, and *Schizosaccharomyces pombe*), nematodes (*Caenorhabditis elegans*), Xenopus laevis oocytes, and animal cells (e.g., CHO cells, COS cells, and HEK293 cells). The methods used to transfer the expression vectors described above into host cells, i.e., the transformation methods, are not particularly limited, and known methods such as electroporation, calcium phosphate methods, liposome methods, and DEAE dextran methods may be used.
[0156] After transforming the host cell with the expression vector, the host cell can be cultured, propagated, or multiplied to produce the CRISPR nuclease peptide. Following expression, the host cell can be collected, and the CRISPR nuclease peptide can be purified from the culture or other sources using conventional methods (e.g., filtration, centrifugation, cell disruption, gel filtration chromatography, ion exchange chromatography, etc.).
[0157] A variety of methods can be used to determine the production levels of mature CRISPR nuclease peptides in host cells. Such methods include, but are not limited to, methods using, for example, polyclonal or monoclonal antibodies that are specific to proteins or to tagging as described elsewhere herein. Exemplary methods include, but are not limited to, enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (MA), fluorescence immunoassay (FIA), and fluorescence activated cell sorting (FACS). These and other assays are well known in the art (see, for example, Maddox et al., Journal of Experimental Medicine 158:1211
[1983] ).
[0158] This disclosure provides methods for expressing CRISPR nuclease peptides (and optionally gRNAs in the gene editing systems disclosed herein) in vivo. Such methods may include providing a polynucleotide encoding a CRISPR nuclease peptide to host cells in a subject (e.g., a human subject), wherein the polynucleotide encodes a CRISPR nuclease peptide derived from the cells that expresses the CRISPR nuclease peptide.
[0159] II. Gene Editing Systems
[0160] In some respects, this disclosure provides a gene editing system with enhanced gene editing efficiency. The gene editing system comprises any CRISPR nuclease polypeptide or nucleic acid encoding a CRISPR nuclease as disclosed herein, derived from nuclease A, nuclease K, or nuclease M, and one or more guide RNAs (gRNAs) or nucleic acids encoding gRNAs.
[0161] CRISPR nuclease
[0162] In some embodiments, the gene editing systems disclosed herein comprise CRISPR nuclease polypeptides as provided herein, such as any reference CRISPR nuclease A, nuclease K, or nuclease M or variants thereof, for example, comprising one or more arginine substitutions and / or lysine substitutions (e.g., arginine substitution), one or more nickase mutations, additional mutations such as N-terminal truncation (if applicable), or combinations thereof. See the above disclosure. Such protein components may form complexes with gRNAs in the same gene editing system. Alternatively, the gene editing system comprises nucleic acids encoding CRISPR nuclease polypeptides. In some cases, the nucleic acid may be an expression vector (e.g., a viral vector) for generating the encoded nuclease polypeptide in a host cell. In some cases, the expression vector may further comprise coding sequences for generating one or more gRNAs of the gene editing system.
[0163] Guide RNA
[0164] The gene editing systems disclosed herein further comprise one or more gRNAs or nucleic acids encoding such gRNAs. As used herein, the terms “RNA guide,” “RNA guide sequence,” or “guide RNA (gRNA)” refer to an RNA molecule or modified RNA molecule that facilitates the targeting of the CRISPR nucleases described herein to a genomic site of interest. For example, an RNA guide may be a molecule comprising a spacer sequence and a scaffold sequence. The spacer sequence recognizes (e.g., binds to) a site in a non-PAM strand complementary to a target sequence in the PAM strand, for example, and is designed to be complementary to a specific nucleic acid sequence. The scaffold sequence contains a nuclease-binding sequence for binding to a CRISPR nuclease. In some embodiments, the scaffold is an RNA sequence.
[0165] In some cases, the gRNA disclosed herein may further include adapter sequences, 5' end protection fragments and / or 3' end protection fragments, or combinations thereof.
[0166] (i) Interval subsequence
[0167] As used herein, the terms "spacer" and "spacer sequence" (also known as DNA-binding sequence) are part of an RNA guide for the RNA equivalent of a target sequence (DNA sequence). A spacer contains a sequence capable of binding to a non-PAM strand through base pairing at a site complementary to the target sequence (which is in the PAM strand). Such spacers are also referred to as being specific to the target sequence. In some cases, the spacer may be at least 75% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%) identical to the target sequence, except for RNA-DNA sequence differences. In some cases, the spacer may be 100% identical to the target sequence, except for RNA-DNA sequence differences.
[0168] The gene editing system disclosed herein comprises one or more gRNAs, each containing a spacer sequence and a scaffold sequence that are specific to a target sequence at a genomic site of interest, the scaffold sequence being recognizable by a CRISPR nuclease peptide contained in the gene editing system.
[0169] The target sequence associated with nuclease A or its variants as disclosed herein may be adjacent to the 5'-NGG-3' PAM (e.g., upstream of it or at 5'), where N represents any nucleotide.
[0170] The target sequence associated with nuclease K or its variants as disclosed herein may be adjacent to the 5'-NGG-3' PAM (e.g., upstream of it or at 5'), where N represents any nucleotide.
[0171] The target sequence associated with nuclease M can be adjacent to the 5'-WTAAH-3' PAM (e.g., upstream of it or at 5'), where W is A or T, and H is A, C, or T. In some instances, the PAM can be 5'-TTAAA-3'.
[0172] As used herein, the term "protospacer adjacent motif" or "PAM sequence" refers to a DNA sequence adjacent to the target sequence. In some embodiments, the PAM sequence is essential for binding CRISPR nucleases and / or indel activity. In a double-stranded DNA molecule, the strand containing the PAM motif is referred to as the "PAM strand," and the complementary strand is referred to as the "non-PAM strand." The gRNA binds to a site in the non-PAM strand that is complementary to the target sequence disclosed herein, and the PAM sequence as described herein is present in the PAM strand. The PAM motif may be located upstream of the target sequence.
[0173] As used herein, the term "proximity" refers to a nucleotide or amino acid sequence that is very close to another nucleotide or amino acid sequence. In some embodiments, if no nucleotide separates the two sequences, the nucleotide sequence is adjacent to the other nucleotide sequence (i.e., directly adjacent). In some embodiments, if a small number of nucleotides separate the two sequences (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides), the nucleotide sequence is adjacent to the other nucleotide sequence. In some embodiments, if the two sequences are separated by about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides, the first sequence is adjacent to the second sequence. In some embodiments, if the two sequences are separated by at most 2, 5, 8, 10, 12, or 15 nucleotides, the first sequence is adjacent to the second sequence. In some embodiments, if two sequences are separated by 2-5 nucleotides, 4-6 nucleotides, 4-8 nucleotides, 4-10 nucleotides, 6-8 nucleotides, 6-10 nucleotides, 6-12 nucleotides, 8-10 nucleotides, 8-12 nucleotides, 10-12 nucleotides, 10-15 nucleotides, or 12-15 nucleotides, then the first sequence is adjacent to the second sequence.
[0174] In specific instances, the spacer is specific to the target sequence at the genomic site of interest, with the target sequence immediately adjacent to the PAM motif. In other specific instances, the target sequence and the PAM motif may have small gaps of less than 5 nucleotides (e.g., 1, 2, 3, 4, or 5 nucleotides).
[0175] The length of the spacer sequence disclosed herein can be from about 15 nucleotides to about 30 nucleotides. For example, the length of the spacer can be from about 15 nucleotides to about 20 nucleotides, from about 15 nucleotides to about 25 nucleotides, from about 20 nucleotides to about 25 nucleotides, or from about 20 nucleotides to about 30 nucleotides. In some embodiments, the spacer in the gRNA can typically be designed to be between 15 and 25 nucleotides in length (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, and 25 nucleotides) and complementary to the specific target sequence. In some embodiments, the spacer sequence can be designed to be between 18 and 22 nucleotides in length (e.g., 20 nucleotides).
[0176] In some embodiments, the spacer sequence may have at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.5% sequence identity with the target sequence as described herein, and may be able to bind to the complementary region of the target sequence via base pairing.
[0177] In some embodiments, the spacer sequence contains only RNA bases. In some embodiments, the spacer sequence contains DNA bases (e.g., the spacer contains at least one thymine). In some embodiments, the spacer sequence contains both RNA and DNA bases (e.g., the DNA-binding sequence contains at least one thymine and at least one uracil).
[0178] (ii) Stent sequence
[0179] The scaffold sequence in gRNA can also be recognized by CRISPR nuclease peptides in gene editing systems.
[0180] In some cases, the scaffold sequence can be recognized by nuclease A (SEQ ID NO: 1) or any engineered variant thereof, as disclosed herein. Such scaffold sequences may include the following SEQ ID NO: 2.
[0181] GUUACAGUUAAGGCUCUUUGGAAACAAAGAAGCCUUAAUUGUAAAACGCCUAUAUGGUAAAGUGAUGUACGUUUGGGUAUAUAUCGCCAGCCUGAACCUCUACGCCAGAAAUGGCAGCUUUAUCAUGGGUUAGGACGAUAUUUAAAAAACUUUCUGCGGCUUGCUUACUUUAGUAAGCUUUGUGGCUGAGGCAGAAUUCCUU (SEQ ID NO: 2)
[0182] Figure 3A The predicted secondary structure of the reference stent sequence of SEQ ID NO: 2 is provided.
[0183] In other cases, the scaffold sequence recognizable by nuclease A (SEQ ID NO: 1) or any engineered variant thereof may be a variant derived from SEQ ID NO: 2. Such variant scaffold sequences may contain at least 80% (e.g., at least 85%, 90%, 95%, 98%, or higher) the same nucleotide sequence as SEQ ID NO: 2. Alternatively or additionally, variant scaffold sequences may contain deletions, nucleotide substitutions, or combinations thereof. The binding of the variant CRISPR nuclease peptide to the variant scaffold sequence is increased compared to the scaffold of SEQ ID NO: 2. In some instances, the variant scaffold may be a fragment of SEQ ID NO: 2 as disclosed herein, or a variant thereof. For example, the length of the variant scaffold used for the gRNA provided herein may be in the range of 100-150.
[0184] In some embodiments, the scaffold sequence recognizable by nuclease A (SEQ ID NO: 1) or any engineered variant thereof may be a truncated variant of SEQ ID NO: 2. Any truncated variants provided herein may be approximately 110-140 nucleotides in length.
[0185] In some instances, truncated variant scaffold sequences recognizable by nuclease A (SEQ ID NO: 1) or any engineered variant thereof may have a 3' truncation relative to SEQ ID NO: 2. Such a 3' truncation can remove… Figure 3A One or more stem structures of P6a, P6b, and P6c depicted in the diagram. In a specific instance, the 3' truncation may include a deletion within residues 143-202 of SEQ ID NO: 2, for example, a deletion of the 143-202 segment of SEQ ID NO: 2. See also Figure 3B .
[0186] Alternatively or additionally, the variant scaffold sequence may be a truncated variant of SEQ ID NO: 2, said truncated variant having an internal truncation relative to SEQ ID NO: 2. For example, such a truncated variant may contain Figure 3A The stem structure P1 depicted in the diagram may be entirely or partially deleted. In some instances, the truncated variant may contain deletions of residues 14-20, residues 25-32, or combinations thereof of SEQ ID NO: 2. In specific instances, the truncated variant may contain deletions of fragments 14-20 and 25-32 of SEQ ID NO: 2. See, for example, Figure 3B and 3C In some cases, truncated variants may include deletions in the loop consisting of residues 77-86 of SEQ ID NO: 2 (see [link to relevant documentation]). Figure 3C This could lead to Figure 3A The stem structure P5 depicted in the figure is shortened. In some instances, the truncated variant may contain a deletion within residues 81-85 of SEQ ID NO: 2 (e.g., a deletion of the 81-85 fragment).
[0187] In some instances, the variant scaffold sequence recognizable by nuclease A (SEQ ID NO: 1) or any engineered variant thereof may comprise a combination of one or more internal truncations, such as the 3' truncation and internal truncation disclosed herein. Alternatively or additionally, the variant scaffold sequence may comprise one or more nucleotide variations relative to the corresponding residues in SEQ ID NO: 2.
[0188] In one specific instance, the scaffold sequence recognizable by nuclease A (SEQ ID NO: 1) or any engineered variant thereof comprises the nucleotide sequence of SEQ ID NO: 27 (e.g., constitutes thereof). In another specific instance, the scaffold sequence comprises SEQ ID NO: 28 (e.g., constitutes thereof). The length of such scaffold sequences can be approximately 115-130 nt.
[0189] In some cases, the scaffold sequence can be recognized by nuclease K (SEQ ID NO: 65) or any engineered variant thereof, as disclosed herein. Such scaffold sequences may include the following SEQ ID NO: 79.
[0190] GUUACAGUUAAGGCUCUGAAAAGAGCCUUAAUUGUAAAACGCCUAUACAGUGAAGGGAUAUACGCUUGGGUUUGUCCAGCCUGAGCCUCUAUGCCAGAAAUGGCGCCUUCAUCGUGGGUUAGGACAUUUAAAUUUAAAAACUAUUCAGCACUGUUUGCUCUUGUCAGCUUGGUGGCAGA (SEQ ID NO: 79)
[0191] In other cases, the scaffold sequence recognizable by nuclease K (SEQ ID NO: 65) or any engineered variant thereof may be a variant derived from SEQ ID NO: 79. Such variant scaffold sequences may contain at least 80% (e.g., at least 85%, 90%, 95%, 98%, or higher) the same nucleotide sequence as SEQ ID NO: 79. Alternatively or additionally, variant scaffold sequences may contain deletions, nucleotide substitutions, or combinations thereof. The binding of variant CRISPR nuclease peptides to variant scaffold sequences is increased compared to the scaffold of SEQ ID NO: 79. In some instances, the variant scaffold may be a fragment of SEQ ID NO: 79 as disclosed herein, or a variant thereof. For example, the length of variant scaffold sequences used for the gRNA provided herein may be in the range of 100-150 nt.
[0192] In some cases, the scaffold sequence can be recognized by nuclease M (SEQ ID NO: 80) or any engineered variant thereof, as disclosed herein. Such scaffold sequences may include the following SEQ ID NO: 94.
[0193] GUCAACUACCCCCGUCUAAAGACGGAGGCAUGAGGUUUCGUAACCAAGUGUUGUACCUGCGGGUACAGUAGUUGAACAGGCGGCGAUGCGGCUGGGCACUCCAGGAUGCCACUCCCAGUCCCGGACACUGCCGACGAGCCGCAUCAAGCCGGGGGAGACCAACCGGCUAACGAUAGCCGAGCAAUUACCUAAAAAGAGGUGCAAAGGAAAUGGUAU (SEQ ID NO: 94)
[0194] In other cases, the scaffold sequence recognizable by nuclease M (SEQ ID NO: 80) or any engineered variant thereof may be a variant derived from SEQ ID NO: 94. Such variant scaffold sequences may contain at least 80% (e.g., at least 85%, 90%, 95%, 98%, or higher) the same nucleotide sequence as SEQ ID NO: 94. Alternatively or additionally, variant scaffold sequences may contain deletions, nucleotide substitutions, or combinations thereof. The binding of the variant CRISPR nuclease peptide to the variant scaffold sequence is increased compared to the scaffold of SEQ ID NO: 94. In some instances, the variant scaffold may be a fragment of SEQ ID NO: 94 as disclosed herein, or a variant thereof. For example, the length of the variant scaffold used for the gRNA provided herein may be in the range of 150-200 nt.
[0195] In any gRNA disclosed herein, the scaffold sequence may be located at the 3' end of the spacer sequence. In some cases, the scaffold sequence and the spacer sequence are directly linked. In other cases, the scaffold sequence and the spacer sequence are linked via nucleotide linkers.
[0196] Nucleic acid modification
[0197] Any guide RNA or encoding nucleic acid (e.g., mRNA encoding a CRISPR nuclease polypeptide) in the gene editing systems disclosed herein may include one or more modifications.
[0198] Exemplary modifications may include any modification to sugars, nucleobases, nucleoside bonds (e.g., to the phosphate ester / phosphodiester bond / phosphodiester backbone), and any combination thereof. Some exemplary modifications provided herein are described in detail below.
[0199] The nucleic acid sequence of the gRNA or any component encoding the composition may include any available modifications such as to sugars, nucleobases, or nucleoside bonds (e.g., to phosphate linkages, phosphodiester bonds, or the phosphodiester backbone). One or more atoms of the pyrimidine nucleobase may be replaced or substituted with an optionally substituted amino group, an optionally substituted thiol group, an optionally substituted alkyl group (e.g., methyl or ethyl), or a halogen (e.g., chlorine or fluorine). One or more atoms of the purine nucleobase may be replaced or substituted with an optionally substituted amino group, an optionally substituted thiol group, an optionally substituted alkyl group (e.g., methyl or ethyl), or a halogen (e.g., chlorine or fluorine). In some embodiments, the modification (e.g., one or more modifications) is present in each of the sugar and nucleoside bonds. In some embodiments, the nucleic acid sequence of the gRNA or any component encoding the composition may include a base-degrading sites (i.e., positions without purines or pyrimidines). Modifications may be modifications of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA), threonucleic acid (TNA), glycol nucleic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), or hybrids thereof. This article describes other modifications.
[0200] In some embodiments, modifications may include chemical modifications or cell-induced modifications. For example, Lewis and Pan describe some non-limiting examples of intracellular RNA modifications in “RNA modifications and structures cooperateto guide RNA-protein interactions” in Nature Reviews Molecular Cell Biology, 2017, 18:202-210.
[0201] Different sugar modifications, nucleotide modifications, and / or nucleotide inter-bonds (e.g., backbone structure) can be present at different positions in the sequence. Those skilled in the art will understand that nucleotide analogs or other modifications can be located at any position in the sequence such that the function of the sequence is not substantially diminished. The sequence may include about 1% to about 100% of modified nucleotides (relative to the total nucleotide content, or relative to one or more types of nucleotides, i.e., any one or more of A, G, U, or C) or any intermediate percentage (e.g., 1% to 20%, 1% to 25%, 1% to 50%, 1% to 60%, 1% to 70%, 1% to 80%, 1% to 90%, 1% to 95%, 10% to 20%, 10% to 25%, 10% to 50%, 10% to 60%, 10% to 70%, 10% to 80%, 10% to 90%, 10% to 95%, 10% to ... 100%, 20% to 25%, 20% to 50%, 20% to 60%, 20% to 70%, 20% to 80%, 20% to 90%, 20% to 95%, 20% to 100%, 50% to 60%, 50% to 70%, 50% to 80%, 50% to 90%, 50% to 95%, 50% to 100%, 70% to 80%, 70% to 90%, 70% to 95%, 70% to 100%, 80% to 90%, 80% to 95%, 80% to 100%, 90% to 95%, 90% to 100%, and 95% to 100%.
[0202] In some embodiments, sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions at one or more ribonucleotides of the sequence, as well as backbone modifications, may include modifications or substitutions of phosphodiester bonds. Specific examples of sequences include, but are not limited to, sequences comprising a modified backbone or sequences without natural internucleotide bonds, such as internucleotide modifications including modifications or substitutions of phosphodiester bonds. In addition, sequences having a modified backbone include sequences without phosphorus atoms in the backbone. For the purposes of this application, and as sometimes referred to in the art, modified RNAs without phosphorus atoms in their internucleotide backbones may also be considered oligonucleotides. In certain embodiments, the sequence will comprise ribonucleotides having phosphorus atoms in their internucleotide backbones.
[0203] The modified sequence backbone may include, for example, thiophosphates, chiral thiophosphates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkylphosphonates, such as 3'-alkylene phosphonates and chiral phosphonates, hypophosphonates, phosphoramidites, such as 3'-aminophosphatides and aminoalkylphosphamides, thiocarbonylphosphatides, thiocarbonylalkylphosphonates, thiocarbonylalkyl phosphate triesters, and borane phosphates having normal 3'-5' bonds, 2'-5' linked analogs of these esters, and esters having reverse polarity, wherein adjacent pairs of nucleoside units are linked in 3'-5' to 5'-3' or 2'-5' to 5'-2' configurations. Various salts, mixed salts, and free acid forms are also included. In some embodiments, the sequence may be negatively or positively charged.
[0204] Modified nucleotides that can be incorporated into the sequence can be modified at nucleoside internucleotides (e.g., phosphate backbones). In this document, the phrases “phosphate ester” and “phosphate diester” are used interchangeably in the context of a polynucleotide backbone. The backbone phosphate ester group can be modified by replacing one or more oxygen atoms with different substituents. Further, modified nucleosides and nucleotides can include large-scale substitution of the unmodified phosphate ester portion with another nucleoside internucleotide as described herein. Examples of modified phosphate ester groups include, but are not limited to, thiophosphates, selenophosphates, boranophosphates, boranophosphate esters, hydrogen phosphates, phosphoramide esters, phosphate diamide esters, alkyl or aryl phosphonates, and triphosphate esters. In dithiophosphates, both non-linked oxygen atoms are replaced by sulfur. The phosphate ester linker can also be modified by replacing the linking oxygen with nitrogen (bridged phosphoramide ester), sulfur (bridged thiophosphate), and carbon (bridged methylene phosphonate).
[0205] The α-thiosubstituted phosphate moiety imparts stability to RNA and DNA polymers via non-natural thiophosphate backbone bonds. Thiophosphate DNA and RNA exhibit increased nuclease resistance and thus a longer half-life in the cellular environment.
[0206] In specific embodiments, the modified nucleosides include α-thio-nucleosides (e.g., 5'-O-(1-thiophosphate)-adenosine, 5'-O-(1-thiophosphate)-cytidine (α-thiocytidine), 5'-O-(1-thiophosphate)-guanosine, 5'-O-(1-thiophosphate)-uridine, or 5'-O-(1-thiophosphate)-pseuuridine).
[0207] This document describes other nucleoside bonds that can be used according to the present invention, including nucleoside bonds that do not contain phosphorus atoms.
[0208] In some embodiments, the sequence may include one or more cytotoxic nucleosides. For example, cytotoxic nucleosides may be incorporated into the sequence, such as through bifunctional modifications. Cytotoxic nucleosides may include, but are not limited to, adenosine arabinoside, 5-azacytidine, 4'-thio-cytarabine, cyclopentenylcytosine, cladribine, clofarabine, cytarabine, cytosine arabinoside, 1-(2-C-cyano-2-deoxy-β-D-arabinose-furanopentosyl)-cytosine, decitabine, 5-fluorouracil, and fludarabine. Examples of cytarabine include fluxuridine, gemcitabine, tegafur, and combinations of uracil, tegafur ((RS)-5-fluoro-1-(tetrahydrofuran-2-yl)pyrimidin-2,4(1H,3H)-dione), troxacitabine, tezacitabine, 2'-deoxy-2'-methylenecytidine (DMDC), and 6-mercaptopurine. Other examples include fludarabine phosphate, N4-behenyl-1-β-D-arabinose-furanose cytosine, N4-octadecyl-1-β-D-arabinose-furanose cytosine, N4-palmitoyl-1-(2-C-cyano-2-deoxy-β-D-arabinose-furanose)cytosine, and P-4055 (cytarabine 5'-trans oleate).
[0209] In some embodiments, the sequence includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, poly-A sequence, methylation, acylation, phosphorylation, methylation of lysine and arginine residues, acetylation, and nitrosylation of thiol and tyrosine residues, etc.). One or more post-transcriptional modifications can be any post-transcriptional modification, such as any of the more than one hundred different nucleoside modifications identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucleic Acid Research 27: 196-197). In some embodiments, the first isolated nucleic acid comprises messenger RNA (mRNA). In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of: pyridine-4-ketoribonucleotide, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyluridine, 1-carboxymethyl-pseudouridine, 5-propynyluridine, 1-propynyl-pseudouridine, 5-taurylmethyluridine, 1-taurylmethyl-pseudouridine, 5-taurylmethyl-2-thio- Urate, 1-taurylmethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-denitro-pseudouridine, 2-thio-1-methyl-1-denitro-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of: 5-aza-cytidine, pseudocytidine, 3-methylcytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudocytidine, pyrrolo-cytidine, pyrrolo-pseudocytidine, 2-thio-cytidine, 2-thio-5-methylcytidine, 4-thio-pseudocytidine, 4-thio-1-methyl- Pseudoisocytidine, 4-thio-1-methyl-1-denitro-pseudoisocytidine, 1-methyl-1-denitro-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine.In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of: 2-aminopurine, 2,6-diaminopurine, 7-deadenine, 7-deaden-8-azaadenine, 7-deaden-2-aminopurine, 7-deaden-8-aza-2-aminopurine, 7-deaden-2,6-diaminopurine, 7-deaden-8-aza-2,6-diaminopurine, 1-methyladenosine, N6 -Methyl adenosine, N6-isopentenyl adenosine, N6-(cis-hydroxyisopentenyl) adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl) adenosine, N6-glycylcarbamoyl adenosine, N6-threonylcarbamoyl adenosine, 2-methylthio-N6-threonylcarbamoyl adenosine, N6,N6-dimethyl adenosine, 7-methyl adenosine, 2-methylthio-adenosine, and 2-methoxy-adenosine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of: inosine, 1-methyl-inosine, woyoside, woyoside, 7-deazoguanosine, 7-deazo-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deazo-guanosine, 6-thio-7-deazo-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methyl-guanosine, N2-methyl-guanosine, N2,N2-dimethyl-guanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.
[0210] The sequence may or may not be uniformly modified along the entire length of the molecule. For example, one or more or all types of nucleotides (e.g., naturally occurring nucleotides, purines or pyrimidines, or any one or all of A, G, U, C, I, pU) may or may not be uniformly modified in the sequence or in its given predetermined sequence region. In some embodiments, the sequence includes pseudouridine. In some embodiments, the sequence includes inosine, which may help the immune system characterize the sequence as endogenous relative to viral RNA. The incorporation of inosine may also mediate an increase in RNA stability / a decrease in degradation. See, for example, Yu, Z. et al., (2015) RNA editing by ADAR1 marks dsRNA as “self”. Cell Research 25, 1283-1284, which is incorporated in its entirety by reference.
[0211] In some embodiments, any RNA sequence described herein may contain end modifications (e.g., 5' or 3' modifications). In some embodiments, the end modifications are chemical modifications. In some embodiments, the end modifications are structural modifications. See the disclosure herein.
[0212] When the gene editing systems disclosed herein contain nucleic acids encoding CRISPR nucleases, such nucleic acid molecules may contain any modifications disclosed herein (if applicable).
[0213] III. Gene editing methods
[0214] Any gene editing system can be used to modify (edit) a target nucleic acid, which can be a gene site of interest, such as a gene site that needs to be edited, for example, to repair gene mutations, introduce protective mutations, or introduce modifications to regulate gene expression.
[0215] The gene editing systems and compositions disclosed herein are suitable for editing and introducing edits into a variety of target sequences. In some embodiments, the target sequence is a DNA molecule, such as a DNA locus (referred to herein as a target sequence or target sequence).
[0216] For gene editing systems containing nuclease A (SEQ ID NO: 1) or its engineered variants, the target sequence is adjacent to the 5'-NGG-3' PAM motif, where N refers to any nucleotide. In some cases, the PAM motif is located 3' (downstream) of the target sequence.
[0217] For gene editing systems containing nuclease K (SEQ ID NO: 65) or an engineered variant thereof, the target sequence is adjacent to the 5'-NGG-3' PAM motif, where N refers to any nucleotide. In some cases, the PAM motif is located 3' (downstream) of the target sequence.
[0218] For gene editing systems containing nuclease M (SEQ ID NO: 80) or an engineered variant thereof, the target sequence is adjacent to a 5'-WTAAH-3' PAM motif, where W is A or T, and H is A, C, or T. In one instance, the PAM is 5'-TTAAA-3'. In some cases, the PAM motif is located 3' (downstream) of the target sequence.
[0219] In some embodiments, the target nucleic acid is a genomic site within the cell. In some cases, the target nucleic acid in which gene editing occurs may be located in a protein-coding region. Alternatively, the target nucleic acid may be located in a regulatory region, such as a promoter, enhancer, or 5' or 3' untranslated region. In other cases, the target nucleic acid may be in a non-coding gene, such as a transposon, miRNA, tRNA, ribosomal RNA, ribozyme, or lincRNA.
[0220] A. Gene editing
[0221] Any gene editing system disclosed herein can be used to edit a target gene of interest, such as a gene involved in a disease (e.g., a genetic disease). In some embodiments, the target gene may be a gene involved in the immune response of a subject. For example, the target gene may be an immune checkpoint gene or a member of the tumor necrosis factor receptor superfamily. Gene editing may occur in exons (e.g., in coding regions). Alternatively, gene editing may occur in introns or regulatory elements (e.g., promoters, enhancers, repressor elements, etc.). In some cases, gene editing may result in a reduction or elimination of the expression of the target gene. In other cases, gene editing may result in an enhancement of the expression of the target gene (e.g., disruption of repressor factors).
[0222] In some respects, this paper provides methods for introducing at least one edit into a target nucleic acid (e.g., a genomic site of interest, such as any target gene disclosed herein) using the gene editing system described herein.
[0223] As used herein, the term "editing" refers to one or more modifications introduced into a target nucleic acid, such as a nucleotide sequence at a genomic site of interest. Editing can occur within a target sequence as defined herein. Alternatively, editing can occur outside a target sequence (e.g., adjacent to a target sequence). Editing can be one or more substitutions, one or more insertions, one or more deletions, or a combination thereof.
[0224] The term "deletion" refers to the loss of one or more nucleotides in a nucleic acid sequence relative to a reference sequence. No specific process is implied in how the sequence containing the deletion is prepared. For example, the sequence containing the deletion can be synthesized directly from individual nucleotides. In other embodiments, the deletion is performed by providing and then modifying a reference sequence. The nucleic acid sequence can be located in the genome of an organism. The nucleic acid sequence can be located in a cell. The nucleic acid sequence can be a DNA sequence. Deletions can be frameshift mutations or non-frameshift mutations. The deletions described herein refer to insertions of up to several thousand bases.
[0225] "Insertion" refers to a gain of one or more nucleotides in a nucleic acid sequence relative to a reference sequence. No specific process is implied in how the sequence containing the insertion is prepared. For example, the sequence containing the insertion can be synthesized directly from individual nucleotides. In other embodiments, insertion is performed by providing and then modifying a reference sequence. The nucleic acid sequence can be located in the genome of an organism. The nucleic acid sequence can be located in a cell. The nucleic acid sequence can be a DNA sequence. The insertion is a frameshift mutation or a non-frameshift mutation. The insertion described herein refers to an insertion of up to several thousand bases.
[0226] In some embodiments, the gene editing methods disclosed herein can introduce editing (including substitution, insertion, deletion or a combination thereof) into target nucleic acids.
[0227] In some instances, an edit may include at least one substitution, at least one insertion, and / or at least one deletion. In some embodiments, an edit includes at least one substitution, insertion, or deletion. In some embodiments, a substitution, insertion, or deletion is at least 1-500 nucleotides (e.g., 1-10 nucleotides, 10-30 nucleotides, 30-50 nucleotides, 50-100 nucleotides, 100-200 nucleotides, 200-300 nucleotides, 300-400 nucleotides, or 400-500 nucleotides).
[0228] In some instances, editing can occur within about 500 nucleotides of any PAM sequence disclosed herein (e.g., a PAM sequence associated with nuclease A, nuclease K, or nuclease M as disclosed herein). In some embodiments, editing occurs in adjacent PAM sequences, for example, within about 1-500 nucleotides upstream or downstream of the PAM sequence. In some embodiments, editing can occur within about 1-10 nucleotides, 10-30 nucleotides, 30-50 nucleotides, 50-100 nucleotides, 100-200 nucleotides, 200-300 nucleotides, 300-400 nucleotides, or 400-500 nucleotides upstream of the PAM sequence. Alternatively or additionally, editing can occur within approximately 1–10 nucleotides, 10–30 nucleotides, 30–50 nucleotides, 50–100 nucleotides, 100–200 nucleotides, 200–300 nucleotides, 300–400 nucleotides, or 400–500 nucleotides downstream of the PAM sequence.
[0229] In some embodiments, editing begins at the PAM sequence. In some embodiments, editing may begin within approximately 1-30 nucleotides downstream of the PAM. Alternatively, editing may begin within approximately 1-30 nucleotides upstream of the PAM.
[0230] In some embodiments, editing may end within about 1-300 nucleotides upstream of the PAM sequence, for example, within about 1-10 nucleotides, about 10-30 nucleotides, about 30-50 nucleotides, about 50-100 nucleotides, about 100-200 nucleotides, or about 200-300 nucleotides upstream of the PAM sequence. Alternatively, editing may end within about 1-300 nucleotides downstream of the PAM sequence, for example, within about 1-10 nucleotides, about 10-30 nucleotides, about 30-50 nucleotides, about 50-100 nucleotides, about 100-200 nucleotides, or about 200-300 nucleotides downstream of the PAM sequence.
[0231] In some embodiments, the editing may end at the PAM sequence. In some embodiments, the editing may end within approximately 1-30 nucleotides downstream of the PAM. In other embodiments, the editing may end within approximately 1-30 nucleotides upstream of the PAM.
[0232] B. Gene editing in cells
[0233] In some respects, this document provides methods for editing genomic sites of interest (e.g., target genes as disclosed herein) in cells using any of the gene editing systems disclosed herein. To perform this method, the gene editing system may be delivered to or introduced into a cell population. In some cases, cells containing the desired gene edit may be collected and optionally cultured and expanded in vitro.
[0234] The cells described herein can be of various types. In some embodiments, the cells are isolated cells. In some embodiments, the cells are in cell cultures or co-cultures of two or more cell types. In some embodiments, the cells are ex vivo. In some embodiments, the cells are obtained from a living organism and maintained in a cell culture. In some embodiments, the cells are single-celled organisms.
[0235] In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a bacterial cell or a cell derived from bacteria. In some embodiments, the cell is an archaea cell or a cell derived from archaea.
[0236] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a plant cell or derived from plant cells. In some embodiments, the cell is a fungal cell or derived from fungal cells. In some embodiments, the cell is an animal cell or derived from animal cells. In some embodiments, the cell is an invertebrate cell or derived from invertebrate cells. In some embodiments, the cell is a vertebrate cell or derived from vertebrate cells. In some embodiments, the cell is a mammalian cell or derived from mammalian cells. In some embodiments, the cell is a human cell. In some embodiments, the cell is a zebrafish cell. In some embodiments, the cell is a primate cell. In some embodiments, the cell is a rodent cell. In some embodiments, the cell is synthetically prepared and is sometimes referred to as an artificial cell.
[0237] In some embodiments, the cells are derived from cell lines. A variety of cell lines are known in the art for tissue culture. Examples of cell lines include, but are not limited to, HEK293T, MF7, K562, HeLa, CHO, and their transgenic variants. Cell lines can be obtained from a variety of sources known to those skilled in the art (see, for example, the United States Type Culture Collection (ATCC) (Manassas, Virginia)). In some embodiments, the cells are immortalized cells or cells that proliferate indefinitely. In some embodiments, the cells are stem cells, such as totipotent stem cells (e.g., pluripotent), multipotent stem cells, oligopotent stem cells, or unipotent stem cells. In some embodiments, the cells are induced pluripotent stem cells (iPSCs) or iPSC-derived cells. In some embodiments, the cells are mesenchymal stem cells. In some embodiments, the cells are embryonic stem cells. In some embodiments, the cells are hematopoietic stem cells. In some embodiments, the cells are differentiated cells. For example, in some embodiments, the differentiated cells are muscle cells (e.g., hepatocytes), adipocytes (e.g., fat cells), bone cells (e.g., osteoblasts, osteocytes, osteoclasts), blood cells (e.g., monocytes, lymphocytes, neutrophils, eosinophils, basophils, macrophages, erythrocytes, or platelets), nerve cells (e.g., neurons), epithelial cells, immune cells (e.g., lymphocytes, neutrophils, monocytes, or macrophages), liver cells (e.g., hepatocytes), fibroblasts, or sex cells. In some embodiments, the cells are terminally differentiated cells. For example, in some embodiments, terminally differentiated cells are neuronal cells, adipocytes, cardiomyocytes, skeletal muscle cells, epidermal cells, or intestinal cells. In some embodiments, the cells are glial cells. In some embodiments, the cells are pancreatic islet cells, including α cells, β cells, δ cells, or enterochromaffin cells. In some embodiments, the cells are immune cells. In some embodiments, the immune cells are T cells. In some embodiments, the immune cells are B cells. In some embodiments, the immune cells are natural killer (NK) cells. In some embodiments, the immune cells are tumor-infiltrating lymphocytes (TILs). In some embodiments, the cells are mammalian cells, such as human cells, primate cells, or mouse cells. In some embodiments, the mouse cells are derived from wild-type mice, immunosuppressed mice, or disease-specific mouse models. In some embodiments, the cells are cells from living tissues, organs, or organisms.
[0238] In some embodiments, the cells are primary cells. For example, cultures of primary cells may be passaged 0, 1, 2, 4, 5, 10, 15, or more times. In some embodiments, primary cells are collected from an individual using any known method. For example, leukocytes may be collected using separation techniques, leukocyte isolation techniques, density gradient separation, etc. Cells from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, and stomach may be collected via biopsy. Appropriate solutions may be used to disperse or suspend the collected cells. Such solutions are typically balanced salt solutions (e.g., physiological saline, phosphate-buffered saline (PBS), Hank's balanced salt solution, etc.), conveniently supplemented with fetal bovine serum or other naturally occurring factors, bound to a low concentration of an acceptable buffer. Buffer solutions may include HEPES, phosphate buffer, lactate buffer, etc. Cells can be used immediately or stored (e.g., by freezing). Frozen cells can be thawed and reused. Cells can be frozen in DMSO, serum, culture medium buffers (e.g., 10% DMSO, 50% serum, 40% buffered medium) and / or some other common solutions used for preserving cells at freezing temperatures.
[0239] In embodiments in which the gene editing system disclosed herein is introduced into multiple cells, at least about 0.5% of the cells contain the desired edit. In some embodiments, at least about 1% of the cells contain the desired edit. In some embodiments, at least about 2% of the cells contain the desired edit. In some embodiments, at least about 3% of the cells contain the desired edit. In some embodiments, at least about 4% of the cells contain the desired edit. In some embodiments, at least about 5% of the cells contain the desired edit. In some embodiments, at least about 10% of the cells contain the desired edit. In some embodiments, at least about 20% of the cells contain the desired edit. In some embodiments, at least about 30% of the cells contain the desired edit. In some embodiments, at least about 40% of the cells contain the desired edit. In some embodiments, at least about 50% of the cells contain the desired edit.
[0240] Cells carrying the desired gene edits, such as cells produced using any gene-editing system also disclosed herein through the methods disclosed herein, are also within the scope of this disclosure. In some cases, cells modified with the CRISPR nuclease peptides disclosed herein can be used as expression systems for preparing biomolecules. For example, modified cells can be used to produce biomolecules such as proteins (e.g., cytokines, antibodies, antibody-based molecules), peptides, lipids, carbohydrates, nucleic acids, amino acids, and vitamins. In other embodiments, modified cells can be used to produce viral vectors such as lentiviruses, adenoviruses, adeno-associated viruses, and oncolytic virus vectors. In some embodiments, modified cells can be used in cytotoxicity studies. In some embodiments, modified cells can be used as disease models. In some embodiments, modified cells can be used for vaccine production. In some embodiments, modified cells can be used for therapy. For example, in some embodiments, modified cells can be used in cell therapies such as blood transfusions and transplants.
[0241] In some embodiments, cells modified with CRISPR nuclease peptides as disclosed herein can be used to establish new cell lines comprising modified genomic sequences. In some embodiments, the modified cells of this disclosure are modified stem cells (e.g., modified totipotent / pluripotent stem cells, modified multipotent stem cells, modified unipotent stem cells, modified oligopotent stem cells, or modified unipotent stem cells) that differentiate into one or more cell lineages lacking the modified stem cells. This disclosure further provides organisms (e.g., animals, plants, or fungi) comprising the modified cells of this disclosure or derived from the modified cells of this disclosure.
[0242] C. Delivering gene editing systems to cells
[0243] In some embodiments, any gene editing system or its components as disclosed herein may be formulated, including, for example, a carrier, such as a carrier and / or a polymer carrier, such as liposomes or lipid nanoparticles, and delivered to cells (e.g., prokaryotes, eukaryotes, plants, mammals, etc.) by known methods. Such methods include, but are not limited to, transfection (e.g., lipid-mediated cationic polymers, calcium phosphate, dendritic macromolecules); electroporation or other membrane rupture methods (e.g., nuclear transfection); viral delivery (e.g., lentiviruses, retroviruses, adenoviruses, AAVs); microinjection; microparticle bombardment (“gene gun”); fugene; direct acoustic loading; cell squeezing; optical transfection; protoplast fusion; impale infection; magnetic transfection; exogenous body-mediated transfer; lipid nanoparticle-mediated transfer; and any combination thereof.
[0244] In some embodiments, the method comprises delivering one or more nucleic acids (e.g., nucleic acids encoding CRISPR nuclease peptides and / or gRNA, and / or pre-formed ribonucleoproteins) into cells. Exemplary intracellular delivery methods include, but are not limited to, viruses or virus-like agents; chemical-based transfection methods, such as those using calcium phosphate, dendritic macromolecules, liposomes, or cationic polymers (e.g., DEAE-glucan or polyethyleneimine); non-chemical methods, such as microinjection, electroporation, cell extrusion, acoustic perforation, optical transfection, puncture transfection, protoplast fusion, bacterial conjugation, plasmid or transposon delivery; particle-based methods, such as those using gene guns, magnetic transfection or magnetically assisted transfection, particle bombardment; and hybridization methods, such as nuclear transfection. In some embodiments, this application further provides cells generated by such methods, and organisms (e.g., animals, plants, or fungi) containing or generated from such cells. In some embodiments, the compositions of the present invention are further delivered together with agents (e.g., compounds, molecules, or biomolecules) that affect DNA repair or DNA repair mechanisms. In some embodiments, the compositions of the present invention are further delivered together with agents that affect the cell cycle (e.g., compounds, molecules, or biomolecules).
[0245] In some embodiments, a first composition comprising a CRISPR nuclease peptide is delivered to a cell. In some embodiments, a second composition comprising gRNA is delivered to a cell. In some embodiments, the first composition is contacted with the cell before the second composition is contacted. In some embodiments, the first composition is contacted with the cell simultaneously with the second composition. In some embodiments, the first composition is contacted with the cell after the second composition is contacted. In some embodiments, the first composition is delivered by a first delivery method, and the second composition is delivered by a second delivery method. In some embodiments, the first delivery method and the second delivery method are the same. For example, in some embodiments, the first composition and the second composition are delivered by viral delivery. In some embodiments, the first delivery method and the second delivery method are different. For example, in some embodiments, the first composition is delivered by viral delivery, and the second composition is delivered by lipid nanoparticle-mediated transfer, and the second composition is delivered by viral delivery, or the first composition is delivered by lipid nanoparticle-mediated transfer, and the second composition is delivered by viral delivery.
[0246] In some instances, any gene editing system provided herein comprises one or more lipid excipients that associate with components of the gene editing system to facilitate delivery of the gene editing system to host cells. In some cases, one or more lipid excipients may form lipid nanoparticles (LNPs) that associate with or encapsulate components of the gene editing system, such as mRNA molecules encoding CRISPR nuclease peptides and gRNA. Such LNP-containing gene editing systems may come into contact with host cells (e.g., administered to a subject requiring gene editing), and the LNPs may facilitate delivery of the gene editing system to host cells.
[0247] In other instances, the gene-editing systems described herein can be delivered to host cells via viral vector-mediated methods. Such gene-editing systems may comprise a viral vector carrying a transgene encoding a CRISPR nuclease polypeptide and optionally a transgene encoding gRNA. In some cases, the viral vector may be an AAV vector. The viral vector can facilitate the delivery of the transgene to host cells, wherein the transgene produces a CRISPR nuclease polypeptide and optionally gRNA to achieve gene editing of a target gene.
[0248] IV. Therapeutic Applications
[0249] Any gene-editing system disclosed herein, or modified cells produced using such a system, can be used to treat diseases that may benefit from gene editing introduced by the system or carried by the modified cells. For example, the disease may be a genetic condition, and the gene editing repairs a gene mutation associated with the genetic condition. Alternatively, the disease may be related to abnormal gene expression, and the gene editing rescues such abnormal expression.
[0250] In some embodiments, this document provides a method for treating a disease, the method comprising administering any gene-editing system disclosed herein to a subject requiring treatment (e.g., a human patient). The gene-editing system can be delivered to a specific tissue or a specific type of cell requiring gene editing. The gene-editing system may comprise an LNP covering one or more of the components, one or more vectors (e.g., a viral vector) encoding one or more of the components, or a combination thereof. The components of the gene-editing system may be formulated to form a pharmaceutical composition, the pharmaceutical composition may further comprise one or more pharmaceutically acceptable carriers.
[0251] In some embodiments, modified cells produced using any of the gene-editing systems disclosed herein may be administered to a subject requiring treatment (e.g., a human patient). The modified cells may contain substitutions, insertions, and / or deletions as described herein. In some instances, the modified cells may comprise cell lines modified with CRISPR nuclease peptides and gRNAs as disclosed herein. In some cases, the modified cells may be a heterologous population comprising cells with different types of gene edits. Alternatively, the modified cells may comprise a substantially homologous population of cells comprising a specific gene edit (e.g., at least 80% of the cells in the entire population). In some instances, the cells may be suspended in a suitable culture medium.
[0252] In some embodiments, this document provides compositions comprising a gene-editing system or components thereof, or modified cells. Such compositions may be pharmaceutical compositions. Useful pharmaceutical compositions may be prepared, packaged, or marketed in formulations suitable for oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, intralesional, buccal, ocular, intravenous, intra-organ, or other routes of administration. The pharmaceutical compositions disclosed herein may be prepared, packaged, or marketed in bulk in the form of a single unit dose or multiple single unit doses. As used herein, a “unit dose” is a discrete amount of a pharmaceutical composition containing a predetermined number of cells. The number of cells is generally equal to the dose of cells to be administered to a subject or a convenient fraction of such a dose, such as half or one-third of such a dose.
[0253] Formulas of pharmaceutical compositions suitable for parenteral administration may comprise an active agent (e.g., a gene-editing system or a component thereof, or modified cells) in combination with a pharmaceutically acceptable carrier (such as sterile water or sterile isotonic saline). Such formulas may be prepared, packaged, or marketed in a form suitable for bolus or continuous administration. Some injectable formulas may be prepared, packaged, or marketed in unit dosage forms, such as ampoules or multi-dose containers containing preservatives. Some formulas for parenteral administration include, but are not limited to, suspensions, solutions, emulsions in oily or aqueous media, pastes, and implantable, sustained-release, or biodegradable formulas. Some formulas may further comprise one or more additional ingredients, including but not limited to suspending agents, stabilizers, or dispersants.
[0254] Pharmaceutical compositions may be in the form of sterile, injectable aqueous or oily suspensions or solutions. Such suspensions or solutions may be formulated according to known techniques and may contain additional components besides cells, such as dispersants, wetting agents, or suspending agents described herein. Such sterile injectable formulations may be prepared using non-toxic, parenterally acceptable diluents or solvents, such as water or saline. Other acceptable diluents and solvents include, but are not limited to, Ringer's solution, isotonic sodium chloride solution, and fixed oils, such as synthetic monoglycerides or diglycerides. Other available parenterally applicable formulations include those that may contain cells in packaged form, in liposome formulations, or as components of a biodegradable polymer system. Some compositions for sustained release or implantation may contain pharmaceutically acceptable polymers or hydrophobic materials, such as emulsions, ion exchange resins, microsoluble polymers, or microsoluble salts.
[0255] V. Kit and its uses
[0256] This disclosure also provides kits or systems that can be used, for example, to implement the methods described herein. In some embodiments, the kit or system includes a CRISPR nuclease peptide and optionally gRNA. In some embodiments, the kit or system includes a polynucleotide encoding the CRISPR nuclease peptide and optionally gRNA. The gRNA in the kit can be programmed to target a sequence of interest. The CRISPR nuclease peptide and gRNA can be packaged in the same vial or other vessel within the kit or system, or they can be packaged in separate vials or other vessels, and their contents can be mixed prior to use. The kit or system may additionally include, optionally, a buffer and / or instructions for use of the CRISPR nuclease peptide and gRNA.
[0257] In some embodiments, the kit comprises a first composition comprising a CRISPR nuclease peptide as disclosed herein. In some embodiments, the kit comprises a second composition comprising gRNA also disclosed herein. In some embodiments, the first and second compositions are packaged in the same vial. In some embodiments, the first and second compositions are packaged in different vials.
[0258] In some embodiments, the kit can be used for research purposes. For example, in some embodiments, the kit can be used to study gene function.
[0259] General technology
[0260] Unless otherwise indicated, the practice of this disclosure will employ conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, which are within the scope of the art. These techniques are well explained in the following literature: *Molecular Cloning: A Laboratory Manual*, 2nd edition (Sambrook et al., 1989), Cold Spring Harbor Press; *Oligonucleotide Synthesis* (edited by MJ Gait, 1984); *Methods in Molecular Biology*, Humana Press; *Cell Biology: A Laboratory Notebook* (edited by JE, 1989), Academic Press; *Animal Cell Culture* (edited by RI Freshney, 1987); *Introduction to Cell and Tissue Culture* (JP Mather and PE Roberts, 1998), Plenum Press; *Cell and Tissue Culture: Laboratory Procedures* (A. Doyle, JB...). Edited by Griffiths and DG Newell, 1993-8; J. Wiley and Sons Publishing; *Methods in Enzymology* (Academic Press, Inc.); *Handbook of Experimental Immunology* (edited by DM Weir and CC Blackwell); *Gene Transfer Vectors for Mammalian Cells* (edited by JMMiller and MP Calos, 1987); *Experimental Guide to Contemporary Molecular Biology* (FM...Ausubel et al. (eds., 1987); *PCR: The Polymerase Chain Reaction* (Mullis et al. (eds., 1994); *Current Protocols in Immunology* (JEColigan et al. (eds., 1991); *Short Protocols in Molecular Biology* (John Willie & Son Publishing, 1999); *Immunobiology* (CA Janeway and P. Travers, 1997); *Antibodies* (P. Finch, 1997); *Antibodies: a practice approach* (D. Catty. (ed., IRL Press, 1988-1989); *Monoclonal antibodies: a practical approach* (P. Shepherd and C. Dean (eds., 1989)). Oxford University Press, 2000; *Using antibodies: a laboratory manual* (E. Harlow and D. Lane, Cold Spring Harbor Laboratory Press, 1999); *The Antibodies* (edited by M. Zanetti and JD Capra, Harwood Academic Publishers, 1995); *DNA Cloning: A practical Approach*, Volumes I and II (edited by DN Glover, 1985); *Nucleic Acid Hybridization* (edited by BD Hames and SJ Higgins, 1985); *Transcription and Translation* (edited by BD Hames and SJ Higgins, 1984); *Animal Cell Culture* (RIFreshney, ed., (1986); *Immobilized Cells and Enzymes* (IRL Press, 1986); and B. Perbal, *A Practical Guide to Molecular Cloning* (1984); FM Ausubel et al., (eds.).
[0261] Without further elaboration, it is believed that those skilled in the art can utilize the invention to the fullest extent based on the above description. Therefore, the following specific embodiments should be interpreted as illustrative only and do not limit the remainder of this disclosure in any way. All disclosures referenced herein for purposes or subject matter are incorporated herein by reference.
[0262] Example 1: Editing human target genes in HEK293T cells using nuclease A-derived CRISPR nuclease peptides
[0263] This example describes genome editing of the AAVS1, EMX1 and VEGFA genes by nuclease A (SEQ ID NO: 1) or a variant thereof, which was introduced into the HEK293T cell line via lipid-based transient transfection.
[0264] Nuclease A was labeled with the N-terminal SV40 nuclear localization signal (NLS) and the C-terminal nucleoplasmic protein nuclear localization signal (NLS), and its coding sequence was converted into a human codon-optimized DNA sequence. The resulting DNA was synthesized and cloned into a pcDNA3.1 vector (Invitrogen) containing a CMV promoter for expression. The reference sequences and NLS-labeled sequences used are shown in Table 1. The plasmid was purified using the midiprep kit.
[0265] RNA guides were designed and cloned into the pUC19 plasmid following the U6 PolIII promoter and terminated with a 6xpolyT sequence. The RNA guides were designed to be specific for target sequences within the coding exons of AAVS1, EMX1, and VEGFA with a 5'-NGG-3' PAM sequence (the PAM sequence located at the 3' end of the target sequence). See Table 2 for all RNA guide sequences. The U6 PolIII promoter uses +1 G at the start of the transcript (i.e., the 5' end of the RNA) for more efficient transcription, and these sequences were excluded from those described in Table 2. The plasmids were purified using the midiprep kit.
[0266] Table 1. Amino acid sequence of nuclease A and its exemplary variants
[0267]
[0268] Table 2. Target sequences and RNA guide sequences of CRISPR nuclease peptides derived from nuclease A
[0269]
[0270] *The spacers are in uppercase letters, and the support (SEQ ID NO: 2) is in lowercase letters.
[0271] Approximately 16 hours prior to transfection, 25,000 HEK293T cells in DMEM / 10% FBS+Pen / Strep (D10 medium) were seeded into each well of a 96-well plate. On the day of transfection, cells were 50-90% confluent. For each well to be transfected, a mixture of Lipofectamine 2000™ (Thermo Fisher Scientific) and Opti-MEM™ (Thermo Fisher Scientific) was prepared and incubated at room temperature for 5 minutes (Solution 1). After incubation, the Lipofectamine 2000™:Opti-MEM™ mixture was added to a separate mixture containing the CRISPR nuclease plasmid (NLS-tagged), the RNA guide plasmid, and Opti-MEM™ (Solution 2). In the negative control case, the CRISPR nuclease plasmid was excluded. Solutions 1 and 2 were mixed by pipetting and then incubated at room temperature for 25 minutes. After incubation, the mixture of Solutions 1 and 2 was added dropwise to each well of a 96-well plate containing cells. Approximately 72 hours post-transfection, the cells were treated with trypsin by adding TrypLE™ (Thermo Fisher Scientific) to the center of each well and incubated at 37°C for approximately 5 minutes. D10 medium was then added to each well and mixed to resuspend the cells. The resuspended cells were centrifuged for 10 minutes to obtain a pellet, and the supernatant was discarded. The cell pellet was then resuspended in QuickExtract™ buffer (Lucigen®) and incubated at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes.
[0272] Next-generation sequencing (NGS) samples were prepared using two rounds of PCR. Three technical replicas of the reference and each target for each variant were analyzed. The first round (PCR1) was used to amplify target-specific genomic regions. A second round of PCR (PCR2) was performed to add Illumina adapters and indexes. The reactions were then pooled and purified by column purification. Sequencing was performed using a kit such as the 150-cycle NextSeq 500 / 550 medium or high output v2.5 kit.
[0273] For NGS analysis, the indel mapping function uses the sample's FASTQ file, amplicon reference sequence, and forward primer sequence. For each read, the KMER scanning algorithm is used to calculate the edit operations (matches, mismatches, insertions, deletions) between the read and the reference sequence. To remove a small number of primer dimers present in some samples, the first 30 nt of each read is required to match the reference, and reads with more than half of the mapped nucleotide mismatches are also filtered out. Up to 50,000 reads that pass through those filters are used for analysis, and if said reads contain insertions or deletions, they are counted as indel reads. The indel% is calculated by dividing the number of reads containing indels by the number of reads analyzed (up to 50,000 reads that pass through the filters). The minimum QC criterion for the number of reads that pass through the filters is 10,000.
[0274] For each target, the indel ratio was calculated for each sample and its homologous protein-free control; the indel ratio refers to the percentage of NGS reads containing indels. When a CRISPR nuclease is included in the transfection, targets containing a higher percentage of indels indicate the outcome of DNA editing in the cells.
[0275] like Figure 1 As shown, when the nuclease A (SEQ ID NO: 1) plasmid was present, four of the six targets tested exhibited high levels of indel.
[0276] Example 2: The effectiveness of CRISPR nuclease A variants in targeting mammalian genes
[0277] This example describes indel evaluation of mammalian targets using a variant of CRISPR nuclease A (SEQ ID NO: 1) transfected into HEK293T cells.
[0278] Arginine scanning mutagenesis was performed to individually replace each non-arginine residue of nuclease A (SEQ ID NO: 1) with arginine. SEQ ID NO: 1 is referred to herein as the nuclease A reference sequence. This yielded 710 single-arginine substitution variants. The nucleic acids encoding nuclease A and each CRISPR nuclease variant were then individually cloned into the pcda3.1 backbone (Ingenie), and plasmids were maximized for preparation and dilution. The plasmids contained a CMV promoter, a first NLS (KRTADGSEFESPKKKRKV; SEQ ID NO: 3) upstream of the coding sequence, an XTEN linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS; SEQ ID NO: 95), followed by a second NLS (KRPAATKKAGQAKKKK; SEQ ID NO: 4) downstream of the coding sequence. See also Example 1 above.
[0279] The RNA guide sequence and target sequence are shown in Table 3. The RNA guide was cloned into the pUC19 backbone following the U6 PolIII promoter (New England Biolabs). ® The U6 PolIII promoter is terminated with a 6x polyT sequence. For more efficient transcription, the U6 PolIII promoter uses +1 G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences described in Table 3. The plasmid was then prepared and diluted to the maximum extent possible.
[0280] Table 3. Mammalian targets and corresponding crRNAs*
[0281]
[0282] *The corresponding sequences are shown in Table 2 above.
[0283] Following the method described in Example 1 above, HEK293T cells were transfected with a plasmid expressing a CRISPR nuclease peptide and the gRNA provided in this article. Next-generation sequencing was then used to determine the gene editing efficiency, also following the method described in Example 1.
[0284] For each target, the indel ratio for the reference and each variant was calculated, where the indel ratio refers to the percentage of NGS reads containing indels. The indel ratio used for fold change calculation was the average of the two technical replicas. Then, to calculate the fold change in the indel ratio at a specific target, the indel ratio of each variant was divided by the reference indel ratio. Table 4 shows the fold change in the indel ratio for each target and the average for the two targets. The number is relative to the reference nuclease (i.e., without NLS) of SEQ ID NO: 1. For example, at the AAVS1 target, the indel ratio of the I67R variant was 7.07 times that of the reference indel ratio, and at the VEGFA target, the indel ratio of the I67R variant was 4.43 times that of the reference indel ratio. At both targets, the average fold change in the indel ratio of the I67R variant was 5.75 times that of the average fold change in the reference indel ratio.
[0285] As shown in Table 4, the 21 variants with monoarginine substitution (left column) are characterized by an indel ratio that increases by at least 1.5 times relative to the reference indel ratio when the average is taken across two targets (right column).
[0286] Table 4. Changes in Indel Ratio*
[0287]
[0288] *Variant indel ratio / Reference indel ratio
[0289] The following 155 variants were analyzed, and when averaged across two targets, their indel ratios were 1 to 1.5 times that of the reference indel ratio: E206R, K74R, S280R, A319R, I96R, A354R, D323R, F265R, E120R, H442R, V284R, E364R, C653R, A365R, N586R, E294R, E490R, N459R, E460R, E144R, S95R, N571R, S587R, P708R, K163R, G485R, V387R, I711R, W148R, K100R, G7 22R, A132R, Q328R, T201R, K322R, Q277R, N87R, K378R, N758R, I563R, G3 30R, D207R, K86R, G359R, K360R, K49R, T741R, I519R, E248R, E299R, K333 R, Q41R, V12R, E152R, V210R, D53R, Q380R, M105R, K145R, S548R, S761R, K28R, V390R, K17R, D537R, K585R, Q772R, K351R, T255R, K287R, K146R, G6 6R, A441R, E491R, K302R, N340R, T98R, T83R, S752R, E149R, A55R, S650R , K355R, N489R, N457R, I753R, E94R, V774R, K572R, E371R, K566R, N462R, D561R, K366R, L369R, K546R, Y693R, S368R, K773R, Q16R, P771R, G385R, S180R, E11R, K614R, V10R, K422R, E129R, H243R, G209R, Q436R, K770R, K7 37R, K767R, K579R, K764R, G97R, K224R, P709R, L8R, K298R, K750R, Q177 R, D533R, D356R, S170R, T395R, K449R, K252R, Y271R, E251R, W141R, K714 R, T726R, K167R, I370R, K184R, K91R, E418R, D173R, Q295R, E131R, K350 R, N332R, Y9R, A242R, K288R, E142R, N154R, F775R, H600R, L266R, E438R,V749R and S134R. The remaining variants (534 variants) reduce the indel ratio relative to the reference indel ratio (the change in indel ratio is less than 1.0): G488R, E300R, K348R, K424R, P437R, K557R, S729R, K702R, G713R, E128R, V654R, S547R, T233R, A135R, T542R, T766R, K386R, T13R, Q64R, E453R, H293R, L543R, D536R, K75R, K363R, P560R, Q140R, K143R, K467R, P192R, M15R, K123R, L245 R, K5R, T297R, K278R, K76R, S235R, H621R, E367R, K689R, D246R, Y534R, I7 56R, V768R, K268R, L434R, T391R, K139R, S465R, E236R, K6R, K82R, K665R, D 7R, K138R, Y29R, V119R, E137R, N39R, E556R, N379R, C4R, K247R, T151R, A23 8R, S742R, T126R, D446R, V372R, Y471R, E357R, N549R, T392R, Q234R, I130R , L701R, P698R, G538R, D272R, I678R, E710R, N517R, M133R, K613R, D159R, E292R, K626R, E552R, E430R, T562R, G516R, K311R, M594R, D164R, A274R, T5 35R, P122R, T70R, F275R, N677R, K687R, N263R, A747R, L559R, S570R, S155R , G550R, K237R, D52R, K620R, G751R, K254R, T763R, T440R, P575R, N196R, K1 09R, E495R, K188R, K216R, F304R, S769R, V205R, Q269R, L291R, G249R, E633 R, K628R, E239R, K691R, K644R, K712R, K715R, A760R, N667R, P493R, I226R, E176R, G754R, L684R, H260R, D554R, H727R, N276R, S625R, Q637R, N315R, E1 60R, A403R, N3R, V433R, K696R, K202R, K402R, D609R, L200R, F229R, V744R,L429R、E18R、T718R、I607R、T73R、H296R、L374R、L78R、D481R、V748R、Q189R、E147R、S583R、H181R、E728R、E150R、E171R、A174R、D220R、K690R、K532R、Y724R、E124R、S759R、Y717R、H703R、P14R、K707R、A681R、L43R、L716R、C629R、L19R、K204R、G740R、T106R、K125R、T175R、E421R、M121R、K113R、K602R、D136R、G361R、I487R、K193R、E617R、G349R、L326R、G623R、K616R、E610R、I695R、E686R、Q497R、E746R、E597R、H558R、N445R、I168R、K313R、E241R、G688R、Q642R、D518R、V733R、K544R、D93R、H362R、L279R、N187R、S732R、K569R、L214R、D225R、V244R、P308R、N464R、G646R、P231R、L412R、I57R、F565R、W508R、Y504R、T162R、Y190R、E730R、D731R、G573R、H478R、G404R、N668R、Q178R、W622R、D540R、S431R、L221R、I199R、T590R、L501R、M541R、G627R、D685R、G601R、I439R、H112R、N407R、V705R、D454R、A344R、N647R、W476R、T631R、N217R、D632R、E80R、N101R、V539R、L486R、E324R、I335R、N211R、N645R、L110R、L492R、V253R、S498R、F409R、G639R、L447R、F195R、V20R、P555R、K738R、V161R、I223R、T638R、T283R、S608R、Y92R、E514R、P428R、T624R、L58R、V218R、K448R、I179R、M458R、I88R、I183R、A692R、Y577R、I325R、L222R、A45R、E102R、A670R、G89R、F329R、T427R、T574R、E634R、D117R、G432R、L232R、V303R、C381R、W660R、D36R, G281R, I584R, A307R, D604R, G679R, G510R, P327R, N316R, L318R, I262R, V227R, F256R, Y84R, N580R, L443R, V320R, F257R, D723R, V54R, F553R, D 342R、L353R、T212R、M649R、Y116R、G425R、F104R、H157R、Y25R、A503R、I666 R、I314R、I589R、G405R、L230R、T312R、I500R、I290R、N40R、Y44R、P435R、P10 3R, L582R, Y468R, V165R, P213R, C419R, G90R, F99R, F704R, F505R, F648R, L306R, A618R, L463R, I388R, N197R, A522R, P545R, H676R, F285R, Y118R, N669 R, Y734R, N697R, V680R, H520R, L373R, I273R, K513R, L240R, F671R, C203R, E376R, L461R, N394R, L719R, I185R, M479R, V682R, A345R, E603R, T47R, F619 R、A762R、L576R、D578R、A410R、I451R、I494R、L81R、I331R、V408R、C208R、S 530R、Y383R、K423R、M469R、F511R、Y108R、V321R、I270R、V382R、S309R、L651 R、T377R、N674R、F640R、L658R、G31R、V107R、A455R、K672R、W509R、I23R、P7 43R、V42R、T596R、L452R、M169R、H289R、M336R、L630R、T30R、V34R、V35R、G22 R、G659R、L85R、C111R、G115R、V317R、V725R、Y683R、V595R、M50R、L615R、I7 36R、V526R、A529R、I606R、Q662R、A414R、C384R、I406R、V48R、L700R、Y611R、 I466R, H397R, A635R, L523R, T347R, F399R, I527R, G657R, C415R, A525R, V664R, L310R, V182R, D524R, F186R, E337R, V675R, L765R, L483R, L32R, K450RS473R, G499R, P581R, L21R, H521R, D396R, C416R, I398R, Y641R, T673R, L739R, G27R, V413R, M663R, L745R, P400R, C286R, T389R, N339R, A393R, I474R, T502R, N411R, A33R, F341R, S338R, G26R, L528R, G475R, N420R, D24R, Y358R, V334R, A472R, I343R, and S470R.
[0290] Based on this experiment, the following substitutions were selected for further engineering of nuclease A relative to SEQ ID NO: 1: I67R, E63R, D56R, T605R, D568R, E694R, G71R, E706R, and E655R.
[0291] Example 3: The effectiveness of CRISPR nuclease A combinatorial variants in targeting mammalian genes
[0292] This example describes the indel evaluation of mammalian targets using nuclease A variants containing two or more substitutions identified in Example 2 as increasing indel activity. 106 combined nuclease A variants were tested.
[0293] As described in Example 2, each nuclease A variant and RNA guide were cloned (see Table 5). HEK293T cells were further transfected as described in Example 2, followed by NGS analysis. For each target, the indel ratio of nuclease A (SEQ ID NO: 1) and each variant of nuclease A was calculated, where the indel ratio refers to the percentage of NGS reads containing the indel. The indel ratios shown in Table 6 were calculated as the average of two biological replicas, each containing two technical replicas.
[0294] Table 5. Mammalian targets and corresponding crRNAs*
[0295]
[0296] *See Table 2 for the corresponding sequences.
[0297] Table 6. Indel ratio of mammalian targets
[0298]
[0299] As shown in Table 6, all variants except one of the nuclease A variants with amino acid substitution combinations exhibited higher indel activity than nuclease A (SEQ ID NO: 1). When averaged across the three targets, the eight CRISPR nuclease variants resulted in an indel ratio exceeding 0.3, indicating that over 30% of the NGS reads contained indels. The eight nuclease variants of nuclease A contain the following substitution combinations: a) I67R, D568R, and E706R; b) I67R, D56R, and D568R; c) I67R, T605R, and D568R; d) I67R, D56R, and E706R; e) I67R, T605R, and E706R; f) I67R, D568R, and E694R; g) I67R, D56R, E694R, and E706R; and h) I67R, D56R, T605R, and E706R. When averaged across the three targets, the 64 variants of nuclease A resulted in an indel ratio exceeding 0.2, indicating that over 20% of NGS reads contained indels. When the average was taken across three targets, 29 variants of nuclease A resulted in an indel ratio exceeding 0.1, indicating that more than 10% of NGS reads contained indels.
[0300] Based on this experiment, nuclease A variants containing the following substitutions were selected for further testing: I67R, D568R, and E706R. Compared to the reference nuclease A (SEQ ID NO: 1), these nuclease A variants exhibited more than an 8-fold increase in indel activity.
[0301] Example 4 - Editing of RNA-guided nuclease A variants and additional RNA scaffold sequences in HEK293T cells Human target genes
[0302] This example describes genome editing of an exemplary target gene via nuclease A of SEQ ID NO: 1 and nuclease A variants of I67R, D568R, and E706R of SEQ ID NO: 6 when combined with an RNA guide containing an alternative scaffold sequence.
[0303] Unless otherwise specified, RNA guides were designed using the RNA scaffold sequences in Table 7 and cloned into the pUC19 plasmid following the U6 PolIII promoter, terminated with a 6x polyT sequence. The RNA guides were designed to be specific to the AAVS1-T3, EMX1-T7, and / or AAVS1-T3b target sequences (see Example 1 and Table 7). See all RNA guide sequences in Table 8. U6PolIII uses +1 G at the start of the transcript (i.e., the 5' end of the RNA) for more efficient transcription, which is excluded from the sequences described in Table 8. Industrial-grade plasmids were obtained from GenScript.
[0304] Table 7. RNA guide component sequences
[0305]
[0306] Table 8. RNA Guide Sequences
[0307]
[0308] *The spacer subsequences are in uppercase letters, and the scaffold sequence is in lowercase letters.
[0309] The RNA guides in Table 8 were combined with nuclease A or its I67R / D568R / E706R variants and introduced into HEK293T cells via lipid-based transient transfection as described in Example 1. The AAVS1-T3 and EMX1-T7 RNA guides from Example 1 were further used as controls. Genomic DNA was recovered approximately 72 hours post-transfection, and samples were prepared for NGS and analysis as described in Example 1. The percentage of indel-containing NGS reads shown in Table 9 was calculated as an average of two biological replicas, each containing two technical replicas. Unless otherwise stated, the percentage of indel-containing NGS reads shown in Table 10 was calculated as an average of three technical replicas.
[0310] Table 9. Indel activity of nuclease A
[0311]
[0312] Table 10. Indel activities of CRISPR nucleases from I67R / D568R / E706R variants
[0313]
[0314] As shown in Table 9, when paired with a reference nuclease A, the RNA guide containing a truncated reference scaffold exhibited higher indel activity than the RNA guide containing a reference scaffold. As shown in Table 10, the RNA guides containing either the reference scaffold or scaffold 2 each exhibited high indel activity with I67R, D568R, or E706R nuclease A variants. In all cases, in the absence of the RNA guide, a significantly higher average indel percentage was observed compared to the background control containing either reference nuclease A or I67R, D568R, or E706R nuclease A variants.
[0315] Therefore, this example demonstrates that the truncated reference scaffold and scaffold 2 can be recognized by nuclease A (SEQ ID NO: 1) and nuclease A variant (SEQ ID NO: 6).
[0316] Example 5: Engineering and effectiveness of CRISPR nuclease A nickase variants targeting mammalian genes
[0317] This example describes the introduction of a mutation into nuclease A of SEQ ID NO: 1, which disrupts either the HNH or RuvC domain to produce a functional nickase. D396, H397, and N420 were identified as putative catalytic residues of the HNH domain. Positions D24, E337, and D524 were identified as putative catalytic residues of the RuvC domain. These positions were identified by analyzing models of structural regions similar to known HNH and RuvC active sites generated using AlphaFold2 (Jumper et al., Nature 596: 583-9 (2021)) and / or by sequence alignment with other nucleases at previously identified candidate positions. Examples of reference structures used to identify HNH and RuvC active sites are represented by the following PDB IDs: 5h0m, 7eu9, 6ltu, 7odf, 7lys, 8dc2, 4cmp, 4oo8, 7z4j, 5axw, 5b2o, 6kc8, 7utn, 8csz, 8ctl, 8dmb.
[0318] The coding sequence for nuclease A was converted into a codon-optimized DNA sequence for *E. coli*, synthesized, and cloned into a pET-28a(+) vector (Novagen), containing lac and T7 RNA polymerase promoters for gene expression. To test nickase activity, individual alanine mutants were cloned for each of the positions of the putative active sites identified as HNH and RuvC domains. A leucine mutant at position H397 was also cloned. Research-grade plasmids were obtained from GenScript. The engineered nickase sequences are shown in Table 11. In the nucleotide sequences, codons encoding substituted residues are uppercase, bold, and underlined, and in the amino acid sequences, substituted residues are shown in bold and underlined. The putative HNH knockout nickase is expected to cleave the non-target strand but not the target strand. The putative RuvC knockout nickase is expected to cleave the target strand but not the non-target strand.
[0319] Table 11. CRISPR nuclease and nickase sequences
[0320]
[0321] A linear DNA template sequence (IDT) encoding an RNA guide was designed and sequenced, with the T7 promoter upstream of the guide and the T7Te terminator sequence downstream. The RNA guide was designed to be specific to previously tested target sequences within the coding exon of AAVS1 having a 5'-NGG-3' PAM sequence (PAM being the 3' of the target sequence), as described in Example 1. The sequence of the encoded RNA guide and its individual components are shown in Table 12. The T7 promoter uses a +1 G at the start of the transcript (i.e., the 5' end of the RNA) for more efficient transcription. This +1 G is shown in SEQ ID NO: 49.
[0322] Table 12. RNA Sequences
[0323]
[0324] DNA targets were designed and sequenced as synthetic linear DNA fragments. Target sequences, consisting of 10 bases upstream and downstream of AAVS1 and exons, were flanked by unrelated sequences of 100 bases upstream and 200 bases downstream. Additional sequences were added to ensure good separation of cleaved and uncleaved products on a gel. Target and non-target strands were labeled with 5' IR800 and 5' IR700, respectively, using PCR amplification with labeled primers. The DNA target sequences, individual components of the DNA target, and labeled PCR primers are shown in Table 13.
[0325] Table 13. Target gBlock and primer sequences
[0326]
[0327] The cleavage activity of nuclease A (SEQ ID NO: 1) and each of the putative cleavage enzymes was assessed using an in vitro cleavage assay. The cleavage activity was evaluated by in vitro cleavage assays using plasmids encoding proteins of interest from Table 11 and linear DNA templates of AAVS1-T3 sgRNA transcribed from T7 from Table 12, in PURExpress containing the SUPERase•In™ RNase inhibitor (Ingenium Biotech). ® Each peptide was co-expressed individually with the RNA guide in vitro by incubation at 37°C for 2 hours in NEB buffer solution. The unpurified peptide / RNA solution was then diluted in 1X NEB buffer 2 (NEB) containing approximately 1 ng / µl of labeled DNA target amplicon. The solution was then incubated at 37°C for 1 hour. The reaction was stopped by incubation at 37°C with RNase Cocktail™ (Ingenieur; final concentration approximately 1 U / µl) for 15 minutes, followed by incubation at 55°C with proteinase K (NEB; final concentration approximately 0.04 U / µl) for 30 minutes. DNA was then purified using CleanNGS DNA and RNA cleaning magnetic beads (Bulldog Bio).
[0328] Samples were run on 10% TBE urea PAGE gels to separate cleaved and uncleaved products of the target and non-target strands. The gels were imaged using a LI-COR Odysssey M imaging system, and 5' IR800 and 5' IR700 labels on the target and non-target strands of the target DNA substrate were visualized using 800 nm and 700 nm channels. Band intensities were quantified using ImageJ software.
[0329] Gel images are shown in Figure 2A-2C The percentage of cleaved target chain and non-target chain is quantitatively shown in the figure. Figure 2D The text indicates uncutable chains, chains cut with HNH, and chains cut with RuvC. Figure 2A The image is a gel image captured using an 800 nm channel, showing the cleavage of the target chain. Figure 2B The image shows a gel captured using a 700 nm channel, illustrating the cleavage of the non-target chain. Figure 2C It comes from Figure 2A and Figure 2B Superimposed gel images. For example... Figure 2A-2DAs shown, nuclease A (SEQ ID NO: 1) cleaves both the target and non-target strands as expected. Each of the four HNH knockout cleavage constructs (D396A, H397A, H397L, and N420A) showed significantly reduced activity on the target strand, while retaining activity on the non-target strand. Two of the three RuvC knockout cleavage constructs (D24A and E337A) showed significantly reduced activity on the non-target strand, while retaining activity on the target strand. Figure 2A-2D ).
[0330] Therefore, this example demonstrates that the HNH knockout nickase and the RuvC knockout nickase have been successfully engineered.
[0331] Other variant nuclease polypeptides derived from nuclease A are provided in Table 14 below, each of which is also within the scope of this disclosure.
[0332] Table 14. Variant CRISPR nuclease peptides
[0333]
[0334] Example 6. Editing human genes in HEK293T cells using nuclease A-derived CRISPR nuclease peptides.
[0335] This example illustrates gene modification of a human gene using nuclease peptides derived from nuclease A, constructed according to the descriptions provided in Examples 1 and 3, and listed in Tables 1 and 14. Specifically, when combined with an RNA guide, the peptides are used to direct indels into human target genes.
[0336] RNA guides (including EMX1-T3 with a reference scaffold, a truncated reference scaffold, or scaffold 2, EMX1-T7 with scaffold 2, and AAVS-T3 with scaffold 2) are shown in Table 2 of Example 1 or Table 8 of Example 4. Samples were prepared as described in Example 1 for NGS and analysis.
[0337] It is expected that the nuclease polypeptide derived from nuclease A described in this example can generate indels at target sites encoded by RNA guides (such as the RNA guide described in this example) and enter the human gene.
[0338] Example 7 Editing human target genes in HEK293T cells using CRISPR nuclease K or its variants.
[0339] This example describes the genome editing of the AAVS1, EMX1, and VEGFA genes by nuclease K of SEQ ID NO: 65, which was introduced into the HEK293T cell line via lipid-based transient transfection.
[0340] Nuclease K was labeled with the N-terminal SV40 nuclear localization sequence (NLS) and the C-terminal nucleoplasmic protein nuclear localization sequence (NLS), and its coding sequence was converted into a human codon-optimized DNA sequence. The resulting DNA was synthesized and cloned into a pcDNA3.1 vector (Ingenie) containing a CMV promoter for expression. The reference sequences and NLS-labeled sequences used are shown in Table 15. The U6PolIII promoter, which uses +1 G at the start of the transcript (i.e., the 5' end of the RNA) for more efficient transcription, was excluded from the sequences described in Table 16. The plasmid was purified using the midiprep kit.
[0341] RNA guides were designed and cloned into the pUC19 plasmid following the U6 PolIII promoter and terminated with a 6xpolyT sequence. The RNA guides were designed to be specific for target sequences within the coding exons of AAVS1, EMX1, and VEGFA, which contain a 5'-NGG-3' PAM sequence (with the PAM sequence located at the 3' end of the target sequence). See Table 16 for all RNA guide sequences. The plasmids were purified using the midiprep kit.
[0342] Table 15. Amino acid sequences of exemplary CRISPR nuclease peptides derived from nuclease K
[0343]
[0344] Table 16. Target sequences and RNA guide sequences
[0345]
[0346] *The spacers are in uppercase letters, and the support (SEQ ID NO: 79) is in lowercase letters.
[0347] Following the method described in Example 1 above, HEK293T cells were transfected with a plasmid expressing a nuclease peptide derived from nuclease K and the gRNA provided in this paper. The gene editing efficiency was also determined by NGS following the method described in Example 1. The NGS results were then analyzed following the method described in Example 1 above.
[0348] For each target, the indel ratio was calculated for each sample and its homologous protein-free control; the indel ratio refers to the percentage of NGS reads containing indels. When nuclease K was included in the transfection, targets containing a higher percentage of indels indicated the outcome of DNA editing in the cells.
[0349] like Figure 4 As shown, when the nuclease K (SEQ ID NO: 65) plasmid was present, four of the six targets tested exhibited high levels of indel.
[0350] Example 8: The effectiveness of CRISPR nuclease K variants in targeting mammalian genes
[0351] This example describes indel evaluation of mammalian targets using a variant of nuclease K transfected into HEK293T cells.
[0352] Arginine scanning mutagenesis was performed to individually replace each non-arginine residue of reference nuclease K (SEQ ID NO: 65) with arginine. This yielded 682 single-arginine substitution variants. The nucleic acids encoding nuclease K and each of its variants were then individually cloned into the pcda3.1 backbone (Ingenie), and plasmids were prepared and diluted to maximize efficiency. The plasmids contained a CMV promoter, a first NLS (KRTADGSEFESPKKKRKV; SEQ ID NO: 3) upstream of the coding sequence, an XTEN linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS; SEQ ID NO: 95), followed by a second NLS (KRPAATKKAGQAKKKK; SEQ ID NO: 4) downstream of the coding sequence. See also Example 7 above.
[0353] The RNA guide sequence and target sequence are shown in Table 17. The RNA guide was cloned into the pUC19 backbone following the U6 PolIII promoter (New England Biolabs). ® The transcription is terminated with a 6x polyT sequence. U6 PolIII uses +1 G at the beginning of the transcript (i.e., the 5' end of the RNA) for more efficient transcription, which is excluded from the sequences described in Table 17. The plasmid is then prepared and diluted to the maximum extent.
[0354] Table 17. Mammalian targets and corresponding crRNAs.
[0355]
[0356] Following the method described in Example 1 above, HEK293T cells were transfected with a plasmid expressing a nuclease peptide derived from nuclease K and the gRNA provided in this paper. The gene editing efficiency was also determined by NGS following the method described in Example 1. The NGS results were then analyzed following the method described in Example 1 above.
[0357] For each target, the indel ratio for the reference and each variant was calculated, where the indel ratio refers to the percentage of NGS reads containing indels. The indel ratio used for fold change calculation was the average of the two technical replicas. Then, to calculate the fold change in the indel ratio at a specific target, the indel ratio of each variant was divided by the reference indel ratio. Table 4 shows the fold change in the indel ratio for each target and the average for the two targets. The number is relative to the reference nuclease (i.e., without NLS) of SEQ ID NO: 1. For example, at the AAVS1 target, the indel ratio of the D582R variant was 3.68 times that of the reference indel ratio, and at the EMX1 target, the indel ratio of the D582R variant was 31.80 times that of the reference indel ratio. At both targets, the average fold change in the indel ratio of the D582R variant was 17.74 times that of the average fold change in the reference indel ratio.
[0358] As shown in Table 18, the 29 variants with monoarginine substitution (left column) are characterized by an indel ratio that increases by at least 1.5 times relative to the reference indel ratio when the average is taken across two targets (right column).
[0359] Table 18. Changes in Indel Ratio*
[0360]
[0361] *Variant indel ratio / Reference indel ratio
[0362] The following 177 variants were analyzed, and when averaged across two targets, their indel ratios were 1 to 1.5 times that of the reference indel ratio: S548R, S365R, S134R, D729R, E747R, A724R, D581R, D301R, E623R, W127R, T180R, V114R, A395R, E123R, T703R, K68R, S258R, K60R, K324R, A297R, A381R, K402R, D217R, D356R, A98R, E266R, T85R, N218R, K311R, V396R, N400R, K131R, K63 3R, V692R, I565R, E671R, I109R, E110R, A709R, Q27R, K300R, Q125R, H25R, S 214R, T438R, V726R, K605R, D186R, K563R, L495R, P637R, E121R, T741R, K30 6R, E107R, N689R, K521R, L269R, V262R, Q58R, K333R, S79R, K414R, E212R, D 367R, K380R, K549R, D512R, K629R, K276R, P101R, E334R, G730R, M112R, K14R , F420R, L407R, D708R, P685R, I149R, K238R, E594R, E203R, K428R, K603R, W 120R, K697R, E108R, E102R, S745R, G492R, N437R, H244R, E115R, K45R, E538 R, K364R, K427R, S88R, K159R, T308R, V189R, K35R, K328R, D113R, N744R, K6 66R, K602R, Q50R, K49R, Q326R, K72R, E116R, G343R, K543R, E590R, T105R, K1 24R, E103R, K142R, K731R, G464R, A322R, K329R, M701R, P686R, S316R, K534 R, E230R, D242R, M571R, K610R, A111R, K664R, K62R, Q687R, E349R, P684R, K 344R, E342R, E431R, V184R, K679R, G337R, G727R, E150R, L193R, K146R, K26 5R, A419R, C606R, S275R, E705R, S511R, L624R, K338R, K240R, D221R, T519R,P469R, K104R, N133R, H239R, S678R, P210R, H145R, K466R, I584R, and K422R.
[0363] The remaining variants (483 variants) reduce the indel ratio relative to the reference indel ratio (the change in indel ratio is less than 1.0): G357R, K216R, N644R, E586R, M100R, D424R, D199R, D545R, K443R, G527R, S106R, G529R, K739R, K562R, K620R, H274R, L200R, E155R, T272R, S73R, S627R, N69R, E345R, K597R, K667R, K122R, K233R, P305R, Q119R, P415R, T346R, S369R, T130 R, K488R, K579R, K613R, N525R, S736R, E2R, E185R, K523R, K673R, E614R, Q2 13R, S412R, K220R, K61R, K668R, K280R, K591R, P171R, N175R, K714R, K118R , K138R, K289R, T587R, N370R, F208R, K593R, T732R, L661R, I540R, S728R, I 268R, E277R, G528R, K467R, K226R, Y95R, N318R, P688R, E533R, K526R, I710R , E215R, E126R, G604R, D513R, K181R, A743R, D254R, I619R, G188R, E723R, E 278R, K742R, E399R, K195R, L224R, P340R, K247R, N310R, S693R, G564R, N77 R, N227R, I158R, Y250R, C577R, N435R, N622R, T154R, K546R, D225R, G509R, A147R, E156R, T695R, N255R, D531R, K256R, Y670R, K92R, E471R, L390R, Y660 R, T738R, G52R, D251R, K621R, I411R, G363R, M518R, H535R, A253R, V735R, P 539R, E707R, L179R, T567R, V725R, P537R, E408R, N654R, S532R, K167R, E49 0R, E78R, Q248R, Q473R, N204R, K172R, E4R, I245R, T241R, G578R, N645R, G5 47R, T608R, V625R, F282R, E663R, F694R, E129R, I43R, L257R, C618R, E416R,S560R、E461R、E335R、Y524R、E302R、E162R、I476R、A350R、G656R、E139R、Q168R、N423R、Y510R、E574R、G550R、V197R、G616R、D143R、K183R、A222R、N465R、G514R、T601R、K348R、D440R、L384R、E231R、H271R、F599R、F174R、A41R、G665R、G690R、I178R、G706R、D609R、Y480R、D517R、V704R、G75R、T59R、G486R、I418R、K508R、K715R、A615R、D457R、Y76R、Y336R、L5R、T56R、L304R、S462R、H680R、P675R、E293R、I211R、L536R、N385R、I417R、F209R、G151R、F542R、F530R、L29R、E611R、S409R、T721R、H372R、V655R、H91R、G327R、I436R、W484R、L89R、V232R、I205R、D662R、G382R、T494R、D658R、V141R、G339R、T80R、V652R、H136R、G717R、Y561R、A463R、G403R、A433R、L64R、N651R、N166R、Y588R、A371R、Q157R、D432R、H653R、T405R、L520R、N196R、L696R、Y169R、K401R、L425R、P192R、P406R、S585R、L352R、N442R、A285R、V313R、F387R、V281R、I74R、A388R、Y454R、V140R、G410R、I682R、N190R、F263R、Y477R、I566R、D96R、I643R、M626R、V647R、N493R、V516R、L44R、S551R、D22R、T290R、N557R、T635R、L677R、I366R、L351R、A669R、F377R、E354R、A31R、F681R、N294R、K291R、G383R、E580R、C264R、G634R、N317R、F617R、A323R、V249R、L331R、F307R、D39R、K426R、V659R、F235R、E66R、P522R、G259R、N646R、D700R、H267R、S474R、L201R、C394R、V161R, G636R, N674R, F648R, L288R, Y15R, L296R, P286R, N32R, T261R, H447R, P413R, A612R, V6R, I386R, Q639R, T82R, T325R, F236R, I 292R, C359R, D38R, K649R, I575R, M640R, V295R, I164R, V40R, S446R, N176R, Y711R, Y97R, T191R, D320R, V583R, V298R, T650R, N389R, I 252R, I641R, V657R, I360R, A153R, I722R, L284R, I321R, V206R, I376R, V28R, M148R, C187R, Y30R, A737R, V391R, N398R, F83R, E315R, I 202R, L628R, V309R, L219R, V144R, M421R, F481R, V21R, L553R, M455R, A505R, P378R, I303R, P720R, I713R, A498R, Y70R, G12R, C182R, V 312R, V86R, L499R, S287R, A19R, M314R, T573R, L740R, L7R, T355R, I299R, L18R, I503R, Y361R, V20R, Y87R, Y11R, G94R, F441R, N26R, F 487R, M445R, L716R, F319R, L468R, H375R, V223R, G8R, A448R, D374R, V439R, A501R, V502R, M36R, L559R, F165R, G17R, L607R, V702R, A5 95R, L34R, K489R, L592R, S506R, W452R, F596R, C362R, L504R, T16R, L430R, S449R, D555R, H496R, D10R, H554R, G13R, G451R, A479R, L67R, T478R, P558R, D500R, I429R, I470R, Y444R, A392R, V572R, I450R, H497R, C393R, C397R, G475R, I9R, W485R, L71R, A90R, L459R, and A81R.
[0364] Based on this experiment, the following variants were selected for further engineering of nuclease K relative to SEQ ID NO: 65: D582R, D46R, I631R, F570R, I53R, G42R, E576R, T734R, C630R, E683R, S719R, and E128R.
[0365] Example 9: The effectiveness of combining CRISPR nuclease K variants in targeting mammalian genes
[0366] This example describes the indel evaluation of mammalian targets using nuclease variants of nuclease K, which contain two or more substitutions identified in Example 8 as increasing indel activity. Seventy-one combined nuclease K variants were tested.
[0367] As described in Example 8, each nuclease K variant and RNA guide were cloned (see Table 19). HEK293T cells were further transfected as described in Example 8, followed by NGS analysis. For each target, the indel ratio of the reference nuclease K (SEQ ID NO: 65) and each nuclease K variant was calculated, where the indel ratio refers to the percentage of NGS reads containing the indel. The indel ratios shown in Table 20 were calculated as the average of two biological replicas, each containing two technical replicas.
[0368] Table 19. Mammalian targets and corresponding crRNAs
[0369]
[0370] Table 20. Indel ratio of mammalian targets
[0371]
[0372] As shown in Table 20, all but three of the combined nuclease K variants exhibited higher indel activity than the reference nuclease K (SEQ ID NO: 65). When averaged across the three targets, two nuclease K variants resulted in an indel ratio exceeding 0.2, indicating that over 20% of the NGS reads contained indels. These two nuclease K variants contained the following substitutions: a) D582R, I53R, and T734R, and b) D582R and G42R. When averaged across the three targets, 42 nuclease K variants resulted in an indel ratio exceeding 0.1, indicating that over 10% of the NGS reads contained indels.
[0373] Based on this experiment, nuclease K variants containing the following substitutions were selected for further testing: D582R and G42R. Compared to the reference nuclease K (SEQ ID NO: 65), these nuclease K variants exhibited more than an 8-fold increase in indel activity.
[0374] Example 10: Nuclease M-mediated editing of human target sequences in HEK293T cells
[0375] This example describes genome editing of the AAVS1, EMX1, and VEGFA genes by nuclease M (SEQ ID NO: 80), which was introduced into the HEK293T cell line via lipid-based transient transfection.
[0376] Nuclease M was labeled with the N-terminal SV40 nuclear localization sequence and the C-terminal nucleoplasmic protein nuclear localization sequence, and converted into a human codon-optimized DNA sequence. This sequence was then synthesized and cloned into a pcDNA3.1 vector (Ingenie) containing the CMV promoter for expression. The reference sequences and NLS-labeled sequences used are shown in Table 21. The U6 PolIII promoter, which uses +1 G at the start of the transcript (i.e., the 5' end of the RNA) for more efficient transcription, was excluded from the sequences described in Table 22. The plasmid was purified using the midiprep kit.
[0377] RNA guides were designed and cloned into the pUC19 plasmid following the U6 PolIII promoter and terminated with a 6xpolyT sequence. The RNA guides were designed to be specific for target sequences within the coding exons of AAVS1, EMX1, and VEGFA, which contain a 5'-WTAAH-3' PAM sequence (with the PAM sequence located at the 3' end of the target sequence). See Table 22 for all RNA guide sequences. The plasmids were purified using the midiprep kit.
[0378] Table 21. Amino acid sequences
[0379]
[0380] Table 22. Target sequences and RNA guide sequences
[0381]
[0382] *The spacers are in uppercase letters, and the support (SEQ ID NO: 94) is in lowercase letters.
[0383] Following the method described in Example 1 above, HEK293T cells were transfected with a plasmid expressing a CRISPR nuclease peptide and the gRNA provided in this paper. Next-generation sequencing was also used to determine gene editing efficiency, following the method described in Example 1. The NGS results were then analyzed, following the method described in Example 1 above.
[0384] For each target, the indel ratio was calculated for each sample and its homologous protein-free control; the indel ratio refers to the percentage of NGS reads containing indels. When nuclease M was included in the transfection, targets containing a higher percentage of indels indicated DNA editing outcomes in the cells.
[0385] like Figure 5 As shown, when the nuclease M (SEQ ID NO: 80) plasmid was present, three of the six targets tested exhibited high levels of indel.
[0386] Example 11: The effectiveness of variant CRISPR nucleases in targeting mammalian genes
[0387] This example describes indel evaluation of mammalian targets using a CRISPR nuclease variant derived from nuclease M (SEQ ID NO: 80) transfected into HEK293T cells.
[0388] Arginine scanning mutagenesis was performed to selectively replace non-arginine residues of nuclease M (SEQ ID NO: 80) with arginine. This yielded 450 single-arginine substitution variants. The nucleic acids encoding the reference CRISPR nuclease and each CRISPR nuclease variant were then individually cloned into the pcDNA3.1 backbone (Invitrogen). TMThe plasmid was prepared and diluted to the maximum extent possible. The plasmid contained a CMV promoter, a first NLS (KRTADGSEFESPKKKRKV; SEQ ID NO: 3) upstream of the coding sequence, an XTEN adapter (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS; SEQ ID NO: 95), followed by a second NLS (KRPAATKKAGQAKKKK; SEQ ID NO: 4) downstream of the coding sequence. See also Example 10 above.
[0389] The RNA guide sequence and target sequence are shown in Table 23. The RNA guide was cloned into the pUC19 backbone following the U6 PolIII promoter (New England Biolabs). ® The transcription is terminated with a 6x polyT sequence. U6 PolIII uses +1 G at the beginning of the transcript (i.e., the 5' end of the RNA) for more efficient transcription, which is excluded from the sequences described in Table 23. The plasmid is then prepared and diluted to the maximum extent.
[0390] Table 23. Mammalian targets and corresponding crRNAs*
[0391]
[0392] *: Sequence information is shown in Table 22 above.
[0393] Following the method described in Example 1 above, HEK293T cells were transfected with a plasmid expressing a CRISPR nuclease peptide and the gRNA provided herein. Gene editing efficiency was determined by NGS, following the method described in Example 10. The NGS results were also analyzed, following the method described in Example 10 above.
[0394] The indel ratio for the reference and each variant was calculated, where the indel ratio refers to the percentage of NGS reads containing indels. The indel ratio used for fold change calculation was the average of the two technical replicas. Then, to calculate the fold change in the indel ratio, the indel ratio of each variant was divided by the indel ratio of the reference. Table 4 shows the fold change in the indel ratio. The numbers are relative to the reference nuclease M (i.e., without NLS) of SEQ ID NO: 80.
[0395] As shown in Table 24, the eight variants of nuclease M with monoarginine substitution (left column) are characterized by increasing the indel ratio by at least 1.5 times relative to the reference indel ratio.
[0396] Table 24. Changes in Indel Ratio*
[0397]
[0398] *Variant indel ratio / Reference indel ratio
[0399] 173 variants of the following nuclease M were analyzed, and their indel ratios were 1-1.5 times that of the reference indel ratio: G211R, D52R, T54R, K237R, E207R, A327R, K283R, A133R, K212R, E282R, K126R, L412R, K365R, K236R, S46R, F57R, A328R, V112R, L312R, N106R, Y141R, K143R, K277R, T414R, V116R, K10R, I102R, V81R, E27R, Y256R, E75R, G132R, G215R, V245R, A353R, Y 454R, G333R, V270R, E315R, H18R, T73R, L171R, S310R, T138R, K7R, V33R, V1 28R, L232R, K26R, I180R, V430R, Y49R, G351R, T294R, I356R, L219R, A225R, A335R, T398R, P11R, P182R, S79R, K140R, S253R, V4R, Q378R, Y367R, K436R, K257R, L487R, D386R, V32R, W130R, A318R, L167R, T50R, I137R, S470R, K28R, S249R, K127R, K29R, K110R, K309R, P37R, P447R, S6R, L24R, D443R, Y343R, G 9R, D194R, N218R, L25R, Y362R, L20R, K119R, N340R, V177R, Y321R, H370R, N 172R, T314R, I488R, K111R, L347R, Y408R, E146R, I359R, L242R, P131R, I19 5R, A492R, M301R, S337R, F174R, P136R, G402R, K145R, A451R, I5R, E474R, Y2 22R, S331R, L198R, T39R, N217R, T489R, K330R, E324R, F446R, L280R, N118R , S438R, D45R, V69R, W158R, Q400R, Y3R, Q227R, S339R, I296R, F122R, L12R, Y364R, G417R, S295R, I183R, C485R, T160R, I239R, E407R, S426R, I472R, Y1 21R, Y78R, A163R, M13R, P361R, T15R, Y44R, V434R, F38R, V2R, T169R, F149R,P125R, P144R, C375R, I317R, F374R, N151R, P14R, L186R, Q64R, C263R, C266R, N86R, M93R, Y431R, L334R, L308R, L297R, Y479R, Q366R, G252R, G60R, T162R, I389R, H244R, H271R, D376 R, P161R, A377R, V444R, K384R, V89R, A30R, A115R, A55R, L166R, T120R, C450R, E189R, T67 R, A147R, V460R, V21R, H267R, A68R, D58R, I66R, C234R, L181R, I300R, I40R, Y323R, T84R, N 258R, S65R, G191R, H165R, T59R, I142R, A139R, E307R, S463R, H243R, S80R, L471R, L188R, F465R, G56R, H170R, L428R, T439R, A421R, N466R, S404R, I53R, I425R, Y173R, L82R, F464R, D380R, G261R, L304R, S342R, K336R, L262R, G322R, L382R, L405R, A433R, A391R, A346R, G449R, V452R, V461R, G468R, N393R, C231R, F193R, D341R, T325R, Y383R, M345R, H338R, and A344R.
[0400] The remaining variants of nuclease M (269 variants) reduce the indel ratio relative to the nuclease M indel ratio (the fold change in the indel ratio is less than 1.0).
[0401] Based on this experiment, variants of nuclease M (SEQ ID NO: 80) with one or more of the following substitutions exhibited enhanced activity: E88R, S95R, L92R, E401R, E83R, N371R, P481R, and A373R.
[0402] Other embodiments
[0403] All features disclosed in this specification can be combined in any combination. Each feature disclosed in this specification can be replaced by alternative features for the same, equivalent, or similar purposes. Therefore, unless otherwise expressly stated, each disclosed feature is merely an example of an equivalent or similar feature in a general series.
[0404] From the above description, those skilled in the art can readily determine the essential characteristics of the present invention, and various changes and modifications can be made to adapt it to various uses and conditions without departing from the spirit and scope of the invention. Therefore, other embodiments are also within the scope of the claims.
[0405] Equivalent solution
[0406] Although several embodiments of the invention have been described and illustrated herein, those skilled in the art will readily conceive of various other methods and / or structures for performing the functions described herein and / or obtaining one or more of the results and / or advantages, and each of such variations and / or modifications is considered to be within the scope of the embodiments of the invention described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications used in the teachings of this invention. Those skilled in the art will recognize or be able to determine many equivalents of the specific embodiments of the invention described herein using only conventional experiments. Therefore, it should be understood that the foregoing embodiments are presented by way of example only, and that the embodiments of the invention may be practiced in ways different from those specifically described and claimed within the scope of the appended claims and their equivalents. The embodiments of the invention disclosed herein relate to each individual feature, system, article of manufacture, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits and / or methods is included within the scope of this disclosure if such features, systems, articles, materials, kits and / or methods do not contradict each other.
[0407] All definitions defined and used herein should be understood as controls over dictionary definitions, definitions incorporated by reference in other documents, and / or the general meaning of the definition terms.
[0408] All references, patents, and patent applications disclosed herein are incorporated by reference to their respective subjects, and in some cases, may cover the entire document.
[0409] Unless explicitly stated otherwise, as used herein in the specification and claims, the indefinite articles “a / an” and “a type” shall be understood to mean “at least one type”.
[0410] As used herein in the specification and claims, the phrase “and / or” should be understood to mean “any one or both” of the elements so combined, that is, elements that exist together in some cases and separately in others. Multiple elements listed with “and / or” should be interpreted in the same way, that is, “one or more elements” of the elements so combined. In addition to the elements specifically identified by the “and / or” clause, other elements may optionally exist, whether related to or unrelated to those specifically identified. Thus, as a non-limiting example, when used in conjunction with open-ended language such as “comprising,” a reference to “A and / or B” may in one embodiment refer only to A (optionally including elements other than B); in another embodiment, only to B (optionally including elements other than A); in yet another embodiment, both A and B (optionally including other elements); and so on.
[0411] As used herein in the specification and claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” should be interpreted as inclusive, i.e., including a plurality of elements or at least one element in a list of elements, but also including more than one element and optionally additional unlisted items. Only explicitly indicating the opposite terms such as “only one” or “exactly one” or, when used in claims, “consisting of…” means including a plurality of elements or exactly one element in a list of elements. In general, when preceded by exclusive terms such as “any one of…”, “one of…”, “only one of…” or “exact one of…”, the term “or” as used herein should be interpreted only to indicate a single alternative (i.e., “one or the other, but not both”). When used in claims, “consisting substantially of…” should have the ordinary meaning as used in the field of patent law.
[0412] As used herein in the specification and claims, the phrase "at least one" relating to a list of one or more elements should be understood to mean at least one element selected from any one or more elements in the element list, but not necessarily including at least one element from every element specifically listed in the element list, and does not exclude any combination of elements in the element list. This definition also allows for the optional presence of elements other than those specifically identified in the element list referred to by the phrase "at least one," whether related to or unrelated to those specifically identified elements. Thus, as a non-limiting example, in one embodiment, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") may mean optionally including more than one A, with no B (and optionally including elements other than B); in another embodiment, it may mean optionally including more than one B, with no A (and optionally including elements other than A); in yet another embodiment, it may mean optionally including at least one more than one A, and optionally including more than one B (and optionally including other elements); and so on.
[0413] It should also be understood that, unless expressly stated to the contrary, in any method claimed herein that includes more than one step or action, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are described.
Claims
1. A CRISPR nuclease polypeptide comprising a RuvC nuclease domain and a HNH nuclease domain, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 90% identical to SEQ ID NO: 1; optionally wherein the CRISPR nuclease polypeptide is a variant comprising at least one mutation relative to SEQ ID NO:
1.
2. The CRISPR nuclease polypeptide of claim 1, wherein the CRISPR nuclease polypeptide is the variant comprising the at least one mutation, the at least one mutation comprising: (a) one or more arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 1, optionally one or more arginine substitutions; (b) one or more nickase mutations in the HNH nuclease domain or in the RuvC nuclease domain of SEQ ID NO: 1; (c) an N-terminal truncation; (d) a C-terminal truncation; or (e) a combination of (a), (b), (c), and / or (d).
3. The CRISPR nuclease polypeptide of claim 1 or claim 2, wherein the at least one mutation comprises (a) in the bridge helix (BH) domain, in the phospho-lock loop (PLL) domain, in the wedge (WED) domain, in the PAM interaction (PID) domain, or a combination thereof.
4. The CRISPR nuclease polypeptide of claim 2, wherein at least one mutation comprises (a), and wherein the one or more arginine substitutions and / or lysine substitutions, optionally one or more arginine substitutions, are at one or more of positions D56, E59, G60, E63, I67, G71, S564, D568, T605, E655, E694, and E706 in SEQ ID NO:
1.
5. The CRISPR nuclease polypeptide of claim 4, wherein the CRISPR nuclease polypeptide comprises arginine substitutions and / or lysine substitutions at the following positions relative to SEQ ID NO: 1: a) I67, D568, and E706; b) I67, D56, and D568; c) I67, T605, and D568; d) I67, D56, and E706; e) I67, T605, and E706; f) I67, D568, and E694; g) I67, D56, E694, and E706; h) I67, D56, T605, and E706; or i) I67, D568, W593, and E706.
6. The CRISPR nuclease polypeptide of any one of claims 1-5, wherein the CRISPR nuclease polypeptide has an N-terminal truncation relative to SEQ ID NO: 1, optionally wherein the N-terminal truncation is a deletion within residues 1-15 of SEQ ID NO: 1; preferably wherein the N-terminal truncation is a deletion of residues 1-14 or residues 1-15 of SEQ ID NO:
1.
7. The CRISPR nuclease of any one of claims 2-6, wherein the CRISPR nuclease polypeptide contains up to 20 arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 1, optionally up to 20 arginine substitutions; optionally wherein the CRISPR nuclease polypeptide contains up to 15 arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 1, optionally up to 15 arginine substitutions.
8. The CRISPR nuclease of claim 7, wherein the CRISPR nuclease polypeptide comprises the following combinations of arginine substitutions: a) I67R, D568R, and E706R; b) I67R, D56R, and D568R; c) I67R, T605R, and D568R; d) I67R, D56R, and E706R; e) I67R, T605R, and E706R; f) I67R, D568R, and E694R; g) I67R, D56R, E694R, and E706R; h) I67R, D56R, T605R, and E706R; or i) I67R, D568R, W593R, and E706R.
9. The CRISPR nuclease of claim 8, wherein the CRISPR nuclease polypeptide comprises the arginine substitutions I67R, D568R, and E706R.
10. The CRISPR nuclease of any one of claims 2-9, wherein the at least one mutation in the CRISPR nuclease polypeptide comprises (b) comprising one or more mutations at positions H397, D396, N420, D24, E337, and / or D524 of SEQ ID NO:
1.
11. The CRISPR nuclease of claim 10, wherein the at least one mutation comprises a mutation at position H397, optionally wherein the mutation is an amino acid substitution of H397A or H397L.
12. The CRISPR nuclease polypeptide of claim 2, wherein the variant comprises: (a) the one or more nickase mutations in the HNH nuclease domain, optionally at positions D396, H397, and / or N420 of SEQ ID NO: 1; optionally wherein the nickase mutation is at position H397; and (b) the one or more arginine substitutions and / or lysine substitutions, optionally at positions I67, D568, and E706 of SEQ ID NO:
1.
13. The CRISPR nuclease polypeptide of claim 2, wherein the variant comprises: (a) the one or more nickase mutations in the HNH nuclease domain, optionally at positions D396, H397, and / or N420 of SEQ ID NO: 1; (b) one or more arginine substitutions and / or lysine substitutions, optionally at positions I67, D568, and E706 of SEQ ID NO: 1; and (c) the N-terminal truncation within residues 1-15 of SEQ ID NO: 1, optionally wherein the N-terminal truncation is the deletion of residues 1-14 or 1-15 of SEQ ID NO:
1.
14. The CRISPR nuclease polypeptide of claim 13, wherein the variant comprises: (a) the nickase mutation of H397A; (b) the arginine substitutions of I67R, D568R, and E706R; and (c) the N-terminal truncation comprising a deletion of residues 1-14 or 1-15 of SEQ ID NO:
1.
15. The CRISPR nuclease polypeptide of claim 1, wherein the CRISPR nuclease polypeptide is listed in Table 1, Table 11, or Table 14.
16. The CRISPR nuclease polypeptide of any one of claims 1-15, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 95% identical to SEQ ID NO:
1.
17. The CRISPR nuclease polypeptide of claim 16, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 98% identical to SEQ ID NO:
1.
18. The CRISPR nuclease polypeptide of any one of claims 1-17, which is a fusion polypeptide further comprising one or more functional fragments.
19. The CRISPR nuclease polypeptide of claim 18, wherein the one or more functional fragments comprise one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof.
20. A nucleic acid comprising a nucleotide sequence encoding the CRISPR nuclease polypeptide of any one of claims 1-19.
21. The nucleic acid of claim 20, wherein the nucleic acid is messenger RNA (mRNA).
22. The nucleic acid of claim 20, wherein the nucleic acid is an expression vector, wherein the nucleotide sequence encoding the CRISPR nuclease polypeptide is operably linked to a promoter; optionally wherein the expression vector is a viral vector.
23. A host cell comprising the nucleic acid of claim 21 or claim 22.
24. A gene editing system comprising: (a) a CRISPR nuclease polypeptide or a first nucleic acid encoding a CRISPR nuclease, wherein the CRISPR nuclease polypeptide is as described in any one of claims 1-19; and (b) a guide RNA (gRNA) or a second nucleic acid encoding the gRNA, wherein the gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease and a spacer specific for a target sequence in a genomic site of interest, wherein the target sequence is adjacent to a protospacer adjacent motif (PAM).
25. The gene editing system of claim 24, wherein the scaffold sequence comprises a nucleotide sequence that is at least 85% identical to SEQ ID NO: 2; optionally wherein the scaffold sequence comprises one or more deletions, one or more nucleotide substitutions, or a combination thereof, as compared to SEQ ID NO:
2.
26. The gene editing system of claim 25, wherein the scaffold sequence is a truncated variant of SEQ ID NO: 2 having a 3' truncation relative to SEQ ID NO: 2, and wherein the truncated variant is about 110-140 nt in length, optionally wherein the 3' truncation comprises a deletion within residues 143-202 of SEQ ID NO:
2.
27. The gene editing system of claim 26, wherein the truncated variant further comprises one or more deletions within residues 10-40 of SEQ ID NO: 2; optionally wherein the deletions comprise residues 14-20 and / or 25-32 of SEQ ID NO:
2.
28. The gene editing system of claim 26 or claim 27, wherein the truncated variant further comprises one or more mutations relative to SEQ ID NO: 2; optionally wherein the one or more mutations comprise a deletion within residues 82-85 of SEQ ID NO:
2.
29. The gene editing system of claim 26, wherein the scaffold sequence comprises the nucleotide sequence of SEQ ID NO: 27 or SEQ ID NO: 28; and wherein the scaffold sequence is about 115-130 nt in length.
30. The gene editing system of claim 29, wherein in the gene editing system: (a) the CRISPR nuclease polypeptide comprises the amino acid sequence of SEQ ID NO: 1; and the scaffold comprises the nucleotide sequence of SEQ ID NO: 27; (b) the CRISPR nuclease polypeptide is a variant of SEQ ID NO: 1 comprising mutations at positions I67, D568, and E706 of SEQ ID NO: 1, which are optionally the arginine substitutions I67R, D568R, and E706R, and the scaffold comprises the nucleotide sequence of SEQ ID NO: 28; or (c) the CRISPR nuclease polypeptide is a variant of SEQ ID NO: 1 comprising a deletion within residues 1-15 of SEQ ID NO: 1; optionally a deletion of residues 1-14 or 1-15 of SEQ ID NO: 1; and the scaffold comprises the nucleotide sequence of SEQ ID NO:
28.
31. The gene editing system of any one of claims 24-30, wherein the target sequence is adjacent to the PAM of 5'-NGG-3', wherein N represents any nucleotide.
32. The gene editing system of any one of claims 24-31, further comprising one or more lipid excipients associated with element (a) and / or element (b) of the gene editing system; optionally wherein the one or more lipid excipients form a lipid nanoparticle associated with or encapsulating the element (a) and / or the element (b) of the gene editing system.
33. A method of gene editing comprising delivering the gene editing system of any one of claims 24-32 to a host cell to edit a genomic site targeted by the gRNA of the gene editing system.
34. A guide RNA comprising a spacer sequence and a scaffold sequence, wherein the scaffold sequence is a variant of SEQ ID NO: 2 comprising one or more deletions, one or more nucleotide substitutions, or a combination thereof, as compared to SEQ ID NO: 2, and wherein the scaffold sequence is recognizable by the CRISPR nuclease polypeptide of any one of claims 1-19.
35. The guide RNA of claim 34, wherein the scaffold sequence is a truncated variant of SEQ ID NO: 2 having a 3' truncation relative to SEQ ID NO: 2, and wherein the truncated variant is about 110-140 nt in length.
36. The guide RNA of claim 35, wherein the 3' truncation comprises a deletion within residues 143-202 of SEQ ID NO:
2.
37. The guide RNA of claim 35 or claim 36, wherein the truncated variant further comprises one or more deletions within residues 10-40 of SEQ ID NO:
2.
38. The guide RNA of claim 37, wherein the deletions comprise residues 14-20 and / or 25-32 of SEQ ID NO:
2.
39. The guide RNA of claim 37 or claim 38, wherein the truncated variant further comprises one or more mutations relative to SEQ ID NO: 2; optionally the one or more mutations comprise a deletion within residues 81-85 of SEQ ID NO:
2.
40. The guide RNA of claim 35, wherein the scaffold sequence comprises the nucleotide sequence of SEQ ID NO: 27 or SEQ ID NO: 28; and wherein the scaffold is about 115-130 nt in length.
41. A CRISPR nuclease polypeptide comprising a RuvC nuclease domain and an HNH nuclease domain, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 90% identical to SEQ ID NO: 65; optionally wherein the CRISPR nuclease polypeptide is a variant comprising at least one mutation relative to SEQ ID NO:
65.
42. The CRISPR nuclease polypeptide of claim 41, wherein the CRISPR nuclease polypeptide is the variant comprising at least one mutation comprising: (a) one or more arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 65, optionally one or more arginine substitutions; (b) one or more nickase mutations in the HNH nuclease domain or in the RuvC nuclease domain of SEQ ID NO: 65; or (c) a combination of (a) and (b).
43. The CRISPR nuclease polypeptide of claim 42, wherein the CRISPR nuclease polypeptide comprises a bridge helix (BH) domain, a wedge (WED) domain, and a PAM interacting (PID) domain, and wherein the CRISPR nuclease polypeptide is the variant comprising the one or more arginine substitutions and / or lysine substitutions of (a) in the BH domain, in the WED domain, in the PID domain, or a combination thereof.
44. The CRISPR nuclease polypeptide of claim 43, wherein the one or more arginine substitutions and / or lysine substitutions are at one or more of positions G42, D46, I53, F83, E128, E541, F570, E576, D582, C630, I631, E683, S719, T734 in SEQ ID NO:
65.
45. The CRISPR nuclease polypeptide of claim 44, wherein the CRISPR nuclease polypeptide comprises arginine substitutions and / or lysine substitutions at the following positions relative to SEQ ID NO: 65: (i) D582 and G42; or (ii) D582, I53, and T734R.
46. The CRISPR nuclease polypeptide of claim 45, wherein the CRISPR nuclease polypeptide comprises the following arginine substitutions: a) D582R and G42R; or b) D582R, I53R, and T734R.
47. The CRISPR nuclease polypeptide of any one of claims 42-46, wherein the CRISPR nuclease polypeptide contains up to 20 arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 65, optionally up to 20 arginine substitutions; optionally wherein the CRISPR nuclease polypeptide contains up to 15 arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 65, optionally up to 15 arginine substitutions.
48. The CRISPR nuclease polypeptide of any one of claims 42-47, wherein the CRISPR nuclease polypeptide comprises the one or more nickase mutations of (b) at positions D374, H375, and / or N398 in SEQ ID NO:
65.
49. The CRISPR nuclease polypeptide of claim 48, wherein the nickase mutation is at position H375; optionally wherein the mutation is an amino acid substitution of H375A.
50. The CRISPR nuclease polypeptide of any one of claims 41-49, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 95% identical to SEQ ID NO:
65.
51. The CRISPR nuclease polypeptide of claim 50, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 98% identical to SEQ ID NO:
65.
52. The CRISPR nuclease polypeptide of any one of claims 41-51, which is a fusion polypeptide comprising one or more additional functional elements.
53. The CRISPR nuclease polypeptide of claim 52, wherein the one or more additional functional elements comprise one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof.
54. A nucleic acid comprising a nucleotide sequence encoding the CRISPR nuclease polypeptide of any one of claims 41-53.
55. The nucleic acid of claim 54, wherein the nucleic acid is an expression vector in which the nucleotide sequence encoding the CRISPR nuclease polypeptide is operably linked to a promoter; optionally wherein the expression vector is a viral vector.
56. The nucleic acid of claim 54, wherein the nucleic acid is a messenger RNA (mRNA).
57. A host cell comprising the nucleic acid of any one of claims 54-56.
58. A gene editing system comprising: (a) the CRISPR nuclease polypeptide of any one of claims 41-53 or a first nucleic acid encoding the CRISPR nuclease polypeptide; and (b) a guide RNA (gRNA) or a second nucleic acid encoding the gRNA, wherein the gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease polypeptide and a spacer sequence specific for a target sequence within a genomic locus of interest, wherein the target sequence is adjacent to a protospacer adjacent motif (PAM).
59. The gene editing system of claim 58, wherein the scaffold sequence comprises a nucleotide sequence that is at least 85% identical to SEQ ID NO:
79.
60. The gene editing system of claim 59, wherein the scaffold sequence is a fragment of SEQ ID NO:
79.
61. The gene editing system of any one of claims 58-60, wherein the scaffold sequence comprises one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to SEQ ID NO:
79.
62. The gene editing system of any one of claims 58-61, wherein the target sequence is located upstream of the PAM of 5'-NGG-3', wherein N represents any nucleotide.
63. The gene editing system of any one of claims 58-62, further comprising one or more lipid excipients associated with element (a) and / or element (b) of the gene editing system; optionally wherein the one or more lipid excipients form a lipid nanoparticle associated with or encapsulating the element (a) and / or the element (b) of the gene editing system.
64. A method of gene editing comprising delivering the gene editing system of any one of claims 58-63 to a host cell to edit a genomic site targeted by the gRNA of the gene editing system.
65. A nuclease polypeptide comprising a RuvC nuclease domain and a HNH nuclease domain, wherein the nuclease polypeptide comprises an amino acid sequence that is at least 90% identical to SEQ ID NO: 80; optionally wherein the nuclease polypeptide is a variant of SEQ ID NO: 80 comprising at least one mutation relative to SEQ ID NO:
80.
66. The nuclease polypeptide of claim 65, wherein the nuclease polypeptide is the variant comprising the at least one mutation, the at least one mutation comprising: (a) one or more arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 80, optionally one or more arginine substitutions; (b) one or more nickase mutations in the HNH nuclease domain or in the RuvC nuclease domain of SEQ ID NO: 80; or (c) a combination of (a) and (b).
67. The nuclease polypeptide of claim 66, wherein the at least one mutation comprises (a) in a bridge helix (BH) domain, in a nucleic acid recognition (REC) domain, in a phospho- lock loop (PLL) domain, in a wedge (WED) domain, in a PAM interaction (PID) domain, in a nuclease domain, or a combination thereof.
68. The nuclease polypeptide of claim 66 or claim 67, wherein the nuclease polypeptide contains at most 20 arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 80, optionally at most 20 arginine substitutions.
69. The nuclease polypeptide of claim 68, wherein the nuclease polypeptide contains at most 15 arginine substitutions and / or lysine substitutions relative to SEQ ID NO: 80, optionally at most 15 arginine substitutions.
70. The nuclease polypeptide of any one of claims 67-69, wherein the one or more arginine substitution and / or lysine substitution, optionally an arginine substitution, is at position E88, S95, L92, E401, E83, N371, P481, and / or A373 in SEQ ID NO:
80.
71. The nuclease polypeptide of any one of claims 66-70, wherein the at least one mutation comprises (b), and wherein the one or more nickase mutation is at one or more of positions D58, E189, D341, H243, H244, H267, R329, and / or H338 in SEQ ID NO:
80.
72. The nuclease polypeptide of any one of claims 65-71, wherein the nuclease polypeptide comprises an amino acid sequence that is at least 95% identical to SEQ ID NO:
80.
73. The nuclease polypeptide of claim 72, wherein the nuclease polypeptide comprises an amino acid sequence that is at least 98% identical to SEQ ID NO:
80.
74. The nuclease polypeptide of any one of claims 65-73, wherein the nuclease polypeptide is a fusion polypeptide further comprising one or more functional fragments heterologous to the nuclease portion in the fusion polypeptide.
75. The nuclease polypeptide of claim 74, wherein the one or more functional fragments comprise one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof.
76. A nucleic acid comprising a nucleotide sequence encoding the nuclease polypeptide of any one of claims 65-75.
77. The nucleic acid of claim 76, wherein the nucleic acid is an expression vector in which the nucleotide sequence encoding the nuclease is operably linked to a promoter; optionally wherein the expression vector is a viral vector.
78. The nucleic acid of claim 77, wherein the nucleic acid is a messenger RNA (mRNA).
79. A host cell comprising the nucleic acid of any one of claims 76-78.
80. A gene editing system comprising: (a) a nuclease polypeptide or a first nucleic acid encoding a nuclease, wherein the nuclease polypeptide is as described in any one of claims 65-75; and (b) a guide RNA (gRNA) or a second nucleic acid encoding the gRNA, wherein the gRNA comprises a scaffold sequence recognizable by the nuclease polypeptide and a spacer sequence specific for a target sequence in a genomic site of interest, wherein the target sequence is adjacent to a protospacer adjacent motif (PAM).
81. The gene editing system of claim 80, wherein the scaffold sequence comprises a nucleotide sequence that is at least 85% identical to SEQ ID NO:
94.
82. The gene editing system of claim 81, wherein the scaffold sequence is a fragment of SEQ ID NO:
94.
83. The gene editing system of claim 82, wherein the scaffold sequence comprises one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to SEQ ID NO:
94.
84. The gene editing system of any one of claims 80-83, wherein the PAM is 5'-WTAAH-3', wherein W is A or T, and H is A, C, or T; optionally wherein the PAM is 5'-TTAAA-3'.
85. The gene editing system of any one of claims 80-84, further comprising one or more lipid excipients associated with element (a) and / or element (b) of the gene editing system; optionally wherein the one or more lipid excipients form a lipid nanoparticle associated with or encapsulating the element (a) and / or the element (b) of the gene editing system.
86. A method of gene editing, comprising delivering the gene editing system of any one of claims 80-85 to a host cell to edit a genomic site targeted by the gRNA of the gene editing system.
Citation Information
Cited By
Cas protein, fusion protein, corresponding gene editing system and application thereof
CN121780486A
A Cas protein, a fusion protein, its corresponding gene editing system and applications
CN121780486B