CRISPR nuclease polypeptides and gene editing systems containing them
Patent Information
- Application Number
- JP2026513622
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2024-08-30
- Publication Date
- 2026-09-03
Smart Images

Figure 2026530077000001_ABST
Abstract
Description
Technical Field
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS The present application claims the benefit of the filing dates of U.S. Provisional Application No. 63 / 580,193 filed on September 1, 2023, U.S. Provisional Application No. 63 / 580,173 filed on September 1, 2023, U.S. Provisional Application No. 63 / 553,840 filed on February 15, 2024, and U.S. Provisional Application No. 63 / 638,552 filed on April 25, 2024. Each of the priority applications is incorporated herein by reference in their entireties.
[0002] SEQUENCE LISTING This application contains a Sequence Listing which has been filed electronically in XML format and is incorporated herein by reference in its entirety. The XML copy, created on August 27, 2024, is named 063586-524001WO_SeqList_ST26.xml and has a size of 330.5 bytes. Background Art
[0003] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes, collectively known as CRISPR-Cas or CRISPR / Cas systems, are adaptive immune systems in archaea and bacteria that protect specific species against exogenous genetic elements.
[0004] A CRISPR-Cas system typically comprises a CRISPR nuclease and one or more RNA components that direct the CRISPR nuclease to a target genomic site for gene editing. It is an object of the present invention to develop efficient CRISPR nucleases to improve gene editing efficiency. Summary of the Invention
[0005] This disclosure provides CRISPR nucleases and gene editing systems containing them for use in genetic modification of target genes. In some embodiments, the CRISPR nucleases are engineered CRISPR nuclease polypeptides that exhibit advantageous enzymatic activity (e.g., high indel activity and / or DNA cleavage activity). Therefore, gene editing systems containing them, as provided herein, are expected to exhibit superior effects when used in gene editing.
[0006] Accordingly, this specification provides a CRISPR nuclease polypeptide derived from the reference CRISPR nuclease of SEQ ID NO: 1 (e.g., the manipulated CRISPR nuclease polypeptide of SEQ ID NO: 1), a gene editing system comprising the CRISPR nuclease polypeptide and a guide RNA targeting a genomic site of interest, and the use of the gene editing system for modifying a genomic site of interest in a host cell. In some cases, the CRISPR nuclease polypeptide disclosed herein is a manipulated CRISPR nuclease polypeptide. In some cases, the manipulated CRISPR nuclease polypeptide is an arginine and / or lysine substitution variant of the reference nuclease, which may exhibit enhanced bioactivity such as indel activity. In other cases, the manipulated CRISPR nuclease polypeptide may be a nickase variant of the reference nuclease, the variant nuclease having reduced or eliminated nuclease activity in one of the nuclease domains relative to the reference nuclease, thereby exhibiting nickase activity. In other instances, the manipulated CRISPR nuclease polypeptide may contain one or more mutations that result in reduced PAM recognition stringency relative to the reference enzyme. In some examples, the manipulated CRISPR nuclease polypeptide may contain arginine / lysine substitutions, nickase mutations, and / or mutations that result in less PAM recognition stringency.
[0007] In some embodiments, the Disclosure provides an engineered CRISPR nuclease polypeptide which is a variant of the reference CRISPR nuclease described as SEQ ID NO: 1. The reference nuclease of SEQ ID NO: 1 comprises a RuvC nuclease domain and an HNH nuclease domain. With respect to the reference CRISPR nuclease, the engineered CRISPR nuclease polypeptide comprises (i) one or more mutations located in the HNH nuclease domain or in the RuvC nuclease domain that reduce or eliminate the nuclease activity thereof, optionally one or more of which are located in the HNH nuclease domain; (ii) one or more arginine and / or lysine substitutions, optionally one or more arginine substitutions; or (iii) a combination of (i) and (ii).
[0008] In some embodiments, the manipulated CRISPR nuclease polypeptide may contain one or more mutations in the HNH nuclease domain at positions D844, H845, and / or N868 relative to SEQ ID NO: 1. In some examples, the mutation is at position H845. In some cases, one or more mutations in the HNH nuclease domain may be, for example, amino acid residue substitutions at one or more of positions D844, H845, and N868. In one example, the mutation at D844 is an amino acid substitution, e.g., D844A, D844G, D844L, or D844S. In another example, the mutation at H845 is an amino acid substitution, e.g., H845A, H845G, H845L, or H845S. In one specific example, the manipulated CRISPR nuclease polypeptide contains a mutation at position H845 (e.g., H845A) relative to SEQ ID NO: 1 (e.g., containing the amino acid sequence of SEQ ID NO: 32). In yet another example, the mutation at N868 is an amino acid substitution, e.g., N868A, N868G, N868L, or N868S.
[0009] In some embodiments, the CRISPR nuclease polypeptides disclosed herein may contain one or more nickase mutations (e.g., amino acid residue substitutions) in the RuvC nuclease domain at positions D10, E763, and / or D991, for example, relative to SEQ ID NO: 1. In some examples, the mutation may be at position E763 or D991.
[0010] Alternatively or in addition, the manipulated CRISPR nuclease polypeptide may contain a bridge helix (BH) domain, a nucleic acid recognition (REC) domain, a phosphate-locked loop (PLL) domain, a wedge (WED) domain, and a PAM interaction (PID) domain, and may contain one or more arginine and / or lysine substitutions (e.g., arginine substitutions) in the BH domain, REC domain, PLL domain, WED domain, PID domain, or a combination thereof. In some embodiments, the manipulated CRISPR nuclease polypeptide may contain up to 20 arginine and / or lysine substitutions relative to the reference CRISPR nuclease. For example, the manipulated CRISPR nuclease polypeptide may contain up to 15 arginine and / or lysine substitutions relative to the reference CRISPR nuclease. In specific examples, one or more arginine and / or lysine substitutions may be located at positions K736, L784, Q812, N813, I857, and / or A919 in SEQ ID NO: 1. For example, in some embodiments, the CRISPR nuclease includes the I857R substitution.
[0011] In a specific example, the manipulated CRISPR nuclease polypeptide has arginine and / or lysine substitutions at the following positions relative to SEQ ID NO: 1: (a) I857, L784, and K736 (e.g., I857R, L784R, K736R), (b) I857, A919, and K736 (e.g., I857R, A919R, K736R), (c) I857, N813, and L784 (e.g., I857R, N813R, L784R), (d) I857, L784, and A919 (e.g., I857R, L784R, A919R), (e) I857, N813, and K736 (e.g., I857R, N813R, K736R), (f) I857 and N813 (e.g., I857R, N813R), (g) L784, A919, and K736 (e.g., L784R, A919R, K736R), (h) I857R and L784 (e.g., I857R and L784R), or (i) I857 and A919 (e.g., I857R and A919R) may contain these. More specifically, the manipulated CRISPR nuclease polypeptide may contain arginine-substituted I857R, L784R, and K736R relative to SEQ ID NO: 1.
[0012] In some embodiments, the manipulated CRISPR nuclease polypeptides disclosed herein may contain one or more mutations that enhance the double-stranded nuclease activity of the reference CRISPR nuclease into which the mutation is introduced. Examples of such CRISPR nuclease polypeptides are provided in Table 8 below.
[0013] Alternatively or in addition, the manipulated CRISPR nuclease polypeptides disclosed herein may, or may further, contain one or more mutations to reduce PAM recognition stringency. In some cases, one or more mutations to reduce PAM recognition stringency may be located at positions D61, A68, H494, L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345, and / or T1347 of SEQ ID NO: 1. In some examples, such mutations may include (i) one or more arginine and / or lysine substitutions, optionally arginine substitutions, at positions D61, A68, H494, L1117, G1227, S1327, A1332, and / or T1347 of SEQ ID NO: 1; (ii) one or more amino acid substitutions at positions D1144, S1145, E1228, R1343, and / or R1345 of SEQ ID NO: 1; or (iii) a combination of (i) and (ii). In specific examples, one or more amino acid substitutions in (ii) may optionally include D1144L, S1145W, E1228Q, R1343P, R1345V, and / or R1345Q for SEQ ID NO: 1.
[0014] In specific examples, engineered CRISPR nuclease polypeptides with less PAM recognition stringency may include the following combinations of mutations: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and A68R for SEQ ID NO: 1. In other specific examples, engineered CRISPR nuclease polypeptides with less PAM recognition stringency may include the following combinations of mutations: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R for SEQ ID NO: 1. In further specific examples, engineered CRISPR nuclease polypeptides with less PAM recognition stringency may include the following combinations of mutations: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and H494R for SEQ ID NO: 1. Other exemplary engineered CRISPR nuclease polypeptides can be seen in Table 9, each of which is within the scope of this disclosure.
[0015] The engineered CRISPR nuclease polypeptides disclosed herein recognize the 5'-NDR-3' PAM sequence, where N represents A, C, G, or U, D represents A, G, or T, and R represents G or A. In some cases, engineered CRISPR nuclease polypeptides with reduced PAM recognition stringency, such as those disclosed herein, may recognize the 5'-NGN-3' PAM sequence. See Example 6 below. In some examples, the PAM is 5'-NRG-3' or 5'-NRR-3', where N and R are as defined herein. In some specific examples, the PAM is 5'-NGG-3', where N represents any nucleotide. In other specific examples, the PAM may be 5'-TGC-3' or 5'-GGA-3'.
[0016] Any of the CRISPR nuclease polypeptides may include any of the arginine and / or lysine substitutions disclosed herein, any of the nickase mutations in either the HNH or RuvC nuclease domains disclosed herein, any of the mutations resulting in reduced PAM recognition stringency, or a combination thereof. For example, a CRISPR nuclease polypeptide may include (a) one or more nickase mutations in the HNH nuclease domain at positions D844, H845, and / or N868 relative to SEQ ID NO: 1 (e.g., the mutation is at position H845), and (b) one or more arginine and / or lysine substitutions at positions I857, L784, and K736 relative to SEQ ID NO: 1. For example, in some cases, a CRISPR nuclease polypeptide may contain (e.g., consist of) a nickase mutation at position H845 (e.g., H845A mutation) and an arginine and / or lysine substitution at position I857 (e.g., I857R substitution) relative to SEQ ID NO: 1. In other cases, a CRISPR nuclease polypeptide may contain, or further contain, one or more mutations that result in reduced PAM recognition stringency (e.g., positions L1117, D1144, G1227, E1228, A1332, R1345, and / or T1347 of SEQ ID NO: 1, and optionally one or more positions D61, A68, and H494 of SEQ ID NO: 1).
[0017] In some embodiments, the manipulated CRISPR nuclease polypeptides provided herein may have an amino acid sequence identical to at least 90% of SEQ ID NO: 1. In some examples, the manipulated CRISPR nuclease polypeptides may have an amino acid sequence identical to at least 95% of SEQ ID NO: 1. In specific examples, the manipulated CRISPR nuclease polypeptides may have an amino acid sequence identical to at least 98% of SEQ ID NO: 1.
[0018] Any of the manipulated CRISPR nuclease polypeptides may include any of the arginine and / or lysine substitutions disclosed herein, any of the mutations in either the HNH or RuvC nuclease domain disclosed herein, or a combination thereof.
[0019] In some embodiments, the manipulated CRISPR nuclease polypeptides disclosed herein may be fusion polypeptides, which may further comprise one or more functional fragments. In some embodiments, one or more functional fragments may comprise one or more nuclear localization signals (NLS), one or more peptide linkers, or a combination thereof. One or more NLS may be located at the N-terminus, C-terminus, or both.
[0020] In another embodiment, the disclosure features a nucleic acid comprising a nucleotide sequence encoding an engineered CRISPR nuclease polypeptide as disclosed herein. In some embodiments, the nucleic acid is an expression vector, and the nucleotide sequence encoding the engineered CRISPR nuclease polypeptide is operably ligated to a promoter. In other embodiments, the nucleic acid is messenger RNA (mRNA). Host cells comprising nucleic acids encoding an engineered CRISPR nuclease polypeptide as disclosed herein are also within the scope of the disclosure.
[0021] In addition, the disclosure features a gene editing system comprising (a) a CRISPR nuclease polypeptide or a first nucleic acid encoding a CRISPR nuclease. The CRISPR nuclease polypeptide has an amino acid sequence at least 90% identical to the reference CRISPR nuclease described as Sequence ID No. 1, and (b) a guide RNA (gRNA), or a second nucleic acid encoding a gRNA. The gRNA comprises a scaffold recognizable by the manipulated CRISPR nuclease polypeptide and a spacer sequence specific to a target sequence within a genomic site of interest. The target sequence is upstream of a protospacer motif (PAM).
[0022] In some embodiments, the CRISPR nuclease polypeptide is any of the engineered CRISPR nuclease polypeptides disclosed herein.
[0023] In some embodiments, the scaffold may comprise a nucleotide sequence that is at least 85% identical to SEQ ID NO: 2. In some examples, the scaffold comprises SEQ ID NO: 2. Alternatively, the scaffold comprises one or more deletions, one or more nucleotide substitutions, or a combination thereof, relative to SEQ ID NO: 2.
[0024] Further, the present disclosure provides a gene editing system, comprising delivering a gene editing system host cell as disclosed herein to edit a genomic site targeted by the gRNA of the gene editing system. In some embodiments, the host cell is cultured in vitro. In other embodiments, the host cell is located within a subject in need of gene editing, and the gene editing system is administered to the subject via a suitable route.
[0025] The details of one or more embodiments of the present invention are set forth in the description below. Other features or advantages of the present invention will be apparent from the following drawings and the detailed description of several embodiments, and also from the appended claims.
[0026] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure, which can be better understood by reference to the drawings in combination with the detailed description of the specific embodiments presented herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] [Figure 1] FIG. 1 is a diagram showing the gene editing efficacy of reference CRISPR nuclease SEQ ID NO: 1 against exemplary target genes AAVS1, EMX1, and VEGFA. [Figure 2A]Includes gel images showing quantification of nuclease activity. These are gel images captured using a 700 nm channel, showing in vitro cleavage of the target strand of the target DNA substrate (labeled with IR700 dye at the 5' end) by a reference CRISPR nuclease, putative HNH knockout nickasase, or putative RuvC knockout nickasase. [Figure 2B] Includes gel images showing quantification of nuclease activity. These are gel images captured using an 800 nm channel, showing in vitro cleavage of the target strand of the target DNA substrate (labeled with IR800 dye at the 5' end) by a reference CRISPR nuclease, putative HNH knockout nickasase, or putative RuvC knockout nickasase. [Figure 2C] Includes gel images showing quantification of nuclease activity. These are overlay images captured using the 700 nm and 800 nm channels in Figures 2A and 2B. [Figure 2D] Includes gel images showing quantification of nuclease activity. Quantification of percentages of cleaved target and non-target DNA produced by the reference CRISPR nuclease, putative HNH knockout nickasase, and putative RuvC knockout nickasase tested in Example 3. [Modes for carrying out the invention]
[0028] CRISPR nuclease polypeptides derived from the reference CRISPR nuclease of Sequence ID No. 1 are provided herein. Such CRISPR nuclease polypeptides may include the reference nuclease disclosed herein or variants thereof. Variant CRISPR nuclease polypeptides may contain one or more mutations (e.g., arginine and / or lysine substitutions, optionally arginine substitutions) relative to the reference CRISPR nuclease. Alternatively or in addition, variant CRISPR nuclease polypeptides may contain one or more mutations in either the RuvC nuclease domain or the HNH nuclease domain. Such mutations (i.e., nickase mutations) may reduce or eliminate the nuclease activity of either the RuvC or HNH nuclease domain, resulting in a variant exhibiting nickase activity. As used herein, the term “nickase” refers to an enzyme that cleaves one strand of double-stranded DNA at a specific recognition nucleotide sequence (e.g., a target sequence disclosed herein). Nickase can interact with one strand of a DNA double helix to produce a DNA molecule that is cleaved (also known as nicked) on one strand. In some embodiments, nickase is a variant of a CRISPR nuclease containing an inactivated HNH domain. In some embodiments, nickase is a variant of a CRISPR nuclease containing an inactivated RuvC domain. In other embodiments, a variant CRISPR nuclease polypeptide may contain one or more mutations that result in reduced PAM recognition stringency compared to SEQ ID NO: 1, compared to a counterpart CRISPR nuclease polypeptide that does not have such mutations. Any of the variant CRISPR nuclease polypeptides disclosed herein may share high sequence homology (e.g., at least 85% sequence identity) with respect to a reference CRISPR nuclease.
[0029] The variant CRISPR nuclease polypeptides provided herein are expected to have advantageous characteristics compared to reference CRISPR nucleases, such as exhibiting nickas activity and / or higher nuclease activity. Therefore, the variant CRISPR nuclease polypeptides disclosed herein are expected to exhibit improved gene editing, such as higher efficacy and precision in gene editing involving strand replacement, compared to reference CRISPR nucleases.
[0030] Alternatively, or in addition, a CRISPR nuclease polypeptide may be a fusion polypeptide comprising a CRISPR nuclease (e.g., SEQ ID NO: 1 or a variant thereof) and one or more additional functional fragments, e.g., those described herein (e.g., nuclease localization signals or NLS). In addition to the above-mentioned advantageous features, the fusion polypeptide may have additional functionalities that can be attributed to its fusion partner.
[0031] Accordingly, this disclosure provides CRISPR nuclease polypeptides derived from the reference CRISPR nuclease of SEQ ID NO: 1 (SEQ ID NO: 1 or its variants), a gene editing system comprising them, and a gene editing method using them.
[0032] I. CRISPR nuclease polypeptide As used herein, the term “CRISPR nuclease” refers to an RNA-guided effector capable of binding to nucleic acids and introducing single-strand or double-strand breaks. CRISPR nucleases typically comprise multiple functional domains, such as nuclease domains (e.g., RuvC and HNH), bridge helix (BH) domains, nucleic acid recognition (REC) domains, phosphate-locked loop (PLL) domains, wedge domains (WED), PAM interaction domains (PID), or combinations thereof. As used herein, the term “domain” refers to a characteristic functional and / or structural unit of a polypeptide. In some cases, functional domains may be linear. In other cases, functional domains may be discontinuous and stereostructural. In some embodiments, a domain may comprise a conserved amino acid sequence across different CRISPR nucleases.
[0033] The reference CRISPR nuclease of SEQ ID NO: 1 (see Table 1 below) is a CRISPR nuclease containing both a RuvC nuclease domain (located at residues 1-59, 722-771, and 927-1101 of SEQ ID NO: 1) and an HNH domain (located at residues 772-926 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate DNA strand cleavage to a 5'-NRG-3' PAM motif, where N represents any nucleotide and R represents A or G. In some embodiments, the PAM motif is 5'-NGG-3'. Positions D10, E763, and D991 are considered active sites in the RuvC domain, and positions D844, H845, and N868 are considered active sites in the HNH domain. In addition to the nuclease domain, the reference CRISPR nuclease of SEQ ID NO: 1 also contains a BH domain (residues 60-93 of SEQ ID NO: 1), a REC domain (residues 94-721 of SEQ ID NO: 1), a PLL domain (residues 1102-1148 of SEQ ID NO: 1), a WED domain (residues 1149-1208 of SEQ ID NO: 1), and a PID domain (residues 1209-1378 of SEQ ID NO: 1).
[0034] (A) Variants of CRISPR nuclease polypeptides The variant CRISPR nuclease polypeptides provided herein are derived from the reference CRISPR nuclease of SEQ ID NO: 1 by introducing one or more mutations into the reference CRISPR nuclease to modulate (e.g., enhance or reduce) one or more activities of the nuclease. As used herein, the term “variant CRISPR nuclease polypeptide” refers to a CRISPR nuclease polypeptide that, compared to the reference CRISPR nuclease (SEQ ID NO: 1), contains changes, e.g., substitutions, insertions, deletions, and / or fusions at one or more residue positions.
[0035] A variant CRISPR nuclease polypeptide may contain one or more mutations (e.g., arginine substitutions) compared to the reference CRISPR nuclease. Alternatively or in addition, a variant CRISPR nuclease polypeptide may contain one or more mutations in either the RuvC nuclease domain or the HNH nuclease domain. Such mutations may reduce or eliminate the nuclease activity of either the RuvC or HNH nuclease domain, resulting in a variant exhibiting nickase activity. A variant CRISPR nuclease polypeptide may share high sequence homology (e.g., at least 85% sequence identity) with respect to the reference CRISPR nuclease.
[0036] As used herein, the term “nickase” refers to an enzyme that cleaves one strand of a double-stranded DNA at a specific recognition nucleotide sequence (e.g., a target sequence disclosed herein). A nickase may interact with one strand of a DNA double helix to produce a DNA molecule that is cleaved (also known as nicked) on one strand. In some embodiments, the nickase is a variant of a CRISPR nuclease containing an inactivated HNH domain. In some embodiments, the nickase is a variant of a CRISPR nuclease containing an inactivated RuvC domain. In other embodiments, the variant CRISPR nuclease polypeptide may contain one or more mutations that result in reduced PAM recognition stringency compared to SEQ ID NO: 1, compared to a counterpart CRISPR nuclease polypeptide that does not have such mutations. The variant CRISPR nuclease polypeptide may share high sequence homology (e.g., at least 85% sequence identity) with respect to a reference CRISPR nuclease.
[0037] In some embodiments, the variant CRISPR nuclease polypeptides provided herein comprise one or more mutations in either the RuvC nuclease domain or the HNH nuclease domain (e.g., the HNH nuclease domain) to reduce or eliminate nuclease activity compared to reference CRISPR nuclease SEQ ID NO: 1, and / or comprise one or more arginine and / or lysine substitutions to improve nuclease characteristics suitable for use in gene editing.
[0038] The variant CRISPR nuclease polypeptides provided herein are expected to exhibit one or more regulated activities (e.g., enhanced or reduced) relative to a reference CRISPR nuclease. As used herein, the term “activity” refers to biological activity. In some embodiments, activity includes enzymatic activity, e.g., catalytic capacity of an effector. For example, activity may include nuclease activity, such as double-stranded nuclease activity. In some embodiments, activity includes nickasing activity. For example, a variant CRISPR nuclease polypeptide may substantially cleave only one strand of a target DNA double-stranded DNA.
[0039] In some embodiments, the activity includes binding activity, e.g., binding of an effector (e.g., a CRISPR nuclease) to an RNA guide and / or target nucleic acid. In some examples, the variant CRISPR nuclease polypeptides disclosed herein have enhanced binding to a congeneral guide RNA (gRNA) compared to a reference CRISPR nuclease, for example, having binding activity that is at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2x, 2x, 5x, 10x, or more greater than that of the reference CRISPR nuclease. A congeneral gRNA refers to a gRNA that has a scaffold recognizable by a CRISPR nuclease.
[0040] In some cases, the variant CRISPR nuclease polypeptides disclosed herein have enhanced enzymatic activity relative to a reference CRISPR nuclease, for example, having at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2x, 2x, 5x, 10x, or more greater enzymatic activity than that of the reference CRISPR nuclease. In other cases, the variant CRISPR nuclease polypeptides disclosed herein have reduced enzymatic activity relative to a reference CRISPR nuclease (e.g., enzymatic activity for cleaving both strands of a target DNA double helix), for example, having at least 20%, 30%, 40%, 50%, 60%, or 70% lower enzymatic activity than that of the reference CRISPR nuclease. In some cases, the reduced enzymatic activity is achieved by reducing the nuclease activity of the RuvC domain. In other cases, reduced enzyme activity is achieved by decreasing or lowering the nuclease activity of the HNH domain.
[0041] In some cases, the variant CRISPR nuclease polypeptides disclosed herein have enhanced indel activity compared to a reference CRISPR nuclease. As used herein, the term “indel activity” refers to the ability of a CRISPR nuclease to introduce indels (insertions / deletions) into a sequence (e.g., a genomic target). For example, in some embodiments, a CRISPR nuclease introduces a double-strand break into a sequence (e.g., a genomic target within a cell), and an indel is created via a DNA repair mechanism.
[0042] In some embodiments, the variant CRISPR nuclease polypeptides provided herein share high sequence homology with a reference CRISPR nuclease. For example, a variant CRISPR nuclease polypeptide may contain at least 70% (e.g., at least 80%, 85%, 90%, 95%, or more) identical amino acid sequences to SEQ ID NO: 1. In some cases, a variant CRISPR nuclease polypeptide may contain at least 90% identical amino acid sequences to SEQ ID NO: 1. In some cases, a variant CRISPR nuclease polypeptide may contain at least 95% identical amino acid sequences to SEQ ID NO: 1. In other cases, a variant CRISPR nuclease polypeptide may contain at least 97% (e.g., 98%, 99%, 99.5%, or more) identical amino acid sequences to SEQ ID NO: 1.
[0043] The "percent identity" (also known as sequence identity) of two nucleic acids or two amino acid sequences is determined using the algorithm of Karlin and Altschul Proc.Natl.Acad.Sci.USA 87:2264-68, 1990, as modified in Karlin and Altschul Proc.Natl.Acad.Sci.USA 90:5873-77, 1993. Such an algorithm is incorporated into the NBLAST and XBLAST programs (version 2.0) of Altschul, et al. J.Mol.Biol.215:403-10, 1990. BLAST nucleotide search can be performed using the NBLAST program, score=100, word length=12 to obtain nucleotide sequences homologous to the nucleic acid molecule of the present invention. BLAST protein search can be performed using the XBLAST program, score=50, word length=3 to obtain amino acid sequences homologous to the protein molecule of the present invention. If a gap exists between two sequences, gap BLAST can be used as described in Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997. When using the BLAST and Gapped BLAST programs, the initial settings parameters for each program (e.g., XBLAST and NBLAST) can be used.
[0044] The variant CRISPR nuclease polypeptides provided herein may contain one or more modifications to the reference CRISPR nuclease of SEQ ID NO: 1, such as one or more amino acid residue substitutions, one or more deletions, one or more insertions, fusions, or combinations thereof. In some cases, the modifications may be introduced within the BH domain, PLL domain, WED domain, PID domain, or combinations thereof.
[0045] In some embodiments, the variant CRISPR nuclease polypeptides provided herein may contain one or more arginine substitutions, one or more lysine substitutions, or a combination thereof, relative to SEQ ID NO: 1. “Arginine substitution” or “lysine substitution” refers to replacing a non-arginine or non-lysine residue in SEQ ID NO: 1 with an arginine or lysine residue. In some examples, the variant CRISPR nuclease polypeptide may contain up to 20 arginine and / or lysine substitutions, e.g., up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine and / or lysine substitutions. In specific examples, the variant CRISPR nuclease polypeptide may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine and / or lysine substitutions.
[0046] In some cases, one or more of the substituted arginine residues may be replaced by conserved amino acid residues, such as lysine or histidine. In some embodiments, the variant CRISPR nuclease polypeptides provided herein may comprise one or more arginine substitutions, one or more lysine substitutions, or a combination thereof.
[0047] In some cases, arginine and / or lysine substitutions (e.g., arginine substitutions) can be located in the BH domain, PLL domain, WED domain, PID domain, or a combination thereof. In some cases, variant CRISPR nuclease polypeptides have one or more arginine and / or lysine substitutions at the positions I79, E331, Y348, S473, F501, I581, D720, A730, G731, Q741, V752, M753, Q809, Q840, Q849, S872, S898, E982, K918, D985, Y986, Y1015, E1037, K1091, S1094, It may contain one or more of the following: P1096, N1099, T1104, E1105, I1106, T1108, L1117, K1131, I1147, E1179, M1205, P1208, E1214, A1226, Q1230, A1236, P1238, F1241, L1281, D1284, F1285, A1292, N1295, K1298, G1329, A1333, K1344, S1348, Q1360, and I1370. In specific examples, arginine and / or lysine substitutions (e.g., combinations of arginine substitutions) can be located at one or more of the following positions in SEQ ID NO: 1: K736, L784, N813, Q812, I857, and A919.
[0048] In some specific examples, the variant CRISPR nuclease polypeptides for SEQ ID NO: I79R, E331R, Y348R, S473R, F501R, I581R, D720R, A730R, G731R, Q741R, V752R, M753R, Q809R, Q840R, Q849R, S872R, S898R, E982R, K918R, D985R, Y986R, Y1015R, E1037R, K1091R, S1094R, P1096R, N1099R, It may contain one or more arginine substitutions of T1104R, E1105R, I1106R, T1108R, L1117R, K1131R, I1147R, E1179R, M1205R, P1208, E1214R, A1226R, Q1230R, A1236R, P1238R, F1241R, L1281R, D1284R, F1285R, A1292R, N1295R, K1298R, G1329R, A1333R, K1344R, S1348R, Q1360R, and I1370R.
[0049] In some cases, variant CRISPR nuclease polypeptides may contain one or more arginine substitutions at one or more of the positions described above. Examples include I857R, N813R, L784R, A919R, Q812R, or combinations thereof. In other cases, variant CRISPR nuclease polypeptides may contain one or more lysine substitutions at one or more of the positions described above. Examples include I857K, N813K, L784K, A919K, Q812K, or combinations thereof.
[0050] In other examples, variant CRISPR nuclease polypeptides have arginine and lysine substitution combinations, such as I79, E331, Y348, S473, F501, I581, D720, A730, G731, Q741, V752, M753, Q809, Q840, Q849, S872, S898, E982, K918, D985, Y986, Y1015, E1037, K1091, S1094, P1 May be contained in 096, N1099, T1104, E1105, I1106, T1108, L1117, K1131, I1147, E1179, M1205, P1208E1214, A1226, Q1230, A1236, P1238, F1241, L1281, D1284, F1285, A1292, N1295, K1298, G1329, A1333, K1344, S1348, Q1360, and / or I1370. In specific examples, variant CRISPR nuclease polypeptides may contain combinations of arginine and lysine substitutions (e.g., combinations of arginine substitutions) in SEQ ID NO: I857, K736, L784, N813, Q812, I857, and / or A919.
[0051] In some cases, variant CRISPR nuclease polypeptides may contain one or more arginine substitutions at one or more of the aforementioned positions. Examples include I857R, N813R, L784R, K736R, A919R, Q812R, or combinations thereof.
[0052] In specific examples, the manipulated CRISPR nuclease polypeptide may contain arginine and / or lysine substitutions at the following positions relative to SEQ ID NO: 1: (a) I857, L784, and K736; (b) I857, A919, and K736; (c) I857, N813, and L784; (d) I857, L784, and A919; (e) I857, N813, and K736; (f) I857 and N813; (g) L784, A919, and K736; (h) I857 and L784; or (i) I857 and A919. In some cases, the manipulated CRISPR nuclease polypeptide may contain arginine substitutions at any of the positional combinations in SEQ ID NO: 1. In one specific example, the manipulated CRISPR nuclease polypeptide may contain the arginine substitutions I857R, L784R, and K736R relative to SEQ ID NO: 1. Other examples of arginine and / or lysine substitutions can be seen in Table 7 below.
[0053] Alternatively, or in addition, the variant CRISPR nuclease polypeptides provided herein may contain one or more mutations within either the RuvC or HNH nuclease domain to reduce or eliminate the nuclease activity of a target domain, thereby producing a variant with nickase activity. Such mutations (nickase activity) may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation within either the RuvC or HNH nuclease domain is an amino acid substitution, and the substituted amino acid residue of the amino acid substitution is not a conserved substitution of the native amino acid residue at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue except K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue except R. Bases for conserved amino acid residue substitutions are provided herein.
[0054] In some cases, one or more nickase mutations may occur within the HNH nuclease domain, for example, in D844, H845, and / or N868 of SEQ ID NO: 1. In some cases, the mutation may be an amino acid residue substitution, and the native amino acid residue in SEQ ID NO: 1 may be replaced by an amino acid residue of a different type than the native residue. For example, a positively charged residue may be replaced by an uncharged amino acid residue, or vice versa. In some cases, the amino acid residue substitution in D844 may be D844G, D844A, D844L, or D844S. In one specific example, the mutation may be D844A. In another example, the amino acid residue substitution in H845 may be H845G, H845A, H845I, H845L, H845M, H845V, or H845S. In one example, the manipulated CRISPR nuclease is a nickase mutant (e.g., consisting of) a residue substitution at position H845 (e.g., H845A) relative to SEQ ID NO: 1. In one specific example, the mutation at position H845 could be H845A. Alternatively or in addition, the amino acid residue substitution could be at position N868, e.g., N868G, N868A, N868L, or N868S. In one example, the mutation at position N868 is N868A.
[0055] In some cases, one or more mutations may be associated with the RuvC nuclease domain in SEQ ID NO: 1, for example, at positions D10, E763, D991, or a combination thereof (e.g., positions E763 and / or D991). In some cases, the mutations may be amino acid residue substitutions, and native amino acid residues in SEQ ID NO: 1 may be replaced by amino acid residues of a different type than the native residues. For example, a positively charged residue may be replaced by an uncharged amino acid residue, or vice versa. In some cases, the amino acid residue substitution at D10 may be D10G, D10A, D10L, or D10S. In some cases, the amino acid residue substitution at E763 may be E763G, E763A, E765L, or E763S. Alternatively or in addition, the amino acid residue substitution at D991 may be D991G, D991A, D991L, or D991S.
[0056] In some examples, the variant CRISPR nuclease polypeptides provided herein may be nickase variants, which include one or more mutations in a single nuclease domain (e.g., an HNH nuclease domain at position H845, e.g., H845A). Exemplary nickase variants are provided in Table 4 below. Such nickase variants may further include one or more arginine or lysine substitutions (e.g., arginine substitutions) to enhance certain features, such as indel activity. For example, exemplary arginine and / or lysine substitutions at one or more of the positions I857R, K736, L784, N813, Q812, I857, and A919 of SEQ ID NO: 1 (e.g., arginine substitutions at positions I857, L784, and K736) are provided herein. In some examples, CRISPR nuclease polypeptides may contain (or consist of) nickase mutations at position H845 (e.g., H845A mutation) and arginine and / or lysine substitutions at position I857 (e.g., I857R substitution).
[0057] In some cases, the variant CRISPR nuclease polypeptides disclosed herein exhibit enhanced double-stranded nuclease activity. Examples of such CRISPR nuclease polypeptides are provided in Table 8 below, each of which is within the scope of this disclosure.
[0058] In some embodiments, the manipulated CRISPR nuclease polypeptides disclosed herein may contain one or more mutations that reduce PAM recognition stringency compared to a reference CRISPR nuclease. In some examples, one or more mutations that reduce PAM recognition stringency may be located at positions L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345, and / or T1347 of SEQ ID NO: 1. In some examples, one or more mutations may include (i) one or more arginine and / or lysine substitutions, optionally arginine substitutions, at positions L1117, G1227, S1327, A1332, and / or T1347 of SEQ ID NO: 1; (ii) one or more amino acid substitutions at positions D1144, S1145, E1228, R1343, and / or R1345 of SEQ ID NO: 1; or (iii) a combination of (i) and (ii).
[0059] In some cases, such variant CRISPR nuclease polypeptides exhibit mutations (e.g., amino acid residue substitutions) at positions D61 (e.g., D61R or D61K), L1117 (e.g., L1117R or L1117K), D1144 (e.g., D1144V, D1144A, D1144G, or D1144S), S1145 (e.g., S1145W, S1145Y, or S1145F), G1227 (e.g., G1227K or G1227R), E 1228 (e.g., E1228F, E1228Y, or E1228W), A1327 (e.g., A1327R, or A1332K), A1332 (e.g., A1332R, or A1332K), R1345 (e.g., R1345Q, or R1345N), R1345 (e.g., R1345V, R1345A, R1345G, or R1345S, or R1345Q, R1345N), T1347 (e.g., T1347R, or T1347K), or combinations thereof may be contained.
[0060] In a specific example, one or more amino acid substitutions in (ii) may include D1144L, S1145W, E1228Q, R1343P, R1345V, and / or R1345Q for SEQ ID NO: 1. Alternatively, the substituted amino acid residues at one or more of positions D1114, S1145, E1228, R1343, and R1345 may be conserved substitutions of L, W, Q, P, V, and Q, respectively. For example, a substitution at position D1114 may be D1114M, D1114I, or D1114V. A substitution at position S1145 may be S1145F or S1145Y. A substitution at position E1228 may be E1228N. Substitutions at position R1345 may be R1345M, R1345I, R1345L, or R1345N. Specific examples of such variant CRISPR nuclease polypeptides are provided in Table 9, each of which is within the scope of this disclosure.
[0061] In some specific examples, the manipulated CRISPR nuclease polypeptide contains (or consists of) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, and T1347R relative to SEQ ID NO: 1. Such a CRISPR nuclease polypeptide has the amino acid sequence described in SEQ ID NO: 130.
[0062] In some examples, the manipulated CRISPR nuclease polypeptides disclosed herein (e.g., those having mutations at positions L1117, D1144, G1227, E1228, A1332, R1343, R1345, and / or T1347) for SEQ ID NO: 130, have mutations at positions D61, A68, H494, L64, S410, T67, Q849, G1110, F501, T659, L784, Y516, G55, E1037, N57, D720, A919, A1294, Q812, N700, H657, T73, Q899, I857, K751, D3 27, I581, D462, E331, A589, D471, I699, N1295, T470, I1147, E130, S473, A353, K40, K334, A60, S1348, K3 67, A1118, K31, Q349, K341, Q83, K585, Q840, G660, K527, G727, Y42, L1281, L122, Q123, T1108, E41, K113 1, K30, S872, I1206, D1132, K460, L80, E459, K1182, M696, K918, K126, N721, Q809, K1091, K736, K783, N4 This may further include one or more single arginine substitutions (e.g., a single arginine substitution) in 98, K723, H1119, F463, L594, D472, K744, E365, G595, K45, Y348, K964, S1181, N813, D407, S839, Y658, E586, G754, A730, Y1015, D903, A1333, S461, or H1359.
[0063] In some specific examples, the manipulated CRISPR nuclease polypeptide contains L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and A68R relative to SEQ ID NO: 1. In other specific examples, the manipulated CRISPR nuclease polypeptide contains L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R relative to SEQ ID NO: 1. In yet another specific example, the manipulated CRISPR nuclease polypeptide contains L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and H494R relative to SEQ ID NO: 1.
[0064] In some cases, variant CRISPR nuclease polypeptides exhibiting less stringent recognition of PAM sequences may recognize 5'-NGN-3', 5'-NRN-3', or 5'-NYN-3' PAM sequences, where N represents any nucleotide, R represents A or G, and Y represents C or T.
[0065] In some cases, the variant CRISPR nuclease polypeptide may contain one or more of the mutations disclosed herein (e.g., one or more arginine and / or lysine substitutions, one or more nickase mutations, and / or one or more mutations resulting in less PAM recognition stringency).
[0066] Any of the variant CRISPR nuclease polypeptides disclosed herein may share at least 90% (e.g., 95%, 97%, 98%, 99%, 99.5%, or more) sequence identity with SEQ ID NO: 1.
[0067] In some cases, variant CRISPR nuclease polypeptides may contain one or more conserved amino acid residue substitutions, in addition to mutations in the HNH or RuvC nuclease domain, arginine / lysine substitutions, and / or mutations that result in reduced PAM recognition stringency.
[0068] As used herein, “conservative amino acid substitution” refers to an amino acid substitution that does not alter the relative charge or size properties of the protein in which the substitution is made. Variants can be prepared by methods for altering polypeptide sequences known to those skilled in the art, for example, as found in references summarizing such methods, e.g., Molecular Cloning: A Laboratory Manual, J. Sambrook, et al., eds., Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989, or Current Protocols in Molecular Biology, FMAusubel, et al., eds., John Wiley & Sons, Inc., New York. Conservative amino acid substitutions include substitutions made with amino acids in the following groups: (a) M, I, L, V; (b) F, Y, W; (c) K, R, H; (d) A, G; (e) S, T; (f) Q, N; and (g) E, D.
[0069] Exemplary CRISPR nuclease polypeptides for use in the gene editing systems provided herein are disclosed in Tables 1, 4, 8, and 9 below, each of which is within the scope of this disclosure.
[0070] In some embodiments, a variant CRISPR nuclease polypeptide may be a fusion polypeptide comprising a CRISPR nuclease and one or more additional functional moieties. As used herein, the terms “fusion” and “fused” refer to the linkage of at least two nucleotides or protein molecules. For example, “fusion” and “fused” may refer to the linkage of at least two polypeptide domains that are naturally encoded by distinct genes. The fusion may be an N-terminal fusion, a C-terminal fusion, or an intramolecular fusion. In some embodiments, the domains are transcribed and translated to produce a single polypeptide. In some cases, the CRISPR nuclease moiety in the fusion polypeptide may be the reference CRISPR nuclease of SEQ ID NO: 1. Alternatively, the CRISPR nuclease moiety in the fusion polypeptide may be a variant CRISPR nuclease derived from SEQ ID NO: 1, such as those disclosed herein.
[0071] Exemplary additional functional segments that may be included within the fusion polypeptide include peptide tags, fluorescent proteins, base editing domains, DNA methylation domains, histone residue modification domains, localization factors, transcription modifiers, photo-gate regulators, chemoinducible factors, chromatin visualization factors, or combinations thereof.
[0072] In some embodiments, additional functional components may include a nuclear localization signal (NLS), a nuclear export signal (NES), or a combination thereof. In some examples, the fusion polypeptide may contain an NLS, which may be located at either the N-terminus or the C-terminus. In specific examples, the fusion polypeptide may contain a first NLS located at the N-terminus and a second NLS located at the C-terminus. The first and second NLS fragments may be identical. Alternatively, the two NLS fragments may be different. In some embodiments, the fusion polypeptide may contain the NLS near the N-terminus and / or C-terminus (e.g., within about one, two, three, four, or five amino acids from the first or last amino acid of the CRISPR nuclease). In some embodiments, the fusion polypeptide may contain the NLS within the mobile loop of the CRISPR nuclease.
[0073] In some embodiments, the additional functional component may be a mobile peptide linker, such as an XTEN peptide linker or a G / S-rich peptide linker. An example of such a peptide linker is provided in Example 1 below.
[0074] In some embodiments, the gene editing systems provided herein may comprise a CRISPR nuclease polypeptide, which may form a ribonucleoprotein (RNP) complex with a congeneral guide RNA. As used herein, the term “complex” refers to a grouping of two or more molecules. In some embodiments, the complex comprises polypeptides and nucleic acid molecules that interact with each other (e.g., bound, in contact, or adhered).
[0075] In other embodiments, the gene editing systems provided herein may include nucleic acids encoding CRISPR nuclease polypeptides. In some examples, the nucleotide sequences encoding CRISPR nuclease polypeptides described herein can be codon-optimized for use in specific host cells or organisms. For example, nucleic acids can be codon-optimized for any non-human eukaryote, which includes mice, rats, rabbits, dogs, livestock, or non-human primates. Codon usage tables are readily available, for example, in the "Codon Usage Database" available on the worldwide website kazusa.orjp / codon / , and these tables can be adapted in several formats. See Nakamura et al. Nucl. Acids Res. 28:292 (2000), which is incorporated herein by reference in its entirety. Computer algorithms for codon-optimizing specific sequences for expression in specific host cells, e.g., Gene Forge (Aptagen, Jacobus, PA), are also available. In some examples, the nucleic acids encoding the CRISPR nuclease polypeptides disclosed herein may be mRNA molecules that can be codon-optimized. Exemplary codon-optimized nucleotide sequences encoding exemplary CRISPR nuclease polypeptides can be found in Tables 1 and 4 below, any of which are within the scope of this disclosure.
[0076] In some cases, gene editing systems may include vectors encoding CRISPR nuclease polypeptides (e.g., viral vectors, such as AAV vectors, AdV vectors, or retroviral vectors).
[0077] B. Preparation of variant CRISPR nuclease polypeptides The variant CRISPR nuclease polypeptides disclosed herein can be prepared by conventional methods or by methods disclosed herein. For example, variant CRISPR nuclease polypeptides can be prepared by culturing host cells capable of producing nuclease polypeptides, such as bacterial or mammalian cells, isolating the nuclease polypeptides thus produced, and optionally purifying the nuclease polypeptides. The variant CRISPR nuclease polypeptides thus prepared can form complexes with gRNA.
[0078] Variant CRISPR nuclease polypeptides can also be prepared by in vitro coupled transcription-translation systems and, optionally, by complexing with gRNA. The bacteria that can be used for the preparation of variant CRISPR nuclease polypeptides are not particularly limited, as long as they are capable of producing variant CRISPR nuclease polypeptides. Some non-limiting examples of bacteria include E. coli cells, as described herein.
[0079] Unless otherwise noted, all compositions, complexes, and polypeptides provided herein are prepared with reference to their activity levels, excluding impurities that may be present in commercially available sources, such as residual solvents or by-products. Enzymatic component weights are based on total active protein. Unless otherwise indicated, all percentages and ratios are calculated by weight. Unless otherwise indicated, all percentages and ratios are calculated based on the total composition. In exemplary compositions, enzyme levels are expressed by the weight of pure enzyme in the total composition, and components are expressed by the weight of the total composition unless otherwise specified.
[0080] (i) Vector This disclosure provides vectors for expressing variant CRISPR nuclease polypeptides. In some embodiments, the vectors disclosed herein include a nucleotide sequence encoding a variant CRISPR nuclease polypeptide. In some embodiments, the vectors include a Pol II promoter or a Pol III promoter.
[0081] The expression of native or synthetic polynucleotides is typically achieved by operably ligating a polynucleotide encoding a variant CRISPR nuclease polypeptide to a promoter and incorporating the construct into an expression vector. The expression vector is not particularly limited as long as it contains a polynucleotide encoding a variant CRISPR nuclease polypeptide and may be suitable for replication and incorporation in eukaryotic cells.
[0082] A typical expression vector contains transcription and translation terminators, start sequences, and promoters useful for the expression of a desired polynucleotide. For example, plasmid vectors containing recognition sequences for RNA polymerase (e.g., pSP64, pBluescript) can be used. Vectors containing retroviruses, such as those derived from lentiviruses, are suitable tools for achieving long-term gene transfer, as they allow for the long-term stable integration of the transgene and its proliferation in daughter cells. Examples of vectors include expression vectors, replication vectors, probe-generating vectors, and sequencing vectors. Expression vectors can be supplied to cells in the form of viral vectors.
[0083] Viral vector technology is well-known in the art and is described in various virology and molecular biology manuals. Useful viruses as vectors include, but are not limited to, phage viruses, retroviruses, adenoviruses, adeno-associated viruses, herpesviruses, and lentiviruses. Generally, a suitable vector contains a functional replication origin, promoter sequence, convenient restriction endonuclease site, and one or more selectable markers in at least one organism.
[0084] The type of vector is not particularly limited, and any vector that can be expressed in host cells can be appropriately selected. More specifically, depending on the type of host cell, a promoter sequence is appropriately selected to ensure the expression of a polypeptide(s) from a polynucleotide, and this promoter sequence and polynucleotide are inserted into one of various plasmids or other materials for preparing the expression vector.
[0085] Additional promoter elements, such as enhancing sequences, regulate the frequency of transcription initiation. Typically, these are located 30–110 bp upstream of the start site, although some promoters have recently been shown to also contain functional elements downstream of the start site. Depending on the promoter, individual elements appear to be able to activate transcription either cooperatively or independently.
[0086] Furthermore, this disclosure should not be limited to the use of constitutive promoters. Inducible promoters are also intended as part of this disclosure. The use of inducible promoters provides a molecular switch that can turn on the expression of a polynucleotide sequence that is operatively ligated when such expression is desired, or turn off the expression when expression is not desired. Examples of inducible promoters include, but are not limited to, metallothione promoters, glucocorticoid promoters, progesterone promoters, and tetracycline promoters.
[0087] The introduced expression vector may also contain either or both a selectable marker gene or a reporter gene to facilitate the identification and selection of expressing cells from a population of cells intended for transfection or infection via the viral vector. In other embodiments, the selectable marker may be supported on a separate DNA fragment and used in a co-transfection procedure. Both the selectable marker and reporter gene may be flanked by appropriate transcriptional regulatory sequences to enable expression in host cells. Examples of such markers include dihydrofolate reductase genes and neomycin resistance genes for eukaryotic cell culture, as well as tetracycline resistance genes and ampicillin resistance genes for E. coli and other bacterial cultures. The use of such selectable markers allows for confirmation that the polynucleotide encoding the polypeptide(s) of the present invention has been transferred into host cells and subsequently expressed without failure.
[0088] The preparation methods using recombinant expression vectors are not particularly limited and include methods using plasmids, phages, or cosmids.
[0089] (ii) Method of expression This disclosure includes a method for protein expression, comprising translating a variant CRISPR nuclease polypeptide as described herein.
[0090] In some embodiments, the host cells described herein are used to express the variant CRISPR nuclease polypeptide. The host cells are not particularly limited, and a variety of known cells can be preferably used. Specific examples of host cells include bacteria, e.g., E. coli, yeast (budding yeast Saccharomyces cerevisiae and fission yeast Schizosaccharomyces pombe), nematodes (Caenorhabditis elegans), Xenopus laevis oocytes, and animal cells (e.g., CHO cells, COS cells, and HEK293 cells). The method for transferring the expression vector into the host cells, i.e., the transformation method, is not particularly limited, and known methods, e.g., electroporation, calcium phosphate method, liposome method, and DEAE dextran method, can be used.
[0091] After the host has been transformed with the expression vector, the host cells can be cultured, cultivated, or propagated for the production of the variant CRISPR nuclease polypeptide. After expression, the host cells can be collected, and the variant CRISPR nuclease polypeptide can be purified from the culture or other source according to conventional methods (e.g., filtration, centrifugation, cell disruption, gel filtration chromatography, ion exchange chromatography, etc.).
[0092] Various methods can be used to determine the level of production of mature variant CRISPR nuclease polypeptides in host cells. Such methods include, but are not limited to, the use of protein-specific polyclonal or monoclonal antibodies, or labeled tags as described elsewhere herein. Exemplary methods include, but are not limited to, enzyme-linked immunosorbent assays (ELISA), radioimmunoassays (RIA), fluorescence immunoassays (FIA), and fluorescence-activated cell sorting (FACS). These and other assays are well known in the art (see, for example, Maddox et al., J.Exp.Med.158:1211
[1983] ).
[0093] This disclosure provides a method for in vivo expression of a variant CRISPR nuclease polypeptide (and optionally, gRNA in a gene editing system disclosed herein). Such a method may include providing a polyribonucleotide encoding the variant CRISPR nuclease polypeptide to a host cell in a subject (e.g., a human subject), the polyribonucleotide encoding the variant CRISPR nuclease polypeptide that causes the cell to express the variant CRISPR nuclease polypeptide.
[0094] II. Gene Editing Systems In some embodiments, the disclosure provides a gene editing system having enhanced gene editing efficiency. The gene editing system comprises one of the variant CRISPR nuclease polypeptides disclosed herein, or a nucleus encoding a CRISPR nuclease, and one or more guide RNAs (gRNAs) or nucleic acids (may be multiple) encoding gRNAs.
[0095] A. CRISPR nuclease In some embodiments, the gene editing systems disclosed herein include variant CRISPR nuclease polypeptides provided herein, for example, variant CRISPR nuclease polypeptides comprising one or more arginine or lysine substitutions, one or more mutations in one of the nuclease domains, for example, in the HNH nuclease domain, or a combination thereof. See above disclosure. Such protein components may form complexes with gRNAs in the same gene editing system.
[0096] Alternatively, the gene editing system includes a nucleic acid encoding a CRISPR nuclease polypeptide. In some cases, the nucleic acid may be an expression vector (e.g., a viral vector) for producing the encoded nuclease polypeptide in a host cell. In some cases, the expression vector may further include coding sequences for producing one or more gRNAs of the gene editing system. In other cases, the nucleic acid encoding the CRISPR nuclease polypeptide may be a messenger RNA (mRNA) molecule. In some cases, the mRNA molecule may further include coding sequences for the gRNAs of the gene editing system.
[0097] B. Guide RNA The gene editing systems disclosed herein further comprise one or more gRNAs or encoding nucleic acids. As used herein, the terms “RNA guide,” “RNA guide sequence,” or “guide RNA (gRNA)” refer to an RNA molecule or modified RNA molecule that facilitates the CRISPR nuclease described herein to target a genomic site of interest. For example, an RNA guide may be a molecule comprising a spacer sequence and a scaffold sequence. The spacer sequence recognizes (e.g., binds to) a site on a non-PAM strand that is complementary to the target sequence on the PAM strand, for example, a site on a non-PAM strand designed to be complementary to a specific nucleic acid sequence. The scaffold sequence contains a nuclease-binding sequence for binding to the CRISPR nuclease. In some embodiments, the scaffold is an RNA sequence.
[0098] In some cases, the gRNAs disclosed herein may further include a linker sequence, a 5' end and / or a 3' end protection fragment, or a combination thereof.
[0099] (i) Spacer array As used herein, the terms “spacer” and “spacer sequence” (also known as DNA-binding sequence) refer to a portion (DNA sequence) of an RNA guide that is the RNA equivalent of the target sequence. A spacer contains a sequence that can bind to the non-PAM strand via base pairing at a site complementary to the target sequence (which is located on the PAM strand). Such spacers are also known to be specific to the target sequence. In some cases, a spacer may be at least 75% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%) identical to the target sequence, except for the RNA-DNA sequence difference. In some cases, a spacer may be 100% identical to the target sequence, except for the RNA-DNA sequence difference.
[0100] The gene editing systems disclosed herein comprise one or more gRNAs, each comprising a spacer sequence for targeting a genomic site of interest and a scaffold recognizable by a variant CRISPR nuclease polypeptide contained in the gene editing system. The target sequence can be adjacent to a 5'-NDR-3' protospacer-adjacent motif (PAM), where N represents any nucleotide, D represents A, G, or T, and R represents G or A. In some cases, the PAM is 5'-NRG-3', where N represents any nucleotide and R represents A or G. In some cases, the PAM is 5'-NRR-3', where N represents any nucleotide and R represents A or G. Alternatively, the PAM may be 5'-NGN-3', where N represents any nucleotide (A, C, G, or U). In some specific examples, the PAM motif is 5'-NGG-3', where N represents any nucleotide. The PAM motif is located on the 3' end of the target sequence. As used herein, the terms “protospacer-adjacent motif” or “PAM sequence” refer to a DNA sequence adjacent to the target sequence. In some embodiments, the PAM sequence is required for CRISPR nuclease binding and / or indel activity. In a double-stranded DNA molecule, the strand containing the PAM motif is referred to as the “PAM strand,” and the complementary strand is referred to as the “non-PAM strand.” The gRNA binds to a site on the non-PAM strand that is complementary to the target sequence disclosed herein, and the PAM sequence described herein is present on the PAM strand. The PAM motif may be located upstream of the target sequence.
[0101] As used herein, the term “adjacent to” means that a nucleotide or amino acid sequence is in close proximity to another nucleotide or amino acid sequence. In some embodiments, if there are no nucleotides separating the two sequences, the nucleotide sequences are adjacent to (i.e., directly adjacent to) another nucleotide sequence. In some embodiments, a nucleotide sequence is adjacent to another nucleotide sequence if a small number of nucleotides (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides) separate the two sequences. In some embodiments, a first sequence is adjacent to a second sequence if the two sequences are separated by about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides. In some embodiments, the first sequence is adjacent to the second sequence if the two sequences are separated by up to 2 nucleotides, up to 5 nucleotides, up to 8 nucleotides, up to 10 nucleotides, up to 12 nucleotides, or up to 15 nucleotides. In some embodiments, the first sequence is adjacent to the second sequence if the two sequences are separated by 2-5 nucleotides, 4-6 nucleotides, 4-8 nucleotides, 4-10 nucleotides, 6-8 nucleotides, 6-10 nucleotides, 6-12 nucleotides, 8-10 nucleotides, 8-12 nucleotides, 10-12 nucleotides, 10-15 nucleotides, or 12-15 nucleotides. In a specific example, the target sequence (corresponding to the spacer sequence) is directly adjacent to the PAM motif.
[0102] Spacers disclosed herein may have a length of about 15 to about 30 nucleotides. For example, a spacer may have a length of about 15 to about 20 nucleotides, about 15 to about 25 nucleotides, about 20 to about 25 nucleotides, or about 20 to about 30 nucleotides. In some embodiments, spacers in gRNA may be designed to generally have a length of 15 to 25 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, and 25) and to be complementary to a specific target sequence. In some embodiments, the spacer sequence may be designed to have a length of 18 to 22 nucleotides (e.g., 20 nucleotides).
[0103] In some embodiments, the spacer sequence may have at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.5% sequence identity with the target sequence described herein, and may bind to a complementary region of the target sequence via base pairing.
[0104] In some embodiments, the spacer sequence contains only RNA bases. In some embodiments, the spacer sequence contains DNA bases (e.g., the spacer contains at least one thymine). In some embodiments, the spacer sequence contains both RNA bases and DNA bases (e.g., the DNA-binding sequence contains at least one thymine and at least one uracil).
[0105] (ii) Scaffolding arrangement Scaffold sequences in gRNAs are similarly recognizable by variant CRISPR nuclease polypeptides in gene editing systems. In some cases, the scaffold sequence contains SEQ ID NO: 2, which is a cognate scaffold for the reference CRISPR nuclease SEQ ID NO: 1. GUUUUAGAGCUGUGCUGAAAAGCACAGCACGUUAAAAUAAGGCAGUGAUUGAAAAAUCCAGUCCGUAUUCAGCUUGAAAAAGUGAGCACCGAAUCGGUGCUU(Sequence ID 2)
[0106] In other instances, the scaffold sequence may be a variant derived from SEQ ID NO: 2. Such a variant scaffold sequence may contain nucleotide sequences identical to SEQ ID NO: 2 by at least 80% (e.g., at least 85%, 90%, 95%, 98%, or more). Alternatively or in addition, the variant scaffold sequence may include deletions, nucleotide substitutions, or combinations thereof. The variant CRISPR nuclease polypeptide may have increased binding to the variant scaffold sequence compared to the SEQ ID NO: 2 scaffold. In some examples, the variant scaffold may be a fragment of SEQ ID NO: 2 disclosed herein or a variant thereof. For example, the variant scaffold for use in gRNA provided herein may have a length in the range of 100–150.
[0107] In gRNA, the scaffold may be located at the 3' end of the spacer. In some cases, the scaffold and spacer are directly linked. In other cases, the scaffold and spacer may be linked via a nucleotide linker.
[0108] C. Modification of nucleic acids Any of the RNA components in the gene editing systems disclosed herein, such as editing template RNA, gRNA, and RT donor RNA, may include one or more modifications.
[0109] Exemplary modifications may include modifications to sugars, nucleic acid bases, nucleoside bonds (e.g., binding phosphate / phosphodiester bond / phosphodiester skeleton), and any combination thereof. Some of the exemplary modifications provided herein are described in detail below.
[0110] Any of the gRNAs or nucleic acid sequences encoding components of the composition may include any useful modifications, e.g., modifications to sugars, nucleic acid bases, or nucleoside bonds (e.g., binding phosphate / phosphodiester bond / phosphodiester skeleton). One or more atoms of pyrimidine nucleic acid bases may be replaced, or may be replaced, with optionally substituted aminos, optionally substituted thiols, optionally substituted alkyls (e.g., methyl or ethyl), or halos (e.g., chloro or fluoro). One or more atoms of purine nucleic acid bases may be replaced, or may be replaced, with optionally substituted aminos, optionally substituted thiols, optionally substituted alkyls (e.g., methyl or ethyl), or halos (e.g., chloro or fluoro). In certain embodiments, modifications (e.g., one or more modifications) are present in each of the sugar and nucleoside bonds. Modifications may include modifications from ribonucleic acid (RNA) to deoxyribonucleic acid (DNA), threose nucleic acid (TNA), glycol nucleic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), or hybrids thereof. Additional modifications are described herein.
[0111] In some embodiments, modifications may include chemical or cell-inducible modifications. For example, some non-limiting examples of intracellular RNA modifications are described by Lewis and Pan in “RNA modifications and structures cooperate to guide RNA-protein interactions” from Nat Reviews Mol Cell Biol, 2017, 18:202-210.
[0112] Different sugar modifications, nucleotide modifications, and / or nucleoside bonds (e.g., skeletal structures) can be present at various positions in the sequence. It will be understood by those skilled in the art that nucleotide analogs or other modifications may be located at any position in the sequence such that the function of the sequence is not substantially diminished. The sequence may contain approximately 1% to approximately 100% modified nucleotides (relative to the total nucleotide content, or to one or more types of nucleotides, i.e., A, G, U, or C), or any intervening percentage (e.g., 1% to over 20%, 1% to 25%, 1% to 50%, 1% to 60%, 1% to 70%, 1% to 80%, 1% to 90%, 1% to 95%, 10% to 20%, 10% to 25%, 10% to 50%, 10% to 60%, 10% to 70%, 10% to 80%, 10% to 90%, 10% to 90%, 10% to 10% It may contain modified nucleotides of the following percentages: 95%, 10%-100%, 20%-25%, 20%-50%, 20%-60%, 20%-70%, 20%-80%, 20%-90%, 20%-95%, 20%-100%, 50%-60%, 50%-70%, 50%-80%, 50%-90%, 50%-95%, 50%-100%, 70%-80%, 70%-90%, 70%-95%, 70%-100%, 80%-90%, 80%-95%, 80%-100%, 90%-95%, 90%-100%, and 95%-100%.
[0113] In some embodiments, sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions in one or more ribonucleotides of a sequence may include modifications or substitutions of phosphodiester bonds, as well as skeletal modifications. Specific examples of sequences include, but are not limited to, sequences containing a modified skeleton or nucleoside modifications, such as modifications or substitutions of phosphodiester bonds, which are unnatural nucleoside-to-nucleoside bonds. Sequences having a modified skeleton include, among other things, those that do not have a phosphorus atom in the skeleton. For the purposes of this application and as is often referred to in the art, modified RNAs that do not have a phosphorus atom in the nucleoside-to-nucleoside skeleton can also be considered oligonucleosides. In certain embodiments, the sequence comprises ribonucleotides, each having a phosphorus atom in its nucleoside-to-nucleoside skeleton.
[0114] The modified sequence skeleton may include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkylphosphonates, such as 3'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramides, such as 3'-aminophosphoramides and aminoalkylphosphoramides, thionophosphoramides, thionoalkyl phosphonates, thionoalkyl phosphotriesters, as well as boranophosphates having the usual 3'-5' bond, their 2'-5' bond analogues, and those having reverse polarity with adjacent pairs of nucleoside units bonded from 3'-5' to 5'-3' or 2'-5' to 5'-2'. Also included are various salts, mixed salts, and free acid forms. In some embodiments, the sequences may be negatively or positively charged.
[0115] Modified nucleotides that can be incorporated into a sequence can be modified on nucleoside bonds (e.g., phosphate backbone). In this specification, the terms “phosphate” and “phosphodiester” are used interchangeably in the context of polynucleotide backbones. A phosphate group on a backbone can be modified by substituting one or more oxygen atoms with different substituents. Furthermore, modified nucleosides and nucleotides can include extensive substitutions of the unmodified phosphate moiety at other nucleoside bonds, as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramides, phosphorodiamidates, alkyl or arylphosphonates, and phosphotryesters. Phosphorodithioates have both sulfur-substituted and unbound oxygen atoms. Phosphate linkers can also be modified by substituting bound oxygen with nitrogen (bridged phosphoramide), sulfur (bridged phosphorothioate), and carbon (bridged methylene phosphonate).
[0116] To confer stability to RNA and DNA polymers via non-natural phosphorothioate backbone binding, α-thio-substituted phosphate moieties are provided. Phosphothioate DNA and RNA exhibit increased nuclease resistance and subsequently have a longer half-life in the cellular environment.
[0117] In specific embodiments, the modified nucleoside includes alpha-thio-nucleosides (e.g., 5'-O-(1-thiophosphate)-adenosine, 5'-O-(1-thiophosphate)-cytidine (α-thiocytidine), 5'-O-(1-thiophosphate)-guanosine, 5'-O-(1-thiophosphate)-uridine, or 5'-O-(1-thiophosphate)-psoidouridine).
[0118] Other nucleoside bonds that may be used in the present invention, including nucleoside bonds that do not contain phosphorus atoms, are described herein.
[0119] In some embodiments, the sequence may contain one or more cytotoxic nucleosides. For example, cytotoxic nucleosides may be incorporated into the sequence, such as through bifunctional modification. Cytotoxic nucleosides may include, but are not limited to, adenosine arabinoside, 5-azacitidine, 4'-thio-aracitidine, cyclopentenylcytosine, cladribine, clofarabine, cytarabine, cytosine arabinoside, 1-(2-C-cyano-2-deoxy-beta-D-arabino-pentofuranosyl)-cytosine, decitabine, 5-fluorouracil, fludarabine, phloxuridine, gemcitabine, combinations of tegafur and uracil, tegafur((RS)-5-fluoro-1-(tetrahydrofuran-2-yl)pyrimidine-2,4(1H,3H)-dione), troxacitabine, tezacitabine, 2'-deoxy-2'-methylidencytidine (DMDC), and 6-mercaptopurine. Additional examples include fludarabine phosphate, N4-behenoyl-1-beta-D-arabinofuranosylcytosine, N4-octadecyl-1-beta-D-arabinofuranosylcytosine, N4-palmitoyl-1-(2-C-cyano-2-deoxy-beta-D-arabino-pentofuranosyl)cytosine, and P-4055 (cytarabine 5'-elaidic acid ester).
[0120] In some embodiments, the sequence includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, poly-A sequence, methylation, acylation, phosphorylation, methylation and acetylation of lysine and arginine residues, and nitrosylation of thiol and tyrosine residues). One or more post-transcriptional modifications can be any of the more than 100 different nucleoside modifications identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucleoside Acids Res 27: 196-197). In some embodiments, the first isolated nucleic acid includes messenger RNA (mRNA). In some embodiments, mRNA is pyridine-4-onribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-psoiduridine, 2-thio-psoiduridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyluridine, 1-carboxymethyl-psoiduridine, 5-propynyluridine, 1-propynyl-psoiduridine, 5-taurinomethyluridine, 1-taurinomethyl-psoiduridine, 5-taurinomethyl-2-thiouridine, 1-taurinomethyl-4-thiouridine, 5 It comprises at least one nucleoside selected from the group consisting of -methyluridine, 1-methylpsoiduridine, 4-thio-1-methylpsoiduridine, 2-thio-1-methylpsoiduridine, 1-methyl-1-deaz-psoiduridine, 2-thio-1-methyl-1-deaz-psoiduridine, dihydrouridine, dihydropsoiduridine, 2-thio-dihydrouridine, 2-thio-dihydropsoiduridine, 2-methoxyuridine, 2-methoxy-4-thiouridine, 4-methoxypsoiduridine, and 4-methoxy-2-thiopsoiduridine.In some embodiments, mRNA is 5-aza-cytidine, pseudoisocytidine, 3-methylcytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-psoidisocytidine, pyrrolo-cytidine, pyrrolo-psoidisocytidine, 2-thiocytidine, 2-thio-5-methylcytidine, 4-thio-psoidisocytidine, 4-thio-1-methyl-psoidisocytidine, 4- It comprises at least one nucleoside selected from the group consisting of thio-1-methyl-1-deazal-psoidisocytidine, 1-methyl-1-deazal-psoidisocytidine, zebralin, 5-aza-zebralin, 5-methyl-zebralin, 5-aza-2-thio-zebralin, 2-thio-zebralin, 2-methoxy-cytidine, 2-methoxy-5-methylcytidine, 4-methoxy-psoidisocytidine, and 4-methoxy-1-methyl-psoidisocytidine. In some embodiments, mRNA is 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyiso It comprises at least one nucleoside selected from the group consisting of pentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine.In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of inosine, 1-methylinosine, waiosin, waibutosin, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.
[0121] The sequence may or may not be uniformly modified along the entire length of the molecule. For example, one or more or all types of nucleotides (e.g., spontaneous nucleotides, purines, or pyrimidines, or one or more or all of A, G, U, C, I, pU) may or may not be uniformly modified in the sequence or in a given given sequence region. In some embodiments, the sequence contains pseudouridine. In some embodiments, the sequence contains inosine, which may assist the immune system in characterizing the sequence as endogenous antiviral RNA. Inosine incorporation may also mediate improved RNA stability / reduced degradation. See, for example, Yu, Z. et al. (2015) RNA editing by ADAR1 marks dsRNA as “self”. Cell Res. 25, 1283-1284, which is incorporated in its entirety by reference.
[0122] In some embodiments, any RNA sequence described herein, for example, an edited template RNA, may include terminal modifications (e.g., 5'-terminal modifications or 3'-terminal modifications). In some embodiments, the terminal modifications are chemical modifications. In some embodiments, the terminal modifications are structural modifications. See the disclosures herein for further details.
[0123] When the gene editing systems disclosed herein include nucleic acids encoding CRISPR nucleases and / or RT polypeptides, such as mRNA molecules, such nucleic acid molecules may contain any of the modifications disclosed herein, where applicable.
[0124] III. Gene Editing Methods Any of the gene editing systems may be used to genetically modify (edit) a target nucleic acid, which can be a gene site of interest, such as a gene site where gene editing is needed, for example, to repair a gene mutation, to introduce a protective mutation, or to introduce a modification to regulate gene expression.
[0125] The gene editing systems and compositions disclosed herein are applicable to editing and the introduction of editing into various target sequences. In some embodiments, the target sequence is a DNA molecule, e.g., a DNA locus (referred herein to as the target sequence or on-target sequence). The target sequence is adjacent to a PAM motif. In some cases, the PAM motif is 5' relative to the target sequence. In some embodiments, the target nucleic acid is a genomic site in a cell. In some cases, the target nucleic acid in which gene editing occurs may be located within a protein-coding region. Alternatively, the target nucleic acid may be located within a regulatory region, e.g., a promoter, enhancer, or 5' or 3' untranslated region. In other cases, the target nucleic acid may be located in a non-coding gene, e.g., a transposon, miRNA, tRNA, ribosomal RNA, ribozyme, or lincRNA.
[0126] A. Gene editing Any of the gene editing systems disclosed herein may be used to edit a target gene of interest, for example, a gene involved in a disease (e.g., a genetic disorder). In some embodiments, the target gene may be one involved in an immune response in the subject. For example, the target gene may be an immune checkpoint gene or a member of the tumor necrosis factor receptor superfamily. Gene editing may occur in exons (e.g., in coding regions). Alternatively, gene editing may occur in introns or regulatory elements (e.g., promoters, enhancers, inhibitory elements, etc.). In some cases, gene editing may result in a reduction or elimination of the expression of the target gene. In other cases, gene editing may result in an enhancement of the expression of the target gene (e.g., destruction of an inhibitor).
[0127] In some embodiments, methods are provided herein for introducing at least one edit into a target nucleic acid (e.g., a genomic site of interest, e.g., a genomic site in any of the target genes disclosed herein) using a gene editing system described herein.
[0128] As used herein, the term “editing” refers to the introduction of one or more modifications within a nucleotide sequence in a target nucleic acid, for example, within a nucleotide sequence at a genomic site of interest. Editing may occur within a target sequence as defined herein. Alternatively, editing may occur outside a target sequence (for example, adjacent to a target sequence). Editing may consist of one or more substitutions, one or more insertions, one or more deletions, or a combination thereof.
[0129] A deletion refers to the loss of one or more nucleotides in a nucleic acid sequence relative to a reference sequence. No specific process for constructing a sequence containing a deletion is suggested. For example, a sequence containing a deletion can be synthesized directly from individual nucleotides. In other embodiments, the deletion is made by providing a reference sequence and then modifying it. Nucleic acid sequences can be found within the genome of an organism. Nucleic acid sequences can be found within cells. Nucleic acid sequences can be DNA sequences. Deletions can be frameshift mutations or non-frameshift mutations. As used herein, a deletion refers to an insertion of up to several kilobases.
[0130] An insertion refers to the acquisition of one or more nucleotides in a nucleic acid sequence relative to a reference sequence. No specific process for how to construct a sequence containing an insertion is implied. For example, a sequence containing an insertion can be synthesized directly from individual nucleotides. In other embodiments, the insertion is performed by providing a reference sequence and then modifying it. Nucleic acid sequences can be found within the genome of an organism. Nucleic acid sequences can be found within cells. Nucleic acid sequences can be DNA sequences. Insertions can be frameshift mutations or non-frameshift mutations. As described herein, an insertion refers to an insertion of up to several kilobases.
[0131] In some embodiments, the gene editing methods disclosed herein can introduce indels (e.g., insertions, deletions, or combinations thereof) into a gene site. In some examples, the indels may occur within approximately 1 to 10 nucleotides, 10 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 100 to 200 nucleotides, 200 to 300 nucleotides, 300 to 400 nucleotides, or 400 to 500 nucleotides upstream of the PAM sequence. Alternatively or in addition, the indels may occur within approximately 1 to 10 nucleotides, 10 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 100 to 200 nucleotides, 200 to 300 nucleotides, 300 to 400 nucleotides, or 400 to 500 nucleotides downstream of the PAM sequence.
[0132] In some embodiments, the indel may begin in the PAM sequence. In some embodiments, the indel may begin within approximately 1 to 30 nucleotides downstream of the PAM. Alternatively, the indel may begin within approximately 1 to 30 nucleotides upstream of the PAM.
[0133] B. Gene editing in cells In some embodiments, methods are provided herein for editing a target genomic site in a cell (e.g., a target gene disclosed herein) using one of the gene editing systems disclosed herein. To carry out this method, the gene editing system can be delivered to or introduced into a population of cells. In some cases, cells containing the desired gene edit can be collected, optionally cultured in vitro, and grown.
[0134] The cells described herein can be a variety of cells. In some embodiments, the cells are isolated cells. In some embodiments, the cells are in cell culture or co-culture of two or more cell types. In some embodiments, the cells are ex vivo. In some embodiments, the cells are obtained from living organisms and maintained in cell culture. In some embodiments, the cells are single-celled organisms.
[0135] In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are bacterial cells or derived from bacterial cells. In some embodiments, the cells are archaeal cells or derived from archaeal cells.
[0136] In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are plant cells or derived from plant cells. In some embodiments, the cells are fungal cells or derived from fungal cells. In some embodiments, the cells are animal cells or derived from animal cells. In some embodiments, the cells are invertebrate cells or derived from invertebrate cells. In some embodiments, the cells are vertebrate cells or derived from vertebrate cells. In some embodiments, the cells are mammalian cells or derived from mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the cells are zebrafish cells. In some embodiments, the cells are primate cells. In some embodiments, the cells are rodent cells. In some embodiments, the cells are synthetically produced and are often referred to as artificial cells.
[0137] In some embodiments, the cells are derived from cell lines. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, HEK293T, MF7, K562, HeLa, CHO, and their transgenic variants. Cell lines are available from various sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC) (Manassas, Va.)). In some embodiments, the cells are immortal or immortalized cells. In some embodiments, the cells are stem cells, e.g., totipotent stem cells (e.g., universally pluripotent), pluripotent stem cells, duplexitic stem cells, oligopotent stem cells, or unipotent stem cells. In some embodiments, the cells are induced pluripotent stem cells (iPSCs) or derived from iPSCs. In some embodiments, the cells are mesenchymal stem cells. In some embodiments, the cells are embryonic stem cells. In some embodiments, the cells are hematopoietic stem cells. In some embodiments, the cells are differentiated cells. For example, in some embodiments, differentiated cells are muscle cells (e.g., myocytes), fat cells (e.g., adipocytes), bone cells (e.g., osteoblasts, osteocytes, osteoclasts), blood cells (e.g., monocytes, lymphocytes, neutrophils, eosinophils, basophils, macrophages, erythrocytes, or platelets), nerve cells (e.g., neurons), epithelial cells, immune cells (e.g., lymphocytes, neutrophils, monocytes, or macrophages), liver cells (e.g., hepatocytes), fibroblasts, or sex cells. In some embodiments, the cells are terminally differentiated cells. For example, in some embodiments, terminally differentiated cells are neuronal cells, adipocytes, cardiomyocytes, skeletal muscle cells, epidermal cells, or intestinal cells. In some embodiments, the cells are glial cells. In some embodiments, the cells are islet cells, which include alpha cells, beta cells, delta cells, or enterochromaffin cells. In some embodiments, the cells are immune cells. In some embodiments, the immune cells are T cells.In some embodiments, the immune cells are B cells. In some embodiments, the immune cells are natural killer (NK) cells. In some embodiments, the immune cells are tumor-infiltrating lymphocytes (TILs). In some embodiments, the cells are mammalian cells, such as human cells or primate cells or mouse cells. In some embodiments, the mouse cells are derived from wild-type mice, immunosuppressed mice, or disease-specific mouse models. In some embodiments, the cells are living tissue, organs, or cells within an organism.
[0138] In some embodiments, the cells are primary cells. For example, a culture of primary cells can be passaged 0, 1, 2, 4, 5, 10, or 15 or more times. In some embodiments, primary cells are collected from an individual by any known method. For example, leukocytes can be collected by apheresis, leukocyte apheresis, density gradient separation, etc. Cells from tissues, such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, etc., can be collected by biopsy. A suitable solution can be used to disperse or suspend the collected cells. Such solutions can generally be equilibrium salt solutions (e.g., ordinary saline, phosphate-buffered saline (PBS), Hanks equilibrium salt solution, etc.) conveniently complemented with fetal bovine serum or other spontaneous factors along with an acceptable buffer at a low concentration. Buffers can include HEPES, phosphate buffer, lactate buffer, etc. Cells can be used immediately or stored (e.g., by freezing). Frozen cells may be thawed and reused. Cells can be frozen in DMSO, serum, culture buffer (e.g., 10% DMSO, 50% serum, 40% buffered medium), and / or several other common solutions used to store cells at freezing temperatures.
[0139] In embodiments in which the gene editing system disclosed herein is introduced into a plurality of cells, at least about 0.5% of the cells contain the desired editing. In some embodiments, at least about 1% of the cells contain the desired editing. In some embodiments, at least about 2% of the cells contain the desired editing. In some embodiments, at least about 3% of the cells contain the desired editing. In some embodiments, at least about 4% of the cells contain the desired editing. In some embodiments, at least about 5% of the cells contain the desired editing. In some embodiments, at least about 10% of the cells contain the desired editing. In some embodiments, at least about 20% of the cells contain the desired editing. In some embodiments, at least about 30% of the cells contain the desired editing. In some embodiments, at least about 40% of the cells contain the desired editing. In some embodiments, at least about 50% of the cells contain the desired editing.
[0140] Cells possessing desired gene editing, for example, cells produced using any of the gene editing systems also disclosed herein by the methods disclosed herein, are also within the scope of this disclosure. In some cases, cells modified with the CRISPR nuclease polypeptides disclosed herein may be useful as expression systems for producing biomolecules. For example, modified cells may be useful for producing biomolecules, such as proteins (e.g., cytokines, antibodies, antibody-based molecules), peptides, lipids, carbohydrates, nucleic acids, amino acids, and vitamins. In other embodiments, modified cells may be useful in the production of viral vectors, such as lentiviruses, adenoviruses, adeno-associated viruses, and oncolytic viral vectors. In some embodiments, modified cells may be useful in cytotoxicity studies. In some embodiments, modified cells may be useful as disease models. In some embodiments, modified cells may be useful in vaccine production. In some embodiments, modified cells may be useful in therapeutics. For example, in some embodiments, modified cells may be useful in cell therapies, such as infusions and transplants.
[0141] In some embodiments, cells modified with variant CRISPR nuclease polypeptides such as those disclosed herein may be useful for establishing novel cell lines containing modified genomic sequences. In some embodiments, the modified cells of the Disclosure are modified stem cells (e.g., modified totipotent / pluripotent stem cells, modified pluripotent stem cells, modified duplexipotent stem cells, modified oligopotent stem cells, or modified unipotent stem cells) which differentiate into one or more cell lineages containing modified stem cell deletions. The Disclosure further provides organisms (e.g., animals, plants, or fungi) containing or produced therefrom the modified cells of the Disclosure.
[0142] Delivery of gene editing systems to C. cells In some embodiments, the gene editing systems disclosed herein or any of their components may be formulated, for example, comprising a carrier, e.g., a carrier and / or polymer carrier, e.g., liposomes or lipid nanoparticles, and delivered to cells (e.g., prokaryotes, eukaryotes, plants, mammals, etc.) by known methods. Such methods include, but are not limited to, transfection (e.g., lipid-mediated cationic polymers, calcium phosphate, dendrimers), electroporation or other methods of membrane disruption (e.g., nucleofection), viral delivery (e.g., lentiviruses, retroviruses, adenoviruses, AAVs), microinjection, particle gunning ("gene gun"), fugene, direct sonic loading, cell squeezing, phototransfection, protoplast fusion, imparefection, magnetofection, exome-mediated transfer, lipid nanoparticle-mediated transfer, and any combination thereof.
[0143] In some embodiments, the method involves delivering one or more nucleic acids (e.g., variant CRISPR nuclease polypeptides and / or gRNAs and / or nucleic acids encoding pre-formed ribonucleoproteins) to cells. Exemplary intracellular delivery methods include, but are not limited to, viruses or viral drugs; chemical-based transfection methods, e.g., using calcium phosphate, dendrimers, liposomes, or cationic polymers (e.g., DEAE-dextran or polyethyleneimine); non-chemical methods, e.g., microinjection, electroporation, cell squeezing, sonoporation, phototransfection, imparefection, protoplast fusion, bacterial conjugation, plasmid or transposon delivery; particle-based methods, e.g., particle bombardment, magnetofection or magnetic-assisted transfection, using particle guns; and hybrid methods, e.g., nucleofection. In some embodiments, the application further provides cells produced by such methods, and organisms (e.g., animals, plants, or fungi) containing or produced therefrom such cells. In some embodiments, the composition of the present invention is further delivered together with a drug (e.g., a compound, molecule, or biomolecule) that affects DNA repair or DNA repair mechanisms. In some embodiments, the composition of the present invention is further delivered together with a drug (e.g., a compound, molecule, or biomolecule) that affects the cell cycle.
[0144] In some embodiments, a first composition containing a variant CRISPR nuclease polypeptide is delivered to cells. In some embodiments, a second composition containing gRNA is delivered to cells. In some embodiments, the first composition contacts cells before the second composition contacts the cells. In some embodiments, the first composition contacts cells at the same time as the second composition contacts the cells. In some embodiments, the first composition contacts cells after the second composition has contacted the cells. In some embodiments, the first composition is delivered by a first delivery method, and the second composition is delivered by a second delivery method. In some embodiments, the first delivery method is the same as the second delivery method. For example, in some embodiments, the first and second compositions are delivered via viral delivery. In some embodiments, the first delivery method is different from the second delivery method. For example, in some embodiments, the first composition is delivered by viral delivery and the second composition is delivered by lipid nanoparticle-mediated transfer, and the second composition is delivered by viral delivery, or the first composition is delivered by lipid nanoparticle-mediated transfer and the second composition is delivered by viral delivery.
[0145] IV. Therapeutic uses Any gene editing system disclosed herein, or any modified cells produced using such a gene editing system, may be used to treat diseases that may be benefits from gene editing introduced by the gene editing system or carried by the modified cells. For example, the disease may be a genetic disease, and gene editing repairs gene mutations associated with the genetic disease. Alternatively, the disease may be associated with abnormal gene expression, and gene editing rescues such abnormal expression.
[0146] In some embodiments, methods for treating a disease are provided herein, comprising administering one of the gene editing systems disclosed herein to a subject in need of treatment (e.g., a human patient). The gene editing system may be delivered to a specific tissue or a specific cell type in which gene editing is required. The gene editing system may include an LNP comprising one or more of its components, one or more vectors (e.g., viral vectors) encoding one or more of its components, or a combination thereof. The components of the gene editing system may be formulated to form a pharmaceutical composition, which may further include one or more pharmaceutically acceptable carriers.
[0147] In some embodiments, modified cells produced using any of the gene editing systems disclosed herein may be administered to subjects in need of treatment (e.g., human patients). Modified cells may include substitutions, insertions, and / or deletions described herein. In some examples, modified cells may include cell lines modified with variant CRISPR nuclease polypeptides and gRNAs disclosed herein. In some cases, modified cells may be a heterogeneous population, which may include cells with different types of gene editing. Alternatively, modified cells may include a substantially homogeneous cell population (e.g., at least 80% of the cells in the whole population), which may include one specific gene editing. In some examples, cells may be suspended in a suitable culture medium.
[0148] In some embodiments, gene editing systems or components thereof, or compositions comprising modified cells are provided herein. Such compositions may be pharmaceutical compositions. Useful pharmaceutical compositions may be prepared, packaged, or marketed in formulations suitable for oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, intrafocal, buccal, ocular, intravenous, intraorganic, or other routes of administration. Pharmaceutical compositions of this disclosure may be prepared, packaged, and / or marketed in bulk as single unit doses or as multiple single unit doses. As used herein, “unit dose” is a distinct amount of a pharmaceutical composition containing a predetermined number of cells. The number of cells is generally equal to the dose of cells that would be administered to a subject, or a convenient portion of such a dose, for example, half or one-third of such a dose.
[0149] Formulations of pharmaceutical compositions suitable for parenteral administration may include an activator (e.g., a gene editing system or its components, or modified cells) combined with a pharmaceutically acceptable carrier, such as sterile water or sterile isotonic saline. Such formulations may be prepared, packaged, or sold in forms suitable for bolus or serial administration. Some injectable formulations may be prepared, packaged, or sold in unit dosage forms, for example, in ampoules or multi-dose containers containing preservatives. Some formulations for parenteral administration include, but are not limited to, suspensions, solutions, emulsions in oily or aqueous vehicles, pastes, and embedded sustained-release formulations or biodegradable formulations. Some formulations may further include, but are not limited to, one or more additional components, including suspending agents, stabilizers, or dispersants.
[0150] Pharmaceutical compositions may be in the form of sterile, injectable aqueous or oily suspensions or solutions. These suspensions or solutions may be formulated by known techniques and may contain, in addition to cells, additional components, such as dispersants, wetting agents, or suspending agents as described herein. Such sterile, injectable formulations may be prepared using non-toxic, parenterally acceptable diluents or solvents, such as water or saline. Other acceptable diluents and solvents include, but are not limited to, Ringer's solution, isotonic sodium chloride solution, and fixative oils such as synthetic monoglycerides or diglycerides. Other useful parenterally administered formulations include those that may contain cells in packaged form, within liposome preparations, or as components of biodegradable polymer systems. Some compositions for sustained release or embedding may contain pharmaceutically acceptable polymers or hydrophobic materials, such as emulsions, ion exchange resins, sparingly soluble polymers, or sparingly soluble salts.
[0151] V. Kits and their use This disclosure also provides kits or systems that can be used, for example, to carry out the methods described herein. In some embodiments, the kit or system comprises a variant CRISPR nuclease polypeptide and, optionally, a gRNA. In some embodiments, the kit or system comprises a variant CRISPR nuclease polypeptide and, optionally, a polynucleotide encoding the gRNA. The gRNA of the kit can be designed to target a sequence of interest. The variant CRISPR nuclease polypeptide and the gRNA can be packaged in the same vial or other container within the kit or system, or in separate vials or other containers, and their contents can be mixed before use. In addition, the kit or system may optionally include a buffer, and / or instructions on how to use the variant CRISPR nuclease polypeptide and the gRNA.
[0152] In some embodiments, the kit comprises a first composition comprising a variant CRISPR nuclease polypeptide as disclosed herein. In some embodiments, the kit comprises a second composition comprising a gRNA also disclosed herein. In some embodiments, the first and second compositions are packaged in the same vial. In some embodiments, the first and second compositions are packaged in different vials.
[0153] In some embodiments, the kit may be useful for research purposes. For example, in some embodiments, the kit may be useful for studying gene function.
[0154] General technology The practices described herein will employ conventional techniques within the scope of the art, including molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, unless otherwise specified. Such techniques are described in Molecular Cloning: A Laboratory Manual, second edition (Sambrook, et al., 1989) Cold Spring Harbor Press, Oligonucleotide Synthesis (MJ Gait, ed. 1984), Methods in Molecular Biology, Humana Press, Cell Biology: A Laboratory Notebook (JECellis, ed., 1989) Academic Press, Animal Cell Culture (RIFreshney, ed. 1987), Introduction to Cell and Tissue Culture (JP Mather and PE Roberts, 1998) Plenum Press, Cell and Tissue Culture: Laboratory Procedures (A. Doyle, JBGriffiths, and DG Newell, eds. 1993-8) J. Wiley and Sons, Methods in Enzymology (Academic Press, Inc.), Handbook of Experimental Immunology (DMWeir and CCBlackwell, eds.): Gene Transfer Vectors for Mammalian Cells (JMMiller and MPCalos, eds., 1987), Current Protocols in Molecular Biology (FMAusubel, et al. eds. 1987); PCR: The Polymerase Chain Reaction, (Mullis, et al., eds. 1994), Current Protocols in Immunology (JEColigan et al., eds., 1991), Short Protocols in Molecular Biology (Wiley and Sons, 1999), Immunobiology (C.A. Janeway and P. Travers, 1997), Antibodies (P. Finch, 1997), Antibodies: a practice approach (D. Catty., ed., IRL Press, 1988-1989), Monoclonal antibodies: a practical approach (P. Shepherd and C. Dean, eds., Oxford University Press, 2000), Using antibodies: a laboratory manual (E. Harlow and D. Lane (Cold Spring Harbor Laboratory Press, 1999), The Antibodies (M. Zanetti and J.D. Capra, eds. Harwood Academic Publishers, 1995), DNA Cloning: A practical Approach, Volumes I and II (D.N. Glover ed. 1985), Nucleic Acid Hybridization (B.D. Hames & S.J. Higgins eds. (1985)), Transcription and Translation (B.D. Hames & S.J. Higgins, eds. (1984)), Animal Cell Culture (R.I. Freshney, ed. (1986)); Immobilized Cells and Enzymes (IRL Press, (1986)), and are sufficiently described in documents such as B. Perbal, A practical Guide To Molecular Cloning (1984), F.M. Ausubel et al. (eds.).
[0155] Without further detail, those skilled in the art will likely be able to make the most of the present invention based on the above description. Accordingly, the following specific embodiments should be construed as merely illustrative and in no way limit the remainder of this disclosure. All publications referenced herein are incorporated by reference for the purposes or subjects referenced herein. [Examples]
[0156] The following embodiments are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit the scope of the present disclosure. It will be understood that, by their exemplary nature, other procedures, methods, or techniques known to those skilled in the art may be used as alternatives.
[0157] Example 1: CRISPR nuclease-mediated editing of human target genes in HEK293T cells This example describes genome editing of exemplary target genes, including the AAVS1, EMX1, and VEGFA genes, by introducing the CRISPR nuclease of SEQ ID NO: 1 into HEK293T cell lines via lipid-based transient transfection.
[0158] CRISPR nucleases were tagged with the N-terminal SV40 nuclear localization signal (NLS) and the C-terminal XTEN linker immediately upstream of the nucleoplasmin NLS. Their coding sequences were converted to human codon-optimized DNA sequences, synthesized, and cloned into a pcDNA3.1 vector (Invitrogen) containing a CMV promoter for expression. The reference and NLS-tagged sequences used are shown in Table 1. Plasmids were purified using a midiprep kit. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5]
[0159] RNA guides were designed and cloned into the pUC19 plasmid, followed by the U6 PolIII promoter, and terminated with a 6x polyT sequence. The RNA guides were designed to be specific to target sequences within AAVS1, EMX1, and VEGFA, which have a 5'-NGG-3'PAM sequence (the PAM sequence is located on the 3' end of the target sequence). For more efficient transcription, the U6 PolIII promoter used +1 G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences described here. See Table 2 for all RNA guide sequences. Plasmids were purified using a midiprep kit. [Table 2]
[0160] Approximately 16 hours prior to transfection, 25,000 HEK293T cells were seeded in DMEM / 10% FBS+Pen / Strep (D10 medium) into each well of a 96-well plate. On the day of transfection, the cells were 70–90% confluent. For each well to be transfected, a mixture of Lipofectamine 2000® (ThermoFisher Scientific) and Opti-MEM® (ThermoFisher Scientific) was prepared and incubated at room temperature for 5 minutes (Solution 1). After incubation, the Lipofectamine 2000®:Opti-MEM® mixture was added to a separate mixture containing a CRISPR nuclease plasmid (NLS tagged), an RNA guide plasmid, and Opti-MEM® (Solution 2). In the case of a negative control, the CRISPR nuclease plasmid was excluded. Solutions 1 and 2 were mixed by pipetting up and down, and then incubated at room temperature for 25 minutes. After incubation, the mixture of solutions 1 and 2 was added dropwise to each well of a 96-well plate containing cells. Approximately 72 hours after transfection, the cells were trypsinized by adding TrypLE® (Thermo Fisher Scientific) to the center of each well and incubating at 37°C for approximately 5 minutes. D10 medium was then added to each well and mixed to resuspend the cells. The resuspended cells were centrifuged for 10 minutes to obtain a pellet, and the supernatant was discarded. The cell pellet was then resuspended in QuickExtract® buffer (Lucigen®), and the cells were incubated at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes.
[0161] Next-generation sequencing (NGS) samples were prepared by two rounds of PCR. Three technical replicates were analyzed for each target, reference, and variant. The first round (PCR1) was used to amplify specific genomic regions in a target-dependent manner. The second round of PCR (PCR2) was performed to add the Illumina adapter and index. The reaction products were then pooled and purified by column purification. Sequencing was performed using a 150-cycle NextSeq 500 / 550 intermediate or high-power v2.5 kit or a 200-cycle NovaSeq 6000 SP or S1 reagent kit v1.5.
[0162] For NGS analysis, the indel mapping function used the sample's fastq file, amplicon reference sequence, and forward primer sequence. For each read, the KMER scanning algorithm was used to calculate the editing behavior (match, mismatch, insertion, deletion) between the read and the reference sequence. To remove small amounts of primer dimers present in some samples, the first 30 nt of each read was required to match the reference, and reads with more than half of the mapping nucleotides being mismatched were filtered out. Up to 50,000 reads that passed these filters were used for analysis, and if a read contained an insertion or deletion, it was counted as an indel read. The QC criterion for the minimum number of filtered reads was 10,000.
[0163] For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for each sample and its related protein-free control. The higher percentage of indels in targets when CRISPR nucleases were included in transfection demonstrated the DNA editing results in cells.
[0164] As shown in Figure 1, each of the six targets tested demonstrated a larger level of indels than was observed when the CRISPR nuclease plasmid was present. Therefore, this example demonstrates that the CRISPR nuclease of Sequence ID No. 1 edited human genes.
[0165] Example 2: Effect of variant CRISPR nuclease for targeting mammalian genes This example describes an indel evaluation for an exemplary mammalian target using a CRISPR nuclease variant transfected into HEK293T cells.
[0166] Arginine scanning mutagenesis was performed to individually substitute selected non-arginine residues with arginine in the reference CRISPR nuclease (SEQ ID NO: 1). SEQ ID NO: 1 is referred to herein as the reference sequence. This resulted in 372 single-arginine substitution variants. The reference CRISPR nuclease variants and the nucleic acids encoding each CRISPR nuclease variant were then individually cloned into a pcDNA3.1 backbone (Invitrogen™) to prepare plasmids, which were then diluted. The plasmids contained a CMV promoter, a first NLS upstream of the coding sequence (KRTADGSEFESPKKKRKV, SEQ ID NO: 3), an XTEN linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS, SEQ ID NO: 8), and a second NLS downstream of the coding sequence (KRPAATKKAGQAKKKK, SEQ ID NO: 4). See also Example 1 above.
[0167] Exemplary RNA guides for VEGFA-T6 and EMX1-T7 were used in this study. Details of these gRNAs are provided in Table 2 above. The RNA guides were cloned into a pUC19 backbone (New England Biolabs®). Plasmids were purified and diluted using a maxi-prep kit. Cells were transfected, and samples were prepared for NGS as described in Example 1. Indel ratios, which refer to the percentage of NGS reads containing indels, were calculated for the reference and for each variant. The indel ratio used for fold change calculations was the average of two technical replicates. Then, to calculate the fold change in indel ratio, the indel ratio for each variant was divided by the indel ratio for the reference. Table 3 shows the fold change in indel ratio for each target tested. The numbering is relative to the reference nuclease (i.e., without NLS) of Sequence ID No. 1.
[0168] As shown in Table 3, six of the 372 variants with a single arginine substitution (left column) were characterized as resulting in an increase of at least 1.5 times the indel ratio compared to the reference indel ratio when averaged across two targets (right column). [Table 3]
[0169] Fifty-five variants with a single arginine substitution were analyzed to have indel ratios 1 to 1.5 times that of the reference indel ratio: G1329R, Q741R, P1238R, A1236R, Q1230R, E1214R, A730R, Q849R, S473R, D985R, M753R, K918R, I1106R, M1205R, Y1015R, Q1360R, K1344R, F501R, K1091R, G731R, N1295R, F1241R, V752R, N1099 R, E1179R, D720R, F1285R, I1370R, E1037R, E982R, S1094R, S872R, P1096R, A1226R, Q840R, I1147R, L1117R, E331R, T1108R, K1298R, Y986R, S898R, A1333R, P1208R, E1105R, Q809R, L1281R, A1292R, K1131R, I581R, I79R, D1284R, T1104R, Y348R, and S1348R. The remaining variants with a single arginine substitution (311 variants) resulted in a reduced indel ratio relative to the reference indel ratio (a multiplier change in indel ratios less than 1.0).
[0170] The following variants, which exhibited at least a 1.5-fold increase in the indel ratio relative to the reference, were selected and used to manipulate the combined variants: I857R, N813R, L784R, K736R, A919R, and Q812R.
[0171] Example 3: Manipulation and effects of CRISPR nickerse variants for targeting mammalian genes This example describes introducing a mutation into the CRISPR nuclease of SEQ ID NO: 1 that disrupts either the HNH or RuvC domain to produce a functional nickase. D844, H845, and N868 were identified as putative catalytic residues in the HNH domain. D10, E763, and D991 were identified as putative catalytic residues in the RuvC domain. These positions were identified by analyzing structural regions similar to known HNH and RuvC active sites using a model generated with AlphaFold2 (Jumper et al., Nature 596:583-9 (2021)) and / or by performing sequence alignment with other nucleases in which candidate sites had been previously identified. Examples of reference structures used to identify HNH and RuvC active sites are represented by the following Protein Databank (PDB) identifiers: 5h0m, 7eu9, 6ltu, 7odf, 7lys, 8dc2, 4cmp, 4oo8, 7z4j, 5axw, 5b2o, 6kc8, 7utn, 8csz, 8ctl, and 8dmb.
[0172] The coding sequences of CRISPR nucleases were converted to E. coli codon-optimized DNA sequences, synthesized, and cloned into a pET-28a(+) vector (Novagen) containing lac and T7 RNA polymerase promoters for gene expression. To test nickase activity, individual alanine variants were cloned for each of the positions identified as putative active site residues in the HNH and RuvC domains. A leucine variant was also cloned for position H845. Research-grade plasmids were received from GenScript. The manipulated nickase sequences are shown in Table 4. Codons encoding substituted residues are shown in underlined bold capital letters in the nucleotide sequence, and substituted residues are shown in underlined bold letters in the amino acid sequence. The putative HNH knockout nickase was expected to cleave non-target strands but not target strands. The putative RuvC knockout nickase was expected to cleave target strands but not non-target strands. Table 4-1 Table 4-2 Table 4-3 Table 4-4 Table 4-5 Table 4-6 Table 4-7 Table 4-8 Table 4-9 Table 4-10 Table 4-11 Table 4-12 Table 4-13 Table 4-14
[0173] A linear DNA template encoding an RNA guide was designed, with a T7 promoter upstream and a T7Te terminator sequence downstream. The RNA guide was designed to be specific to the previously tested target sequence described in Example 1 and Table 2 above, located within the coding exon of EMX1 having a 5'-NGG-3'PAM sequence (where PAM is the 3' end of the target sequence). For more efficient transcription, the T7 promoter used +1 G at the start of the transcript (i.e., the 5' end of the RNA), as shown for Sequence ID No. 45. The sequences of the encoded RNA guide and its individual components are shown in Table 5. [Table 5]
[0174] The DNA target was designed and ordered as a synthetic linear DNA fragment. The target sequence from EMX1 and the 10 upstream and downstream bases within the exon was adjacent to 200 upstream and 100 downstream bases of an unrelated sequence. An extra sequence was added, thereby allowing for good separation of the cleaved and uncleaved products on the gel. The target and non-target strands were labeled with 5'IR700 and 5'IR800 labels, respectively, via PCR amplification using labeled primers. The sequences of the DNA target, the individual components of the DNA target, and the labeled PCR primers are shown in Table 6. [Table 6]
[0175] The cleavage activity of the CRISPR nuclease (SEQ ID NO: 1) and putative nickase was evaluated using an in vitro cleavage assay. Each polypeptide was individually co-expressed in vitro with an RNA guide by incubating a plasmid encoding the target protein from Table 4 and a linear DNA template of T7 transcription EMX1-T2 sgRNA from Table 5 in PURExpress® solution (NEB) containing SUPERase·In® RNase Inhibitor (Invitrogen) for 2 hours at 37°C. The unpurified polypeptide / RNA solution was then diluted in a solution of 1x NEB Buffer 2 (NEB) containing approximately 1 ng / μl of labeled DNA target amplicon. The solution was then incubated at 37°C for 1 hour. The reaction was then stopped by incubation with RNase Cocktail (Invitrogen, final concentration approximately 1 U / μl) at 37°C for 15 minutes, followed by incubation with proteinase K (NEB, final concentration approximately 0.04 U / μl) at 55°C for 30 minutes. The DNA was then purified using CleanNGS DNA & RNA Clean-Up Magnetic Beads (Bulldog Bio).
[0176] The cleaved and uncleaved products of the target and non-target strands were separated by running the sample on a 10% TBE-urea PAGE gel. The gel was imaged using a LI-COR Odyssey M imaging system with 700 nm and 800 nm channels to visualize 5'IR700 and 5'IR800 labeling in the target and non-target strands of the target DNA substrate. Band intensities were quantified using ImageJ software.
[0177] Gel images are shown in Figures 2A–2C, and the quantification of the percentage of cleaved target and non-target chains is shown in Figure 2D. Uncleaved chain, HNH-cleaved chain, and RuvC-cleaved chain are shown. Figure 2A is a gel image captured using a 700 nm channel, showing cleavage of the target chain. Figure 2B is a gel image captured using an 800 nm channel, showing cleavage of the non-target chain. Figure 2C is an overlay of gel images from Figures 2A and 2B. As shown in Figures 2A–2D, the reference CRISPR nuclease (SEQ ID NO: 1) cleaved both the target and non-target chains, as expected. Three of the four HNH knockout nickase constructs (H845A, H845L, and N868A) showed significantly reduced activity in the target chain while retaining activity in the non-target chain. Each of the three RuvC knockout nickase constructs (D10A, E763A, and D991A) exhibited significantly reduced activity in the non-target chain while retaining activity in the target chain (Figures 2A-2D).
[0178] Therefore, this example demonstrates the successful manipulation of HNH knockout nickases and RuvC knockout nickases.
[0179] Example 4: Efficacy of combined CRISPR nuclease variants for targeting mammalian genes This example describes indel evaluation for mammalian targets using CRISPR nuclease variants containing two or more substitutions identified in Example 2 as increasing indel activity. Thirty-five combinations of CRISPR nuclease variants were tested.
[0180] Each CRISPR nuclease variant and RNA guide was cloned as described in Example 2. Exemplary RNA guides for VEGFA-T6 and EMX1-T7 were used in this study. Details of these gRNAs are provided in Table 2 above. HEK293T cells were further transfected as described in Example 2, followed by NGS analysis. For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for the reference CRISPR nuclease (SEQ ID NO: 1) and for each variant CRISPR nuclease. The indel ratios shown in Table 7 were calculated as the average of two biological replicas, each containing two technical replicas. [Table 7]
[0181] As shown in Table 7, each CRISPR nuclease variant with amino acid substitution combinations exhibited higher indel activity than the reference CRISPR nuclease (SEQ ID NO: 1). Nine CRISPR nuclease variants yielded indel ratios greater than 0.25 when averaged across both targets, indicating that more than 25% of NGS reads contained indels. These nine CRISPR nuclease variants included the following substitution combinations: a) I857R, L784R, K736R; b) I857R, A919R, K736R; c) I857R, N813R, L784R; d) I857R, L784R, A919R; e) I857R, N813R, K736R; f) I857R, N813R; g) L784R, A919R, K736R; h) I857R, L784R; and i) I857R, A919R. Eight of the CRISPR nuclease variants yielded indel ratios of 0.2–0.24 when averaged across both targets, indicating that 20%–24% of NGS reads contained indels. The 18 CRISPR nuclease variants yielded indel ratios of 0.1–0.19 when averaged across both targets, indicating that 10–19% of NGS reads contained indels. The average indel ratio across both targets exceeded the average indel ratio of the reference for all variants tested.
[0182] Based on this experiment, the highest-performing CRISPR nuclease variants, including substituted I857R, L784R, and K736R, were selected for further testing. These CRISPR nuclease variants exhibited a 2.5-fold increase in indel activity compared to the reference CRISPR nuclease.
[0183] Example 5: Design of additional CRISPR nucleases This example describes the design of an additional CRISPR nuclease having desired bioactivity.
[0184] Manipulate additional variants of the CRISPR nickase described herein. Replace each individual amino acid residue of the CRISPR nickase having the H845A substitution from Example 3 with the remaining 19 available amino acids (excluding the H845A residue). Clone the nucleic acid encoding the CRISPR nickase variant into a pcDNA3.1 vector (Invitrogen) containing the CMV promoter. Introduce the expression vector into host cells for the expression of the CRISPR nickase variant. Purify the CRISPR nickase variant expressed in the host cells and evaluate their nickase activity according to the procedure provided in Example 3 above.
[0185] Additional CRISPR nuclease variants are engineered to increase double-stranded nuclease activity. The variants in Table 8 are cloned and evaluated as described in Example 2. [Table 8-1] [Table 8-2] [Table 8-3] [Table 8-4] [Table 8-5] [Table 8-6] [Table 8-7] [Table 8-8] [Table 8-9] [Table 8-10] Table 8-11 Table 8-12 Table 8-13 Table 8-14 Table 8-15 Table 8-16 Table 8-17 Table 8-18 Table 8-19 Table 8-20 Table 8-21 Table 8-22 Table 8-23 Table 8-24 Table 8-25 Table 8-26 Table 8-27 Table 8-28 Table 8-29 Table 8-30 Table 8-31 Table 8-32 Table 8-33 Table 8-34 Table 8-35 Table 8-36 Table 8-37 Table 8-38 Table 8-39 Table 8-40 Table 8-41 Table 8-42 Table 8-43 Table 8-44 [Table 8-45]
[0186] Additional CRISPR nuclease variants are engineered and evaluated for their ability to recognize less stringent PAM sequences. The variants listed in Table 9 are cloned and evaluated as described in Example 2 using a target sequence adjacent to a 5'-NGN-3', 5'-NRN-3', or 5'-NYN-3' PAM sequence, wherein N represents any nucleotide, R represents G or A, and Y represents C or T. [Table 9-1] [Table 9-2] [Table 9-3]
[0187] Example 6: Effect of CRISPR Nuclease Variants with Relaxed PAM Stringency for Targeting Exemplary Mammalian Genes This example describes indel assessment for exemplary mammalian targets using CRISPR nuclease variants with relaxed PAM requirements transfected into HEK293T cells.
[0188] Arginine scanning mutagenesis was performed to individually substitute selected non-arginine residues of the CRISPR nuclease variant of SEQ ID NO: 130 with arginine. This resulted in 372 single arginine substitution variants. The variants were cloned and evaluated as described in Example 2 using target sequences flanking 5'-NGN-3' PAM sequences summarized in Table 10. [Table 10]
[0189] HEK293T cells were further transfected as described in Example 2, followed by NGS analysis. The indel activity of the CRISPR nuclease variant of SEQ ID NO: 130 is shown in Table 11. The data in Table 11 are the average of 10 control samples, each of which had two biological replicas and two technical replicas. [Table 11]
[0190] Next, for each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for the variant CRISPR nuclease of SEQ ID NO: 130 and for each variant CRISPR nuclease. Then, to calculate the multiplier change in the indel ratio, the indel ratio for each variant was divided by the indel ratio for the variant CRISPR nuclease of SEQ ID NO: 130. The indel ratio used for the multiplier change calculation was the average of the two technical replicates. As shown in Table 12, three of the 372 variants with a single arginine substitution (left column) were characterized as resulting in at least a twofold increase in the indel ratio compared to the indel ratio for the variant CRISPR nuclease of SEQ ID NO: 130 when averaged across the two targets (right column). [Table 12]
[0191] Eleven variants with a single arginine substitution were analyzed, assuming they had indel ratios 1.5 to 2 times the reference indel ratio: L64R, S410R, T67R, Q849R, G1110R, F501R, T659R, L784R, Y516R, G55R, and E1037R. 92 variants exhibited indel ratios 1 to 1.4 times the reference indel ratio: N57R, D720R, A919R, A1294R, Q812R, N700R, H657R, T73R, Q899R, T1347R, I857R, K751R, D327R, I581R, D462R, E331R, A589R, D471R, I699R. N1295R, T470R, I1147R, E130R, S473R, A353R, K40R, K334R, A60R, S1348R, K367R, A1118R, K 31R, Q349R, K341R, Q83R, K585R, Q840R, G660R, K527R, G727R, Y42R, L1281R, L122R, Q123R, T 1108R, E41R, K1131R, K30R, S872R, I1206R, D1132R, K460R, L80R, E459R, K1182R, L1117R, M 696R, K918R, K126R, N721R, G1227R, Q809R, K1091R, K736R, A1332R, K783R, N498R, K723R, E 1228R, H1119R, F463R, L594R, D472R, K744R, E365R, G595R, K45R, Y348R, K964R, S1181R, N813R, D407R, S839R, Y658R, E586R, G754R, A730R, Y1015R, D903R, A1333R, S461R, and H1359R. The remaining variants with a single arginine substitution (266 variants) resulted in a decrease in the indel ratio compared to the variant CRISPR nuclease of SEQ ID NO: 130 (indel ratio multiplier change less than 1.0).
[0192] Therefore, this embodiment demonstrates that the CRISPR nuclease variant of SEQ ID NO: 130 is an active nuclease capable of editing a target sequence adjacent to 5'-NGN-3'PAM (where N represents A, C, G, or U), and that certain further arginine substitutions (e.g., D61R, A68R, and / or H494R) increase nuclease activity.
[0193] Other Embodiments All features disclosed herein can be combined in any combination. Each feature disclosed herein may be replaced by an alternative feature that serves the same, equivalent, or similar purpose. Thus, unless otherwise expressly stated, each feature disclosed is merely an example of a general set of equivalent or similar features.
[0194] From the above description, those skilled in the art will readily grasp the essential features of this disclosure and can make various modifications and alterations to the present invention to suit various uses and conditions without departing from its spirit and scope. Accordingly, other embodiments are also within the scope of the claims.
[0195] Equal portions While several embodiments of the present invention are described and illustrated herein, various other means and / or structures for carrying out the functions described herein and / or obtaining one or more of the results and / or benefits will be readily conceivable to those skilled in the art, and each of such variations and / or modifications will be considered to fall within the scope of the embodiments of the present invention described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials and arrangements described herein are intended to be illustrative, and that actual parameters, dimensions, materials and / or arrangements will depend on the specific application or multiple application in which the teachings of the present invention are used. Those skilled in the art will be able to recognize or grasp many equivalents to the specific embodiments of the present invention described herein by means of conventional experimentation alone. Therefore, it should be understood that the embodiments described herein are presented merely as examples, and embodiments of the present invention may be specifically described or practiced outside the scope of the appended claims and their equivalents. Embodiments of the present invention in this disclosure cover each individual feature, system, article, material, kit and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present invention as long as such features, systems, articles, materials, kits, and / or methods are not inconsistent with each other.
[0196] All definitions defined and used herein should be understood to take precedence over dictionary definitions, definitions incorporated by reference in documents, and / or the ordinary meanings of the defined terms.
[0197] All references, patents, and patent applications disclosed herein are incorporated by reference with respect to the subject matter they cite, which may in some cases encompass the entire document.
[0198] As used in this specification and the claims, the indefinite articles "a" or "an" should be understood to mean "at least one", unless a contrary intention is clearly indicated.
[0199] As used in this specification and the claims, the phrase "and / or" should be understood to mean "either or both" of the elements so conjoined, that is, elements that are present conjointly in some cases and separately in other cases. Multiple elements listed with "and / or" should be construed in the same fashion, that is, "one or more" of the elements so conjoined. Other elements besides those specifically identified by the "and / or" clause may optionally be present, regardless of whether they are related or unrelated to those specifically identified elements. Therefore, as a non-limiting example, a reference to "A and / or B", when used in conjunction with an open-ended expression such as "comprising", can refer to, in one embodiment, only A (optionally including elements other than B), in another embodiment, only B (optionally including elements other than A), and in yet another embodiment, both A and B (optionally including other elements), and the like.
[0200] Where used herein and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as inclusive, that is, inclusion of several elements or a list of elements and at least one additional item not enumerated by choice, but also including two or more. Only terms that clearly indicate the opposite meaning, such as “one of” or “exactly one of,” or, when used in the claims, “consisting of,” would refer to the inclusion of just one element from several elements or a list of elements. In general, where used herein, the term “or” shall be interpreted only as indicating exclusive substitution (i.e., “one or the other, but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “one of,” or “exactly one of.” Where used in the claims, “essentially consisting of” has its usual meaning as used in the field of patent law.
[0201] As used herein in this specification and in the claims, the phrase “at least one” should be understood to mean at least one element selected from any one or more elements in a list of elements, but not necessarily including at least one of each element specifically enumerated in the list of elements or all elements, nor excluding any combination of elements in the list of elements. This definition also allows for the existence of elements other than those specifically identified in the list of elements to which the phrase “at least one” refers, whether related to or unrelated to those specifically identified elements, at the discretion of the definition. Therefore, as a non-restrictive example, “at least one of A and B” (or equivalently, “at least one of A or B,” or equivalently, “at least one of A and / or B”) could mean, in one embodiment, at least one A (and optionally including elements other than B) in which B is absent and optionally two or more are included; in another embodiment, at least one B (and optionally including elements other than A) in which A is absent and optionally two or more are included; in yet another embodiment, at least one A and optionally two or more are included, and at least one B (and optionally including other elements), and so on.
[0202] Furthermore, unless explicitly stated otherwise, it should be understood that in any method claimed herein that includes two or more steps or actions, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are enumerated.
Claims
1. A modified CRISPR nuclease polypeptide, wherein the modified CRISPR nuclease polypeptide is a variant of the reference CRISPR nuclease described as Sequence ID No. 1, comprising a RuvC nuclease domain and an HNH nuclease domain, and the modified CRISPR nuclease polypeptide is, (i) One or more nickase mutations in the HNH nuclease domain or the RuvC nuclease domain that reduce or eliminate the nuclease activity thereof, wherein the one or more mutations are located in the HNH nuclease domain, (ii) One or more arginine and / or lysine substitutions, optionally one or more arginine substitutions, (iii) One or more mutations to reduce PAM recognition stringency, or A modified CRISPR nuclease polypeptide, including combinations of (iv), (i), (ii), and / or (iii).
2. The manipulated CRISPR nuclease polypeptide according to claim 1, wherein the CRISPR nuclease polypeptide comprises one or more mutations in the HNH nuclease domain at positions D844, H845, and / or N868 relative to SEQ ID NO: 1, and optionally the mutation is located at position H845.
3. (a) The mutation in D844 is an amino acid substitution of D844A, D844G, D844L, or D844S, (b) The mutation in H845 is an amino acid substitution of H845A, H845G, H845L, or H845S, (c) The modified CRISPR nuclease polypeptide according to claim 2, wherein the mutation in N868 is an amino acid substitution of N868A, N868G, N868L, or N868S.
4. The manipulated CRISPR nuclease polypeptide according to claim 1, wherein the CRISPR nuclease polypeptide comprises one or more mutations in the RuvC nuclease domain at positions D10, E763, and / or D991 relative to SEQ ID NO: 1, and optionally the mutations are located at position E763 or D991.
5. The manipulated CRISPR nuclease polypeptide according to any one of claims 1 to 4, wherein the manipulated CRISPR nuclease polypeptide comprises a bridge helix (BH) domain, a nucleic acid recognition (REC) domain, a phosphate-locked loop (PLL) domain, a wedge (WED) domain, and a PAM interaction (PID) domain, wherein one or more arginine and / or lysine substitutions, optionally an arginine substitution, is located in the BH domain, the REC domain, the PLL domain, the WED domain, the PID domain, or a combination thereof.
6. The manipulated CRISPR nuclease polypeptide according to any one of claims 1 to 5, wherein the manipulated CRISPR nuclease polypeptide contains up to 20 arginine and / or lysine substitutions relative to the reference CRISPR nuclease, and optionally, the manipulated CRISPR nuclease polypeptide contains up to 15 arginine and / or lysine substitutions relative to the reference CRISPR nuclease, preferably, one or more arginine and / or lysine substitutions are located at positions K736, L784, Q812, N813, I857, and / or A919, wherein this is optionally I857R.
7. The manipulated CRISPR nuclease polypeptide according to claim 6, wherein the CRISPR nuclease polypeptide contains at least two arginine and / or lysine substitutions relative to the reference CRISPR nuclease, the at least two arginine and / or lysine substitutions being located at positions K736, L784, Q812, N813, I857, and / or A919.
8. The CRISPR nuclease polypeptide is subjected to arginine and / or lysine substitutions relative to the reference CRISPR nuclease at the following positions: (a) I857, L784, and K736, (b) I857, A919, and K736, (c) I857, N813, and L784, (d) I857, L784, and A919, (e) I857, N813, and K736, (f) I857 and N813, (g) L784, A919, and K736, (h) I857 and L784, and (i) The manipulated CRISPR nuclease polypeptide according to claim 7, which is contained in I857 and A919.
9. The CRISPR nuclease polypeptide is subjected to the following arginine substitutions relative to the reference CRISPR nuclease: (a) I857R, L784R, and K736R, (b) I857R, A919R, and K736R, (c) I857R, N813R, and L784R, (d) I857R, L784R, and A919R, (e) I857R, N813R, and K736R, (f) I857R and N813R, (g) L784R, A919R, and K736R, (h) I857R and L784R, and (i) Including I857R and A919R, The modified CRISPR nuclease polypeptide according to claim 8, wherein the CRISPR nuclease polypeptide optionally comprises the arginine substitution of (a).
10. The modified CRISPR nuclease polypeptide according to claim 1, comprising, with respect to SEQ ID NO: 1, a nickase mutation at position H845, optionally H845A, and an arginine or lysine substitution at position I857, optionally I857R.
11. The manipulated CRISPR nuclease polypeptide according to any one of claims 1 to 10, wherein the CRISPR nuclease polypeptide comprises one or more mutations for reducing the PAM recognition stringency, and optionally, the one or more mutations are located at positions D61, A68, H494, L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345, and / or T1347 of SEQ ID NO:
1.
12. The one or more mutations described above (i) One or more arginine and / or lysine substitutions at positions D61, A68, H494, L1117, G1227, S1327, A1332, and / or T1347 of SEQ ID NO: 1, optionally arginine substitutions. (ii) One or more amino acid substitutions at positions D1144, S1145, E1228, R1343, and / or R1345 of SEQ ID NO: 1, optionally D1144L, S1145W, E1228Q, R1343P, R1345V, and / or R1345Q, or The manipulated CRISPR nuclease polypeptide according to claim 11, comprising a combination of (iii)(i) and (ii).
13. The following mutation combinations for sequence number 1: (i) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and A68R, (ii) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R, or (iii) The operated CRISPR nuclease polypeptide according to claim 12, comprising L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and H494R.
14. The manipulated CRISPR nuclease polypeptide according to any one of claims 1 to 13, wherein the manipulated CRISPR nuclease polypeptide comprises an amino acid sequence identical to that of SEQ ID NO: 1 by at least 90%.
15. The manipulated CRISPR nuclease polypeptide according to claim 14, wherein the manipulated CRISPR nuclease polypeptide comprises at least 95% the same amino acid sequence as SEQ ID NO:
1.
16. The manipulated CRISPR nuclease polypeptide according to claim 15, wherein the manipulated CRISPR nuclease polypeptide comprises at least 98% the same amino acid sequence as SEQ ID NO:
1.
17. The manipulated CRISPR nuclease polypeptide according to claim 1, wherein the manipulated CRISPR nuclease polypeptide is one of those listed in Table 4, Table 8, or Table 9.
18. An operated CRISPR nuclease polypeptide according to any one of claims 1 to 17, which is a fusion polypeptide further comprising one or more functional fragments.
19. The modified CRISPR nuclease polypeptide according to claim 18, wherein the one or more functional fragments comprise a nuclear localization signal (NLS), a peptide linker, or a combination thereof.
20. The manipulated CRISPR nuclease polypeptide according to claim 19, wherein the NLS is fused to the manipulated CRISPR nuclease polypeptide at its N-terminus and / or C-terminus.
21. A nucleic acid comprising a nucleotide sequence encoding an manipulated CRISPR nuclease polypeptide according to any one of claims 1 to 20.
22. The nucleic acid according to claim 21, wherein the nucleic acid is an expression vector, and the nucleotide sequence encoding the manipulated CRISPR nuclease polypeptide is operably linked to a promoter.
23. The nucleic acid according to claim 22, which is messenger RNA (mRNA).
24. A host cell comprising the nucleic acid described in claim 21 or 22.
25. It is a gene editing system, (a) A CRISPR nuclease polypeptide or a first nucleic acid encoding the CRISPR nuclease, wherein the CRISPR nuclease polypeptide contains an amino acid sequence that is at least 90% identical to the reference CRISPR nuclease described as Sequence ID No. 1, (b) A gene editing system comprising a guide RNA (gRNA) or a second nucleic acid encoding the gRNA, wherein the gRNA comprises a scaffold recognizable by an engineered CRISPR nuclease polypeptide and a spacer sequence specific to a target sequence in a genomic region of interest, and the target sequence is located upstream of a protospacer motif (PAM) and comprises the gRNA or the second nucleic acid encoding the gRNA.
26. The gene editing system according to claim 25, wherein the CRISPR nuclease polypeptide is an engineered CRISPR nuclease polypeptide according to any one of claims 1 to 20.
27. The gene editing system according to claim 25 or 26, wherein the scaffold comprises a nucleotide sequence identical to at least 85% of Sequence ID No.
2.
28. The gene editing system according to claim 27, wherein the scaffold includes Sequence ID No.
2.
29. The gene editing system according to claim 27, wherein the scaffold comprises one or more deletions, one or more nucleotide substitutions, or a combination thereof, compared to SEQ ID NO:
2.
30. The gene editing system according to any one of claims 25 to 29, wherein the PAM is 5'-NDR-3' or 5'-NGN-3', where N represents any nucleotide, D represents A, G, or T, and R represents G or A, and optionally the PAM is 5'-NRG-3' or 5'-NRR-3', where N and R are as defined herein, preferably the PAM is 5'-NGG-3', and N represents any nucleotide.
31. A gene editing method comprising delivering the gene editing system according to any one of claims 25 to 30 to a host cell to edit a genomic site targeted by the gRNA of the gene editing system.
32. The gene editing method according to claim 31, wherein the host cells are cultured in vitro.
33. The gene editing method according to claim 32, wherein the host cell is located within the target that requires gene editing.