Engineered cas polypeptides having improved editing performance and uses thereof
Engineered Cas12f polypeptides with targeted amino acid substitutions enhance indel activity and adaptability, addressing limitations in CRISPR/Cas systems for efficient gene editing within AAV vectors, thereby facilitating clinical applications.
Patent Information
- Application Number
- PCT/KR2025/010508
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-07-16
- Publication Date
- 2026-01-22
AI Technical Summary
Existing CRISPR/Cas systems, particularly Cas12f polypeptides, exhibit limited indel activity in eukaryotic cells and are constrained by AAV vector packaging limits, posing challenges for clinical applications.
Engineering Cas12f polypeptides with specific amino acid substitutions to enhance indel activity and adaptability to a broader range of PAM sequences, enabling efficient gene editing within AAV vectors.
The engineered Cas12f polypeptides demonstrate significantly improved indel activity in eukaryotic cells, facilitating effective gene editing and clinical applications by overcoming packaging constraints.
Smart Images

Figure KR2025010508_22012026_PF_FP_ABST
Abstract
Description
Engineered CAS polypeptides with improved editing performance and uses thereof
[0001] This application claims priority to Korean Patent Application No. 2024-0094066, filed July 16, 2024, the disclosure of which is incorporated herein by reference in its entirety.
[0002] Sequence list
[0003] This application includes a sequence listing, which has been submitted electronically in XML format and is incorporated herein by reference in its entirety. A copy of said sequence listing, dated July 16, 2025, is named KC24119.SEQ.xml and is 906 kilobytes in size.
[0004] The present invention relates to Cas polypeptides engineered to have improved gene editing performance compared to wild-type Cas polypeptides and to uses thereof.
[0005] The CRISPR / Cas system is a promising gene editing technology that edits and improves the genetic information of organisms by inserting, deleting, substituting, or modifying specific DNA in the genome, and is being actively researched. This system is classified into Class 1 and Class 2 based on the type, number, and domain structure of the Cas polypeptides contained in the effector complex. Class 1 includes Types I, III, and IV Cas nucleolytic proteins, while Class 2 includes Types II, V, and VI Cas nucleolytic proteins (Koonin et al., 2017; Makarova et al., 2020). The effector complex of the Class 2 CRISPR / Cas system is characterized by being composed of a single, large protein subdivided into multiple domains.
[0006] Recently, small Cas polypeptides, such as Cas12f (Cas14) and Cas12j (CasΦ), classified as class 2, type V CRISPR nucleases, have been identified from archaea and large bacteriophages (Harrington et al., 2018). These small effectors consist of approximately 400–700 amino acid residues and contain a RuvC domain. The discovery of the small protein Cas12f has raised expectations for the development of gene editing tools using adeno-associated virus (AAV) vectors with limited packaging capacity (approximately 4.7 kb). AAV vectors have been approved by the US FDA due to their safety, durability, and amenability to mass production. However, limitations have been reported, such as Cas12f effectors exhibiting only single-stranded DNA cleavage activity and extremely low indel activity in eukaryotic cells (Karvelis et al., 2020).
[0007] By engineering the natural guide RNA of archaeal Un1Cas12f1, we have demonstrated that the average indel activity in eukaryotic cells can be increased by more than 800-fold. This successfully transformed small Cas polypeptides, such as the Cas12f effector, into efficient gene editing tools (Kim, DY et al., 2021). Most CRISPR / Cas systems (e.g., CRISPR / SpCas9) are too large to be delivered via AAV vectors, which limit their payloads to less than 4.7 kb, making clinical applications challenging. In contrast, the Cas12f polypeptide and engineered guide RNA (TaRGET editing system) developed by the inventors are small enough to be packaged within an AAV vector along with other components, making them advantageous for clinical applications.
[0008] Additionally, the present inventors developed a variant of the Cas12f polypeptide capable of recognizing an extended range of PAM sequences and implemented it as an adenine base editor (ABE) system (Kim, DY et al., 2023).
[0009] From the perspective of clinical application of the CRISPR / Cas system, reducing off-target effects and improving gene editing efficiency remain critical challenges.
[0010] [Prior Art Literature]
[0011] [Patent Document]
[0012] (Patent Document 1) KR 10-2023-0076820 A
[0013] (Patent Document 2) WO 2022-075816 A1
[0014] [Non-patent literature]
[0015] (Non-patent Document 1) Plagens et al., DNA and RNA interference mechanisms by CRISPR-Cas surveillance complexes. FEMS Microbiol Rev. 39(3):442-463(2015)
[0016] (Non-patent Document 2) Nishimasu, H., & Nureki, O. Structures and mechanisms of CRISPR RNA-guided effector nucleases. Curr Opin Struct Biol, 43, 68-78 (2017)
[0017] (Non-patent Document 3) Koonin, EV. et al., Mobile genetic elements and evolution of CRISPR-Cas system; All the way there and back, Genome Biol. Evol., Vol. 9, No. 10, 2812-2825(2017)
[0018] (비특허문헌 4) Makarova, KS. et al., Evolutionary classification of the CRISPR-Cas system: a burst of class 2 and derived variants, Nat. Rev. Microbiol., Vol. 18, 67-83(2020)
[0019] (비특허문헌 5) Harrington, LB. et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science, Vol. 362, 839-842(2018)
[0020] (비특허문헌 6) Karvelis, T. et al., PAM recognition by miniature CRISPR-Cas12f nucleases triggers programmable double-stranded DNA target cleavage, Nucleic Acids Research, Vol. 48, No. 9, 5016-5023(2020)
[0021] (비특허문헌 7) Kim, D. Y. et al., Efficient CRISPR editing with a hypercompact Cas12f1 and engineered guide RNAs delivered by adeno-associated virus. Nat. Biotechnol.40, 94-102(2021)
[0022] (비특허문헌 8) Kim, D. Y. et al., Author Correction: Hypercompact adenine base editors based on a Cas12f variant guided by engineered RNA. Nature Chemical Biology, 19(3), 389-389(2023)
[0023] The present invention aims to solve one or more of the problems of the above-mentioned prior art.
[0024] An object of the present invention is to provide engineered small Cas polypeptides having improved editing performance.
[0025] Another object of the present invention is to provide engineered small Cas polypeptides that can be programmed to exhibit improved editing performance through site-specific mutations.
[0026] Another object of the present invention is to provide a nucleic acid encoding an engineered small Cas polypeptide.
[0027] Another object of the present invention is to provide a CRISPR / Cas system, composition, kit, or delivery system comprising an engineered small Cas polypeptide or a nucleic acid encoding the same.
[0028] Another object of the present invention is to provide a vector or vector system comprising a nucleic acid encoding an engineered small Cas polypeptide.
[0029] Another object of the present invention is to provide a cell, cell line or organism comprising an engineered small Cas polypeptide, a CRISPR / Cas system, a nucleic acid, a vector, a vector system or a delivery system.
[0030] Another object of the present invention is to provide a method for modulating, modifying or editing one or more target nucleic acids using engineered small Cas polypeptides.
[0031] The purpose of the present invention is not limited to the purposes mentioned above. The purpose of the present invention will become more apparent from the following description and may be realized by the means and combinations thereof set forth in the claims.
[0032] A representative configuration of the present invention to achieve the above purpose is as follows.
[0033] In one aspect of the present invention, an engineered Cas polypeptide is provided comprising an amino acid sequence that is modified compared to a wild-type Cas12f1 polypeptide. Specifically, the modified amino acid sequence comprises one or more amino acid substitutions compared to the amino acid sequence of a wild-type Cas12f1 polypeptide, wherein the one or more amino acid substitutions are based on SEQ ID NO: 1.
[0034] (i) I365, A368, H399, A402, A406, C433, I435, A436 and D437;
[0035] (ii) I65, E125, I126, L133, K135, I140, Y141, K155, S164, H167, D171, T175, E179, L180, D216, L222, Q229, T231, R250, F266, E269, K305, Q312, T313, S322, G325, M331, L332, D337, I341, P347, K358, S359, L361, F369, K407, T414, E418, S420, E421, C433, I435, I440, N442, Q448, N451, R457, E459, K479, H508, K530 and E555;
[0036] (iii) K31, N32, T33, K39, N47, P221, L222, K224, Q225, Q229, Y230, T231, E234, N237, K245, I246, R250, Q252, K254, K255, E256, I257, K259, K265, F266, D267, E269, Q270, Q272, I278, L282, K288, K295, D296, E297, K304, K305, N308, D310, Q312, T313, K319, G321, S322, K323, G325, M331, L332, N333 and D337;
[0037] (iv) S48, E50, K53, D57, E58, N61, I65, L67, N70, D72, T89 and A120;
[0038] (v) Q124, E125, I126, L133, K135, I140, Y141, I154, K155, A162, S163, S164, E166, H167, D171, T175, R176, A178, E179, L180, I186, L190, L200, N205, M206, G209, S215, D216 and F218;
[0039] (vi) I341, D342, K343, D346 and P347;
[0040] (vii) T502, K505, H508, L509, N511, E516, K519, K527, K530 and E535; and
[0041] (viii) comprises an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of K358, S359, L361, I365, A368, F369, K385, H399, A402, A406, K407, T414, E418, S420, E421, C433, I435, A436, D437, I440, N442, T446, Q448, N451, S454, R457, E459, I465, K479, K501, L543, L550, K551, S552, E555 and E556.
[0042] In some implementations, one or more amino acid substitutions are
[0043] (i) I365V, A368S, H399L, A402S, A406S, C433K, I435V, I435L, A436T and D437N;
[0044] (ii) I65S, I65E, E125S, E125N, E125M, I126M, L133V, L133S, L133M, L133C, K135R, I140V, I140M, Y141F, K155R, K155Y, S164D, H167K, H167R, D171E, D171N, D171R, T175R, E179A, L180N, D216E, L222F, Q229G, Q229K, T231E, T231K, R250K, F266C, E269D, K305R, Q312P, Q312K, T313I, T313V, S322K, G325C, M331F, L332V, D337R, I341K, P347K, K358R, S359N, S359H, L361A, F369N, F369D, F369G, F369W, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435L, I440L, N442H, N442M, Q448V, Q448N, N451S, R457K, E459K, K479M, H508C, H508Y, K530R, more
[0045] (iii) K31R, N32Q, T33S, K39R, N47Q, P221T, P221S, P221C, P221G, P221A, P221V, P221I, P221M, P221F, L222F, L222I, K224R, Q225R, Q225H, Q225D, Q225E, Q225N, Q225Y, Q225K, Q225F, Q225C, Q225V, Q225P, Q225S, Q225T, Q225A, Q229G, Q229K, Q229N, Q229R, Q229H, Q229C, Q229P, Y230E, Y230R, Y230H, Y230K, Y230D, Y230N, Y230Q, Y230F, Y230C, Y230W, T231E, T231K, T231P, E234D, N237Q, K245R, I246L, R250K, Q252N, K254R, K255R, E256D, I257L, K259R, K265R, F266C, D267E, E269D, Q270R, Q272N, Q272R, I278L, L282I, K288R, K295R, D296E, E297D, K304R, K305R, N308Q, D310E, Q312K, Q312P, Q312N T313I, T313V, K319R, G321K, S322K, K323R, G325C, M331F, L332V, N333Q and D337R;
[0046] (iv) S48T, E50D, K53R, D57E, E58D, N61Q, I65E, I65L, I65S, L67I, N70Q, D72E, T89S 및 A120T;
[0047] (v) Q124N, E125M, E125N, E125S, I126M, L133C, L133M, L133S, L133V, K135R, I140M, I140V, Y141F, I154L, K155R, K155Y, A162D, A162Q, A162P, A162V, S163R, S163K, S163N, S163M, S163F, S163Y, S163G, S163P, S163A, S163L, S164D, E166D, H167K, H167R, D171N, D171R, D171E, T175R, R176K, A178S, E179A, L180N, I186L, L190I, L200F, N205H, M206H, G209R, S215H, D216E and F218H;
[0048] (vi) I341K, D342E, K343R, D346N and P347K;
[0049] (vii) T502S, K505R, H508C, H508Y, L509I, N511Q, E516D, K519R, K527R, K530R and E535D; and
[0050] (viii) K358R, S359H, S359N, L361A, I365V, A368S, F369D, F369G, F369N, F369W, K385R, H399L, A402S, A406S, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435V, I435L, A436T, D437N, I440L, N442H, N442M, T446S, Q448N, Q448V, N451S, S454G, R457K, E459K, I465R, K479M, K501R, It may include one or more selected from the group consisting of L543I, L550I, K551R, S552T, E555D, E555R and E556D.
[0051] In some embodiments, one or more amino acid substitutions are (ii) I65, E125, I126, L133, K135, I140, Y141, K155, S164, H167, D171, T175, E179, L180, D216, L222, Q229, T231, R250, F266, E269, K305, Q312, T313, S322, G325, M331, L332, D337, I341, P347, K358, S359, L361, F369, K407, T414, E418, S420, E421, C433, I435, I440, N442, Q448, N451, It may include an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of R457, E459, K479, H508, K530, and E555.
[0052] In some embodiments, one or more amino acid substitutions are (ii) I65S, I65E, E125S, E125N, E125M, I126M, L133V, L133S, L133M, L133C, K135R, I140V, I140M, Y141F, K155R, K155Y, S164D, H167K, H167R, D171E, D171N, D171R, T175R, E179A, L180N, D216E, L222F, Q229G, Q229K, T231E, T231K, R250K, F266C, E269D, K305R, Q312P, Q312K, T313I, It may include at least one selected from the group consisting of T313V, S322K, G325C, M331F, L332V, D337R, I341K, P347K, K358R, S359N, S359H, L361A, F369N, F369D, F369G, F369W, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435L, I440L, N442H, N442M, Q448V, Q448N, N451S, R457K, E459K, K479M, H508C, H508Y, K530R, and E555R.
[0053] In some embodiments, the one or more amino acid substitutions may comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of (i) I365, A368, H399, A402, A406, C433, I435, A436, and D437.
[0054] In some embodiments, the one or more amino acid substitutions may comprise one or more selected from the group consisting of (i) I365V, A368S, H399L, A402S, A406S, C433K, I435V, I435L, A436T, and D437N.
[0055] In some embodiments, the one or more amino acid substitutions may comprise amino acid substitutions at residues corresponding to any one selected from I365 / A368, H399 / A402 / A406, C433 / I435 / A436 / D437, and any combination thereof.
[0056] In some embodiments, the one or more amino acid substitutions may include an amino acid substitution selected from I365V / A368S; H399L / A402S / A406S; C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N; and any combination thereof.
[0057] In some embodiments, one or more amino acid substitutions may include I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N.
[0058] In some embodiments, one or more amino acid substitutions may include I365V / A368S; H399L / A402S / A406S; and C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N.
[0059] In some embodiments, the one or more amino acid substitutions may further comprise one or more other amino acid substitutions in one or more domains selected from the group consisting of WED, ZF, REC, TNB, Linker, and RuvC domains.
[0060] In some implementations, one or more amino acid substitutions are
[0061] (iii) K31, N32, T33, K39, N47, P221, L222, K224, Q225, Q229, Y230, T231, E234, N237, K245, I246, R250, Q252, K254, K255, E256, I257, K259, K265, F266, D267, E269, Q270, Q272, I278, L282, K288, K295, D296, E297, K304, K305, N308, D310, Q312, T313, K319, G321, S322, K323, G325, M331, L332, N333 D337;
[0062] (iv) S48, E50, K53, D57, E58, N61, I65, L67, N70, D72, T89 및 A120;
[0063] (v) Q124, E125, I126, L133, K135, I140, Y141, I154, K155, A162, S163, S164, E166, H167, D171, T175, R176, A178, E179, L180, I186, L190, L200, N205, M206, G209, S215, D216 및 F218;
[0064] (vi) I341, D342, K343, D346 및 P347;
[0065] (vii) T502, K505, H508, L509, N511, E516, K519, K527, K530 및 E535; 및
[0066] (viii) may further comprise a substitution at a residue corresponding to one or more residues selected from the group consisting of K358, S359, L361, I365, A368, F369, K385, H399, A402, A406, K407, T414, E418, S420, E421, C433, I435, A436, D437, I440, N442, T446, Q448, N451, S454, R457, E459, I465, K479, K501, L543, L550, K551, S552, E555 and E556.
[0067] In some implementations, one or more amino acid substitutions are
[0068] (iii) K31R, N32Q, T33S, K39R, N47Q, P221T, P221S, P221C, P221G, P221A, P221V, P221I, P221M, P221F, L222F, L222I, K224R, Q225R, Q225H, Q225D, Q225E, Q225N, Q225Y, Q225K, Q225F, Q225C, Q225V, Q225P, Q225S, Q225T, Q225A, Q229G, Q229K, Q229N, Q229R, Q229H, Q229C, Q229P, Y230E, Y230R, Y230H, Y230K, Y230D, Y230N, Y230Q, Y230F, Y230C, Y230W, T231E, T231K, T231P, E234D, N237Q, K245R, I246L, R250K, Q252N, K254R, K255R, E256D, I257L, K259R, K265R, F266C, D267E, E269D, Q270R, Q272N, Q272R, I278L, L282I, K288R, K295R, D296E, E297D, K304R, K305R, N308Q, D310E, Q312K, Q312P, Q312N T313I, T313V, K319R, G321K, S322K, K323R, G325C, M331F, L332V, N333Q and D337R;
[0069] (iv) S48T, E50D, K53R, D57E, E58D, N61Q, I65E, I65L, I65S, L67I, N70Q, D72E, T89S 및 A120T;
[0070] (v) Q124N, E125M, E125N, E125S, I126M, L133C, L133M, L133S, L133V, K135R, I140M, I140V, Y141F, I154L, K155R, K155Y, A162D, A162Q, A162P, A162V, S163R, S163K, S163N, S163M, S163F, S163Y, S163G, S163P, S163A, S163L, S164D, E166D, H167K, H167R, D171N, D171R, D171E, T175R, R176K, A178S, E179A, L180N, I186L, L190I, L200F, N205H, M206H, G209R, S215H, D216E and F218H;
[0071] (vi) I341K, D342E, K343R, D346N and P347K;
[0072] (vii) T502S, K505R, H508C, H508Y, L509I, N511Q, E516D, K519R, K527R, K530R and E535D; and
[0073] (viii) K358R, S359H, S359N, L361A, I365V, A368S, F369D, F369G, F369N, F369W, K385R, H399L, A402S, A406S, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435V, I435L, A436T, D437N, I440L, N442H, N442M, T446S, Q448N, Q448V, N451S, S454G, R457K, E459K, I465R, K479M, K501R, It may additionally include one or more selected from the group consisting of L543I, L550I, K551R, S552T, E555D, E555R and E556D.
[0074] In some embodiments, the one or more amino acid substitutions may further comprise a substitution at a residue corresponding to one or more residues selected from the group consisting of K135, Y141, K305, and M331.
[0075] In some embodiments, one or more amino acid substitutions may additionally include K135R / Y141F / K305R / M331F.
[0076] In some embodiments, the one or more amino acid substitutions may further comprise a substitution at a residue corresponding to one or more residues selected from the group consisting of L67, Q124, K155, and E556.
[0077] In some embodiments, one or more amino acid substitutions may additionally include L67I / Q124N / K155R / E556D.
[0078] In some embodiments, the one or more amino acid substitutions may further comprise a substitution at a residue corresponding to one or more residues selected from the group consisting of T175, G209, K358, and S454.
[0079] In some embodiments, one or more amino acid substitutions may additionally include T175R / G209R / K358R / S454G.
[0080] In some embodiments, the one or more amino acid substitutions may comprise any one selected from the group consisting of amino acid substitutions listed in Table 9.
[0081] In some embodiments, the engineered Cas polypeptide can comprise, consist of, or consist essentially of a sequence having at least 80% identity to an amino acid sequence selected from the group consisting of amino acid sequences listed in Tables 10 to 19.
[0082] In some embodiments, the engineered Cas polypeptide can comprise, consist of, or consist essentially of an amino acid sequence selected from the group consisting of SEQ ID NOs: 201 to 384 and 410 to 606.
[0083] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions:
[0084] L67I / Q124N / K135R / Y141F / K155R / K305R / M331F / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K530R / E556D (M1010);
[0085] L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1120);
[0086] L67I / Q124N / K135R / Y141F / K155R / N205H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1223);
[0087] L67I / Q124N / K135R / Y141F / K155R / M206H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1224);
[0088] L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1225);
[0089] Q272R / N205H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1231);
[0090] Q272R / M206H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1232);
[0091] Q272R / Q270R / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1233);
[0092] Q225P / Q229C / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM622);
[0093] Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM625);
[0094] Q229R / Y230W / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM628);
[0095] K230H / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM638);
[0096] K230C / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM646);
[0097] S163R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM674);
[0098] S163K / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM676);
[0099] T175R / G209R / K358R / S454G / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM714);
[0100] D171R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM715);
[0101] T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM716);
[0102] I465R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM722);
[0103] G209R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM724);
[0104] D171R / T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM727);
[0105] T175R / G209R / S454G / K385R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM728);
[0106] D171R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM729);
[0107] T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM730);
[0108] I465R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM731);
[0109] G209R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM732);
[0110] D171R / T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM733);
[0111] K358R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM734); 또는
[0112] T175R / G209R / K358R / S454G / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM735).
[0113] In some implementations, the one or more substitutions may comprise between 2 and 40 amino acid substitutions.
[0114] In some embodiments, the engineered Cas polypeptide can comprise, consist of, or consist essentially of an amino acid sequence having at least 80% identity to any one amino acid sequence selected from the group consisting of SEQ ID NOs: 201 to 207, 284 to 384, and 410 to 606.
[0115] In some embodiments, the engineered Cas polypeptide can comprise, consist of, or consist essentially of an amino acid sequence having at least 80% identity to any one amino acid sequence selected from the group consisting of SEQ ID NOs: 380 to 384 and 470 to 472.
[0116] In some embodiments, the engineered Cas polypeptide may be part of a fusion protein.
[0117] In some embodiments, a fusion protein comprising an engineered Cas polypeptide of the present invention is provided.
[0118] In some embodiments, the engineered Cas polypeptides exhibit improved indel efficiency, PAM sequence recognition range, and / or editing window.
[0119] In one aspect of the present invention, based on sequence number 1
[0120] (i) I365 / A368,
[0121] (ii) H399 / A402 / A406,
[0122] (iii) C433 / I435 / A436 / D437,
[0123] (iv) K135 / Y141 / K305 / M331,
[0124] (v) L67 / Q124 / K155 / E556,
[0125] (vi) T175 / G209 / K358 / S454, or
[0126] (vii) Engineered Cas polypeptides are provided comprising amino acid substitutions at residues corresponding to any combination of these.
[0127] In some implementations, based on sequence number 1
[0128] (i) I365V / A368S,
[0129] (ii) H399L / A402S / A406S,
[0130] (iii) C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N,
[0131] (iv) K135R / Y141F / K305R / M331F,
[0132] (v) L67I / Q124N / K155R / E556D,
[0133] (vi) T175R / G209R / K358R / S454G, or
[0134] (vii) Engineered Cas polypeptides are provided comprising amino acid substitutions corresponding to any combination of these.
[0135] In some embodiments, an isolated nucleic acid encoding an engineered Cas polypeptide of the present invention is provided.
[0136] In some embodiments, a vector is provided comprising a nucleic acid sequence encoding an engineered Cas polypeptide of the present invention.
[0137] In some implementations, the vector may be a non-viral vector or a viral vector.
[0138] In some embodiments, the non-viral vector may be plasmid DNA, mRNA (transcript), or PCR amplicon.
[0139] In some embodiments, the viral vector may be selected from the group consisting of a retrovirus, a lentivirus, an adenovirus, an adeno-associated virus, a vaccinia virus, a poxvirus, and a herpes simplex virus.
[0140] In one aspect of the present invention, a vector system is provided comprising one or more vectors, wherein the one or more vectors comprise (a) a first nucleic acid construct operably linked to a nucleotide encoding an engineered Cas polypeptide of the present invention; and (b) a second nucleic acid construct operably linked to a nucleotide encoding a first guide nucleic acid comprising a guide sequence complementary to a target sequence, wherein the nucleic acid constructs of (a) and (b) are contained in the same or different vectors.
[0141] In some embodiments, the vector system further comprises a third nucleic acid construct operably linked to a nucleotide encoding a second guide nucleic acid comprising a guide sequence complementary to the target sequence, wherein the second guide nucleic acid targets the same or different target sequence as the first guide nucleic acid, and the third nucleic acid construct can be contained in a vector that is the same or different from the first nucleic acid construct or the second nucleic acid construct.
[0142] In some embodiments, the first guide nucleic acid may comprise a scaffold region sequence of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 151 to 153.
[0143] In some embodiments, the first guide nucleic acid may additionally comprise a U-rich tail at the 3' end of the guide sequence. Specifically, the sequence of the U-rich tail is represented as 5'-(UmV)nUo-3', wherein V is each independently A, C, or G, m and o are integers ranging from 1 to 20, and n may be an integer ranging from 0 to 5.
[0144] In some embodiments, the first guide nucleic acid may comprise a nucleic acid sequence selected from the group consisting of SEQ ID NO: 151 to SEQ ID NO: 153.
[0145] In one aspect of the present invention, a system is provided comprising an engineered Cas polypeptide of the present invention or a nucleic acid encoding the polypeptide; and one or more guide nucleic acids or nucleic acids encoding the guide nucleic acids.
[0146] In some embodiments, the guide nucleic acid comprises i) a first sequence comprising the nucleotide sequence 5'-GGAAUGCAAC-3' and ii) a second sequence that hybridizes to a target sequence of a target nucleic acid, wherein the first sequence may be located 5' to the second sequence.
[0147] In some embodiments, the guide nucleic acid is an engineered guide nucleic acid, wherein the engineered guide nucleic acid comprises, in a 5' to 3' direction, first, second, third and fourth stem-loop regions, a tracrRNA-crRNA complementary region and a guide sequence, compared to an unmodified guide RNA, (1) deletion of part or all of the first stem-loop region; (2) deletion of part or all of the second stem-loop region; (b) deletion of part or all of the tracrRNA-crRNA complementary region; (c) substitution of one or more Uracils (U) with A, G or C when three or more, four or more or five or more consecutive Us are present within the tracrRNA-crRNA complementary region; and (d) addition of a U-rich tail to the 3'-end of the crRNA sequence (wherein the sequence of the U-rich tail is represented as 5'-(UmV)nUo-3', wherein V is each independently A, C or G, m and o are integers ranging from 1 to 20, and n is an integer ranging from 0 to 5).
[0148] In some embodiments, the guide nucleic acid may comprise a scaffold region sequence of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 13, SEQ ID NO: 149 to SEQ ID NO: 186, and SEQ ID NO: 196.
[0149] In some embodiments, the guide nucleic acid may comprise a scaffold region sequence of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 151 to 153.
[0150] In some embodiments, the guide nucleic acid may comprise a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 156, 183, and 196.
[0151] In one aspect of the present invention, an engineered Cas polypeptide
[0152] L67I / Q124N / K135R / Y141F / K155R / K305R / M331F / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K530R / E556D (M1010);
[0153] L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1120);
[0154] L67I / Q124N / K135R / Y141F / K155R / N205H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1223);
[0155] L67I / Q124N / K135R / Y141F / K155R / M206H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1224);
[0156] L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1225);
[0157] Q272R / N205H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1231);
[0158] Q272R / M206H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1232);
[0159] Q272R / Q270R / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1233);
[0160] Q225P / Q229C / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM622);
[0161] Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM625);
[0162] Q229R / Y230W / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM628);
[0163] K230H / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM638);
[0164] K230C / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM646);
[0165] S163R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM674);
[0166] S163K / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM676);
[0167] T175R / G209R / K358R / S454G / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM714);
[0168] D171R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM715);
[0169] T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM716);
[0170] I465R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM722);
[0171] G209R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM724);
[0172] D171R / T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM727);
[0173] T175R / G209R / S454G / K385R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM728);
[0174] D171R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM729);
[0175] T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM730);
[0176] I465R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM731);
[0177] G209R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM732);
[0178] D171R / T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM733);
[0179] K358R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM734); 또는
[0180] T175R / G209R / K358R / S454G / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM735) 포함하고,
[0181] An editing system is provided, wherein the guide nucleic acid comprises a scaffold region sequence of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 151 to 153.
[0182] In some implementations, the system comprises two or more guide nucleic acids, each of which can independently comprise a spacer sequence complementary to the same or different target sequence.
[0183] In one aspect of the present invention, a composition is provided comprising an engineered Cas polypeptide of the present invention or a nucleic acid encoding the polypeptide, and one or more guide nucleic acids comprising a guide sequence complementary to a target sequence or a nucleic acid encoding the guide nucleic acid.
[0184] In some embodiments, the engineered Cas polypeptide and guide nucleic acid can form a ribonucleoprotein complex.
[0185] In one aspect of the present invention, an isolated cell comprising an engineered Cas polypeptide, vector, vector system, system, or composition described herein is provided.
[0186] In some embodiments, the cell may be a eukaryotic cell.
[0187] In one aspect of the present invention, a method of regulating or editing a target nucleic acid in a cell is provided, comprising contacting the cell with an engineered Cas polypeptide, vector, vector system, system, or composition described herein.
[0188] In some embodiments, the cell may be a eukaryotic cell.
[0189] In some embodiments, the methods of the present invention can be performed in vitro or ex vivo.
[0190] In some embodiments, a fusion protein comprising an engineered Cas polypeptide of the present invention and an effector domain may be provided.
[0191] The present invention provides engineered small Cas polypeptides with enhanced editing performance. The invention is based, at least in part, on the identification of amino acid positions and modifications in Cas12f polypeptides, such as UnCas12f1 or CWCas12f1, that can improve editing performance, such as indel efficiency, expansion of recognizable PAM sequences, and expansion of the editing window. Accordingly, the present invention provides programmable Cas12f protein variants that exhibit improved editing performance compared to conventional CRISPR / Cas12f editing systems. It has been confirmed that Cas12f variants of the present invention comprising one or more amino acid substitutions at specific residue positions exhibit higher indel efficiency compared to the wild type. The Cas12f variants of the present invention can exhibit improved editing performance for a wider range of target sequences by recognizing a wider range of PAM sequences compared to the wild type.
[0192] In addition, the ultra-small CRISPR / Cas12f system ("TaRGET") comprising the engineered Cas12f polypeptide of the present invention and the engineered guide RNA described herein is small enough to be packaged into AAV, a validated vehicle for in vivo delivery, and thus can be utilized as a therapeutic for various diseases.
[0193] Figures 1a and 1b illustrate the amino acid sequences of UnCas12f1 and CWCas12f1, respectively.
[0194] Figure 2 depicts modification sites MS1 to MS5 in an engineered guide RNA (engineered gRNA) for the Cas12f system (MS, modification site).
[0195] Figures 3a to 3c are graphs showing the indel efficiency for target sequences containing the canonical PAM sequence of the M1120 variant. Figure 3a is a graph showing the indel efficiency of CWCas12f1 and the M1120 variant for 31 target sequences. Figure 3b is a graph showing the indel efficiency of CWCas12f1 and the M1220 variant for the F3, F5, and R10 targets, and Figure 3c is a graph showing the indel efficiency of the M1120 variant and the M1223 to M1225 variants for the R10 target sequence.
[0196] Figure 4 is a graph showing the indel efficiency for target sequences containing non-circular PAM sequences of the HS1 / 5 / 8 variants and the M1120 variant. NT represents the untreated group.
[0197] Figures 5a and 5b show results showing that the M1120 variant and the M1223 to M1225 variants recognize non-circular PAM sequences.
[0198] Figures 6a and 6b show results showing that the base editing module (TaRGET-ABE) constructed using the M1120 variant has an expanded base editing window in various targets (DY10 and Inter22).
[0199] Figures 7a and 7b are graphs showing the excellent indel efficiency of exemplary variants of the present invention.
[0200] The detailed description of the present invention described below will be described with reference to specific embodiments in which the present invention may be practiced (if any), with reference to the drawings, but the present invention is not limited thereto, but is defined only by the appended claims to the full scope equivalent to or equivalent to what the claims describe. It should be understood that the various embodiments / embodiments of the present invention, while different from each other, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be changed from one embodiment / embodiment to another, or multiple embodiments / embodiments may be combined, without departing from the spirit and scope of the present invention. Technical and scientific terms used herein, unless otherwise defined, have the same meaning as commonly used in the art to which the present invention belongs. For the purpose of interpreting this specification, the following definitions will apply, and terms expressed in the singular should be construed to also refer to the plural (i.e., at least one), unless the context makes it inappropriate.
[0201] I. Definition
[0202] The term "about" as used herein refers to the conventional error range known to those skilled in the art for the value modified by the term. Furthermore, unless otherwise specified, all numbers, values, and / or expressions expressing ingredients, conditions, compositions, amounts, and so forth used herein are to be understood as being modified by the term "about" because such numbers are approximations that inherently reflect, among other things, the various uncertainties of measurement that arise in obtaining such values. Specifically, "about" can mean an amount, level, value, number, frequency, percent, dimension, size, amount, weight, or length that varies by ±30%, ±25%, ±20%, ±15%, ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, or ±1% with respect to the reference amount, level, value, number, frequency, percent, dimension, size, amount, weight, or length.
[0203] In this specification, when a composition is said to "contain" or "include" a component, this does not exclude other components unless specifically stated to the contrary, but rather means that it may further include other components.
[0204] The terms "polynucleotide" and "nucleic acid" are used interchangeably herein and refer to a polymer of nucleotides of any length, such as ribonucleotides or deoxyribonucleotides, in single- or double-stranded form. Unless otherwise specified, this includes known analogs of naturally occurring nucleotides that can function in a manner similar to naturally occurring nucleotides. It also includes nucleic acid-like structures with synthetic backbones, as well as amplified products. For example, polynucleotides may include natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromuridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), inserted The nucleic acid may comprise bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages) in the form of single-stranded or multi-stranded DNA or RNA, genomic DNA, cDNA, or DNA-RNA hybrids. The terms "polynucleotide" and "nucleic acid" are to be understood to encompass single-stranded (e.g., sense or antisense) and double-stranded forms, as applicable to the embodiments described.
[0205] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein and refer to a polymer of amino acids of any length, which may include genetically encoded or non-genetically encoded amino acids, chemically or biochemically synthesized, modified, or derivatized amino acids, and polypeptides having modified peptide backbones.
[0206] The term "oligonucleotide" refers to a series of nucleotides or their analogs. Oligonucleotides can be obtained by a number of methods, including, for example, chemical synthesis, restriction enzyme digestion, or PCR. As will be appreciated by those skilled in the art, the length (i.e., the number of nucleotides) of an oligonucleotide will often vary depending on the intended function or use of the oligonucleotide. Throughout this specification, when an oligonucleotide is represented by a series of letters (in the case of DNA, selected from the four base letters A, C, G, and T, representing adenosine, cytidine, guanosine, and thymidine, respectively; in the case of RNA, selected from the four base letters A, C, G, and U, representing adenosine, cytidine, guanosine, and uracil, respectively), this is represented in 5' to 3' order from left to right. In certain embodiments, the sequence of an oligonucleotide comprises one or more degenerate residues.
[0207] The terms "wild type" or "naturally occurring" are used interchangeably herein and refer to a nucleic acid, protein, cell, or organism that is found in nature.
[0208] The terms "parent polypeptide," "parent sequence," "reference polypeptide," or "reference sequence" are used interchangeably herein and may refer to a native polypeptide (e.g., a wild-type Cas12f1 polypeptide) that has been modified to generate an engineered Cas polypeptide of the invention. In some embodiments, the parent sequence may refer to the sequence of UnCas12f1 (SEQ ID NO: 5) or CWCas12f1 (SEQ ID NO: 1).
[0209] The term "nuclease" refers to a polypeptide capable of cleaving a phosphodiester bond between nucleotide subunits of a nucleic acid, and the term "endonuclease" refers to a polypeptide capable of cleaving a phosphodiester bond within a polynucleotide chain.
[0210] The terms "engineered," "mutated," or "modified" are used interchangeably herein and refer to a protein or polypeptide that has been altered such that one or more amino acid residues differ from the original amino acid sequence (e.g., a parent sequence of SEQ ID NO: 1 or 5). In some embodiments, an engineered protein or polypeptide may be altered to include truncation, deletion, insertion, substitution, or a combination thereof at one or more amino acid residues to improve editing performance, such as indel efficiency, PAM sequence recognition range, or editing window.
[0211] The term "variant" refers to a protein or polypeptide that has been engineered, mutated, or altered. Examples of such variants include engineered versions of Cas polypeptides with alterations in one or more amino acid residues at positions critical to editing performance, such as indel efficiency, PAM sequence recognition range, or editing window.
[0212] The term "performance" refers to biological activity. In some embodiments, performance includes the editing performance of a CRISPR / Cas12f system complexed with a guide RNA toward a target nucleic acid, such as targeting ability, indel efficiency, PAM sequence recognition range, editing window, etc.
[0213] The term "indel" refers to the insertion or deletion of one or more nucleotide bases in a nucleic acid. Such insertions or deletions can cause frame-shift mutations in the coding region of a gene. Indel efficiency can be calculated using methods commonly used in the industry. For example, indel efficiency can be calculated by measuring the frequency of truncation, deletion, insertion, substitution, or a combination thereof that occurs at the target site due to gene editing. After extracting genomic DNA from edited cells, the target site is amplified by PCR and sequenced through next-generation sequencing (NGS). In the analyzed sequence, the number of reads containing indels (edited sequences) compared to the wild-type sequence is divided by the total number of reads to calculate the indel efficiency (%). For example, indel efficiency can be calculated as follows: Indel efficiency (%) = (Number of reads containing indels / Total number of reads) χ 100. Indel analysis can be performed with tools such as CRISPResso2, ICE, and TIDE, but is not limited thereto.
[0214] In some embodiments, the engineered Cas polypeptides of the present invention having enhanced editing performance may refer to polypeptides having improved indel efficiency compared to a wild-type Cas polypeptide. Specifically, the engineered Cas polypeptides of the present invention may exhibit improved indel efficiency for a target nucleic acid compared to a wild-type or naturally occurring parent polypeptide when expressed and purified under the same conditions. Improved indel efficiency may refer to an improvement in the targeting ability of the Cas polypeptide, its ability to bind to a target nucleic acid, its ability to maintain binding to a target nucleic acid, its cleavage efficiency, or any combination thereof. In some embodiments, the engineered Cas polypeptides of the invention can exhibit an indel efficiency that is increased by at least about 1.1-fold, about 1.2-fold, about 1.3-fold, about 1.4-fold, about 1.5-fold, about 1.6-fold, about 1.7-fold, about 1.8-fold, about 1.9-fold, about 2.0-fold, about 2.5-fold, about 3.0-fold, about 3.5-fold, about 4.0-fold, about 4.5-fold, about 5.0-fold, about 5.5-fold, about 6.0-fold, about 6.5-fold, about 7.0-fold, about 7.5-fold, about 8.0, or about 8.5-fold or more compared to a wild-type polypeptide.
[0215] The term "protospacer-adjacent motif" or "PAM" refers to a unique sequence determined by a Cas protein that is adjacent to a target sequence to which a complex comprising an effector (e.g., an (endo)nuclease) and a guide RNA binds. In some embodiments, the PAM may be a sequence required for the enzymatic activity of the (endo)nuclease. When the target nucleic acid is a double-stranded molecule, one strand may be a sequence referred to as the "PAM strand" that includes the target sequence adjacent to the PAM, and the other complementary strand may be a sequence referred to as the "non-PAM strand." As used herein, the term "adjacent" may include cases where the guide RNA of the complex specifically binds, interacts, or associates with the target sequence immediately adjacent to the PAM. There may be no nucleotides between the target sequence and the PAM. Additionally, the term "contiguous" includes cases where there are some nucleotides (e.g., 1, 2, 3, 4, or 5 nucleotides) between the target sequence to which the guide RNA binds and the PAM. In some embodiments, the PAM sequence recognized by the wild-type Cas12f of the present invention can be a 5'-TTTR-3' sequence, wherein R can be A or G. In some embodiments, the engineered Cas12f of the present invention can be a variant having an expanded PAM sequence recognized by the wild-type Cas12f. For example, the engineered Cas12f having an expanded recognizable PAM sequence can be a variant that recognizes a non-circular PAM sequence in addition to the circular PAM sequence, thereby improving the editing performance of the Cas polypeptide.
[0216] In some embodiments, the engineered Cas polypeptide of the present invention can recognize a non-canonical PAM sequence, such as 5'-CCTG-3', 5'-GTTG-3', 5'-TAGG-3', 5'-TTCG-3', 5'-ATAG-3', 5'-TCTG-3', 5'-ATTA-3', 5'-TTTA-3', 5'-CTTA-3', 5'-TTTG-3', 5'-TCTA-3', or 5'-CTCA-3', in addition to the canonical PAM 5'-TTTR-3' sequence (wherein R is A or G) recognized by wild-type Cas12f (e.g., UnCas12f1 or CWCas12f1).
[0217] The term "editing window" refers to the target range of base editing by a base editing system (e.g., a CRISPR / Cas-ABE or CRISPR / Cas-CBE module). For example, the target base of the editing window may be proximal to the PAM sequence or may be at any distance. In some embodiments, the engineered Cas polypeptide of the present invention may have an expanded editing window compared to wild-type Cas12f (e.g., UnCas12f1 or CWCas12f1).
[0218] The terms "guide nucleic acid," "guide RNA," and "gRNA" refer to any nucleic acid (e.g., RNA) molecule that directs or facilitates targeting of a molecule, such as a nuclease, gene editing protein, or nucleic acid degrading protein, to a target nucleic acid. For example, the guide nucleic acid may be a guide RNA that recognizes (or binds to) a target nucleic acid by forming a complex with an engineered Cas polypeptide of the present invention. The guide nucleic acid may be designed to be complementary to the target strand (e.g., the non-PAM strand) of the target nucleic acid sequence, and the complementary binding to the target nucleic acid sequence is sufficient to result in sequence-specific binding. In some embodiments, a naturally occurring guide nucleic acid may comprise one or more stem / loop or optimized secondary structures. A stem / loop structure may refer to a secondary structure that forms a hairpin structure as a result of base pairing that occurs in a single-stranded nucleic acid (e.g., DNA or RNA). “Engineered guide RNA” refers to a guide RNA that has undergone an artificial modification to the structure of a guide RNA that naturally exists in nature.
[0219] The terms "target gene" or "target nucleic acid" basically refer to a cellular gene or nucleic acid that is the target of gene regulation, modification, or editing. The terms "target gene" and "target nucleic acid" may be used interchangeably and may refer to the same entity. Unless otherwise specified, a target gene or target nucleic acid may refer to either a native gene or nucleic acid of a target cell, cell line, or organism, or an exogenous gene or nucleic acid, and is not particularly limited as long as it can be the target of gene regulation, modification, or editing. A target gene or target nucleic acid may be single-stranded DNA, double-stranded DNA, and / or RNA.
[0220] The term "target sequence" refers to a sequence to which a guide nucleic acid specifically recognizes or binds. In some embodiments, a DNA-binding sequence (e.g., a spacer sequence) of a guide RNA binds to a target sequence. In some embodiments, the target sequence is a sequence contained within a target gene or a target nucleic acid sequence, and refers to a sequence that complementarily binds to a spacer sequence (guide sequence) contained in a guide RNA. A target nucleic acid may comprise the target sequence and optionally additional coding or non-coding sequences. In some embodiments, the target sequence may be a segment of DNA adjacent to a PAM motif (on the PAM strand). The target sequence may be located at the 3' end of the PAM motif or the 5' end of the PAM motif, depending on the Cas polypeptide that recognizes the PAM motif, as is well known in the art. In some embodiments, the target sequence to which the guide RNA binds may refer to the non-PAM strand including the protospacer, but may also refer to all double-stranded DNA of the target sequence targeted by the guide RNA described herein, as will be apparent from the context throughout the specification.
[0221] The term "substantially complementary" or "complementary" means that a polynucleotide (e.g., a spacer sequence of a guide RNA) has a certain level of complementarity to a target sequence. In some embodiments, a certain level of complementarity is such that the polynucleotide and the Cas polypeptide that forms a complex with the polynucleotide can hybridize to the target sequence with sufficient affinity to allow the polynucleotide to act on (e.g., cleave) the target sequence.
[0222] The term "off-target" refers to unintended or unexpected binding, cleavage, and / or editing of a nucleic acid region by a gene editing system. For example, a nucleic acid region is an off-target region if it differs by 1, 2, 3, 4, 5, 6, 7, or more nucleotides from the nucleic acid region to which it is intended or expected to be bound, cleaved, and / or edited. As another example, a nucleic acid region is an off-target region if it has the same nucleotide sequence as the nucleic acid region to which it is intended or expected to be cleaved and / or edited, but at a different location in the genome.
[0223] The term "domain" refers to a functional and / or structural unit of a polypeptide. In some embodiments, a Cas polypeptide (e.g., UnCas12f1 or a variant thereof) may comprise a wedge (WED) domain, a zinc finger (ZF) domain, a recognition (REC) domain, a linker domain, a linker-loop (RuvC) domain, or a target nucleotide binding (TNB) domain. In some embodiments, a Cas polypeptide (e.g., CWCas12f1 or a variant thereof) may further comprise an auxiliary domain.
[0224] The term "fusion protein" refers to a protein formed by combining two or more distinct proteins or fragments thereof, and may be in the form of an effector domain, such as a transcriptional regulatory domain or an epigenetic modification domain, attached, linked, or fused, depending on the desired purpose. In some embodiments, two or more proteins or protein fragments thereof may be connected by a linker or spacer.
[0225] The term "isolated" may mean a polynucleotide, polypeptide, or cell that is in an environment other than the environment in which the polynucleotide, polypeptide, or cell naturally occurs.
[0226] The term “construct” or “vector” means a recombinant nucleic acid, typically recombinant DNA, created for the purpose of expressing and / or propagating a particular nucleotide sequence(s) or for use in the construction of other recombinant nucleotide sequences.
[0227] The term “operably linked” in gene expression technology refers to the juxtaposition of a particular component (e.g., a functional element, a regulatory element, a control element, etc.) with another component in a relationship that allows these components to function in their intended manner. A regulatory and / or control element (e.g., a promoter, an enhancer, an intron, a polyadenylation signal, a Kozak consensus sequence, an internal ribosome entry site (IRES), a splice acceptor, a 2A sequence, and / or an origin of replication, etc.) that is “operably linked” to a functional element is associated in such a way that expression and / or activity of the functional element is achieved under conditions compatible with the regulatory and / or control elements. For example, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence.
[0228] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues with similar side chains within a protein. A conservative amino acid substitution does not substantially alter the structure, function, or both of a polypeptide. Conservative amino acid substitutions are provided in Table 1 under the heading "Preferred Substitutions" and are further described below with respect to amino acid side chain classes (1) through (6). In some embodiments, a conservative amino acid substitution refers to a substitution at an amino acid position that does not significantly affect performance (e.g., indel efficiency, recognizable PAM sequence, editing window). For example, an engineered Cas polypeptide comprising an amino acid sequence having at least 80% identity to a sequence means a sequence having at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid identity compared to a parent sequence.
[0229]
[0230] Amino acids can be grouped according to their common side chain properties as follows: (1) hydrophobic: norleucine, Met, Ala, Val, Leu, Ile;
[0231] (2) Neutral hydrophilic: Cys, Ser, Thr, Asn, Gln;
[0232] (3) Acidic: Asp, Glu;
[0233] (4) Basic: His, Lys, Arg;
[0234] (5) Residues affecting chain orientation: Gly, Pro;
[0235] (6) Aromatic: Trp, Tyr, Phe.
[0236] The term "identity" refers to the overall relatedness between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, nucleic acid molecules or polypeptide molecules may be considered "substantially identical" to each other if the sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. Calculating the percent identity of sequences can be performed by aligning two sequences for optimal comparison purposes, wherein gaps can be introduced into one or both of the first and second sequences for optimal alignment, and non-identical sequences can be ignored for comparison purposes. Comparing the sequences to be compared and determining the percent identity can be accomplished using a mathematical algorithm. As is well known in the art, amino acid or nucleic acid sequences can be compared using any of a variety of algorithms available in commercial computer programs, including BLASTN for nucleotide sequences and BLASTP, Gapped BLAST, and PSI-BLAST for amino acids.
[0237] The term "subject" is used interchangeably with "subject," "subject," or "patient," and can be a mammal, such as a primate (e.g., a human), a companion animal (e.g., a dog, a cat, etc.), a livestock animal (e.g., a cow, a pig, a horse, a sheep, a goat, etc.), or a laboratory animal (e.g., a rat, a mouse, a guinea pig, etc.), in need of prevention or treatment of a disease or condition. In some embodiments, the subject is a human.
[0238] The term "treatment" generally refers to obtaining a desired pharmacological and / or physiological effect. Such effect is therapeutic in that it partially or completely cures a disorder, disease, and / or other unwanted or undesirable condition. Desirable therapeutic effects include, but are not limited to, preventing the occurrence or recurrence of a disorder or disease, improving symptoms, reducing any direct or indirect pathological consequences of the disorder or disease, preventing metastasis, slowing the progression of the disorder or disease, improving or palliating the state of the disorder or disease, and remission or improved prognosis. Preferably, "treatment" may refer to medical intervention for an already present disorder or disease.
[0239] The term "administration" means providing an active ingredient to a subject to achieve a preventive or therapeutic purpose.
[0240] II. Engineered Cas Polypeptides
[0241] The present invention is based, at least in part, on the discovery that CRISPR-Cas effector polypeptides can exhibit improved editing performance compared to wild-type effector polypeptides through site-specific amino acid mutations. The inventors have surprisingly determined that CRISPR-Cas effector polypeptides can be programmed to exhibit significantly improved editing performance through one or more of the site-specific amino acid mutations described herein, or a combination thereof. Accordingly, in one aspect of the present invention, a CRISPR-Cas effector polypeptide is provided that is engineered to have improved editing performance. The engineered Cas polypeptide exhibits significantly improved gene editing efficiency compared to a wild-type Cas polypeptide (a reference Cas polypeptide) and can function in a eukaryotic cell. Furthermore, the engineered Cas polypeptide can be a variant that has an expanded target range (a recognizable PAM sequence), an expanded editing window, or both, compared to a wild-type Cas polypeptide.
[0242] In some embodiments, the engineered Cas polypeptide of the present invention is an isolated or purified polypeptide.
[0243] In some embodiments, the wild-type Cas polypeptide can be a type VF Cas polypeptide, specifically a Cas12f polypeptide, more specifically a Cas12f1 polypeptide, and more specifically UnCas12f1 or CWCas12f1. The amino acid sequence of UnCas12f1 or CWCas12f1 is depicted in FIG. 1A (SEQ ID NO: 5) and FIG. 1B (SEQ ID NO: 1). In some embodiments, the wild-type Cas polypeptide can comprise or consist of the amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 5. Human codon-optimized nucleic acid sequences encoding the proteins of UnCas12f1 and CWCas12f1 are exemplified by SEQ ID NO: 10 and SEQ ID NO: 6.
[0244] The CWCas12f1 polypeptide comprises, from the N-terminus to the C-terminus, an auxiliary domain (positions 1-28 based on SEQ ID NO: 1), a wedge (WED) domain (positions 29-47 based on SEQ ID NO: 1), a zinc finger (ZF) domain (positions 48-123 based on SEQ ID NO: 1), a recognition (REC) domain (positions 124-220 based on SEQ ID NO: 1), a WED domain (positions 221-340 based on SEQ ID NO: 1), a linker domain (positions 341-349 based on SEQ ID NO: 1), a RuvC domain (positions 350-501 based on SEQ ID NO: 1), a target nucleotide binding (TNB) domain (positions 502-536 based on SEQ ID NO: 1), and a RuvC domain (SEQ ID NO: 1 standard, positions 537-557). For convenience, amino acid positions are indicated based on the CWCas12f1 polypeptide sequence throughout this specification.
[0245] In some embodiments, the engineered Cas polypeptide can comprise a modified amino acid sequence that is at least 70% identical to the amino acid sequence of a wild-type Cas12f polypeptide. For example, the engineered Cas polypeptide can be at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of a wild-type Cas12f polypeptide (e.g., the amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 5).
[0246] In some embodiments, the engineered Cas polypeptide can comprise a modification in one or more amino acids relative to the amino acid sequence of a wild-type Cas12f polypeptide (e.g., the amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 5), at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 162, 164, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 193, 194,Amino acids 195, 196, 197, 198, 199, or 200 may be altered. In some embodiments, the engineered Cas polypeptide has an amino acid sequence of 1 to 50, 1 to 49, 1 to 48, 1 to 47, 1 to 46, 1 to 45, 1 to 44, 1 to 43, 1 to 42, 1 to 41, 1 to 40, 1 to 39, 1 to 38, 1 to 37, 1 to 36, 1 to 35, 1 to 34, 1 to 33, 1 to 32, 1 to 31, 1 to 30, 1 to 29, 1 to 28, 1 to 27, 1 to 26, 1 to 25, 1 to 24, 1 to 23, 1 to 22, 1 to 21, 1 to 20, 1 to 19, 1 to 18, 1 to 17, 1 to 16, 1 to 15, 1 to 14, 1 to 13, 1 to 12, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 to 2, or 1 amino acid It may include transformations.
[0247] In the present invention, amino acid modifications may include substitution, insertion, deletion, or addition of one or more amino acids or nucleotides in a polypeptide or nucleic acid compared to a reference sequence, or a combination thereof.
[0248] In some embodiments, the amino acid sequence of the engineered Cas polypeptide can comprise one or more amino acid substitutions relative to the amino acid sequence of a wild-type Cas12f polypeptide.
[0249] In some embodiments, the engineered Cas polypeptide can comprise one or more amino acid substitutions in the RuvC domain. In some embodiments, the mutated amino acid sequence of the engineered Cas polypeptide can comprise 1 to 15, 1 to 13, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1, or 2 amino acid substitutions in the RuvC domain.
[0250] In one aspect, an engineered Cas polypeptide is provided that comprises one or more amino acid substitutions relative to the amino acid sequence of a wild-type Cas12f polypeptide. Specifically, the one or more amino acid substitutions comprise SEQ ID NO: 1.
[0251] (i) I365, A368, H399, A402, A406, C433, I435, A436 and D437;
[0252] (ii) I65, E125, I126, L133, K135, I140, Y141, K155, S164, H167, D171, T175, E179, L180, D216, L222, Q229, T231, R250, F266, E269, K305, Q312, T313, S322, G325, M331, L332, D337, I341, P347, K358, S359, L361, F369, K407, T414, E418, S420, E421, C433, I435, I440, N442, Q448, N451, R457, E459, K479, H508, K530 and E555;
[0253] (iii) K31, N32, T33, K39, N47, P221, L222, K224, Q225, Q229, Y230, T231, E234, N237, K245, I246, R250, Q252, K254, K255, E256, I257, K259, K265, F266, D267, E269, Q270, Q272, I278, L282, K288, K295, D296, E297, K304, K305, N308, D310, Q312, T313, K319, G321, S322, K323, G325, M331, L332, N333 D337;
[0254] (iv) S48, E50, K53, D57, E58, N61, I65, L67, N70, D72, T89 및 A120;
[0255] (v) Q124, E125, I126, L133, K135, I140, Y141, I154, K155, A162, S163, S164, E166, H167, D171, T175, R176, A178, E179, L180, I186, L190, L200, N205, M206, G209, S215, D216 및 F218;
[0256] (vi) I341, D342, K343, D346 및 P347;
[0257] (vii) T502, K505, H508, L509, N511, E516, K519, K527, K530 및 E535; 및
[0258] (viii) comprises a substitution at a residue corresponding to one or more residues selected from the group consisting of K358, S359, L361, I365, A368, F369, K385, H399, A402, A406, K407, T414, E418, S420, E421, C433, I435, A436, D437, I440, N442, T446, Q448, N451, S454, R457, E459, I465, K479, K501, L543, L550, K551, S552, E555 and E556.
[0259] In one aspect, an engineered Cas polypeptide is provided that comprises one or more amino acid substitutions relative to the amino acid sequence of a wild-type Cas12f polypeptide. Specifically, the one or more amino acid substitutions comprise SEQ ID NO: 1.
[0260] (i) I365V, A368S, H399L, A402S, A406S, C433K, I435V, I435L, A436T and D437N;
[0261] (ii) I65S, I65E, E125S, E125N, E125M, I126M, L133V, L133S, L133M, L133C, K135R, I140V, I140M, Y141F, K155R, K155Y, S164D, H167K, H167R, D171E, D171N, D171R, T175R, E179A, L180N, D216E, L222F, Q229G, Q229K, T231E, T231K, R250K, F266C, E269D, K305R, Q312P, Q312K, T313I, T313V, S322K, G325C, M331F, L332V, D337R, I341K, P347K, K358R, S359N, S359H, L361A, F369N, F369D, F369G, F369W, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435L, I440L, N442H, N442M, Q448V, Q448N, N451S, R457K, E459K, K479M, H508C, H508Y, K530R, more
[0262] (iii) K31R, N32Q, T33S, K39R, N47Q, P221T, P221S, P221C, P221G, P221A, P221V, P221I, P221M, P221F, L222F, L222I, K224R, Q225R, Q225H, Q225D, Q225E, Q225N, Q225Y, Q225K, Q225F, Q225C, Q225V, Q225P, Q225S, Q225T, Q225A, Q229G, Q229K, Q229N, Q229R, Q229H, Q229C, Q229P, Y230E, Y230R, Y230H, Y230K, Y230D, Y230N, Y230Q, Y230F, Y230C, Y230W, T231E, T231K, T231P, E234D, N237Q, K245R, I246L, R250K, Q252N, K254R, K255R, E256D, I257L, K259R, K265R, F266C, D267E, E269D, Q270R, Q272N, Q272R, I278L, L282I, K288R, K295R, D296E, E297D, K304R, K305R, N308Q, D310E, Q312K, Q312P, Q312N T313I, T313V, K319R, G321K, S322K, K323R, G325C, M331F, L332V, N333Q and D337R;
[0263] (iv) S48T, E50D, K53R, D57E, E58D, N61Q, I65E, I65L, I65S, L67I, N70Q, D72E, T89S 및 A120T;
[0264] (v) Q124N, E125M, E125N, E125S, I126M, L133C, L133M, L133S, L133V, K135R, I140M, I140V, Y141F, I154L, K155R, K155Y, A162D, A162Q, A162P, A162V, S163R, S163K, S163N, S163M, S163F, S163Y, S163G, S163P, S163A, S163L, S164D, E166D, H167K, H167R, D171N, D171R, D171E, T175R, R176K, A178S, E179A, L180N, I186L, L190I, L200F, N205H, M206H, G209R, S215H, D216E and F218H;
[0265] (vi) I341K, D342E, K343R, D346N and P347K;
[0266] (vii) T502S, K505R, H508C, H508Y, L509I, N511Q, E516D, K519R, K527R, K530R and E535D; and
[0267] (viii) K358R, S359H, S359N, L361A, I365V, A368S, F369D, F369G, F369N, F369W, K385R, H399L, A402S, A406S, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435V, I435L, A436T, D437N, I440L, N442H, N442M, T446S, Q448N, Q448V, N451S, S454G, R457K, E459K, I465R, K479M, K501R, Contains one or more amino acid substitutions selected from the group consisting of L543I, L550I, K551R, S552T, E555D, E555R and E556D.
[0268] In one aspect, the engineered Cas polypeptide comprises (ii) I65, E125, I126, L133, K135, I140, Y141, K155, S164, H167, D171, T175, E179, L180, D216, L222, Q229, T231, R250, F266, E269, K305, Q312, T313, S322, G325, M331, L332, D337, I341, P347, K358, S359, L361, F369, K407, T414, E418, S420, E421, C433, I435, I440, N442, Engineered Cas polypeptides are provided comprising amino acid substitutions at one or more residues selected from the group consisting of Q448, N451, R457, E459, K479, H508, K530, and E555, or at a residue corresponding to one residue.
[0269] In some embodiments, the engineered Cas polypeptide comprises (ii) I65S, I65E, E125S, E125N, E125M, I126M, L133V, L133S, L133M, L133C, K135R, I140V, I140M, Y141F, K155R, K155Y, S164D, H167K, H167R, D171E, D171N, D171R, T175R, E179A, L180N, D216E, L222F, Q229G, Q229K, T231E, T231K, R250K, F266C, E269D, K305R, Q312P, Q312K, T313I, It may comprise one or more amino acid substitutions or one amino acid substitution selected from the group consisting of T313V, S322K, G325C, M331F, L332V, D337R, I341K, P347K, K358R, S359N, S359H, L361A, F369N, F369D, F369G, F369W, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435L, I440L, N442H, N442M, Q448V, Q448N, N451S, R457K, E459K, K479M, H508C, H508Y, K530R and E555R.
[0270] In one aspect, the engineered Cas polypeptide can comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of (i) I365, A368, H399, A402, A406, C433, I435, A436, and D437, based on SEQ ID NO: 1.
[0271] In some embodiments, the engineered Cas polypeptide can comprise one or more substitutions selected from the group consisting of (i) I365V, A368S, H399L, A402S, A406S, C433K, I435V, A436T, and D437N, based on SEQ ID NO: 1.
[0272] In some embodiments, the engineered Cas polypeptide may comprise a substitution at a residue corresponding to a residue selected from I365 / A368, H399 / A402 / A406, C433 / I435 / A436 / D437, and any combination thereof, based on SEQ ID NO: 1. As used herein, " / " means "and." For example, the expression "I365 / A368" means that the amino acid modifications are at "I365 and A368."
[0273] In some embodiments, the engineered Cas polypeptide can comprise an amino acid substitution selected from I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N, C433K / I435L / A436T / D437N, and combinations thereof, based on SEQ ID NO: 1. In some embodiments, engineered Cas polypeptides (e.g., Cas12f1, such as UnCas12f1 or CWCas12f1) comprising these amino acid mutation combinations (I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N) exhibited significantly improved indel efficiency compared to wild-type Cas12f.
[0274] In some implementations, the engineered Cas polypeptide comprises SEQ ID NO: 1.
[0275] I365V / A368S,
[0276] H399L / A402S / A406S,
[0277] C433K / I435V or I435L / A436T / D437N,
[0278] I365V / A368S / H399L / A402S / A406S,
[0279] I365V / A368S / C433K / I435V or I435L / A436T / D437N,
[0280] H399L / A402S / A406S / C433K / I435V or I435L / A436T / D437N, or
[0281] It may contain amino acid substitutions of I365V / A368S / H399L / A402S / A406S / C433K / I435V or I435L / A436T / D437N.
[0282] In some embodiments, the engineered Cas polypeptide comprises I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N and / or C433K / I435L / A436T / D437N, and optionally further comprises an amino acid substitution at one or more residues selected from the group consisting of amino acid residues listed in (ii), (iii), (iv), (v), (vi), (vii) and (viii) described herein.
[0283] In some embodiments, the engineered Cas polypeptide may further comprise one or more amino acid substitutions in the RuvC domain in addition to I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N and / or C433K / I435L / A436T / D437N.
[0284] In some embodiments, the engineered Cas polypeptide may further comprise one or more amino acid substitutions in one or more domains selected from the group consisting of RuvC, WED, ZF, REC, Linker, and TNB domains, in addition to I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N, and / or C433K / I435L / A436T / D437N.
[0285] Exemplary amino acid positions and substitutions in the RuvC, WED, ZF, REC, Linker, and TNB domains are listed in Table 2.
[0286] In some implementations, the engineered Cas polypeptide comprises:
[0287] One or more amino acid substitutions in the WED domain, e.g., (iii) K31, N32, T33, K39, N47, P221, L222, K224, Q225, Q229, Y230, T231, E234, N237, K245, I246, R250, Q252, K254, K255, E256, I257, K259, K265, F266, D267, E269, Q270, Q272, I278, L282, K288, K295, D296, E297, K304, K305, N308, D310, Q312, T313, K319, G321, S322, K323, G325, A substitution at a residue corresponding to one or more residues selected from the group consisting of M331, L332, N333 and D337;
[0288] One or more amino acid substitutions in the ZF domain, for example, (iv) a substitution at a residue corresponding to one or more residues selected from the group consisting of S48, E50, K53, D57, E58, N61, I65, L67, N70, D72, T89 and A120;
[0289] One or more amino acid substitutions in the REC domain, for example, (v) a substitution at a residue corresponding to one or more residues selected from the group consisting of Q124, E125, I126, L133, K135, I140, Y141, I154, K155, A162, S163, S164, E166, H167, D171, T175, R176, A178, E179, L180, I186, L190, L200, N205, M206, G209, S215, D216 and F218;
[0290] One or more amino acid substitutions in the linker domain, for example, (vi) a substitution at a residue corresponding to one or more residues selected from the group consisting of I341, D342, K343, D346 and P347;
[0291] one or more amino acid substitutions in the TNB domain, for example, (vii) a substitution at a residue corresponding to one or more residues selected from the group consisting of T502, K505, H508, L509, N511, E516, K519, K527, K530, and E535; and / or
[0292] The RuvC domain may further comprise one or more amino acid substitutions, for example, (viii) a substitution at a residue corresponding to one or more residues selected from the group consisting of K358, S359, L361, I365, A368, F369, K385, H399, A402, A406, K407, T414, E418, S420, E421, C433, I435, A436, D437, I440, N442, T446, Q448, N451, S454, R457, E459, I465, K479, K501, L543, L550, K551, S552, E555, and E556, wherein the amino acid substitution position is based on SEQ ID NO: 1.
[0293] In some implementations, the engineered Cas polypeptide comprises SEQ ID NO: 1.
[0294] (iii) K31R, N32Q, T33S, K39R, N47Q, P221T, P221S, P221C, P221G, P221A, P221V, P221I, P221M, P221F, L222F, L222I, K224R, Q225R, Q225H, Q225D, Q225E, Q225N, Q225Y, Q225K, Q225F, Q225C, Q225V, Q225P, Q225S, Q225T, Q225A, Q229G, Q229K, Q229N, Q229R, Q229H, Q229C, Q229P, Y230E, Y230R, Y230H, Y230K, Y230D, Y230N, Y230Q, Y230F, Y230C, Y230W, T231E, T231K, T231P, E234D, N237Q, K245R, I246L, R250K, Q252N, K254R, K255R, E256D, I257L, K259R, K265R, F266C, D267E, E269D, Q270R, Q272N, Q272R, I278L, L282I, K288R, K295R, D296E, E297D, K304R, K305R, N308Q, D310E, Q312K, Q312P, At least one selected from the group consisting of Q312N, T313I, T313V, K319R, G321K, S322K, K323R, G325C, M331F, L332V, N333Q and D337R;
[0295] (iv) at least one selected from the group consisting of S48T, E50D, K53R, D57E, E58D, N61Q, I65E, I65L, I65S, L67I, N70Q, D72E, T89S and A120T;
[0296] (v) Q124N, E125M, E125N, E125S, I126M, L133C, L133M, L133S, L133V, K135R, I140M, I140V, Y141F, I154L, K155R, K155Y, A162D, A162Q, A162P, A162V, S163R, S163K, S163N, S163M, S163F, S163Y, S163G, S163P, S163A, S163L, S164D, E166D, H167K, H167R, D171N, D171R, D171E, T175R, R176K, A178S, At least one selected from the group consisting of E179A, L180N, I186L, L190I, L200F, N205H, M206H, G209R, S215H, D216E and F218H;
[0297] (vi) one or more selected from the group consisting of I341K, D342E, K343R, D346N and P347K;
[0298] (vii) one or more selected from the group consisting of T502S, K505R, H508C, H508Y, L509I, N511Q, E516D, K519R, K527R, K530R and E535D; and / or
[0299] (viii) K358R, S359H, S359N, L361A, I365V, A368S, F369D, F369G, F369N, F369W, K385R, H399L, A402S, A406S, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435V, I435L, A436T, D437N, I440L, N442H, N442M, T446S, Q448N, Q448V, N451S, S454G, R457K, E459K, I465R, K479M, K501R, It may additionally include one or more substitutions selected from the group consisting of L543I, L550I, K551R, S552T, E555D, E555R and E556D.
[0300] For example, additional amino acid substitutions or combinations thereof are presented in Tables 3 and 9.
[0301] In some embodiments, the engineered Cas polypeptide may comprise amino acid substitutions, such as C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N, and optionally one or more other amino acid substitutions. The optionally additional one or more other amino acid substitutions are set forth, for example, in Tables 3 and 9.
[0302] In some embodiments, the engineered Cas polypeptide can comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of residues K135, Y141, K305, and M331 of SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can further comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of residues K135, Y141, K305, and M331 in addition to the amino acid substitutions described herein. In some embodiments, the engineered Cas polypeptide can comprise an amino acid substitution at a residue corresponding to residues K135, Y141, K305, and M331. In some embodiments, the engineered Cas polypeptide can further comprise an amino acid substitution at a residue corresponding to residues K135, Y141, K305, and M331 in addition to the amino acid substitutions described herein.
[0303] In some embodiments, the engineered Cas polypeptide can comprise one or more amino acid substitutions selected from the group consisting of K135R, Y141F, K305R, and M331F, based on SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can further comprise one or more amino acid substitutions selected from the group consisting of K135R, Y141F, K305R, and M331F, based on SEQ ID NO: 1, in addition to the amino acid substitutions described herein. In some embodiments, the engineered Cas polypeptide can comprise K135R / Y141F / K305R / M331F, based on SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can comprise K135R / Y141F / K305R / M331F, based on SEQ ID NO: 1, in addition to the amino acid substitutions described herein.
[0304] For example, the engineered Cas polypeptide may further comprise amino acid substitutions exemplified in Tables 3 and 9 in addition to the amino acid substitutions K135R / Y141F / K305R / M331F.
[0305] In some embodiments, the engineered Cas polypeptide may comprise amino acid substitutions of I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N, and K135R / Y141F / K305R / M331F based on SEQ ID NO: 1. Throughout the specification, C433K / I435L / A436T / D437N may be substituted for C433K / I435V / A436T / D437N. In one embodiment, Cas variants comprising amino acid substitutions I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N, and K135R / Y141F / K305R / M331F were found to exhibit significantly improved indel efficiency compared to the wild-type Cas polypeptide.
[0306] In some embodiments, the engineered Cas polypeptide can comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of residues L67, Q124, K155, and E556 of SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can further comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of residues L67, Q124, K155, and E556, in addition to the amino acid substitutions described herein. In some embodiments, the engineered Cas polypeptide can comprise an amino acid substitution at a residue corresponding to residues L67, Q124, K155, and E556 of SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can further comprise an amino acid substitution at a residue corresponding to residues L67, Q124, K155, and E556, in addition to the amino acid substitutions described herein.
[0307] In some embodiments, the engineered Cas polypeptide can comprise the amino acid substitutions L67I / Q124N / K155R / E556D relative to SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can further comprise the amino acid substitutions L67I / Q124N / K155R / E556D in addition to the amino acid substitutions described herein relative to SEQ ID NO: 1.
[0308] For example, the engineered Cas polypeptide may further comprise amino acid substitutions exemplified in Tables 3 and 9 in addition to the amino acid substitutions L67I / Q124N / K155R / E556D.
[0309] In some embodiments, the engineered Cas polypeptide can comprise amino acid substitutions of I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N, K135R / Y141F / K305R / M331F, and L67I / Q124N / K155R / E556D relative to SEQ ID NO: 1.
[0310] In some embodiments, the engineered Cas polypeptide can comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of T175, G209, K358, and S454 of SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can further comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of T175, G209, K358, and S454 of SEQ ID NO: 1 in addition to the amino acid substitutions described herein. In some embodiments, the engineered Cas polypeptide can comprise an amino acid substitution at residues T175, G209, K358, and S454 of SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can further comprise an amino acid substitution at residues T175, G209, K358, and S454 of SEQ ID NO: 1 in addition to the amino acid substitutions described herein.
[0311] In some embodiments, the engineered Cas polypeptide can comprise one or more amino acid substitutions selected from the group consisting of T175R, G209R, K358R, and S454G, based on SEQ ID NO: 1. In some embodiments, the engineered Cas polypeptide can further comprise one or more amino acid substitutions selected from the group consisting of T175R, G209R, K358R, and S454G, based on SEQ ID NO: 1, in addition to the amino acid substitutions described herein. In some embodiments, the engineered Cas polypeptide can comprise the amino acid substitutions T175R / G209R / K358R / S454G. In some embodiments, the engineered Cas polypeptide can further comprise the amino acid substitutions T175R / G209R / K358R / S454G, in addition to the amino acid substitutions described herein.
[0312] For example, the engineered Cas polypeptide may further comprise amino acid substitutions exemplified in Tables 3 and 9 in addition to the amino acid substitutions T175R / G209R / K358R / S454G.
[0313] In some embodiments, the engineered Cas polypeptide can comprise amino acid substitutions of I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N, K135R / Y141F / K305R / M331F, L67I / Q124N / K155R / E556D, and T175R / G209R / K358R / S454G relative to SEQ ID NO: 1.
[0314] In one aspect of the present invention, engineered Cas polypeptides are provided that exhibit improved indel efficiency compared to wild type and can recognize an expanded range of PAM sequences, such as amino acid substitution combinations labeled starting with M1010, M1120, M1223, M1224, M1125, M1231, M1232, M1233, and PM of Table 9.
[0315] In one aspect of the present invention, an engineered Cas polypeptide is provided comprising an amino acid substitution at the following position relative to SEQ ID NO: 1:
[0316] (i) I365 / A368,
[0317] (ii) H399 / A402 / A406,
[0318] (iii) C433 / I435 / A436 / D437,
[0319] (iv) K135 / Y141 / K305 / M331,
[0320] (v) L67 / Q124 / K155 / E556,
[0321] (vi) T175 / G209 / K358 / S454, or
[0322] (vii) combinations of these.
[0323] In one aspect of the present invention, an engineered Cas polypeptide is provided comprising the following amino acid substitutions based on SEQ ID NO: 1:
[0324] (i) I365V / A368S,
[0325] (ii) H399L / A402S / A406S,
[0326] (iii) C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N,
[0327] (iv) K135R / Y141F / K305R / M331F,
[0328] (v) L67I / Q124N / K155R / E556D,
[0329] (vi) T175R / G209R / K358R / S454G, or
[0330] (vii) Amino acid substitutions corresponding to any combination of these.
[0331] Through the embodiments of the present invention, it was confirmed that these amino acid substitutions can be effectively used for the production of engineered Cas12f polypeptides with improved editing performance.
[0332] In one aspect of the present invention, an engineered Cas polypeptide is provided comprising the following amino acid substitutions based on SEQ ID NO: 1:
[0333] L67I / Q124N / K135R / Y141F / K155R / K305R / M331F / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K530R / E556D (M1010).
[0334] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0335] L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1120).
[0336] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0337] L67I / Q124N / K135R / Y141F / K155R / N205H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1223).
[0338] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0339] L67I / Q124N / K135R / Y141F / K155R / M206H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1224).
[0340] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0341] L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1225).
[0342] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0343] Q272R / N205H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D31 0E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1231).
[0344] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0345] Q272R / M206H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D31 0E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1232).
[0346] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0347] Q272R / Q270R / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D31 0E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1233).
[0348] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0349] Q225P / Q229C / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM622).
[0350] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0351] Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D31 0E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM625).
[0352] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0353] Q229R / Y230W / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D31 0E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM628).
[0354] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0355] K230H / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM638).
[0356] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0357] K230C / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM646).
[0358] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0359] S163R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM674).
[0360] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0361] S163K / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM676).
[0362] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0363] T175R / G209R / K358R / S454G / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K30 5R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM714).
[0364] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0365] D171R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM715).
[0366] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0367] T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM716).
[0368] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0369] I465R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM722).
[0370] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0371] G209R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM724).
[0372] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0373] D171R / T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N30 8Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM727).
[0374] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0375] T175R / G209R / S454G / K385R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K30 5R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM728).
[0376] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0377] D171R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D31 0E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM729).
[0378] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0379] T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D31 0E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM730).
[0380] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0381] I465R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D31 0E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM731).
[0382] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0383] G209R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D31 0E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM732).
[0384] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0385] D171R / T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM733).
[0386] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0387] K358R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM734).
[0388] In some embodiments, an engineered Cas polypeptide is provided comprising the following amino acid substitutions relative to SEQ ID NO: 1:
[0389] T175R / G209R / K358R / S454G / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM735).
[0390] Exemplary amino acid positions and substitutions of the present invention are shown in Tables 2 and 3.
[0391]
[0392]
[0393]
[0394]
[0395]
[0396]
[0397]
[0398]
[0399]
[0400]
[0401]
[0402]
[0403]
[0404]
[0405]
[0406]
[0407]
[0408]
[0409]
[0410]
[0411]
[0412]
[0413]
[0414]
[0415]
[0416]
[0417]
[0418]
[0419] The expression "based on SEQ ID NO: 1" herein is intended to provide a description of the amino acid mutation positions of the present invention and should not be construed as limiting the amino acid sequence of the engineered Cas polypeptides disclosed herein. The expression is also used to determine the corresponding amino acid residue in another species of Cas effector polypeptide classified in the same Cas effector category as the Cas polypeptide having the amino acid sequence of SEQ ID NO: 1. For example, in another Cas effector polypeptide, the corresponding amino acid residue or position at an amino acid mutation residue or position of the present invention can be determined by aligning the amino acid sequence of the other Cas effector polypeptide with SEQ ID NO: 1. Methods for aligning two sequences are known in the art. For example, the UnCas12f1 polypeptide having the amino acid sequence of SEQ ID NO: 5 includes an amino acid sequence in which 28 amino acids are deleted from the N-terminus of the amino acid sequence of SEQ ID NO: 1, so the corresponding position is obtained by subtracting 28 from the amino acid residue number indicated based on SEQ ID NO: 1.
[0420] In some embodiments, the engineered Cas12f protein comprises any one amino acid substitution or combination of amino acid substitutions listed in Table 3 based on SEQ ID NO: 1, and can be, for example, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 1.
[0421] In some embodiments, the engineered Cas12f protein can comprise one or more amino acid substitutions at positions corresponding to amino acid positions disclosed herein relative to the amino acid sequence of the CWCas12f-v1 protein (SEQ ID NO: 2), the CWCas12f-v2 protein (SEQ ID NO: 3), or the CWCas12f-v3 protein (SEQ ID NO: 4). For example,
[0422] “CWCas12f-v1 protein” MEKRINKIRKKLSADNATKPVSRSGPMAKNTITKTLKLRIVRPYNSAEVEKIVADEKNNREKIALEKNKDKVKEACSKHLKVAAYCTTQVERNACLFCKARKLDDKFYQKLRGQFPDAVFWQEISEIFRQLQKQAAEI YNQSLIELYYEIFIKGKGIANASSVEHYLSDVCYTRAAELFKNAAIASGLRSKIKSNFRLKELKNMKSGLPTTKSDNFPIPLVKQKGGQYTGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQVQKSPKPIS LLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSKIGEKSAWMLNLSIDVPKIDKGVDPSIIGGIDVGVKSPLVCAINNAFSRYSISDNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTE KSERFRKKLIERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNIRLRGFWPYAEMQNKIEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKCNFKENADYNAALNISNPKLKSTKEEP (SEQ ID NO: 2);
[0423] "CWCas12f-v2 단백질" MAGGPGAGSAAPVSSTSSLPLAALNMRVMAKNTITKTLKLRIVRPYNSAEVEKIVADEKNNREKIALEKNKDKVKEACSKHLKVAAYCTTQVERNACLFCKARKLDDKFYQKLRGQFPDAVFWQEISEIFRQLQKQAAEIYNQSLIELYYEIFIKGKGIANASSVEHYLSDVCYTRAAELFKNAAIASGLRSKIKSNFRLKELKNMKSGLPTTKSDNFPIPLVKQKGGQYTGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQVQKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSKIGEKSAWMLNLSIDVPKIDKGVDPSIIGGIDVGVKSPLVCAINNAFSRYSISDNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTEKSERFRKKLIERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNIRLRGFWPYAEMQNKIEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKCNFKENADYNAALNISNPKLKSTKEEP (서열번호 3);
[0424] “CWCas12f-v3 protein” MAGGPGAGSAAPVSSTSSLPLAALNMMAKNTITKTLKLRIVRPYNSAEVEKIVADEKNNREKIALEKNKDKVKEACSKHLKVAAYCTTQVERNACLFCKARKLDDKFYQKLRGQFPDAVFWQEISEIFRQLQKQAAEI YNQSLIELYYEIFIKGKGIANASSVEHYLSDVCYTRAAELFKNAAIASGLRSKIKSNFRLKELKNMKSGLPTTKSDNFPIPLVKQKGGQYTGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQVQKSPKPIS LLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSKIGEKSAWMLNLSIDVPKIDKGVDPSIIGGIDVGVKSPLVCAINNAFSRYSISDNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTE KSERFRKKLIERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNIRLRGFWPYAEMQNKIEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKCNFKENADYNAALNISNPKLKSTKEEP (SEQ ID NO: 4).
[0425] The human codon-optimized nucleic acid sequences encoding the above CWCas12f-v1 protein, CWCas12f-v2 protein and CWCas12f-v3 protein are represented by SEQ ID NO: 7 to SEQ ID NO: 9.
[0426] In some embodiments, the engineered Cas12f protein may further comprise 1 to 600 amino acids at the N-terminus or C-terminus. There is no limitation on the sequence of 1 to 600 amino acids added. For example, the added 1 to 600 amino acids may be an amino acid sequence of SEQ ID NO: 190 or SEQ ID NO: 191. Meanwhile, the engineered Cas12f protein may further comprise a nuclear localization signal (NLS) or a nuclear export signal (NES) sequence at the N-terminus and / or C-terminus. In addition, the NLS sequence or NES sequence may further be included between the added sequence and the Cas12f variant protein.
[0427] In the present invention, engineered Cas polypeptides comprising one or more amino acid substitutions exhibited significantly improved editing performance compared to the parent Cas12f polypeptide. Without being bound by theory, it is believed that the one or more amino acid substitutions included in the engineered Cas polypeptides of the present invention contribute to the improved activity compared to the parent Cas12f polypeptide, such as increased interaction with a target nucleic acid, increased binding affinity for a target nucleic acid, increased binding affinity for a guide RNA, increased on-target specific binding activity, decreased off-target binding activity, decreased dissociation from a target nucleic acid, increased stability, expansion of the recognized PAM sequence, expansion of the editing window, or a combination thereof. In some embodiments, the M1120, M1223, M1224, and M1225 variants exhibited superior indel efficiency not only for the canonical PAM but also for non-canonical PAM sequences, suggesting that they can recognize an expanded range of PAM sequences.
[0428] In some embodiments, the engineered Cas polypeptide comprises or consists of an amino acid sequence having at least 80% sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 201 to 384 or SEQ ID NOs: 410 to 606. The engineered Cas polypeptide can be at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 201 to 384 or SEQ ID NOs: 410 to 606. In some embodiments, the engineered Cas polypeptide may additionally comprise conservative amino acid substitutions.
[0429] In some embodiments, the engineered Cas polypeptides of the invention can be fused with one or more effector domains. For example, the effector domains can include, but are not limited to, polypeptides that can cleave a nucleic acid (e.g., DNA or RNA), edit a nucleotide, modulate transcription, modulate translation, methylate or demethylate a nucleic acid, affect RNA splicing, enable affinity purification or immunoprecipitation (e.g., FLAG, HA, biotin, or HALO tags), and / or enable protein labeling or identification.
[0430] In some embodiments, the engineered Cas polypeptides of the present invention may further comprise one or more amino acid mutations that render them catalytically inactive or substantially devoid of cleavage activity. Furthermore, the engineered Cas polypeptides of the present invention may further comprise one or more amino acid mutations that broaden the editing window. For a discussion of such additional amino acid mutations, see, for example, International Publication No. 2023-282597, the entire disclosure of which is incorporated herein by reference.
[0431] In some embodiments, the additional amino acid mutation that renders the engineered Cas polypeptide catalytically inactive or substantially devoid of cleavage activity can be a substitution at a residue corresponding to one or more residues selected from the group consisting of residues D354, E450, R518, and D538 of SEQ ID NO: 1. In other embodiments, the additional amino acid mutation that widens the editing window in the engineered Cas polypeptide can be a substitution at a residue corresponding to I159 and / or S164 of SEQ ID NO: 1.
[0432] In some embodiments, the additional amino acid mutations that render the engineered Cas polypeptide catalytically dead or substantially devoid of cleavage activity can include one or more selected from the group consisting of D354A, E450A, E450Q, E450L, E450W, R518A, R518Q, R518L, R518W, D538A, D538L, D538V, based on SEQ ID NO: 1. Such additional amino acid mutations that render the engineered Cas polypeptide catalytically dead or substantially devoid of cleavage activity can result in a nuclease activity of less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the wild-type or prior to the additional mutations. In another embodiment, additional amino acid mutations that widen the editing window in the engineered Cas polypeptide can include I159W and / or S164Y relative to SEQ ID NO: 1.
[0433] In some embodiments, engineered Cas polypeptides of the invention comprising additional amino acid mutations that render them catalytically dead or substantially inactive or that provide a widened editing window can be fused to one or more heterologous protein domains or heterologous polypeptides.
[0434] In some embodiments, the heterologous protein domain may be a base-editing domain. A base-editing domain is a domain that has the function of correcting a specific base in a gene to another base. When combined with the high editing performance of the engineered Cas polypeptide of the present invention, a base at a specific site can be corrected to a desired base.
[0435] In some embodiments, the base editing domain can be an adenosine deaminase and / or a cytidine deaminase, and can function as an adenine and / or cytosine base editor by fusion to an engineered Cas polypeptide of the invention comprising additional amino acid mutations that render the base editing domain catalytically dead or substantially devoid of cleavage activity.
[0436] Adenosine deaminase refers to a protein, or a domain that functions to participate in a deamination reaction that targets adenine (A) in RNA / DNA and DNA duplexes, and hydrolyzes the adenine moiety of adenine or an adenine-containing molecule (e.g., adenosine, DNA, RNA) to hypoxanthine or the hypoxanthine moiety of a hypoxanthine-containing molecule (e.g., inosine (I)). Adenosine deaminase is rarely found in higher animals, but is known to be present in small amounts in the muscles of cows, milk, and blood of rats, and in large amounts in the intestines of crayfish and insects. In some embodiments, the adenosine deaminase includes, but is not limited to, tRNA adenosine deaminase (TadA) from Escherichia coli, a mutant of TadA, or a combination thereof.
[0437] Cytidine deaminase is an enzyme protein that targets cytosine (C), deaminating it and converting it to uracil (U). By removing the amine group from cytosine to create uracil, a series of intracellular repair mechanisms convert uracil to thymine (T), ultimately leading to base correction of cytosine (C) to thymine (T). Cytidine deaminases generally act on RNA, but some are known to be able to act on single-stranded DNA (ssDNA) as well (Harris et al., 2002), including but not limited to human activation-induced cytidine deaminase (AID), human APOBEC3G, mouse APOBEC1, APOBEC3A, APOBEC3B, CDA, AID, single-stranded DNA deaminase (Sdd), and lamprey PmCDA1.
[0438] In some embodiments, a dCWCas12f1-ABE-C3.1 module is provided utilizing an inactive variant of CWCas12f1 lacking cleavage activity (e.g., dead M1120). The dCWCas12f1-ABE-C3.1 module was found to have an extended editing window (extension of positions at which base transitions occur) for targets containing the circular PAM sequence (5'-TTTG-3' or 5'-TTTA-3') compared to the control (dCWCas12f1-ABE module).
[0439] In some embodiments, the heterologous polypeptide includes, but is not limited to, a polypeptide that provides an activity that indirectly increases transcription by acting directly on the target DNA or on a polypeptide associated with the target DNA (e.g., a histone or other DNA binding protein). Additional suitable fusion targets include, but are not limited to, a polypeptide that provides methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, SUMOylating activity, deSUMOylating activity, ribosylating activity, deribosylating activity, myristoylating activity, or demyristoylating activity. Additional suitable fusion targets include, but are not limited to, polypeptides that directly provide increased transcription of the target nucleic acid (e.g., transcriptional activators or fragments thereof, proteins that recruit transcriptional activators or fragments thereof, small molecule / drug-responsive transcriptional regulators, etc.). In some embodiments, the heterologous polypeptide may be, but is not limited to, DNMT, TET, KRAB, DHAC, LSD, p300, or moloney murine leukemia virus (M-MLV). In some embodiments, when a reverse transcriptase is fused, the fusion protein may function as a prime editor.
[0440] In some embodiments, engineered Cas polypeptides of the invention comprising additional amino acid mutations that render them catalytically dead or substantially devoid of cleavage activity can also be fused to other proteins or domains, such as Clo51 or FokI nucleases, to generate double-strand breaks.
[0441] In some embodiments, the engineered Cas polypeptides of the invention may comprise additional amino acid mutations that allow them to cleave only one strand of a double-stranded target nucleic acid.
[0442] In some embodiments, the engineered Cas polypeptides of the present invention may include additional amino acid mutations that allow for an expanded targetable region or additional amino acid mutations that allow for appropriate PAM specificity for a target gene. For a discussion of such additional amino acid mutations, see, for example, International Publication No. 2023-229222, the entire disclosure of which is incorporated herein by reference.
[0443] In some embodiments, the engineered Cas polypeptide of the present invention may be part of a fusion protein. In other words, a fusion protein is provided comprising the engineered Cas polypeptide of the present invention and an effector domain.
[0444] In some embodiments, the engineered Cas polypeptides of the present invention can be attached to, linked to, or fused to an effector domain, such as a transcriptional regulatory domain or an epigenetic modification domain. In some embodiments, such an effector domain can be fused to the C-terminus or N-terminus of the engineered Cas polypeptide.
[0445] In some embodiments, the separate proteins (effector domains) included in the fusion protein may be linked (or connected) by a linker or spacer. Specifically, the linker may be a peptide linker comprising 2 to 20 amino acids. The linker may be, for example, GGGS, GGGGS, or two, three, four, or more repeats thereof.
[0446] In some embodiments, the effector domain may comprise an intracellular localization signal, such as, but not limited to, a nuclear localization signal (NLS), a nuclear export signal (NES), or a mitochondrial localization signal.
[0447] III. Guide nucleic acid
[0448] A nucleic acid molecule that binds to the engineered Cas polypeptide of the present invention to form a ribonucleoprotein complex (RNP) and directs the complex to a specific location within a target nucleic acid (e.g., target DNA) is referred to as a guide nucleic acid or guide RNA. In some cases, the guide RNA may comprise DNA bases in addition to RNA bases, resulting in a hybrid DNA / RNA, and the term "guide RNA" is understood to encompass such hybrid DNA / RNA molecules.
[0449] The guide RNA comprises a spacer region (or guide region) comprising a guide sequence hybridizable with a target sequence adjacent to a PAM, and a scaffold region capable of interacting with a Cas polypeptide to form a complex. The guide sequence of the guide RNA is designed to be substantially complementary to its target sequence. Hybridization between the guide sequence and the target sequence allows the Cas polypeptide to be positioned at the target sequence. The guide sequence of the guide RNA is not necessarily required to be perfectly complementary to its target sequence; it is sufficient that the complementarity allows hybridization between the guide sequence and the target sequence. In some embodiments, the complementary bond between the guide sequence and the target sequence may include one or more mismatch bonds, preferably 0 to 7, more preferably 0 to 6, and even more preferably 0 to 5 mismatches. In some embodiments, the guide sequence may comprise or consist of at least 10 and no more than 50, preferably at least 15 and no more than 40 nucleotides.
[0450] In some embodiments, the guide RNA may be an engineered guide RNA. Specifically, the engineered guide RNA can direct the engineered Cas polypeptide described herein to a specific target nucleic acid. More specifically, the engineered guide RNA may have a shorter nucleotide length compared to the wild-type guide RNA of the Cas12f1 polypeptide.
[0451] In some embodiments, the (engineered) guide RNA comprises i) a first sequence comprising the nucleotide sequence 5'-GGAAUGCAAC-3' and ii) a second sequence that hybridizes to a target sequence of a target nucleic acid, wherein the first sequence is located at the 5' side of the second sequence. Through previous studies, the present inventors found that guide RNAs that bind to the Cas12f Cas protein can be engineered to be shorter than the original length to exhibit excellent editing efficiency in eukaryotic cells, and revealed that an essential part of the crRNA repeat sequence is 5'-GGAAUGCAAC-3'. Based on these research results, the present inventors developed a new CRISPR / Cas system, the CRISPR / Cas12f system, and named it the TaRGET (Tiny nuclease augmented RNA-based Genome Editing Technology) system. The CRISPR / Cas12f system is a novel CRISPR / Cas system that was first reported in a previous study (see [Harrington et al., Science, 362, 839-842, 2018)). Despite the advantage of having a remarkably small effector protein size, it was reported that its application to gene editing technology was limited due to its lack or extremely low double-stranded DNA cleavage activity. To overcome this limitation, the present inventors developed an engineered guide RNA with enhanced cleavage activity for double-stranded DNA, enabling its use in gene editing. Detailed descriptions of the engineered guide RNA are provided in Korean Patent Application Nos. 10-2021-0044152, 10-2022-0049512, and 10-2021-0051552; International Application Nos. PCT / KR2021 / 013898, PCT / KR2021 / 013923, and PCT / KR2021 / 013933; and Kim, DY et al.Efficient CRISPR editing with a hypercompact Cas12f1 and engineered guide RNAs delivered by adeno-associated virus. Nat. Biotechnol. 40, 94-102, 2021], the entire disclosure of which is incorporated herein by reference.
[0452] In some embodiments, the (engineered) guide RNA comprises a tracrRNA sequence, a crRNA sequence and a guide sequence, and a scaffold region formed by the tracrRNA sequence and the direct repeat sequence comprises, in the 5' to 3' direction, one or more stem-loop regions (e.g., a first stem-loop region, a second stem-loop region, a third stem-loop region and a fourth stem-loop region), a tracrRNA-crRNA complementarity region (which may also be referred to as a fifth stem region), and optionally (iii) a region comprising at least three, at least four or at least five consecutive uracils (U), and is capable of interacting with a Cas12f polypeptide.
[0453] In some embodiments, the (engineered) guide RNA can comprise a tracrRNA, a crRNA, and a guide sequence in the 5' to 3' end direction. The tracrRNA can comprise a nucleotide sequence of SEQ ID NO: 11 or a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, the tracrRNA may essentially comprise 5'-GGCUGCUUGCAUCAGCCUAAUGUCGAGAAGUGCUUUCUUCGGAAAGUAACCCUCGAAACAAA-3'.
[0454] The crRNA can comprise the nucleotide sequence of SEQ ID NO: 12 or a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, the crRNA can essentially comprise 5'-GGAAUGCAAC-3'.
[0455] In some implementations, the (engineered) guide RNA comprises, from the 5' end to the 3' end, i) a first sequence comprising the nucleotide sequence 5'-GGCUGCUUGCAUCAGCCUAAUGUCGAGAAGUGCUUUCUUCGGAAAGUAACCCUCGAAACAAA-3', ii) a second sequence comprising the nucleotide sequence 5'-GGAAUGCAAC-3', and iii) a third sequence that hybridizes to a target sequence of a target nucleic acid.
[0456] In some embodiments, the (engineered) guide RNA may be a single guide RNA (sgRNA). Specifically, the (engineered) guide RNA may comprise or consist of a nucleotide sequence of SEQ ID NO: 13 or a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0457] Exemplary sequences of tracrRNA, crRNA, and sgRNA are presented in Table 4.
[0458]
[0459] In Table 4 and throughout the specification, the sequence indicated as 'NNNNNNNNNNNNNNNNNNNN' refers to a guide sequence (spacer sequence) having any length (e.g., 15 to 40 nucleotides in length) that can hybridize with a target sequence within a target site. The crRNA of the Cas12f guide RNA may include a direct repeat sequence and a guide sequence, and the direct repeat sequence may be located at the 5' end of the guide sequence. In the present specification, crRNA may refer only to a sequence that forms a homologous bond with a tracrRNA excluding the guide sequence and a sequence including the direct repeat sequence, depending on the context. In addition, the crRNA may be located at the 3' end of the tracrRNA. The scaffold region of the Cas12f guide RNA includes a portion of the tracrRNA and the crRNA. For detailed information on the structure of the Cas12f guide RNA, see Takeda et al., Structure of the miniature type VF CRISPR-Cas effector enzyme, Molecular Cell 81, 1-13 (2021), the entire disclosure of which is incorporated herein by reference.
[0460] In some embodiments, the engineered guide RNA may comprise a modified scaffold region and / or guide sequence. The sequence of the scaffold region is located 5' to the guide sequence.
[0461] In some embodiments, the engineered gRNA comprises a sequence in which one or more nucleotides are substituted, deleted, inserted or added from the wild-type gRNA sequence, and can have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to an unmodified Cas12f gRNA, excluding the guide sequence.
[0462] In some embodiments, the structure of the engineered gRNA and its modifications are described in detail for each of the five modification sites. The modification sites are abbreviated as "MS (modification site)" throughout this specification, and the numbers following "modification site" or "MS" are sequentially assigned according to the experimental engineering flow of each modification site according to one embodiment. However, it does not mean that engineering (modification) at a modification site with a later number necessarily includes engineering (modification) at a modification site with a preceding number. Figure 2 illustrates modification sites MS1 to MS5 included in an engineered guide RNA according to an embodiment of the present invention on an unmodified guide RNA sequence.
[0463] For example, among the detailed regions of the gRNA disclosed herein, the first stem-loop region including modification site 3 (MS3), the second stem-loop region including modification site 5 (MS5), the tracrRNA-crRNA complementarity region (the fifth stem region or the fifth stem-loop region) including modification site 1 (MS1) and modification site 4 (MS4), and the addition of one or more uridines to the 3'-end of the crRNA sequence (MS2) may be defined as regions corresponding to or including regions indicated by single-dotted line boxes with different shades of color in FIG. 2.
[0464] In some embodiments, the engineered guide RNAs disclosed herein can achieve higher editing performance (e.g., improved recognition / cleavage efficiency for target nucleic acids) and be shorter in length than unmodified guide RNAs. The engineered guide RNAs disclosed herein, due to their shorter length, can allocate more space within the packaging limit (about 4.7 kb) of vectors such as adeno-associated virus (AAV) to other components for various purposes or applications (e.g., additional guide RNAs, shRNAs for suppressing specific gene expression, etc.), thereby imparting highly efficient gene editing effects that could not be achieved with existing CRISPR / Cas systems.
[0465] In some embodiments, the engineered guide RNA comprises, compared to an unmodified Cas12f gRNA, e.g., an unmodified Cas12f gRNA comprising (i) one or more stem-loop regions, (ii) a tracrRNA-crRNA complementary region, and optionally (iii) a region comprising at least three, at least four, or at least five consecutive uracils (Us), (a) deletion of part or all of one or more of the stem-loop regions (e.g., the first stem-loop region, the second stem-loop region, or both); (b) deletion of part or all of the tracrRNA-crRNA complementary region; (c) substitution of one or more of the Us with another nucleotide (e.g., A, G, or C), if three, at least four, or at least five consecutive uracils (Us) are present; and (d) may comprise one or more modifications selected from the group consisting of addition of one or more uridines to the 3'-end of the crRNA sequence.
[0466] In some embodiments, the engineered guide RNA may comprise the addition of one or more uridines (e.g., U-rich tail) to the 3'-end of the crRNA sequence. In some embodiments, the sequence of the U-rich tail is 5'-(U m V) n U o -3', where V is independently A, C or G, m and o are integers between 1 and 20, and n is an integer between 0 and 5.
[0467] In some implementations, the engineered guide RNA may comprise a deletion of part or all of one or more stem-loop regions (e.g., the first stem-loop region, the second stem-loop region, or both) compared to the unmodified guide RNA.
[0468] In some implementations, the engineered guide RNA may comprise a deletion of part or all of the tracrRNA-crRNA complementarity region (the fifth stem region or, in the case of sgRNA, the fifth stem-loop region).
[0469] Deletion of part or all of the stem-loop region or complementary region means removal of one or more pairs of nucleotides that form complementary bonds (including mismatched bonds) from the stem region, and if a pair of nucleotides that form complementary bonds (including mismatched bonds) exists to form the stem region, the loop region may remain without being removed.
[0470] In some implementations, the engineered guide RNA may comprise a modified scaffold region comprising a sequence represented by Formula (I).
[0471]
[0472] In equation (I), X a , X b1 , X b2 , X c1 and X c2are each independently composed of 0 to 35 (poly)nucleotides, and Lk is a polynucleotide linker of length 2 to 20 or is absent. [In formula (I), the black solid line represents a chemical bond (e.g., a phosphodiester bond) between nucleotides, and the gray bold line represents a complementary bond between nucleotides.]
[0473] In equation (I), X a , X b1 , X b2 , X c1 or X c2 If it consists of 0 nucleotides, then X a , X b1 , X b2 , X c1 or X c2 It is interpreted to mean that there is no existence.
[0474] X in equation (I) a , X b1 , X b2 , X c1 or X c2 If it consists of 0 nucleotides or is absent, X a , X b1 , X b2 , X c1 or X c2 If there are two or more nucleotides connected through , they are interpreted as being directly connected in some way. For example, in formula (I), X b1 If it consists of 0 nucleotides or is absent, X b1 Nucleotides directly linked to the 5'-end of X b1 The nucleotides directly linked to the 3'-end may be directly linked, for example, by a phosphodiester bond.
[0475] In some implementations, X a may be a (poly)nucleotide that may be absent or may have a stem-loop conformation. In some embodiments, the X amay consist of 0 to 20 (poly)nucleotides.
[0476] In some implementations, X b1 and X b2 may be (poly)nucleotides capable of complementary binding to each other. In some embodiments, X b1 may consist of 0 to 13 (poly)nucleotides, or X b2 may consist of 0 to 14 (poly)nucleotides.
[0477] In some implementations, X c1 and X c2 may be (poly)nucleotides capable of complementary binding to each other. In some embodiments, X c1 may consist of 0 to 28 (poly)nucleotides, or X c2 may consist of 0 to 27 (poly)nucleotides.
[0478] In some embodiments, Lk is a polynucleotide linker of length 2 to 20, length 2 to 15, length 2 to 10, or length 2 to 8, or is absent.
[0479] In some embodiments, the scaffold region of the engineered gRNA can comprise or consist of a sequence represented by Formula (I), or a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity thereto.
[0480] When referring to the scaffold region of the unmodified guide RNA, the first stem-loop region of the scaffold sequence is X in formula (I). a corresponds to or X aIt may be a region including. For example, it may be a region corresponding to the sequence indicated as MS3 in Figure 2. The second stem-loop region of the scaffold sequence is X in formula (I). b1 and X b2 may be an area corresponding to or containing them. For example, X b1 and X b2 The second stem-loop region containing 5'-CCGCUUCAC-X b1 -uuag-X b2 -AGUGAAGGUG―3' sequence. The third stem region of the scaffold sequence may be a region corresponding to or including the 5'-GGCUGCUUGCAUCAGCC-3' sequence in formula (I). The fourth stem-loop region of the scaffold sequence may be a region corresponding to or including the 5'-UCGAGAAGUGCUUUCUUCGGAAAGUAACCCUCGA-3' sequence in formula (I). The tracrRNA-crRNA complementarity region (the fifth stem(-loop) region) of the scaffold sequence may be a region corresponding to or including the 5'-UCGAGAAGUGCUUUCUUCGGAAAGUAACCCUCGA-3' sequence in formula (I). c1 and X c2 It may be an area corresponding to .
[0481] We describe in detail the modifications at each modification site in the engineered gRNA.
[0482] (1) Modification at modification site 1 (MS1)
[0483] This article describes modifications in MS1 (Fig. 2). In some embodiments, a wild-type tracrRNA (e.g., SEQ ID NO: 11) that can be included in a naturally occurring guide RNA may have a sequence containing five consecutive uracils (Us) within the sequence. This poses a problem in that, when attempting to express the wild-type tracrRNA in a cell using a vector or the like, under certain conditions, the sequence acts as a transcription termination signal, causing unintended premature termination of transcription. That is, when the sequence containing the five consecutive Us acts as a transcription termination signal, normal or complete expression of the tracrRNA is suppressed, and normal or complete formation of gRNA is also inhibited, resulting in a decrease in gene editing efficiency.
[0484] Therefore, to solve the above-described problem, the engineered gRNA may be one in which at least one of three, four, or five consecutive Us, preferably four or five Us, of the wild-type tracrRNA (e.g., SEQ ID NO: 11) is artificially modified to another nucleotide, such as A, C, T, or G.
[0485] In some embodiments, an engineered gRNA is provided that comprises a modification in which at least one of the three or more, four or more, or five or more consecutive Us, referred to as MS1, is replaced with a different type of nucleotide. For example, the three or more, four or more, or five or more consecutive Us may be present within the tracrRNA-crRNA complementarity region of the tracrRNA, wherein the three or more, preferably four or more, or five or more consecutive Us may be modified such that no consecutive sequence of three or more, preferably four or more, or five or more Us appears by replacing one or more of the three or more, preferably four or more, or five or more consecutive Us with A, G, or C.
[0486] It is preferred that the sequence within the tracrRNA-crRNA complementary region of the crRNA corresponding to the sequence being modified in this way is also modified. In some embodiments, when the sequence 5'-ACGAA-3' is present within the tracrRNA-crRNA complementary region of the crRNA that forms a partial complementary bond with the sequence 5'-UUUUU-3' within the tracrRNA-crRNA complementary region of the tracrRNA, the sequence may be replaced with 5'-NGNNN-3', wherein N is each independently A, C, G, or U.
[0487] In some implementations, X in the engineered gRNA of formula (I) c1 When there are three or more, four or more or five or more consecutive uracils (U) in the sequence, it may include a variant in which one or more of these Us are replaced with A, G or C. For example, X c1 If the sequence 5'-UUUUU-3' exists in the sequence, the sequence can be replaced with 5'-NNNCN-3', where N is independently A, C, G or U. In a more specific example, X c1 The sequence 5'-UUUUU-3' in the sequence may be replaced by any one nucleic acid sequence selected from the group consisting of the following sequences, but is not limited to the following sequences as long as it does not cause a sequence containing three or more consecutive Us, preferably four or more, or five or more consecutive Us to appear: 5'-UUUCU-3', 5'-GUUCU-3', 5'-UCUCU-3', 5'-UUGCU-3', 5'-UUUCC-3', 5'-GCUCU-3', 5'-GUUCC-3', 5'-UCGCU-3', 5'-UCUCC-3', 5'-UUGCC-3', 5'-GCGCU-3', 5'-GCUCC-3', 5'-GUGCC-3', 5'-UCGCC-3', 5'-GCGCC-3' and 5'-GUGCU-3'.
[0488] In some implementations, X in the engineered gRNA of formula (I)c2 The sequence is X c1 A region in which at least part of the sequence forms a complementary bond (also referred to as a tracrRNA-crRNA complementarity region), wherein X c1 X forming at least one complementary bond with three or more, four or more, or five or more consecutive U's present in the sequence c2 The corresponding sequence within the sequence may also be transformed. For example, X in the above formula (I) c2 If the sequence 5'-ACGAA-3' exists in the sequence, the sequence can be replaced with 5'-NGNNN-3', where N is independently A, C, G or U. In a more specific example, X in formula (I) c1 The sequence 5'-ACGAA-3' within the sequence may be replaced by any one nucleic acid sequence selected from the group consisting of the following sequences, but is not limited to the following sequences: 5'-AGGAA-3', 5'-AGCAA-3', 5'-AGAAA-3', 5'-AGCAU-3', 5'-AGCAG-3', 5'-AGCAC-3', 5'-AGCUA-3', 5'-AGCGA-3', 5'-AGCCA-3', 5'-UGCAA-3', 5'-UGCUA-3', 5'-UGCGA-3', 5'-UGCCA-3', 5'-GGCAA-3', 5'-GGCUA-3', 5'-GGCGA-3', 5'-GGCCA-3', 5'-CGCAA-3', 5'-CGCUA-3', 5'-CGCGA-3', and 5'-CGCCA-3'.
[0489] In some implementations, X of formula (I) c1 When a sequence containing three or more, four or more, or five or more consecutive U's in a sequence is transformed into another sequence, the corresponding (i.e., at least some of which form complementary bonds) X c2 It is preferable that the corresponding nucleotide in the sequence be modified so that it can form a complementary bond with the modified nucleotide. For example, X c1X when the sequence 5'-UUUUU-3' within the sequence is transformed into 5'-GUGCU-3' c2 It is preferred that the sequence 5'-ACGAA-3' within the sequence be modified to 5'-AGCAA-3', but complementary binding is not essential.
[0490] (2) Modification at modification site 2 (MS2)
[0491] This article describes modifications in MS2 (Fig. 2). In some embodiments, the engineered guide RNA (gRNA) may be a gRNA found in nature with a novel structure, such as one or more uridines added to the 3'-end of the crRNA sequence, more specifically, the 3'-end of the spacer sequence included in the crRNA. The 3'-end of the crRNA sequence may be the 3'-end of the guide sequence (spacer). As used herein, the one or more uridines added to the 3'-end are also referred to as a "U-rich tail." Engineered gRNAs comprising one or more uridines or U-rich tails added to the 3'-end serve to enhance the nucleic acid cleavage or indel efficiency of the ultra-small CRISPR / Cas12f system for the target gene or target nucleic acid.
[0492] The term "U-rich tail" as used herein may refer not only to the RNA sequence itself, which is rich in uridine (U), but also to the DNA sequence encoding it, and this may be interpreted appropriately depending on the context. The present inventors have experimentally and in detail elucidated the structure and effects of the U-rich tail sequence, which can be found in Korean Patent No. 10-2455623 and International Application No. PCT / KR2020 / 014961. The disclosures of these documents are incorporated herein in their entirety.
[0493] In some implementations, the U-rich tail sequence can be represented by Ux. Wherein x can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20. As an example, x can be an integer within a range of two numbers selected from the numbers listed above. For example, x can be an integer between 1 and 6. As another example, x can be an integer between 1 and 20. In some implementations, x can be an integer greater than or equal to 20.
[0494] In some implementations, the sequence of the U-rich tail is 5'-(U m V) n U o -3', and in the above formula, V is each independently A, C or G, m and o are integers from 1 to 20, and n can be an integer from 0 to 5. For example, n can be 0, 1 or 2. For example, m and o can each independently be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0495] In some implementations, the sequence of the U-rich tail is 5'-(U m V) n U o In the sequence represented by -3', (i) n is 0, o is an integer from 1 to 6, or (ii) V is each independently A or G, m and o are each independently an integer from 3 to 6, and n is an integer from 1 to 3. The U-rich tail may be. In a specific example, the U-rich tail is 5'-U-3', 5'-UU-3', 5'-UUU-3', 5'-UUUU-3', 5'-UUUUU-3', 5'-UUUUU-3', 5'-UUUUUU-3', 5'-UUURUUU-3', 5'-UUURUUU-3', 5'-UUURUUUU-3', 5'-UUUURUUU-3', 5'-UUUURUUU-3', 5'-UUUURUUU-3', 5'-UUUURUUU-3', 5'-UUUURUUUU-3', 5'-UUUURUUUU-3', 5'-UUUURUUUU-3', 5'-UUUURUUUUU-3' and 5'-UUUURUUUUUU-3', wherein R can be a U-rich tail of A or G. For example, the U-rich tail may be a sequence consisting of or including the sequence 5'-UUUUUUUUUU-3' (SEQ ID NO: 192), 5'-UUAUUUAUUU-3' (SEQ ID NO: 193), 5'-UUUCUAUUUU-3' (SEQ ID NO: 194), or 5'-UUAUGUUUUU-3' (SEQ ID NO: 195).
[0496] In some embodiments, the U-rich tail sequence can comprise a modified uridine repeat sequence in which a non-uridine ribonucleoside (A, C, or G) is included for every 1 to 5 uridine repeats. Such modified uridine repeat sequences are particularly useful when designing vectors expressing engineered crRNAs. In some embodiments, the U-rich tail sequence can be or comprise one or more repeats of UV, UUV, UUUV, UUUUV, and / or UUUUUV, wherein V is one of A, C, and G.
[0497] Additionally, the U-rich tail sequence is expressed as Ux and 5'-(U m V) n U o-3' sequences may be combined. In some implementations, the U-rich tail sequence may be represented by (U)n1-V1-(U)n2-V2-Ux, wherein V1 and V2 are each one of adenine (A), cytidine (C), and guanine (G), n1 and n2 are each an integer from 1 to 4, and x may be an integer from 1 to 20. In addition, the length of the U-rich tail sequence may be 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt. In some implementations, the length of the U-rich tail sequence may be 20 nt or more.
[0498] In some embodiments, when the engineered gRNA is expressed in a cell, the U-rich tail can be expressed as one or more sequences due to early transcription termination. For example, according to some embodiments, when a gRNA intended to include a U-rich tail of the sequence 5'-UUUUAUUUUUU-3' is transcribed in a cell, four or more or five or more Ts can act as termination sequences, so that a gRNA including a U-rich tail such as 5'-UUUUAUUUU-3', 5'-UUUUAUUUUUU-3', or 5'-UUUUAUUUUUU-3' can be generated simultaneously. Therefore, in the present invention, a U-rich tail including four or more Us can be understood to also include a U-rich tail sequence of a shorter length than the intended length.
[0499] In some implementations, the U-rich tail sequence may include additional bases in addition to uridine depending on the actual usage environment and expression environment of the gene editing system of the present invention, for example, the environment inside a eukaryotic cell or a prokaryotic cell.
[0500] (3) Modification at modification site 3 (MS3)
[0501] This article describes modifications in MS3 (Fig. 2). As described above, MS3 is a region (which may be referred to as the first stem-loop region) that comprises some or all of the nucleotides that form a stem-loop structure within the gRNA and effector protein complex, and MS3 may comprise a region that does not interact with the effector protein when the gRNA and effector protein form a complex. Modifications in MS3 include the removal of some or all of the first stem-loop region near the 5'-end of the tracrRNA.
[0502] In some embodiments, the engineered gRNA comprises a variant in which part or all of the first stem-loop region (e.g., the sequence of SEQ ID NO: 14) is deleted.
[0503] In some implementations, the engineered gRNA comprises a variant in which part or all of the first stem-loop region on the tracrRNA is deleted, wherein the part or all of the first stem-loop region that is deleted may be from 1 to 20 nucleotides. Specifically, part or all of the first stem-loop region may be 2 to 20, 3 to 20, 4 to 20, 5 to 20, 6 to 20, 7 to 20, 8 to 20, 9 to 20, 10 to 20, 11 to 20, 12 to 20, 13 to 20, 14 to 20, 15 to 20, 16 to 20, 17 to 20, 18 to 20, 19 or 20 nucleotides.
[0504] In some implementations, the MS3 or first stem-loop region is X of formula (I). a As a site corresponding to the nucleotide indicated by X, by a variant in which part or all of the first stem-loop region is deleted, aIt may be composed of 0 to 35 (poly)nucleotides, and preferably 0 to 20, 0 to 19, 0 to 18, 0 to 17, 0 to 16, 0 to 15, 0 to 14, 0 to 13, 0 to 12, 0 to 11, 0 to 10, 0 to 9, 0 to 8, 0 to 7, 0 to 6, 0 to 5, 0 to 4, 0 to 3, 0 to 2, 1 or 0 (poly)nucleotides.
[0505] In some implementations, X in the scaffold sequence of formula (I) a may comprise a nucleic acid sequence of SEQ ID NO: 14 or may comprise a nucleic acid sequence in which all or part of the sequence is deleted, preferably 1 to 20 nucleotides are deleted from the sequence of SEQ ID NO: 14. For example, the deletion of nucleotides may be a deletion of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 15, 16, 17, 18, 19 or 20 nucleotides randomly from the sequence of SEQ ID NO: 14. As a preferred example, the deletion of the nucleotides may be a deletion of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides sequentially from the 5'-end in the sequence of SEQ ID NO: 14. In this respect, X a The deletion of a nucleotide in may be a deletion of a complementary nucleotide pair. More specifically, X in formula (I) ais 5'-CUUCACUGAUAAAGUGGAGA-3' (SEQ ID NO: 14), 5'-UUCACUGAUAAAGUGGAGA-3' (SEQ ID NO: 15), 5'-UCACUGAUAAAGUGGAGA-3' (SEQ ID NO: 16), 5'-CACUGAUAAAGUGGAGA-3' (SEQ ID NO: 17), 5'-ACUGAUAAAGUGGAGA-3' (SEQ ID NO: 18), 5'-CUGAUAAAGUGGAGA-3' (SEQ ID NO: 19), 5'-UGAUAAAGUGGAGA-3' (SEQ ID NO: 20), 5'-GAUAAAGUGGAGA-3' (SEQ ID NO: 21), 5'-AUAAAGUGGAGA-3' (SEQ ID NO: 22), 5'-UAAAGUGGAGA-3' (SEQ ID NO: 23), 5'-AAAGUGGAGA-3' (SEQ ID NO: 24), 5'-AAGUGGAGA-3', 5'-AGUGGAGA-3', 5'-GUGGAGA-3', 5'-UGGAGA-3', 5'-GGAGA-3', 5'-GAGA-3', 5'-AGA-3', 5'-GA-3' or 5'-A-3', or may comprise or consist of the sequence X a may not exist.
[0506] (4) Modification at modification site 4 (MS4)
[0507] This article describes modifications in MS4 (Fig. 2). MS4 may include part or all of a sequence referred to as a tracrRNA-crRNA complementarity region (which may also be referred to as a fifth stem region), which is located across the 3'-terminus of tracrRNA and the 5'-terminus of crRNA, or, in the case of a single guide RNA, a region where a sequence corresponding to tracrRNA and a sequence corresponding to crRNA form at least a partial complementary bond. In the present invention, the tracrRNA-crRNA complementarity region may include both modification site 1 (MS1) and modification site 4 (MS4). Modifications in MS4 include deletions of part or all of the tracrRNA-crRNA complementarity region. The tracrRNA-crRNA complementarity region may include nucleotides capable of forming complementary bonds with some nucleotides in the tracrRNA and some nucleotides in the crRNA within a complex of the gRNA and the nucleotide-degrading protein, and may include adjacent nucleotides. The tracrRNA-crRNA complementarity region of the tracrRNA may include a region that does not interact with the nucleotide-degrading protein within the complex of the gRNA and the nucleotide-degrading protein.
[0508] In some embodiments, the engineered gRNA comprises a deletion of part or all of the tracrRNA-crRNA complementary region in the tracrRNA, a deletion of part or all of the tracrRNA-crRNA complementary region in the crRNA, or a deletion of part or all of the tracrRNA-crRNA complementary region in both the tracrRNA and the crRNA.
[0509] In some embodiments, the tracrRNA-crRNA complementarity region may comprise the nucleotide sequence of SEQ ID NO: 39 and / or the nucleotide sequence of SEQ ID NO: 58.
[0510] In some implementations, the tracrRNA-crRNA complementarity region may further comprise a linker (e.g., a polynucleotide) connecting the 3'-end of the tracrRNA to the 5'-end of the crRNA.
[0511] In some implementations, the engineered gRNA comprises a variant in which a portion of the tracrRNA-crRNA complementary region is deleted, wherein the portion of the complementary region that is deleted may be from 1 to 54 nucleotides.
[0512] In some embodiments, the engineered gRNA comprises a variant in which the entire tracrRNA-crRNA complementary region is deleted, wherein the entire complementary region deleted may be 55 nucleotides.
[0513] Specifically, part or all of the tracrRNA-crRNA complementarity region may be 3 to 55, 5 to 55, 7 to 55, 9 to 55, 11 to 55, 13 to 55, 15 to 55, 17 to 55, 19 to 55, 21 to 55, 23 to 55, 25 to 55, 27 to 55, 29 to 55, 31 to 55, 33 to 55, 35 to 55, 37 to 55, 39 to 55 or 41 to 55 nucleotides, preferably 42 to 55, 43 to 55, It may be 44 to 55, 45 to 55, 46 to 55, 47 to 55, 48 to 55, 49 to 55, 50 to 55, 51 to 55, 52 to 55, 53 to 55, 54 or 55 nucleotides.
[0514] In some implementations, the MS4 or tracrRNA-crRNA complementary region is X of formula (I). c1 and X c2A region corresponding to or including a polynucleotide indicated by X, wherein the X is a variant in which part or all of the tracrRNA-crRNA complementary region is deleted. c1 and X c2 Each can independently consist of 0 to 35 (poly)nucleotides.
[0515] Preferably, X c1 may be composed of 0 to 28, 0 to 27, 0 to 26, 0 to 25, 0 to 24, 0 to 23, 0 to 22, 0 to 21, 0 to 20, 0 to 19, 0 to 18, 0 to 17, 0 to 16, 0 to 15, 0 to 14, 0 to 13, 0 to 12, 0 to 11, 0 to 10, 0 to 9, 0 to 8, 0 to 7, 0 to 6, 0 to 5, 0 to 4, 0 to 3, 0 to 2, 1 or 0 (poly)nucleotides. Also, preferably, the X c2 may consist of 0 to 27, 0 to 26, 0 to 25, 0 to 24, 0 to 23, 0 to 22, 0 to 21, 0 to 20, 0 to 19, 0 to 18, 0 to 17, 0 to 16, 0 to 15, 0 to 14, 0 to 13, 0 to 12, 0 to 11, 0 to 10, 0 to 9, 0 to 8, 0 to 7, 0 to 6, 0 to 5, 0 to 4, 0 to 3, 0 to 2, 1 or 0 (poly)nucleotides.
[0516] In some implementations, X in the scaffold sequence of formula (I) c1may comprise a nucleic acid sequence of SEQ ID NO: 39 or may comprise a nucleic acid sequence having 1 to 28 nucleotides deleted from the sequence of SEQ ID NO: 39. Preferably, the deletion of nucleotides may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28 nucleotides sequentially removed from the 3'-end of the sequence of SEQ ID NO: 39. More specifically, X c1 silver
[0517] 5'-UUCAUUUUUCCUCUCCAAUUCUGCACAA-3' (SEQ ID NO: 39),
[0518] 5'-UUCAUUUUUCCUCUCCAAUUCUGCACA-3' (SEQ ID NO: 40),
[0519] 5'-UUCAUUUUUCCUCUCCAAUUCUGCAC-3' (SEQ ID NO: 41),
[0520] 5'-UUCAUUUUUCCUCUCCAAUUCUGCA-3' (SEQ ID NO: 42),
[0521] 5'-UUCAUUUUUCCUCUCCAAUUCUGC-3' (SEQ ID NO: 43),
[0522] 5'-UUCAUUUUUCCUCUCCAAUUCUG-3' (SEQ ID NO: 44),
[0523] 5'-UUCAUUUUUCCUCUCCAAUUCU-3' (SEQ ID NO: 45),
[0524] 5'-UUCAUUUUUCCUCUCCAAUUC-3' (SEQ ID NO: 46),
[0525] 5'-UUCAUUUUUCCUCUCCAAUU-3' (SEQ ID NO: 47),
[0526] 5'-UUCAUUUUUCCUCUCCAAU-3' (SEQ ID NO: 48),
[0527] 5'-UUCAUUUUUCCUCUCCAA-3' (SEQ ID NO: 49),
[0528] 5'-UUCAUUUUUCCUCUCCA-3' (SEQ ID NO: 50),
[0529] 5'-UUCAUUUUUCCUCUCC-3' (SEQ ID NO: 51),
[0530] 5'-UUCAUUUUUCCUCUC-3' (SEQ ID NO: 52),
[0531] 5'-UUCAUUUUUCCUCU-3' (SEQ ID NO: 53),
[0532] 5'-UUCAUUUUUCCUC-3' (SEQ ID NO: 54),
[0533] 5'-UUCAUUUUUCCU-3' (SEQ ID NO: 55),
[0534] 5'-UUCAUUUUUCC-3' (SEQ ID NO: 56),
[0535] 5'-UUCAUUUUUC-3' (SEQ ID NO: 57), 5'-UUCAUUUUU-3', 5'-UUCAUUUU-3', 5'-UUCAUUU-3', 5'-UUCAUUU-3', 5'-UUCAUU-3', 5'-UUCAU-3', 5'-UUCA-3', 5'-UUC-3', 5'-UU-3' or 5'-U-3', or X c1 may be non-existent.
[0536] X with some nucleotides removed c1 If there is a region containing three, four, or five or more uracils (U) within the sequence, the modifications in MS1 described above may also be applied. For specific details regarding MS1, refer to the above “(1) Modification at Modification Site 1 (MS1)” section.
[0537] In some implementations, X in the scaffold sequence of formula (I) c2may comprise a nucleic acid sequence of SEQ ID NO: 58 or may comprise a nucleic acid sequence in which 1 to 27 nucleotides are deleted from the sequence of SEQ ID NO: 58. Preferably, the deletion of nucleotides may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 or 27 nucleotides are sequentially removed from the 5'-end of the sequence of SEQ ID NO: 58. More specifically, X c2 Is
[0538] 5'-GUUGCAGAACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 58),
[0539] 5'-UUGCAGAACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 59),
[0540] 5'-UGCAGAACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 60),
[0541] 5'-GCAGAACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 61),
[0542] 5'-CAGAACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 62),
[0543] 5'-AGAACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 63),
[0544] 5'-GAACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 64),
[0545] 5'-AACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 65),
[0546] 5'-ACCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 66),
[0547] 5'-CCCGAAUAGACGAAUGAA-3' (SEQ ID NO: 67),
[0548] 5'-CCGAAUAGACGAAUGAA-3' (SEQ ID NO: 68),
[0549] 5'-CGAAUAGACGAAUGAA-3' (SEQ ID NO: 69),
[0550] 5'-GAAUAGACGAAUGAA-3' (SEQ ID NO: 70),
[0551] 5'-AAUAGACGAAUGAA-3' (SEQ ID NO: 71),
[0552] 5'-AUAGACGAAUGAA-3' (SEQ ID NO: 72),
[0553] 5'-UAGACGAAUGAA-3' (SEQ ID NO: 73),
[0554] 5'-AGACGAAUGAA-3' (SEQ ID NO: 74), 5'-GACGAAUGAA-3' (SEQ ID NO: 75), 5'-ACGAAUGAA-3', 5'-CGAAUGAA-3', 5'-GAAUGAA-3', 5'-AAUGAA-3', 5'-AUGAA-3', 5'-UGAA-3', 5'-GAA-3', 5'-AA-3' or 5'-A-3', or X c2 may not exist.
[0555] X with some nucleotides removed c2 Within the sequence, X c1 If there is a sequence corresponding to a sequence containing three or more Us, or three, four, or five or more Us, the modification in MS1 described above may also be applied. For specific details on MS1, refer to the above “(1) Modification at modification site 1 (MS1)”.
[0556] X in the scaffold sequence of formula (I) c1 and X c2The corresponding regions can be independently subjected to the above-described modifications, but the MS4 or tracrRNA-crRNA complementary region is a region where tracrRNA and crRNA form complementary binding, and must be X to function as a dual guide RNA. c1 and X c2 It is desirable that the positions and numbers of nucleotides missing in each be the same or similar. That is, X c1 and X c2 In order to preserve the complementarity of the sequences, it is preferable to sequentially delete the sequences located at the 3'-end of tracrRNA in MS4 (tracrRNA-crRNA complementarity region), and to sequentially delete the sequences located at the 5'-end of crRNA. In some implementations according to this viewpoint, X c1 and X c2 A deletion in a nucleic acid sequence may be a deletion of one or more complementary nucleotide pairs.
[0557] In some implementations, X in the scaffold sequence of formula (I) c1 3'-end and X c2The 5'-end of the tracrRNA and crRNA can be modified into a single guide RNA (sgRNA) form by being linked with a linker (Lk). The Lk is a sequence that physically or chemically links the tracrRNA and crRNA, and can be a polynucleotide sequence having a length of 1 to 30 nucleotides. In some embodiments, the Lk can be a sequence of 1 to 5, 5 to 10, 10 to 15, 2 to 20, 15 to 20, 20 to 25, or 25 to 30 nucleotides. For example, the Lk can be a 5'-GAAA-3' sequence, but is not limited thereto. As another example, the Lk may be a linker comprising or consisting of the sequence 5'-UUAG-3', 5'-UGAAAA-3', 5'-UUGAAAAA-3', 5'-UUCGAAAGAA-3' (SEQ ID NO: 76), 5'-UUCAGAAAUGAA-3' (SEQ ID NO: 77), 5'-UUCAUGAAAAUGAA-3' (SEQ ID NO: 78), or 5'-UUCAUUGAAAAAUGAA-3' (SEQ ID NO: 79).
[0558] It is possible to use a linker (Lk) to make a single guide RNA (sgRNA), but it is also possible to directly connect the 3'-end of a tracrRNA with some sequences removed from the 3'-end and the 3'-end of a crRNA with some sequences removed from the 5'-end.
[0559] In some embodiments, X in the scaffold sequence of formula (I) c1 and X c2 When connected by a linker, 5'-X as shown in formula (I) c1 -Lk-X c2 -3' can be expressed as, and the above 5'-X c1 -Lk-X c2 -3' is sequence number 80 to sequence number 86 and 5'-Lk-3'(X c1 and X c2 It may be any one nucleic acid sequence selected from the group consisting of, but is not limited to, all of these missing forms.
[0560] (5) Modification at modification site 5 (MS5)
[0561] This article describes modifications in MS5 (Fig. 2). As described above, MS5 corresponds to a region located 3'-endward within tracrRNA, referred to as the second stem-loop region. The second stem-loop region comprises nucleotides that form a stem structure within the guide RNA (gRNA) and nucleic acid editing protein complex, and may include adjacent nucleotides. The stem or stem-loop structure is distinct from the stem contained in the first stem-loop region described above.
[0562] In some implementations, the second stem-loop region may comprise the nucleotide sequence of SEQ ID NO: 25 and / or the nucleotide sequence of SEQ ID NO: 29.
[0563] In some implementations, the MS5 or second stem-loop region is X of formula (I). b1 and X b2 A region comprising a polynucleotide represented by X and an adjacent (poly)nucleotide (including a loop of the 5'-UUAG-3' sequence), wherein a part or all of the second stem-loop region is deleted by a modification b1 and X b2 Each can independently consist of 0 to 35 (poly)nucleotides.
[0564] In some implementations, the engineered gRNA comprises a variant in which part or all of the second stem-loop region is deleted.
[0565] In some implementations, the engineered gRNA comprises a deletion of part or all of the second stem-loop region, wherein the part or all of the second stem-loop region that is deleted may be from 1 to 27 nucleotides. Specifically, part or all of the second stem region comprises 2 to 27, 3 to 27, 4 to 27, 5 to 27, 6 to 27, 7 to 27, 8 to 27, 9 to 27, 10 to 27, 11 to 27, 12 to 27, 13 to 27, 14 to 27, 15 to 27, 16 to 27, 17 to 27, 18 to 27, 19 to 27, 20 to 27, 21 to 27, 22 to 27, 23 to 27, 24 to 27, 25 to 27, It can be 26 or 27 nucleotides.
[0566] Preferably, X of formula (I) b1 may be composed of 0 to 13, 0 to 12, 0 to 11, 0 to 10, 0 to 9, 0 to 8, 0 to 7, 0 to 6, 0 to 5, 0 to 4, 0 to 3, 0 to 2, 1 or 0 (poly)nucleotides. Also, preferably, the X b2 may consist of 0 to 14, 0 to 13, 0 to 12, 0 to 11, 0 to 10, 0 to 9, 0 to 8, 0 to 7, 0 to 6, 0 to 5, 0 to 4, 0 to 3, 0 to 2, 1 or 0 (poly)nucleotides.
[0567] In some implementations, X in the scaffold sequence of formula (I) b1may comprise a nucleic acid sequence of SEQ ID NO: 25 or may comprise a nucleic acid sequence having 1 to 13 nucleotides deleted from the sequence of SEQ ID NO: 25. Preferably, the deletion of the nucleotides may be a sequential removal of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 nucleotides from the 3'-end of the sequence of SEQ ID NO: 25. More specifically, X b1 may comprise or consist of the sequence of 5'-CAAAAGCUGUCCC-3' (SEQ ID NO: 25), 5'-CAAAAGCUGUCC-3' (SEQ ID NO: 26), 5'-CAAAAGCUGUC-3' (SEQ ID NO: 27), 5'-CAAAAGCUGU-3' (SEQ ID NO: 28), 5'-CAAAAGCUG-3', 5'-CAAAAGCU-3', 5'-CAAAAGC-3', 5'-CAAAAG-3', 5'-CAAAA-3', 5'-CAAA-3', 5'-CAA-3', 5'-CA-3' or 5'-C-3', or X b1 may be non-existent.
[0568] In some implementations, X in the scaffold sequence of formula (I) b2 may comprise a nucleic acid sequence of SEQ ID NO: 29 or may comprise a nucleic acid sequence in which 1 to 14 nucleotides are deleted from the sequence of SEQ ID NO: 29. Preferably, the deletion of the nucleotides may be a sequential removal of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 14 nucleotides from the 5'-end of the sequence of SEQ ID NO: 29. More specifically, X b2may comprise or consist of the sequence of 5'-GGGAUUAGAACUUG-3' (SEQ ID NO: 29), 5'-GGAUUAGAACUUG-3' (SEQ ID NO: 30), 5'-GAUUAGAACUUG-3' (SEQ ID NO: 31), 5'-AUUAGAACUUG-3' (SEQ ID NO: 32), 5'-UUAGAACUUG-3' (SEQ ID NO: 33), 5'-UAGAACUUG-3', 5'-AGAACUUG-3', 5'-GAACUUG-3', 5'-AACUUG-3', 5'-ACUUG-3', 5'-CUUG-3', 5'-UUG-3', 5'-UG-3' or 5'-G-3', or X b1 may be non-existent.
[0569] X in the scaffold sequence of formula (I) b1 and X b2 The corresponding regions can be deformed independently, but X is used to preserve the normal stem-loop structure. b1 and X b2 It is desirable that the positions and numbers of nucleotides missing in each be the same or similar. For example, X b1 In case of sequential deletion starting from the 5'-terminal sequence, X b2 In , it is desirable to sequentially delete sequences starting from the 3'-terminal direction. In some implementations according to this viewpoint, X b1 and X b2 A deletion in a nucleic acid sequence may be a deletion of one or more complementary nucleotide pairs.
[0570] In some implementations, X of the scaffold sequence of formula (I) b1 and X b2The sequence of the loop portion connecting is indicated as 5'-UUAG-3', but this can be replaced with other sequences such as 5'-NNNN-3', '5-NNN-3', etc. as needed. Here, N is each independently A, C, G, or U. For example, 5'-NNNN-3' can be 5'-GAAA-3', and the '5-NNN-3' can be 5'-CGA-3'.
[0571] For example, in the scaffold sequence of formula (I), X b1 and X b2 The sequence of the loop portion connecting is 5'-UUAG-3', and the sequence 5'-X in the above formula (I) b1 UUAGX b2 -3' is sequence number 34 to sequence number 38 and 5'-UUAG-3'(X b1 and X b2 It may comprise or consist of any one nucleic acid sequence selected from the group consisting of (all of which are in the form of deletion).
[0572] (6) Examples of gRNAs with modifications applied at modification sites 1 to 5.
[0573] The engineered guide RNA disclosed herein may comprise a modification at two or more of the modification sites 1 (MS1) to 5 (MS5) described above.
[0574] In some embodiments, the engineered guide RNA may comprise one or more modifications selected from the group consisting of (a1) deletion of part or all of the first stem-loop region; (a2) deletion of part or all of the second stem-loop region; (b) deletion of part or all of the tracrRNA-crRNA complementary region; (c) substitution of one or more Us with A, G or C when three or more, four or more or five or more consecutive Us are present within the tracrRNA-crRNA complementary region; and (d) addition of a U-rich tail to the 3'-end of the crRNA sequence. The sequence of the U-rich tail is 5'-(U m V) n U o -3', where V is independently A, C or G, m and o are integers from 1 to 20, and n is an integer from 0 to 5.
[0575] For example, the engineered guide RNA may comprise (d) addition of a U-rich tail to the 3'-end of the crRNA sequence and (c) substitution of one or more Us with A, G or C when three or more, four or more or five or more consecutive Us are present within the tracrRNA-crRNA complementarity region.
[0576] As another example, the engineered guide RNA may comprise (d) addition of a U-rich tail to the 3'-end of the crRNA sequence, (c) substitution of one or more Us with A, G or C when three or more, four or more or five or more consecutive Us are present within the tracrRNA-crRNA complementarity region, and (a1) deletion of part or all of the first stem-loop region.
[0577] As another example, the engineered guide RNA may comprise (d) addition of a U-rich tail to the 3'-end of the crRNA sequence, (a1) deletion of part or all of the first stem-loop region, and (b) deletion of part or all of the tracrRNA-crRNA complementary region, wherein if there are three or more, four or more, or five or more consecutive uracils (U) within the tracrRNA-crRNA complementary region comprising the partial deletion, a substitution of one or more Us with A, G, or C may be further included.
[0578] As another example, the engineered guide RNA may comprise (d) addition of a U-rich tail to the 3'-end of the crRNA sequence, (a1) deletion of part or all of the first stem-loop region, (b) deletion of part or all of the tracrRNA-crRNA complementary region, and (a2) deletion of part or all of the second stem-loop region, wherein if there are three or more, four or more, or five or more consecutive uracils (U) within the tracrRNA-crRNA complementary region comprising the partial deletion, a substitution of one or more Us with A, G, or C may be further included.
[0579] As an example of a tracrRNA having a modification applied at the multiple modification sites (MS) described above, an engineered tracrRNA comprising a nucleotide sequence of SEQ ID NO: 87 to SEQ ID NO: 132 is provided. Specifically, the engineered tracrRNA comprises a nucleotide sequence of SEQ ID NO: 87 (MS1), SEQ ID NO: 88 (MS1 / MS3-1), SEQ ID NO: 89 (MS1 / MS3-2), SEQ ID NO: 90 (MS1 / MS3-3), SEQ ID NO: 91 (MS1 / MS4). * -1), sequence number 92 (MS1 / MS4 * -2), sequence number 93 (MS1 / MS4 *-3), SEQ ID NO: 94 (MS1 / MS5-1), SEQ ID NO: 95 (MS1 / MS5-2), SEQ ID NO: 96 (MS1 / MS5-3), SEQ ID NO: 97 (MS1 / MS3-3 / MS4 * -1), sequence number 98 (MS1 / MS3-3 / MS4 * -2), sequence number 99 (MS1 / MS3-3 / MS4 * -3), sequence number 100 (MS1 / MS4 * -2 / MS5-1), sequence number 101 (MS1 / MS4 * -2 / MS5-2), sequence number 102 (MS1 / MS4 * -2 / MS5-3), SEQ ID NO: 103 (MS1 / MS3-3 / MS5-1), SEQ ID NO: 104 (MS1 / MS3-3 / MS5-2), SEQ ID NO: 105 (MS1 / MS3-3 / MS5-3), SEQ ID NO: 106 (MS1 / MS3-3 / MS4 *-2 / MS5-3), SEQ ID NO: 107 (mature form, MF), SEQ ID NO: 108 (MF / MS3-1), SEQ ID NO: 109 (MF / MS3-2), SEQ ID NO: 110 (MF / MS3-3), SEQ ID NO: 111 (MF / MS4-1), SEQ ID NO: 112 (MF / MS4-2), SEQ ID NO: 113 (MF / MS4-3), SEQ ID NO: 114 (MF / MS5-1), SEQ ID NO: 115 (MF / MS5-2), SEQ ID NO: 116 (MF / MS5-3), SEQ ID NO: 117 (MF / MS5), SEQ ID NO: 118 (MF / MS3-3 / MS4-1), SEQ ID NO: 119 (MF / MS3-3 / MS4-2), SEQ ID NO: 120 (MF / MS3-3 / MS4-3), SEQ ID NO: 121 (MF / MS4-3 / MS5-1), SEQ ID NO: It may comprise or consist of the nucleotide sequence of SEQ ID NO: 122 (MF / MS4-3 / MS5-2), SEQ ID NO: 123 (MF / MS4-3 / MS5-3), SEQ ID NO: 124 (MF / MS4-3 / MS5), SEQ ID NO: 125 (MF / MS3-3 / MS5-1), SEQ ID NO: 126 (MF / MS3-3 / MS5-2), SEQ ID NO: 127 (MF / MS3-3 / MS5-3), SEQ ID NO: 128 (MF / MS3-3 / MS5), SEQ ID NO: 129 (MF / MS3-3 / MS4-3 / MS5-3), SEQ ID NO: 130 (MF / MS3-3 / MS4-1 / MS5), SEQ ID NO: 131 (MF / MS3-3 / MS4-2 / MS5), or SEQ ID NO: 132 (MF / MS3-3 / MS4-3 / MS5).
[0580] As a more specific example, exemplary sequences of engineered tracrRNAs having one or more modifications at one or more modification sites selected from MS1, MS3, MS4, and MS5 are provided in Table 5 below. Such engineered tracrRNAs constitute part of the scaffold sequence of the scaffold region.
[0581]
[0582]
[0583]
[0584] Additionally, an engineered crRNA comprising a nucleotide sequence of SEQ ID NO: 133 to SEQ ID NO: 148 is provided as an example of a crRNA having modifications applied at multiple modification sites (MS).
[0585] Specifically, the engineered crRNA of the present invention comprises SEQ ID NO: 133 (MS1), SEQ ID NO: 134 (MS1 / MS4) * -1), sequence number 135 (MS1 / MS4 * -2), sequence number 136 (MS1 / MS4 * -3), SEQ ID NO: 137 (mature form; MF), SEQ ID NO: 138 (MF / MS4-1), SEQ ID NO: 139 (MF / MS4-2), SEQ ID NO: 140 (MF / MS4-3), SEQ ID NO: 141 (MS1 / MS2), SEQ ID NO: 142 (MS1 / MS2 / MS4 * -1), sequence number 143 (MS1 / MS2 / MS4 * -2), sequence number 144 (MS1 / MS2 / MS4 * -3), may comprise or consist of the nucleotide sequence of SEQ ID NO: 145 (MF / MS2), SEQ ID NO: 146 (MF / MS2 / MS4-1), SEQ ID NO: 147 (MF / MS2 / MS4-2), or SEQ ID NO: 148 (MF / MS2 / MS4-3).
[0586] As some embodiments, exemplary sequences of engineered crRNAs having one or more modifications at one or more of the modification sites selected from MS1, MS2, and MS4 are provided in Table 6 below.
[0587]
[0588] In Table 6, all crRNA sequences, except where necessary, omit the guide sequence (spacer) indication, and the sequence indicated as 'NNNNNNNNNNNNNNNNNNNN' means any guide sequence (spacer) that can hybridize with the target sequence within the target gene. The guide sequence can be appropriately designed by a person skilled in the art according to the desired target gene and / or the target sequence within the target gene as described above, and therefore is not limited to a specific sequence of a specific length.
[0589] In some embodiments, the scaffold region of the engineered gRNA may comprise a tracrRNA comprising or consisting of any one nucleic acid sequence selected from the group consisting of SEQ ID NO: 87 to SEQ ID NO: 132; and a crRNA comprising or consisting of any one nucleic acid sequence selected from the group consisting of SEQ ID NO: 133 to SEQ ID NO: 148. The scaffold region of any one nucleic acid sequence selected from the group consisting of SEQ ID NO: 141 to SEQ ID NO: 148 refers to the remaining region excluding the spacer region (the region indicated as 5'-NNNNNNNNNNNNNNNNNNNN-3' in the nucleic acid sequence) and the U-rich tail region (the region indicated as 5'-UUUUAUUUUU-3' in the nucleic acid sequence) present at the 3'-end of the crRNA.
[0590] In some embodiments, the guide RNA of the present invention may comprise a sequence of a scaffold region of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 149 to 186 and 196. Here, the scaffold region of the nucleic acid sequence refers to the remaining region excluding the spacer region present at the 3'-terminal portion of the crRNA (the region indicated as 5'-NNNNNNNNNNNNNNNNNN-3' in the nucleic acid sequence).
[0591] In some embodiments, when the engineered gRNA of the present invention is in the form of a single guide RNA (sgRNA), the scaffold region of the engineered sgRNA may comprise or consist of any one nucleic acid sequence selected from the group consisting of SEQ ID NOs: 149 to 186 and 196 (provided that the sequence 5'-NNNNNNNNNNNNNNNNNNNN-3', 5'-NNNNNNNNNNNNNNNNNUUUUAUUUU-3' or 5'-NNNNNNNNNNNNNNNNNNNUUUUAUUUUU-3' present at the 3'-end of SEQ ID NOs: 149 to 186 or SEQ ID NO: 196 is excluded).
[0592] For example, the engineered sgRNA can be an sgRNA of SEQ ID NO: 149 comprising a modification in MS1, an sgRNA of SEQ ID NO: 150 comprising a modification in MS1 / MS2, an sgRNA of SEQ ID NO: 151 comprising a modification in MS1 / MS2 / MS3, an sgRNA of SEQ ID NO: 152 comprising a modification in MS2 / MS3 / MS4, or an sgRNA of SEQ ID NO: 153 comprising a modification in MS2 / MS3 / MS4 / MS5. In the nucleic acid sequences of SEQ ID NOs: 149 to 153, the sequence indicated as 5'-NNNNNNNNNNNNNNNNNN-3' refers to a guide sequence.
[0593] In another specific example, the engineered sgRNA comprises SEQ ID NO: 154 (MS1 / MS3-1), SEQ ID NO: 155 (MS1 / MS3-2), SEQ ID NO: 156 (MS1 / MS3-3), SEQ ID NO: 157 (MS1 / MS4 * -1), sequence number 158 (MS1 / MS4 * -2), sequence number 159 (MS1 / MS4 * -3), SEQ ID NO: 160 (MS1 / MS5-1), SEQ ID NO: 161 (MS1 / MS5-2), SEQ ID NO: 162 (MS1 / MS5-3), SEQ ID NO: 163 (MS1 / MS2 / MS4 * -2), sequence number 164 (MS1 / MS3-3 / MS4 *-2), SEQ ID NO: 165 (MS1 / MS2 / MS5-3), SEQ ID NO: 166 (MS1 / MS3-3 / MS5-3), SEQ ID NO: 167 (MS1 / MS4 * -2 / MS5-3), sequence number 168 (MS1 / MS2 / MS3-3 / MS4 * -2), SEQ ID NO: 169 (MS1 / MS2 / MS3-3 / MS5-3), SEQ ID NO: 170 (MS1 / MS2 / MS4 * -2 / MS5-3), sequence number 171 (MS1 / MS3-3 / MS4 * -2 / MS5-3) or sequence number 172 (MS1 / MS2 / MS3-3 / MS4 * -2 / MS5-3) may be an sgRNA comprising or consisting of a nucleotide sequence. In these nucleic acid sequences, the sequence indicated as 5'-NNNNNNNNNNNNNNNNNN-3' indicates a guide sequence.
[0594] Additionally, the sgRNA may be an sgRNA comprising or consisting of the nucleotide sequence of SEQ ID NO: 173, which is a mature form (abbreviated as MF) of the sgRNA.
[0595] In another specific embodiment, an exemplary sgRNA comprising a modification of a nucleic acid sequence in an MF sgRNA is provided. Specifically, the MF sgRNA may be an sgRNA comprising or consisting of a nucleotide sequence of SEQ ID NO: 174 (MS3-1), SEQ ID NO: 175 (MS3-2), SEQ ID NO: 176 (MS3-3), SEQ ID NO: 177 (MS4-1), SEQ ID NO: 178 (MS4-2), SEQ ID NO: 179 (MS4-3), SEQ ID NO: 180 (MS5-1), SEQ ID NO: 181 (MS5-2), SEQ ID NO: 182 (MS5-3), SEQ ID NO: 183 (MS3-3 / MS4-3), SEQ ID NO: 184 (MS3-3 / MS5-3), SEQ ID NO: 185 (MS4-3 / MS5-3), or SEQ ID NO: 186 (MS3-3 / MS4-3 / MS5-3). In these nucleic acid sequences, the sequence indicated as 5'-NNNNNNNNNNNNNNNNNNNN-3' represents the guide sequence.
[0596] In a preferred embodiment, the engineered sgRNA may be comprised of the nucleotide sequence of SEQ ID NO: 151 (Cas12f1 ver3.0), SEQ ID NO: 152 (Cas12f1 ver4.0), or SEQ ID NO: 153 (Cas12f1 ver4.1). In these nucleic acid sequences, the sequence indicated as 5'-NNNNNNNNNNNNNNNNNNNN-3' represents the guide sequence.
[0597] In a preferred embodiment, the engineered sgRNA may comprise the nucleotide sequence of SEQ ID NO: 156 (Cas12f1 ver3.0), SEQ ID NO: 183 (Cas12f1 ver4.0), or SEQ ID NO: 196 (Cas12f1 ver4.1) (excluding the U-rich tail sequence). In these nucleic acid sequences, the sequence indicated as 5'-NNNNNNNNNNNNNNNNNNNN-3' represents the guide sequence.
[0598] The engineered sgRNA is an engineered gRNA having one or more of the five modification sites (MS1, MS2, MS3, MS4, and MS5) as shown in FIG. 2, and the specific sequences are shown in Table 7 below.
[0599]
[0600]
[0601]
[0602]
[0603] In addition, a mature form gRNA was produced in which the variant region MS1 was removed from the canonical gRNA, and the specific sequence is shown in Table 8 below.
[0604]
[0605]
[0606] In some implementations, the (engineered) guide RNA may comprise any one sequence selected from the group consisting of SEQ ID NOs: 13, 149 to 186, and 196.
[0607] In some embodiments, the scaffold region of the (engineered) guide RNA may comprise or consist of any one sequence selected from the group consisting of SEQ ID NOs: 13, 149 to 186, and 196. Here, the scaffold region means the remaining region excluding the spacer region (the region indicated as 5'-NNNNNNNNNNNNNNNNNNNN-3' in the nucleic acid sequence) and / or the U-rich tail region (the region indicated as 5'-UUUUAUUUUUU-3' in the nucleic acid sequence) present at the 3'-terminal portion of the guide RNA.
[0608] In some implementations, the (engineered) guide RNA comprises any one sequence selected from the group consisting of SEQ ID NOs: 13, 149 to 186, and 196, and may further comprise a U-tail sequence disclosed herein at the 3' end of the guide sequence.
[0609] In some implementations, the (engineered) guide RNA may comprise any one sequence selected from the group consisting of SEQ ID NOs: 156, 183, and 196.
[0610] In some embodiments, the scaffold region of the (engineered) guide RNA may comprise or consist of any one sequence selected from the group consisting of SEQ ID NOs: 156, 183, and 196. Here, the scaffold region means the remaining region excluding the spacer region (the region indicated as 5'-NNNNNNNNNNNNNNNNNN-3' in the nucleic acid sequence) present at the 3'-terminal portion of the guide RNA.
[0611] In some embodiments, the (engineered) guide RNA comprises any one sequence selected from the group consisting of SEQ ID NOs: 156, 183, and 196, and may further comprise a U-tail sequence disclosed herein at the 3' end of the guide sequence.
[0612] In some implementations, the (engineered) guide RNA may comprise the sequence of SEQ ID NO: 183.
[0613] In some implementations, the (engineered) guide RNA comprises the sequence of SEQ ID NO: 183 and may further comprise a U-tail sequence disclosed herein at the 3' end of the guide sequence.
[0614] In some implementations, the (engineered) guide RNA may comprise or consist of any one sequence selected from the group consisting of SEQ ID NOs: 151 to 153.
[0615] In some implementations, the (engineered) guide RNA may comprise or consist of the sequence of SEQ ID NO: 152.
[0616] (7) Additional sequence
[0617] The engineered tracrRNA of the present invention may optionally further comprise an additional sequence. In some embodiments, the additional sequence may be located at the 3'-end of the engineered tracrRNA. In some embodiments, the additional sequence may be located at the 5'-end of the engineered tracrRNA. For example, the additional sequence may be located at the 5'-end of the first stem-loop region.
[0618] The additional sequence may be from 1 to 40 nucleotides. In some embodiments, the additional sequence may be any nucleotide sequence or any randomly arranged nucleotide sequence. For example, the additional sequence may be the sequence 5'-AUAAAGGUGA-3' (SEQ ID NO: 187).
[0619] Additionally, the additional sequence may be a known nucleotide sequence. For example, the additional sequence may be a hammerhead ribozyme nucleotide sequence. For example, the nucleotide sequence of the hammerhead ribozyme may be the sequence 5'-CUGAUGAGUCCGUGAGGACGAAACGAGUAAGCUCGUC-3' (SEQ ID NO: 188) or the sequence 5'-CUGCUCGAAUGAGCAAAGCAGGAGUGCCUGAGUAGUC-3' (SEQ ID NO: 189). The above-mentioned sequences are merely examples, and the additional sequence is not limited thereto.
[0620] (8) Chemical modification
[0621] In some embodiments, the engineered tracrRNA or engineered crRNA included in the engineered gRNA may optionally have at least one or more nucleotides chemically modified. The chemical modification may be a modification of various covalent bonds that may occur in the bases and / or sugars of the nucleotides.
[0622] In some embodiments, the chemical modification may be, but is not limited to, methylation, halogenation, acetylation, phosphorylation, phosphorothioate (PS) linkage, locked nucleic acid (LNA), 2'-O-methyl 3'phosphorothioate (MS), or 2'-O-methyl 3'thioPACE (MSP).
[0623] When using an ultra-small gene editing system comprising the engineered gRNA and engineered Cas polypeptide of the present invention, the indel efficiency of a target gene or target nucleic acid in a cell is significantly improved compared to when using a guide RNA found in nature, resulting in large-scale deletion effects. Furthermore, the engineered Cas polypeptide of the present invention has the advantage of having altered PAM sequence specificity, thereby expanding the targetable genome range. Furthermore, the engineered Cas polypeptide of the present invention has the advantage of being able to edit a wider range of target bases due to the expanded editing window.
[0624] Engineered gRNAs can be optimized for length to exhibit high efficiency and thus reduce the cost of gRNA synthesis, secure additional space or capacity when inserted into a viral vector, normal expression of gRNA, increase in operable gRNA expression, increase in stability of gRNA, increase in stability of gRNA and Cas polypeptide complex, induce formation of highly efficient gRNA and Cas polypeptide complex, increase in cleavage efficiency of target nucleic acids by ultra-small gene editing systems including gRNA and Cas polypeptide complex, and expand the targetable genome range.
[0625] Engineered guide RNAs according to embodiments of the present invention may be single guide RNAs or dual guide RNAs. Dual guide RNA refers to a guide RNA composed of two RNA molecules, a tracrRNA and a crRNA. Single guide RNA (sgRNA) refers to a guide RNA in which the 3'-end of the tracrRNA and the 5'-end of the crRNA are connected via a linker.
[0626] In some embodiments, the engineered single guide RNA (sgRNA) may additionally comprise a linker sequence, and the tracrRNA sequence and the crRNA sequence may be connected via the linker sequence. Preferably, the engineered scaffold sequence may comprise the 3'-end of the tracrRNA-crRNA complementary sequence of the tracrRNA and the 5'-end of the tracrRNA-crRNA complementary sequence of the crRNA connected via a linker. More preferably, the tracrRNA-crRNA complementary regions of the tracrRNA and crRNA may be connected at their 3'-end and 5'-end, respectively, via a linker 5'-GAAA-3'. For specific details regarding the linker, refer to the description of Lk in the above-described formula (I).
[0627] In some embodiments, the sequence of the single guide RNA comprises a tracrRNA sequence, a linker sequence, a crRNA sequence, and a U-rich tail sequence sequentially connected from the 5'-end to the 3'-end. A portion of the tracrRNA sequence and all or a portion of the CRISPR RNA repeat sequences included in the crRNA sequence have complementary sequences to each other.
[0628] In addition, the engineered guide RNA according to an embodiment of the present invention may be a dual guide RNA in which the tracrRNA and the crRNA form separate RNA molecules. A portion of the tracrRNA and a portion of the crRNA may have complementary sequences to form a double-stranded RNA. More specifically, in the dual guide RNA, a portion including the 3'-end of the tracrRNA and a portion including the CRISPR RNA repeat sequence of the crRNA may form a double strand. The engineered guide RNA may bind to Cas12f or a variant protein thereof to form a complex of the guide RNA and the protein, and recognize a target sequence complementary to the guide sequence included in the crRNA sequence, thereby editing a target gene or target nucleic acid including the target sequence.
[0629] In some embodiments, the tracrRNA sequence may comprise a complementary sequence having 0 to 20 mismatches with the CRISPR RNA repeat sequence. Preferably, the tracrRNA sequence may comprise a complementary sequence having 0 to 8 or 8 to 12 mismatches with the CRISPR RNA repeat sequence.
[0630] IV. Nucleic Acids
[0631] In another aspect of the present invention, a nucleic acid encoding an engineered Cas polypeptide as described herein or a nucleic acid encoding a guide RNA as described herein is provided. Also, in some embodiments, a nucleic acid encoding a fusion protein comprising an engineered Cas polypeptide as described herein is provided.
[0632] In some embodiments, the nucleic acid may comprise or consist of DNA. In some embodiments, the nucleic acid may comprise or consist of RNA.
[0633] In some embodiments, the nucleic acid may be codon-optimized for expression in a cell or organism (e.g., a cell or organism in which genome editing is desired).
[0634] In some embodiments, the nucleic acid may be a human codon-optimized nucleic acid sequence. For example, a nucleic acid encoding a human codon-optimized CWCas12f1 polypeptide may comprise or consist of the nucleic acid sequence of SEQ ID NO: 6. For example, a nucleic acid encoding a human codon-optimized UnCas12f1 protein may comprise or consist of the nucleic acid sequence of SEQ ID NO: 10.
[0635] V. Vectors or Vector Systems
[0636] In another aspect of the present invention, a vector is provided comprising one or more nucleic acid sequences encoding an engineered Cas polypeptide as described herein and optionally one or more nucleic acid sequences encoding a guide RNA as described herein.
[0637] In some embodiments, a vector system is provided comprising one or more vectors comprising (i) a first nucleic acid construct comprising a nucleotide sequence encoding an engineered Cas polypeptide as described herein; and (ii) a second nucleic acid construct comprising a nucleotide sequence encoding a guide RNA as described herein.
[0638] In some embodiments, each nucleic acid construct may further comprise a promoter operably linked to the nucleotide sequence.
[0639] In some embodiments, the promoter contained in the first nucleic acid construct and the promoter contained in the second nucleic acid construct may be the same or different.
[0640] In some embodiments, the first nucleic acid construct and the second nucleic acid construct may be contained in a single vector.
[0641] In some embodiments, the first nucleic acid construct and the second nucleic acid construct may each be contained in different vectors.
[0642] In some embodiments, the vector system can further comprise (iii) a third nucleic acid construct comprising a nucleotide sequence encoding a guide RNA as described herein. For example, a vector system is provided comprising one or more vectors comprising (i) a first nucleic acid construct comprising a nucleotide sequence encoding an engineered Cas polypeptide as described herein and a first promoter operably linked to the sequence; (ii) a second nucleic acid construct comprising a nucleotide sequence encoding a guide RNA as described herein (a first guide RNA) and a second promoter operably linked to the sequence; and (iii) a third nucleic acid construct comprising a nucleotide sequence encoding a guide RNA as described herein (a second guide RNA) and a third promoter operably linked to the sequence.
[0643] In some embodiments, the first and second guide RNAs may target the same target or different target sequences. For example, a first and second guide RNA targeting the same target may be used to increase the number of expressed copies. For example, a first and second guide RNA targeting different target sequences may be used to remove a desired nucleic acid segment.
[0644] In some implementations, the first, second and third promoters may be the same or different.
[0645] In some embodiments, the first nucleic acid construct, the second nucleic acid construct, and the third nucleic acid construct may be contained in one vector, in any combination of two vectors, or may each be contained in different vectors.
[0646] In some embodiments, the vector system must include one or more regulatory and / or control elements to directly express it within a cell. Specifically, the regulatory and / or control elements may include, but are not limited to, a promoter, an enhancer, an intron, a polyadenylation signal, a Kozak consensus sequence, an internal ribosome entry site (IRES), a splice acceptor, a 2A sequence, and / or an origin of replication. The origin of replication may be, but is not limited to, the f1 origin of replication, the SV40 origin of replication, the pMB1 origin of replication, the adeno origin of replication, the AAV origin of replication, and / or the BBV origin of replication.
[0647] In some embodiments, to express the nucleic acid sequences encoding the components included in the vector system within a cell, a promoter sequence may be operably linked to the sequence encoding each component, thereby enabling activation of an RNA transcription factor within the cell. The promoter sequence may be designed differently depending on the corresponding RNA transcription factor or expression environment, and is not limited as long as it can appropriately express the components of the gene editing system of the present invention within the cell.
[0648] In some embodiments, the promoter sequence may be a promoter that promotes transcription of RNA polymerase (RNA Pol I, Pol II, or Pol III). Specifically, the promoter may be one of a U6 promoter, an EFS promoter, an EF1-α promoter, an H1 promoter, a 7SK promoter, a CMV promoter, an LTR promoter, an Ad MLP promoter, an HSV promoter, an SV40 promoter, a CBA promoter, or an RSV promoter.
[0649] In some embodiments, when the vector sequence includes a promoter sequence, transcription of a sequence operably linked to the promoter is induced by an RNA transcription factor, and a termination signal that induces transcription termination of the RNA transcription factor may be included. The termination signal may vary depending on the type of the promoter sequence. Specifically, when the promoter is a U6 or H1 promoter, the promoter recognizes a sequence of thymidine (T), such as TTTTT (T5) or TTTTTT (T6), as a termination signal.
[0650] The sequence of the engineered guide RNA of the present invention may include a U-rich tail sequence at the 3'-terminus. Accordingly, the sequence encoding the engineered guide RNA includes a T-rich sequence corresponding to the U-rich tail sequence at its 3'-terminus. As described above, some promoter sequences recognize a thymidine (T) sequence, for example, a sequence of five or more thymidines (T) in a row, as a termination signal, and therefore, in some cases, the T-rich sequence may be recognized as a termination signal. Specifically, when the vector sequence provided herein includes a sequence encoding an engineered guide RNA, a sequence encoding a U-rich tail sequence included in the engineered gRNA sequence may be used as a termination signal.
[0651] In some implementations, when a vector sequence comprises a U6 or H1 promoter sequence and a sequence encoding an engineered guide RNA operably linked thereto, a portion of the sequence encoding a U-rich tail sequence included in the guide RNA sequence may be recognized as a termination signal. Specifically, the U-rich tail sequence may comprise a sequence of five or more consecutively linked uridines (U).
[0652] In some embodiments, the vector system may further comprise a nucleic acid sequence encoding an additional expression element that a person skilled in the art desires to express. For example, the additional expression element may be a tag. In another example, the additional expression element may be a herbicide resistance gene, such as glyphosate, glufosinate ammonium, or phosphinothricin, or an antibiotic resistance gene, such as ampicillin, kanamycin, G418, bleomycin, hygromycin, or chloramphenicol. The additional expression element may be expressed independently of the engineered Cas polypeptide and / or guide RNA disclosed herein. Alternatively, the additional expression element may be expressed in conjunction with the engineered Cas polypeptide and / or guide RNA disclosed herein. Additional expression elements may be components that are commonly expressed when expressing the CRISPR / Cas system, and reference is made to known techniques in the art.
[0653] The term "tag" as used herein refers to a functional domain added to facilitate tracking and / or purification of a peptide or protein. For example, tags include, but are not limited to, tag proteins such as histidine (His) tag, V5 tag, FLAG tag, influenza hemagglutinin (HA) tag, Myc tag, VSV-G tag, and thioredoxin (Trx) tag; autofluorescent proteins such as green fluorescent protein (GFP), yellow fluorescent protein (YFP), cyan fluorescent protein (CFP), blue fluorescent protein (BFP), HcRED, and DsRed; and reporter genes such as glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, and luciferase. Additionally, it may also include an NLS sequence that transports material from outside the nucleus into the nucleus through nuclear transport, or an NES sequence that transports material from inside the nucleus out of the nucleus through nuclear transport. The term "tag" includes all meanings that can be recognized by a person skilled in the art and may be appropriately interpreted depending on the context.
[0654] In some implementations, the vector may be designed as a linear or circular vector. If the vector is a linear vector, RNA transcription is terminated at the 3'-end even if the linear vector sequence does not contain a separate termination signal. However, if the vector is a circular vector, RNA transcription is not terminated if the circular vector sequence does not contain a separate termination signal. Therefore, when using a circular vector, a termination signal corresponding to the transcription factor associated with each promoter sequence must be included to express the intended target.
[0655] In some embodiments, the vector may be a viral vector. Specifically, the viral vector may be one or more selected from the group consisting of a retroviral vector (e.g., a foamy viral vector), a lentiviral vector (e.g., a SIN lentiviral vector), an adenoviral vector, an adeno-associated viral vector, a vaccinia viral vector, a poxviral vector, a herpes simplex viral vector, and a phage vector (e.g., a phagemid vector). Preferably, the viral vector may be an adeno-associated viral vector. In some embodiments, the phage may be selected from the group consisting of λgt4λB, λ-charon, λΔz1, and M13.
[0656] In some embodiments, the vector may be a non-viral vector. Specifically, the non-viral vector may be one or more selected from the group consisting of, but not limited to, plasmids, naked DNA, DNA complexes, mRNA (transcripts), and amplicons. For example, the plasmid may be selected from the group consisting of pcDNA series, pSC101, pGV1106, pACYC177, ColE1, pKT230, pME290, pBR322, pUC8 / 9, pUC6, pBD9, pHC79, pIJ61, pLAFR1, pHV14, pGEX series, pET series, and pUC19.
[0657] The term “naked DNA” refers to DNA (e.g., DNA lacking histones) encoding a protein, such as Cas12f of the invention or a variant thereof, cloned into a suitable expression vector (e.g., a plasmid) in an appropriate orientation for expression.
[0658] The term “amplicon,” when used with respect to a nucleic acid, refers to a nucleic acid replication product, wherein the product has a nucleotide sequence that is identical to or complementary to at least a portion of the nucleotide sequence of the nucleic acid. For example, an amplicon can be produced by any of a variety of amplification methods that use a nucleic acid or an amplicon thereof as a template, including polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of the nucleotide sequence (e.g., a concatameric product of RCA).
[0659] In some embodiments, hybrid vectors comprising sequences of one or more viral vectors and / or non-viral vectors can be created and used in appropriate combinations according to methods known in the art.
[0660] In some embodiments, the viral vector or non-viral vector may be delivered by a delivery system such as a liposome, a polymer nanoparticle (e.g., a lipid nanoparticle), an oil-in-water nanoemulsion, or a combination thereof, or may be delivered in viral form.
[0661] In another aspect of the present invention, a recombinant virus or recombinant viral particle produced by the vector system provided in the present invention is provided.
[0662] In some embodiments, the viral vector may be one or more viral vectors selected from the group consisting of, for example, a retroviral vector (e.g., a foamy viral vector), a lentiviral vector (e.g., a SIN lentiviral vector), an adenoviral vector, an adeno-associated viral vector, a vaccinia viral vector, a poxviral vector, a herpes simplex viral vector, and a phage (e.g., a phagemid) vector. Preferably, the viral vector may be an adeno-associated viral vector.
[0663] In some embodiments, the method may comprise introducing into a subject cell a non-viral vector comprising a nucleic acid sequence encoding an engineered Cas polypeptide of the invention and a nucleic acid sequence encoding a guide RNA. For example, the method may comprise introducing into a subject cell a first non-viral vector comprising a nucleic acid sequence encoding an engineered Cas polypeptide of the invention and a second non-viral vector comprising a nucleic acid sequence encoding a guide RNA as described herein.
[0664] In some embodiments, the method may comprise introducing into a subject cell a viral vector comprising a nucleic acid sequence encoding an engineered Cas polypeptide of the invention and a nucleic acid sequence encoding a guide RNA. For example, the method may comprise introducing into a subject cell a first viral vector comprising a nucleic acid sequence encoding an engineered Cas polypeptide of the invention and a second viral vector comprising a nucleic acid sequence encoding a guide RNA.
[0665] The method of delivery of a vector system comprising one or more vectors comprising the engineered Cas polypeptide and / or guide RNA of the present invention is not particularly limited, as long as it can be delivered into cells in an appropriate delivery form. For example, the delivery method may include, but is not limited to, electroporation, a gene gun, sonoporation, magnetofection, and / or transient cell compression or squeezing.
[0666] In some embodiments, the delivery forms and / or delivery methods of the components included in the vector system may be the same or different. For example, the engineered Cas polypeptide or the nucleic acid encoding it may be delivered in a first delivery form (or delivery method), and the guide RNA or the nucleic acid encoding it may be delivered in a second delivery form (or delivery method). The first delivery form and the second delivery form may each be any of the aforementioned delivery forms.
[0667] In some embodiments, the delivery form and / or delivery method of the components included in the vector system can be delivered simultaneously or sequentially with a time difference within the cell. For example, the engineered Cas polypeptide or a nucleic acid encoding it, and the guide RNA or a nucleic acid encoding it, can be delivered simultaneously or sequentially with a time difference. In another example, when two or more guide RNAs are included, the engineered Cas polypeptide or a nucleic acid encoding it, and the two or more guide RNAs or nucleic acids encoding it, can be delivered simultaneously or sequentially with a time difference. If delivered sequentially, the order of delivery can be appropriately selected by a person skilled in the art.
[0668] To efficiently deliver a gene editing system to a target cell or target site via a virus, particularly an adeno-associated virus (AAV), it is crucial to design the nucleotide sequence encoding all components of the editing system to be within the packaging limit of 4.7 kb for AAV. The engineered Cas polypeptide and (engineered) guide RNA of the present invention have the advantage of being sufficiently packaged within an AAV vector even if they contain additional components due to their extremely small size. The engineered Cas polypeptide and (engineered) guide RNA of the present invention can be provided as a single vector system, taking advantage of their extremely small size, or can be provided as a vector system comprising two or more different vectors, depending on the purpose.
[0669] VI. Gene Editing System
[0670] In another aspect of the present invention, a system is provided comprising (i) one or more engineered Cas polypeptides of the present invention or a nucleic acid encoding said polypeptide and (ii) one or more guide RNAs described herein or a nucleic acid encoding said guide RNA.
[0671] In some embodiments, the engineered Cas12f polypeptide and guide RNA can be included in a complex form, e.g., in the form of a ribonucleoprotein particle (RNP). The complex can include a guide RNA and two Cas12f or variant proteins thereof (see Satoru N. Takeda et al., Molecular Cell, 81, 1-13, (2021)). The complex can be formed by an interaction between the guide RNA and the Cas12f molecule.
[0672] In some embodiments, the engineered Cas polypeptide can be any of the engineered Cas polypeptides disclosed herein and described in detail above.
[0673] In some embodiments, the guide RNA can be any of the wild-type guide RNAs or engineered guide RNAs disclosed herein and described in detail above.
[0674] In some embodiments, the system may comprise one or more (e.g., two or more) engineered Cas polypeptides and one or more (e.g., two or more) guide RNAs. In some embodiments, the two or more guide RNAs may target different target sequences.
[0675] In some embodiments, the system may generate indels in a target gene or target nucleic acid as a result of performing a gene editing method. Indels may occur within and / or outside of a portion of the target sequence and / or a portion of the protospacer sequence. An indel refers to a mutation in the nucleotide sequence of a nucleic acid before gene editing, in which some nucleotides are deleted in the middle, any nucleotide is inserted, and / or insertions and deletions are mixed. Generally, when an indel occurs in a target gene or target nucleic acid sequence, the gene or nucleic acid is inactivated. In some embodiments, one or more nucleotides may be deleted and / or added in the target gene or target nucleic acid as a result of performing a gene editing method.
[0676] In some implementations, the system can perform a gene editing method, resulting in base editing within the target gene or target nucleic acid. Unlike indels, which involve deletion or addition of any base within the target gene or target nucleic acid, this refers to a targeted change in one or more specific bases within the nucleic acid. In other words, this involves inducing a pre-designed point mutation at a specific location within the target gene or target nucleic acid. In some implementations, the gene editing method can result in the substitution of one or more bases within the target gene or target nucleic acid with other bases.
[0677] In some embodiments, the system can generate knock-ins within a target gene or target nucleic acid as a result of performing a gene editing method. Knock-ins refer to the insertion of additional nucleic acid sequences within a target gene or target nucleic acid sequence. Knock-ins require a donor comprising additional nucleic acid sequences in addition to the engineered CRISPR / Cas12f1 complex. When the engineered CRISPR / Cas12f1 complex cleaves a target gene or target nucleic acid within a cell, the cleaved target gene or target nucleic acid is repaired. The donor participates in the repair process, allowing the additional nucleic acid sequences to be inserted into the target gene or target nucleic acid. In some embodiments, the gene editing method can additionally comprise delivering a donor into a target cell. The donor comprises an additional target nucleic acid, and the insertion of the additional target nucleic acid into the target gene or target nucleic acid is induced by the donor. When delivering the donor into the target cell, common delivery methods and / or modes, such as electroporation, gene gun, sonoporation, magnetofection, and / or transient cell compression or squeezing, may be used. In another embodiment, nanoparticles may be used when delivering the donor into the target cell. In another embodiment, an expression vector may be used when delivering the donor into the target cell.
[0678] In some embodiments, the system can remove all or part of a target gene or target nucleic acid sequence as a result of performing a gene editing method. Deletion refers to removing a portion of a base sequence within the target gene or target nucleic acid. In some embodiments, the removal of a portion of a base sequence (nucleic acid segment) can be induced by two single-strand breaks or two double-strand breaks in the target nucleic acid sequence.
[0679] VII. Composition
[0680] In another aspect of the present invention, a composition is provided comprising (i) one or more engineered Cas polypeptides of the present invention or a nucleic acid encoding said polypeptide and optionally (ii) one or more guide RNAs described herein or a nucleic acid encoding said guide RNA.
[0681] In another aspect of the present invention, a composition comprising one or more vectors or vector systems described herein is provided.
[0682] In some embodiments, the composition may be a composition for modulating or altering a target sequence within a cell.
[0683] In some embodiments, the composition may be a pharmaceutical composition further comprising a pharmaceutically acceptable carrier and / or excipient.
[0684] In some embodiments, the pharmaceutical composition may be formulated according to the intended route of administration. For example, if the pharmaceutical composition is intended for injection, it may be desirable to use an isotonic agent. Excipients for isotonicity typically include sodium chloride, dextrose, mannitol, sorbitol, and lactose. In some embodiments, isotonic solutions such as phosphate-buffered saline are preferred. In some embodiments, the pharmaceutical composition may further include a stabilizer, such as gelatin and albumin.
[0685] In some embodiments, the pharmaceutical composition may further comprise a pharmaceutically acceptable carrier (e.g., water, saline, ethanol, glycerol, lactose, sucrose, calcium phosphate, gelatin, dextran, agar, pectin, peanut oil, sesame oil, etc.), a diluent, a pharmaceutically acceptable excipient (including a stabilizer), and / or other compounds known in the art.
[0686] In some embodiments, a composition (e.g., a composition comprising one or more vectors included in a vector system described above) may include a gene transfer facilitating agent, such as a lipid, a liposome (including lecithin liposomes or other liposomes known in the art), a DNA-liposome mixture, calcium ions, viral proteins, polyanions, polycations, nanoparticles, or other known gene transfer facilitating agents. Preferably, the gene transfer facilitating agent is a polyanion, a polycation (e.g., poly-L-glutamic acid (LGS)), or a lipid.
[0687] The actual dosage of the composition may vary greatly depending on various factors, such as the type of vector, the target cell, organism, or tissue, the condition of the subject to be treated, the type of transformation / modification sought, etc., and can be selected as an appropriate dosage by a person skilled in the art.
[0688] VIII. Methods for Modifying or Modifying Nucleic Acids
[0689] In another aspect of the present invention, a method of modulating or altering one or more target nucleic acids in a cell is provided, comprising contacting the cell with an engineered Cas polypeptide of the present invention (optionally with a guide RNA described herein); or contacting the cell with a system or composition, vector system, or recombinant virus of the present invention. The engineered Cas polypeptide forms a complex with the guide RNA, and binding of the complex to the target nucleic acid results in alteration of the target nucleic acid. The engineered Cas polypeptide of the present invention exhibits improved gene editing efficiency compared to a wild-type Cas12f polypeptide.
[0690] In some embodiments, each component of the engineered CRISPR / Cas12f1 complex of the present invention can be delivered into a target cell to induce contact of the engineered CRISPR / Cas12f1 complex with a target gene or target nucleic acid.
[0691] In some embodiments, a method of modulating or altering a nucleic acid comprises delivering an engineered Cas polypeptide of the present invention, or a nucleic acid sequence encoding the same, and a guide RNA or a nucleic acid sequence encoding the same, into a target cell. As a result, a complex of the engineered Cas polypeptide and the guide RNA is injected into the target cell, or a complex of the engineered Cas polypeptide and the guide RNA is induced, and a target nucleic acid is cleaved, edited, and / or repaired by the engineered Cas polypeptide and the guide RNA complex. In one aspect, a cell comprising a system or composition, a vector system, or a recombinant virus of the present invention can be provided.
[0692] In some embodiments, nucleic acid modulation or modification may involve nucleic acid cleavage of double-stranded DNA, single-stranded DNA, or a hybrid duplex of DNA and RNA having a target sequence within the target site. Preferably, nucleic acid modulation or modification involves nucleic acid cleavage of double-stranded DNA.
[0693] In some embodiments, modulating or altering a nucleic acid within a cell involves editing a target nucleic acid within the cell. In some embodiments, the editing may involve indel introduction, base correction, gene insertion, or gene deletion.
[0694] In some embodiments, modulating or altering a nucleic acid within a cell comprises modulating the expression of a target nucleic acid within the cell.
[0695] In some embodiments, modulating or altering a nucleic acid within a cell comprises targeting a target nucleic acid within the cell.
[0696] In some embodiments, contacting can be performed in vitro, in vivo, or ex vivo.
[0697] In some embodiments, the cell may be a plant cell, a non-human animal cell, or a human cell. In some embodiments, the cell may be a eukaryotic cell or a prokaryotic cell.
[0698] In some embodiments, the cell may be a stem cell, a blood cell, an immune cell, or part of an organoid.
[0699] In some embodiments, modulating one or more target nucleic acids may comprise activating or inhibiting at least one of the one or more target nucleic acids. In some embodiments, modulating one or more target nucleic acids may comprise editing a primer of at least one of the one or more target nucleic acids. In some embodiments, modulating one or more target nucleic acids may comprise altering the methylation or acetylation of at least one of the one or more target nucleic acids. In some embodiments, modulating one or more target nucleic acids may comprise labeling at least one of the one or more target nucleic acids. In some embodiments, modulating one or more target nucleic acids may comprise nicking at least one of the one or more target nucleic acids.
[0700] IX. Treatment Methods
[0701] In another aspect of the present invention, a method for preventing or treating a disorder or disease, such as a genetic disorder or infectious disease, in a subject is provided. The method may comprise administering to the subject an effective amount of the gene editing system, composition, or vector system described herein. The term "effective amount" refers to an amount sufficient to modulate one or more target nucleic acids associated with the disorder or disease.
[0702] Disorders or conditions suitable for treatment by the methods of the present invention include, but are not limited to, X-linked severe combined immunodeficiency, sickle cell anemia, thalassemia, hemophilia, neoplasms, cancer, age-related macular degeneration, schizophrenia, trinucleotide repeat disorders, fragile X syndrome, prion-related disorders, amyotrophic lateral sclerosis, drug addiction, autism, Alzheimer's disease, Parkinson's disease, cystic fibrosis, blood and coagulation diseases or disorders, inflammation, facioscapulohumeral muscular dystrophy, retinitis pigmentosa, Leber's congenital amaurosis, glaucoma, immune-related diseases or disorders, metabolic diseases and disorders, liver diseases and disorders, kidney diseases and disorders, muscular / skeletal diseases and disorders, nervous system and neurological diseases and disorders, cardiovascular diseases and disorders, pulmonary diseases and disorders, and ocular diseases and disorders.
[0703] Additionally, a method for treating an infection in a subject is provided. In some embodiments, the infectious agent may be a virus. In some embodiments, the gene editing system, composition, or vector system described herein may be administered in an amount sufficient to modulate one or more target nucleic acids associated with the infection.
[0704] In some embodiments, administration can be via delivery systems using viruses, nanoparticles, liposomes, micelles, virosomes, nucleic acid complexes, protein-nucleic acid (RNA) conjugates, and combinations thereof.
[0705] In some embodiments, administration may be via an adeno-associated virus delivery system. For example, administration may be performed using a single adeno-associated virus delivery system or a dual adeno-associated virus delivery system.
[0706] In some embodiments, administration may be oral or parenteral, and may be local or systemic. Administration may be by, but is not limited to, intratumorally, orally, intradermally, transdermally, subcutaneously, intramuscularly, intravenously, intralymphatic, intramedullary, intraperitoneally, buccally, rectal, intraocularly, intravitreally, subretinal, intratracheally, intrapulmonary, intranodally, inhalation, etc. In addition, administration may be systemic, topical, intramural, or localized by the use of an implant that acts to maintain the active dose at the site of implantation.
[0707] In another embodiment, a method is provided for isolating cells (e.g., immune cells or stem cells) from a subject necessary for treating a disease in the subject, applying the gene editing system, composition or vector system described herein to the cells, and administering the resulting cells to the subject.
[0708] All publications mentioned in this specification are incorporated herein by reference to disclose and describe the methods and / or materials cited. Furthermore, any reference to prior art herein is not an admission or suggestion that such prior art is part of the common general knowledge in any jurisdiction, nor does it imply that a person skilled in the art would reasonably be expected to combine such prior art with other prior art.
[0709] The invention provided by this specification is described in more detail below through examples. These examples are intended solely to illustrate the subject matter disclosed by this specification and do not limit the scope of the present invention in any way.
[0710] Example
[0711] Example 1. Engineering of Cas12f polypeptide
[0712] To improve the editing performance of the naturally occurring Cas12f polypeptide (UnCas12f1 or CWCas12f1), Cas12f variants were generated by introducing one or more modifications into specific domains, including RuvC. The mutation sites were selected using analysis of the Cyro-EM structure of Cas12f (PDB ID: 7l49), homology sequence alignment, computational analysis, or a combination of these techniques. Amino acid substitutions were initially selected into three categories: single mutations, Cas12f2 homologous substitutions (hybrid), and computationally designed constructs (cd). Depending on the experimental progress and results, various combinations were tested, such as combinations of single mutation hits (substitutions of two or more amino acid residues), combinations of single mutations and homologous substitutions (hybrid), and combinations of homologous substitutions (hybrid), to select additional mutant candidates.
[0713] The selected variants were generated by mutagenesis using site-directed mutagenesis or Gibson assembly cloning ('insert and backbone PCR' or 'insert and backbone enzyme digestion' or '2 Fragment PCR and enzyme digestion') or Hifi assembly cloning ('insert and backbone PCR' or 'insert and backbone enzyme digestion' or '2 Fragment PCR and enzyme digestion') (see Experimental Example 1).
[0714] Example 2. Construction of an engineered Cas12f editing system.
[0715] Using the CwCas12f1 polypeptide sequence of SEQ ID NO: 1, a human codon-optimized polynucleotide of the Cas12f variant designed in Example 1 was synthesized, and a template polynucleotide for guide RNA was generated according to the target sequence, and a plasmid vector was prepared according to the method described in Experimental Example 3.
[0716] Exemplary Cas variants of the present invention used in the examples and their amino acid substitutions and sequences are presented in Tables 9 to 19.
[0717]
[0718]
[0719]
[0720]
[0721]
[0722]
[0723]
[0724]
[0725]
[0726]
[0727]
[0728]
[0729]
[0730]
[0731]
[0732]
[0733]
[0734]
[0735]
[0736]
[0737]
[0738]
[0739]
[0740]
[0741]
[0742]
[0743]
[0744]
[0745]
[0746]
[0747]
[0748]
[0749]
[0750]
[0751]
[0752]
[0753]
[0754]
[0755]
[0756]
[0757]
[0758]
[0759]
[0760]
[0761]
[0762]
[0763]
[0764]
[0765]
[0766]
[0767]
[0768]
[0769]
[0770]
[0771]
[0772]
[0773]
[0774]
[0775]
[0776]
[0777]
[0778]
[0779]
[0780]
[0781]
[0782]
[0783]
[0784]
[0785]
[0786]
[0787]
[0788]
[0789]
[0790]
[0791]
[0792]
[0793]
[0794]
[0795]
[0796]
[0797]
[0798]
[0799]
[0800]
[0801]
[0802]
[0803]
[0804]
[0805]
[0806]
[0807]
[0808]
[0809]
[0810]
[0811]
[0812]
[0813]
[0814]
[0815]
[0816]
[0817]
[0818]
[0819]
[0820]
[0821]
[0822]
[0823]
[0824]
[0825]
[0826]
[0827]
[0828]
[0829]
[0830]
[0831]
[0832]
[0833]
[0834]
[0835]
[0836]
[0837]
[0838]
[0839]
[0840]
[0841]
[0842]
[0843]
[0844]
[0845]
[0846]
[0847]
[0848]
[0849]
[0850]
[0851]
[0852]
[0853]
[0854]
[0855]
[0856]
[0857]
[0858]
[0859]
[0860]
[0861]
[0862]
[0863] Canonical sgRNA of SEQ ID NO: 13, ge3.0 of SEQ ID NO: 151, ge4.0 of SEQ ID NO: 152, and ge4.1 of SEQ ID NO: 153 were used as guide RNAs (Table 20). The guide RNAs (ge3.0, ge4.0, and ge4.1) were synthesized using a PCR method using oligonucleotides encoding engineered guide RNAs (see Experimental Examples 2 and 3).
[0864]
[0865] In this specification, the sequence indicated as 'NNNNNNNNNNNNNNNNNNNN' in relation to gRNA refers to any guide sequence (spacer sequence) capable of hybridizing with a target sequence within a target gene. The guide sequence can be appropriately designed by those skilled in the art according to the desired target gene and / or target sequence within the target gene, and therefore is not limited to a specific sequence of a specific length.
[0866] Meanwhile, the target sequence for confirming indel efficiency is shown in Table 21.
[0867]
[0868] Example 3. Relative indel analysis of Cas12f variants
[0869] Example 3.1. Measurement of indel efficiency
[0870] The engineered Cas12f editing system of Example 2 was transfected into HEK-293T cells (LentX-293T, Takara) according to the method described in Experimental Example 4. The same plasmid was transfected in three wells each and the experiment was repeated. NGS analysis was performed using cell lysates extracted from cells 72 hours after transfection as a template according to the method described in Experimental Example 5. The relative indel value (relative indel fold change) was derived from the NGS analysis results, and the editing system using the CWCas12f1 polypeptide was used as a control. The relative indel value was derived by comparing the indel (%) of the control CWCas12f1 that was performed in each experiment, and collecting and analyzing the values. Except where otherwise noted, the guide RNA of ge4.0 was used for the relative indel analysis. The indel (%) was calculated through NGS analysis as described in Experimental Example 5.
[0871] Example 3.2. Single amino acid variants
[0872] The indel efficiencies of exemplary variants comprising amino acid substitutions that improve the relative indel efficiency of the Cas12f protein were measured. Exemplary variants and their relative indel efficiencies are presented in Tables 22 and 23. The Cas12f variants were found to have improved indel efficiency compared to the wild type, with indel efficiency improvements ranging from about 1.0-fold to about 2.3-fold.
[0873]
[0874]
[0875]
[0876]
[0877]
[0878] Example 3.3. Variants containing multiple amino acid substitutions in the RuvC domain
[0879] We measured the indel efficiency of exemplary variants containing two or more amino acid substitutions in the RuvC domain. The exemplary Cas12f variants presented in Tables 12 to 15 were confirmed to exhibit significantly enhanced indel efficiency. Referring to Tables 24 to 27, HS1, HS5, and HS8 each significantly increased indel efficiency, and when combined with each other or other amino acid substitutions, the indel efficiency increased by approximately 1.7-fold to approximately 8.6-fold compared to the wild-type. In particular, we confirmed that the combination of HS1, HS5, HS8, and these showed an excellent synergistic effect in indel efficiency. This suggests that the amino acid substitutions in HS1, HS5, and HS8 play a significant role in improving the indel efficiency of the Cas12f protein.
[0880]
[0881]
[0882]
[0883]
[0884]
[0885]
[0886] Example 3.4. Relative indel analysis - Variant I containing multiple amino acid substitutions
[0887] The indel efficiency of exemplary Cas12f variants comprising HS1 / HS5 / HS8 and one or more other amino acid substitutions was measured. These variants comprise HS1 / HS5 / HS8 and K135R / Y141F / K305R / M331F (dM12a36), and optionally further comprise one or more substitutions selected from the group consisting of E555D (RuvC domain), E556D (RuvC domain), N32Q (WED domain), N47Q (WED domain), E50D (ZF domain), N61Q (ZF domain), T89S (ZF domain), E166D (REC domain), D171E (REC domain), R176K (REC domain), D216E (REC domain), N511Q (TNB domain), E516D (TNB domain), and E535D (TNB domain). The relative indel values of these variants are presented in Tables 29 to 31.
[0888]
[0889]
[0890]
[0891]
[0892]
[0893]
[0894] Referring to Tables 29 to 31, even when targeting different target sequences, variants containing multiple substitutions of HS1 / HS5 / HS8 and K135R / Y141F / K305R / M331F exhibited excellent indel effects, with relative indels of more than 2-3 times, in combination with other amino acid substitutions. These results suggest that at least one amino acid position included in HS1 / HS5 / HS8 and K135R / Y141F / K305R / M331F can be utilized to construct Cas12f variants exhibiting excellent indel effects for various target sequences.
[0895] Example 3.5. Relative indel analysis - Variant II containing multiple amino acid substitutions
[0896] The present inventors have identified additional amino acid substitution combination units, L67I / Q124N / K155R / E556D and L67I / Q124N / K155R / K530R / E556D, that can be utilized to generate Cas12f variants exhibiting excellent indel effects. Variants containing multiple amino acid substitutions of HS1 / HS5 / HS8, K135R / Y141F / K305R / M331F, and L67I / Q124N / K155R / E556D or L67I / Q124N / K155R / K530R / E556D were generated and the relative indel effects were measured. The relative indel effects are presented in Table 32.
[0897]
[0898]
[0899]
[0900]
[0901]
[0902]
[0903] Example 3.6. Cas12f variants with improved indel efficiency and expanded range of recognizable PAM sequences.
[0904] We confirmed that the mutants obtained by introducing multiple amino acid substitutions into the wild-type Cas12f protein showed improved indel efficiency compared to the wild-type, and could recognize not only the canonical PAM sequence recognized by the wild-type Cas12f protein but also non-canonical PAM sequences. The indel efficiency for 31 target sequences containing the canonical PAM sequence (5'-TTTG-3' or 5'-TTTA-3') was compared with that of the control CWCas12f1. Referring to Figure 3a, M1120 showed improved indel efficiency compared to the control for most targets (Table 33).
[0905]
[0906] Additionally, indel efficiency was measured for target sequences F3, F5, and F10 containing the circular PAM sequence (5'-TTTG-3' or 5'-TTTA-3').
[0907]
[0908] Referring to Fig. 3b, the M1120 variant exhibited higher indel efficiency than the control in all three target sequences. Meanwhile, compared to the M1120 variant (a variant with 28 amino acid substitutions in the wild-type CWCas12f), the variants (M1223, M1224, M1225) containing two additional amino acid substitutions exhibited increased indel efficiency compared to the M1120 variant for the R10 target sequence containing the circular PAM sequence (5'-TTTA-3') (Fig. 3c).
[0909] Recognition of non-circular PAM sequences (1)
[0910] Referring to Figure 4, the M1120 variant showed an indel efficiency of up to about 0.8% even for targets containing non-canonical PAM sequences. Considering that the indel efficiency of the HS1 / 5 / 8 variants converges to about 0% for target sequences containing non-canonical PAM sequences, it can be seen that the M1120 variant is a variant with an expanded range of recognized PAM sequences. The non-canonical PAMs are 5'-CCTG-3' (Dis3), 5'-GTTG-3' (Dis5), and 5'-TAGG-3' (Dis8), and NT represents the untreated group.
[0911] Additionally, it was confirmed that three mutants (M1223, M1224, and M1225) showed higher editing efficiency for target sequences adjacent to non-circular PAM sequences (PAM sequence for R11 is 5'-TTTG-3', and PAM sequence for L17 is 5'-TTCG-3') than M1120. Referring to Figures 5a and 5b, M1223, M1224, and M1225 all showed higher indel efficiency than M1120, and in particular, M1224 showed the highest indel efficiency.
[0912] Recognition of non-circular PAM sequences (2)
[0913] Additionally, three mutants (M1231, M1232, and M1233) were generated by introducing a single mutation (Q272R) into M1223, M1224, and M1225 as backbones, respectively, and 13 mutants (PM25, PM30, PM44, PM47, PM58, PM141, PM143, PM144, PM145, PM146, PM147, PM149, and PM150) were generated by introducing mutations at one or more positions selected from N205, M206, P221, Q225, Q229, Y230, and Q270 into HS1 / HS5 / HS8 as backbones. For these mutants, indel efficiency was measured for a target sequence containing a different non-circular PAM sequence (L20; its sequence is TGCCGTGCCCCGGGCACTCA, and its PAM sequence is TCTG). ge4.1 was used as the gRNA in this experiment.
[0914] The results are shown in Table 35. The relative indels were calculated using CWCas12f1 as a control. All 16 generated mutants showed improved indel efficiency compared to the wild type, ranging from approximately 1.0-fold to approximately 1.43-fold.
[0915]
[0916]
[0917] Recognition of non-circular PAM sequences (3)
[0918] Additionally, 87 variants (PM211, PM212, PM213, PM214, PM216, PM217, PM218, PM219, PM220, PM233, PM238, PM244, PM246, PM253, PM255, PM256, PM257, PM260, PM261, PM262, PM263, PM264, PM268, PM273, PM277, PM332, PM333, PM334, PM335, PM337, PM339, PM340, PM341, PM350, PM354, PM368, PM383, PM389, PM411, PM427, PM429, PM432, PM434, PM435, PM541, PM542, PM543, PM546, PM547, PM548, PM549, PM550, PM553, PM555, PM558, PM559, PM561, PM563, PM572, PM575, PM576, PM578, PM579, PM580, PM581, PM583, PM584, PM585, PM586, PM587, PM596, PM600, PM608, PM609, PM610, PM611, PM614, PM615, PM617, PM618, PM620, PM622, PM623, PM624, PM625, PM628 and PM629) were additionally created.Among the 87 mutants generated in this way, 6 mutants (PM729, PM730, PM731, PM732, PM733, and PM735) were additionally generated by introducing mutations at one or more positions selected from D171, T175, I465, G209, K358, and S454 using PM218 as the backbone, and 28 mutants (PM630, PM631, PM632, PM638, PM646, PM658, PM663, PM670, PM671, PM674, We additionally generated variants (PM676, PM680, PM683, PM684, PM685, PM687, PM688, PM689, PM692, PM693, PM714, PM715, PM716, PM722, PM724, PM727, PM728, and PM734). For these variants, we measured the indel efficiency for target sequence 1 (Dis8; its sequence is GATAGGTATGAGATATTCAC, and its PAM sequence is TAGG) and target sequence 2 (Dis16; its sequence is GGATAGGTATGAGATATTCA, and its PAM sequence is ATAG) containing different non-circular PAM sequences. ge4.1 was used as the gRNA in this experiment.
[0919] The results are shown in Tables 36 and 37, respectively. Here, the relative indels were calculated using the backbone variants M1225, PM218, or PM625 from which each variant was derived as controls. The generated variants were confirmed to have improved indel efficiency from 0.9 to 7.04 times compared to the control for Dis8 and from 0.9 to 2.67 times compared to the control for Dis16. These results demonstrate that the variants recognize non-circular PAM sequences and exhibit excellent indel efficiency. In addition, PM714, PM728, and PM735 containing amino acid substitutions T175R / G209R / K358R / S454G exhibited excellent indel efficiency.
[0920]
[0921]
[0922]
[0923]
[0924]
[0925]
[0926]
[0927]
[0928]
[0929]
[0930]
[0931]
[0932]
[0933]
[0934]
[0935] Thus, it can be seen that the Cas12f variants of the present invention can exhibit improved editing performance for a wider range of target sequences by recognizing an expanded range of PAM sequences compared to the wild type.
[0936] Measurement of fabrication and editing efficiency of the Cas12f-ABE system
[0937] The TaRGET-ABE module (a system that indicates the conversion efficiency from A to G) was constructed using the M1120 variant. Specifically, the Cas protein of the dCWCas12f1-ABE-C3.1 module (dCWCas12f1 of the present invention includes the D538A mutation) used in the paper (Kim et al. Nat. Chem. Biol. 19,389 (2023)) was specified using GA cloning (Gibson assembly cloning; 'insert and backbone PCR' or 'insert and backbone enzyme digestion' method) and HA cloning (Hifi assembly, NEB; 'insert and backbone PCR' or 'insert and backbone enzyme digestion' method), and then the nucleic acid sequence (SEQ ID NO: 408) encoding the fusion protein of the dCWCas12f1-ABE module was specified, and then the fusion protein was manufactured using dead M1120, which has no cleavage activity, according to the method described in Experimental Example 6. dCWCas12f1-ABE-C3.1 is oriented in the order of NLS, dCWCas12f1 (D538A), TadA, and 8eWQ and connected by a linker. The amino acid sequence of 8eWQ (deaminase) used in the dCWCas12f1-ABE module is as follows: SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGWRQSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 409)
[0938] The editing efficiency was measured using three types of guide RNAs (ge3.0, ge4.0, and ge4.1) for the target DY10 or Inter22 containing a circular PAM sequence (5'-TTTG-3' or 5'-TTTA-3'). Referring to Figures 6A and 6B, it was confirmed that the editing window (the position where the base is switched) was expanded for the M1120-ABE module compared to the CWCas12f1-ABE module.
[0939] Experimental methods and materials
[0940] Experimental Example 1. Mutagenesis
[0941] Mutagenesis for engineering Cas12f polypeptides was performed by one or more of the following methods: site-directed mutagenesis, three types of GA cloning (Gibson assembly cloning), and HA cloning (Hifi assembly cloning).
[0942] Experimental Example 1.1. Site-directed mutagenesis
[0943] Site-specific mutagenesis was performed by preparing a mixture as shown on the left side of Table 38 below and performing PCR under the conditions described on the right side of Table 38 to amplify a plasmid containing the mutation. The PCR was performed using KOD One TM It was performed using PCR Master Mix - Blue (KMM-201, TOYOBO).
[0944]
[0945] After PCR was completed, the PCR product was confirmed by electrophoresis and purified using the GEL & PCR Purification System (GP104-200, Biofact). The final elution was performed in 17 μL.
[0946] After purification of the PCR product, the obtained DNA amplification product was used to prepare a mixture as shown on the left side of Table 39 below, and the Dpn1 (R0176L, NEB) reaction was performed at 37°C for 1 hour. The reaction product was prepared as a mixture as shown on the right side of Table 39 below, and the KLD (M0554S, NEB) reaction was performed at room temperature for 5 minutes. In the KLD reaction mixture, the DNA was directly used as the Dpn1 reaction product.
[0947]
[0948] The entire KLD reaction product was transformed into 50 μl of RBC HIT competent cells (RH617, RBC Bioscience). Specifically, 10 μl of the KLD reaction product was added to the competent cells, vortexed for about 1 second, and incubated on ice for 20 minutes. After incubation, the cells were heat shocked at 42°C for 1 minute and then incubated on ice for 20 minutes. The cells were spread on ampicillin-supplemented medium plates and incubated overnight at 37°C. DNA was extracted from each colony formed on the plates, and sequencing was performed to confirm whether the mutation was correct.
[0949] Experimental Example 1.2. HA cloning (Hifi assembly cloning; insert and backbone PCR) or GA cloning (Gibson assembly cloning; insert and backbone PCR)
[0950] Insert and backbone PCR were performed using mixtures suitable for each construct. The composition of the mixtures and PCR conditions are shown in Table 40 below.
[0951]
[0952] For inserts, 2.5 μl of each of the forward and reverse primers at a concentration of 10 pmole / μl were mixed when preparing the mixture, and for the backbone, 3 μl of each of the forward and reverse primers at a concentration of 1 pmole / μl were mixed. In addition, for the insert, denaturation and annealing were set to 10 s and 5 s, respectively, and 35 cycles were performed, and for the backbone, denaturation and annealing were set to 20 s and 15 s, respectively, and 16 cycles were performed.
[0953] PCR products were confirmed by electrophoresis. Specifically, for the insert, gel extraction was performed using the GEL & PCR Purification System (GP104-200, Biofact). The elution volume was 20 to 40 μl depending on the band intensity on electrophoresis. For the backbone, 1 μl of Dpn1 (R0176L, NEB) and 10 μl of 10X rCutsmart buffer were added, reacted at 37°C for 1 hour, and then purified using the GEL & PCR Purification System (GP104-200, Biofact). The elution volume was 20 to 40 μl depending on the band intensity on electrophoresis. The insert and backbone extracts were quantified using a nanodrop and the molar ratio was calculated according to the base pair to prepare a mixture with insert to backbone ratio of 10 to 1 (total volume 5 μl).
[0954] Next, the insert was diluted to 150 fmole / μl and the backbone to 30 fmole / μl, and a mixture was prepared as shown in Table 41 below, and the NEBuilder® HiFi DNA Assembly Master Mix (E2621X, NEB) reaction was performed at 50°C for 1 hour.
[0955]
[0956] After the reaction, the entire mixture was transformed into 50 μl of RBC HIT competent cells (RH617, RBC Bioscience). Specifically, 10 μl of the mixture was added to the competent cells, vortexed for about 1 second, and incubated on ice for 20 minutes. After incubation, the cells were heat-shocked at 42°C for 1 minute and then incubated on ice for 20 minutes. The cells were plated on a medium plate supplemented with ampicillin and incubated overnight at 37°C. DNA was extracted from each colony formed on the plate, and base sequence analysis was performed to confirm whether the mutation was correct.
[0957] Experimental Example 1.3. HA cloning (Hifi assembly cloning; 2 Fragment PCR and enzyme digestion) or GA cloning (Gibson assembly cloning; 2 Fragment PCR and enzyme digestion)
[0958] The plasmid used as the backbone [px458-TnpB_WT-HA (U6 scaffold removed) or HS1, HS8, etc.] was digested with restriction enzymes. The following combinations of restriction enzymes were used depending on the needs of the modified constructs: XhoI single digestion, AgeI and AflIII double digestion, TfiI single digestion, StuI and BamHI double digestion, BsaI and BamHI double digestion, BsaI and EcoNI double digestion, XhoI and FseI double digestion.
[0959] The mixture was prepared as shown in Table 42 below.
[0960]
[0961] The product was digested with a restriction enzyme at 37°C for 2 hours, the bands were identified by electrophoresis, and then purified using a GEL & PCR Purification System (GP104-200, Biofact). The elution volume was 30 μL.
[0962] Next, a mixture was prepared with the composition shown on the left side of Table 43 below, and PCR was performed under the conditions shown on the right side of Table 43 to synthesize fragments 1 and 2, respectively. Here, fragment 1 used a combination of a restriction enzyme forward primer (restriction enzyme primer_F) and a mutagenic reverse primer (mutagenic primer_R), and fragment 2 used a combination of a mutagenic forward primer (mutagenic primer_F) and a restriction enzyme reverse primer (restriction enzyme primer_R).
[0963]
[0964] After PCR was completed, the PCR product was confirmed by electrophoresis and purified using the GEL & PCR Purification System (GP104-200, Biofact). The elution volume was 30 μL.
[0965] The plasmid and fragment 1 and fragment 2 digested with restriction enzymes were quantified using Nanodrop, and the molar ratio of each quantified DNA was calculated to prepare a mixture so that the ratio of plasmid: fragment 1: fragment 2 was 1:2:2, and the final volume was adjusted to 2.5 to 3.5 μl. To this, an equal volume of NEBuilder® HiFi DNA Assembly Master Mix (E2621X, NEB) was added and reacted at 50°C for 15 to 20 minutes. The entire mixture was transformed into 50 μl of RBC HIT competent cells (RH617, RBC Bioscience). Specifically, 5 to 7 μl of the mixture was added to the competent cells, vortexed for about 1 second, and incubated on ice for 20 minutes. After incubation, the cells were heat-shocked at 42°C for 1 minute and then incubated on ice for 20 minutes. The cells were plated on ampicillin-supplemented media plates and cultured overnight at 37°C. DNA was extracted from each colony formed on the plates and sequenced to confirm whether the mutation was correct.
[0966] Experimental Example 2. Preparation of Guide RNA
[0967] Experimental Example 2.1 Preparation of Guide RNA Forms
[0968] The guide RNA (gRNA) or engineered gRNA of the present invention was prepared by the following method. First, a canonical gRNA or engineered gRNA was chemically synthesized from a pre-designed gRNA to prepare the same, and then a PCR amplicon containing the synthesized gRNA sequence and a T7 promoter sequence was prepared. U-rich tail ligation to the 3'-end of the engineered gRNA was performed using Pfu PCR Master Mix (Biofact) in the presence of sequence-modified primers and a gRNA plasmid vector. The PCR amplicon was prepared by HiGene. TMPurification was performed using the Gel & PCR Purification System (Biofact).
[0969] Using the PCR amplicon as a template, in vitro transcription was performed using NEB T7 polymerase. The in vitro transcription product was treated with DNase I (NEB) and purified using the Monarch RNA Cleanup Kit (NEB), yielding gRNA. Subsequently, a plasmid vector containing the predesigned gRNA sequence and T7 promoter sequence was constructed using the T-blunt plasmid cloning method (Biofact).
[0970] After double-cutting and purifying both ends of the guide RNA sequence containing the T7 promoter sequence in the above vector, in vitro transcription was performed using T7 polymerase (NEB). After treating the in vitro transcription product with DNase I (NEB), it was purified using the Monarch RNA Cleanup Kit (NEB), and gRNA was obtained.
[0971] Experimental Example 2.2. Preparation of Guide RNA Amplicon
[0972] Canonical guide RNAs, Cas12f_ge3.0, Cas12f_ge4.0, and Cas12f_ge4.1 guide RNAs were cloned by synthesizing template DNA (Twist Bioscience) and cloning it into a plasmid vector containing a U6 promoter sequence. If necessary, the vector was used to produce PCR amplicons of the guide RNAs (sgRNA transcription cassette) using a U6-complementary forward primer and a protospacer-complementary reverse primer. The PCR amplicons were purchased from HiGene. TM Purification was performed using the Gel & PCR Purification System (Biofact).
[0973] Experimental Example 3. Design and Preparation of Plasmid Vectors
[0974] To produce a Cas12f variant with improved editing performance of the present invention, a plasmid vector was produced using the Gibson assembly method or the Hifi assembly method, in which the following components were operably linked: 1) a sequence encoding eGFP linked to a chicken β-actin hyprid promoter (CBh) promoter and a self-cleaving T2A peptide (2A), 2) a polynucleotide of a human codon-optimized nucleic acid construct encoding a Cas12f polypeptide or an engineered Cas12f polypeptide, and 3) a canonical gRNA or an engineered gRNA.
[0975] Specifically, oligonucleotides comprising the base sequence of the gene encoding wild-type Cas12f and engineered Cas12f polypeptides, and a nuclear localization signal (NLS) sequence and a linker sequence at the 5'-end and 3'-end, respectively, were synthesized (Bionics), thereby synthesizing polynucleotides of human codon-optimized Cas12f or engineered Cas12f nucleic acid constructs for cleavage of a target gene or target nucleic acid. The polynucleotides of codon-optimized Cas12f or engineered Cas12f nucleic acid constructs were operably linked to a plasmid comprising a sequence encoding eGFP linked to the chicken β-actin hybrid promoter (CBh) promoter and a self-cleaving T2A peptide (2A) and cloned.
[0976] At this time, canonical guide RNAs, Cas12f_ge3.0, Cas12f_ge4.0, and Cas12f_ge4.1 guide RNAs were cloned by synthesizing template DNA (Twist Bioscience) and operably linking it to a plasmid containing a U6 promoter sequence. If necessary, a polynucleotide encoding a guide RNA containing a U6 promoter sequence was additionally cloned into the plasmid containing the polynucleotide of the codon-optimized Cas12f or engineered Cas12f nucleic acid construct.
[0977] Experimental Example 4. Cell Culture and Transfection
[0978] HEK293T (LentX-293T, Takara) cells were cultured at 37°C in 5% CO2 in DMEM medium supplemented with 10% heat-inactivated FBS, 1% penicillin / streptomycin, and 0.1 mM non-essential amino acids. One day before cell transfection, cells were diluted 1 / 100 and subcultured when they reached approximately 80–90% confluency in a 24-well plate. The transfection mixture was prepared as shown in Table 44 below.
[0979]
[0980] A mixture of DNA (vector plasmid + sgRNA transcription cassette), DMEM medium, and FuGENE reagent was vortexed and incubated for 15 minutes. 200 μl of the mixture was treated to cells prepared in a 24-well plate and incubated at 37°C for 72 hours.
[0981] Experimental Example 5. Next-Generation Sequencing (NGS)
[0982] Experimental Example 5.1. NGS Sample Lysis
[0983] Next-generation sequencing (NGS) for indel efficiency analysis was performed as follows. First, the supernatant was removed from the 24-well plate where transfection was completed (72 hours after transfection) according to Experimental Example 4. 200 μl of 1X lysis buffer (0.2 mg / ml Protease K) was added. Next, the 1X lysis buffer and cell mixture were transferred to a 96-well PCR plate and incubated in a PCR machine under the following conditions: 55°C for 2 hours, 95°C for 20 minutes, and stored at 4°C. The composition of 1X lysis buffer was prepared as shown in Table 45 below.
[0984]
[0985] Experimental Example 5.2. NGS PCR (1st, 2nd, and 3rd)
[0986] NGS 1st PCR was performed using cell lysate as the NGS 1st PCR template, and NGS 2nd PCR was performed using the obtained NGS 1st PCR product as the template. In addition, NGS 3rd PCR was performed using the NGS 2nd PCR product as the template.
[0987] NGS 1st, 2nd, and 3rd PCR were prepared by preparing a mixture with the composition shown on the left side of Table 45 below, and PCR was performed under the conditions shown on the right side of Table 46.
[0988]
[0989] The above PCR uses KOD One TM PCR Master Mix -Blue- (KMM-201, Toyobo) was used.
[0990] After completing the third PCR, PCR products were confirmed by electrophoresis. 5 μl of each sample was mixed and purified using a GEL & PCR Purification System (GP104-200, Biofact). Nuclease-free distilled water was used for elution.
[0991] The above extract was quantified using a nanodrop and diluted to a concentration of 15 ng / μl. The final volume was 1 μl times the number of samples. NGS analysis was performed using the iseq100 sequencing system (illumina).
[0992] Experimental Example 6. Preparation of an adenine base editing system (TaRGET-ABE)
[0993] Polynucleotides of base-editing constructs consisting of codon-optimized dCas12f or engineered dCas12f and deaminase were operably linked to a plasmid comprising a chicken β-actin hyprid promoter (CBh) promoter or a chicken β-actin (CBA) promoter, a nuclear localization signal (NLS) sequence at the 5'-terminus and the 3'-terminus, and a sequence encoding eGFP linked to a self-cleaving T2A peptide (2A), or a plasmid comprising a CMV enhancer, a CMV promoter, and a nuclear localization signal (NLS) sequence at the 5'-terminus or the 3'-terminus.
[0994] Experimental Example 7. Statistical Analysis
[0995] Statistical significance was verified using the unpaired t-test using Prism-GraphPad software. A p-value less than 0.05 was considered statistically significant. Data points in box plots and dot plots represent the entire range of values. Error bars for all dot plots and bar plots were drawn using Prism and represent the standard deviation of each data point.
Claims
1. An engineered Cas polypeptide comprising a modified amino acid sequence compared to a wild-type Cas12f1 polypeptide, The modified amino acid sequence comprises one or more amino acid substitutions compared to the amino acid sequence of the wild-type Cas12f1 polypeptide, The above one or more amino acid substitutions are based on SEQ ID NO:
1. (i) I365, A368, H399, A402, A406, C433, I435, A436 and D437; (ii) I65, E125, I126, L133, K135, I140, Y141, K155, S164, H167, D171, T175, E179, L180, D216, L222, Q229, T231, R250, F266, E269, K305, Q312, T313, S322, G325, M331, L332, D337, I341, P347, K358, S359, L361, F369, K407, T414, E418, S420, E421, C433, I435, I440, N442, Q448, N451, R457, E459, K479, H508, K530 and E555; (iii) K31, N32, T33, K39, N47, P221, L222, K224, Q225, Q229, Y230, T231, E234, N237, K245, I246, R250, Q252, K254, K255, E256, I257, K259, K265, F266, D267, E269, Q270, Q272, I278, L282, K288, K295, D296, E297, K304, K305, N308, D310, Q312, T313, K319, G321, S322, K323, G325, M331, L332, N333 and D337; (iv) S48, E50, K53, D57, E58, N61, I65, L67, N70, D72, T89 and A120; (v) Q124, E125, I126, L133, K135, I140, Y141, I154, K155, A162, S163, S164, E166, H167, D171, T175, R176, A178, E179, L180, I186, L190, L200, N205, M206, G209, S215, D216 and F218; (vi) I341, D342, K343, D346 and P347; (vii) T502, K505, H508, L509, N511, E516, K519, K527, K530 and E535; and (viii) comprising an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of K358, S359, L361, I365, A368, F369, K385, H399, A402, A406, K407, T414, E418, S420, E421, C433, I435, A436, D437, I440, N442, T446, Q448, N451, S454, R457, E459, I465, K479, K501, L543, L550, K551, S552, E555 and E556; Engineered Cas polypeptides.
2. In paragraph 1, One or more of the amino acid substitutions mentioned above (i) I365V, A368S, H399L, A402S, A406S, C433K, I435V, I435L, A436T and D437N; (ii) I65S, I65E, E125S, E125N, E125M, I126M, L133V, L133S, L133M, L133C, K135R, I140V, I140M, Y141F, K155R, K155Y, S164D, H167K, H167R, D171E, D171N, D171R, T175R, E179A, L180N, D216E, L222F, Q229G, Q229K, T231E, T231K, R250K, F266C, E269D, K305R, Q312P, Q312K, T313I, T313V, S322K, G325C, M331F, L332V, D337R, I341K, P347K, K358R, S359N, S359H, L361A, F369N, F369D, F369G, F369W, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435L, I440L, N442H, N442M, Q448V, Q448N, N451S, R457K, E459K, K479M, H508C, H508Y, K530R, E555R (iii) K31R, N32Q, T33S, K39R, N47Q, P221T, P221S, P221C, P221G, P221A, P221V, P221I, P221M, P221F, L222F, L222I, K224R, Q225R, Q225H, Q225D, Q225E, Q225N, Q225Y, Q225K, Q225F, Q225C, Q225V, Q225P, Q225S, Q225T, Q225A, Q229G, Q229K, Q229N, Q229R, Q229H, Q229C, Q229P, Y230E, Y230R, Y230H, Y230K, Y230D, Y230N, Y230Q, Y230F, Y230C, Y230W, T231E, T231K, T231P, E234D, N237Q, K245R, I246L, R250K, Q252N, K254R, K255R, E256D, I257L, K259R, K265R, F266C, D267E, E269D, Q270R, Q272N, Q272R, I278L, L282I, K288R, K295R, D296E, E297D, K304R, K305R, N308Q, D310E, Q312K, Q312P, Q312N T313I, T313V, K319R, G321K, S322K, K323R, G325C, M331F, L332V, N333Q and D337R; (iv) S48T, E50D, K53R, D57E, E58D, N61Q, I65E, I65L, I65S, L67I, N70Q, D72E, T89S 및 A120T; (v) Q124N, E125M, E125N, E125S, I126M, L133C, L133M, L133S, L133V, K135R, I140M, I140V, Y141F, I154L, K155R, K155Y, A162D, A162Q, A162P, A162V, S163R, S163K, S163N, S163M, S163F, S163Y, S163G, S163P, S163A, S163L, S164D, E166D, H167K, H167R, D171N, D171R, D171E, T175R, R176K, A178S, E179A, L180N, I186L, L190I, L200F, N205H, M206H, G209R, S215H, D216E and F218H; (vi) I341K, D342E, K343R, D346N and P347K; (vii) T502S, K505R, H508C, H508Y, L509I, N511Q, E516D, K519R, K527R, K530R and E535D; and (viii) K358R, S359H, S359N, L361A, I365V, A368S, F369D, F369G, F369N, F369W, K385R, H399L, A402S, A406S, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435V, I435L, A436T, D437N, I440L, N442H, N442M, T446S, Q448N, Q448V, N451S, S454G, R457K, E459K, I465R, K479M, K501R, Containing at least one selected from the group consisting of L543I, L550I, K551R, S552T, E555D, E555R and E556D Engineered Cas polypeptides.
3. In paragraph 1, wherein said one or more amino acid substitutions are (ii) I65, E125, I126, L133, K135, I140, Y141, K155, S164, H167, D171, T175, E179, L180, D216, L222, Q229, T231, R250, F266, E269, K305, Q312, T313, S322, G325, M331, L332, D337, I341, P347, K358, S359, L361, F369, K407, T414, E418, S420, E421, C433, I435, I440, N442, Q448, N451, R457, Comprising an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of E459, K479, H508, K530 and E555. Engineered Cas polypeptides.
4. In paragraph 3, wherein said one or more amino acid substitutions are (ii) I65S, I65E, E125S, E125N, E125M, I126M, L133V, L133S, L133M, L133C, K135R, I140V, I140M, Y141F, K155R, K155Y, S164D, H167K, H167R, D171E, D171N, D171R, T175R, E179A, L180N, D216E, L222F, Q229G, Q229K, T231E, T231K, R250K, F266C, E269D, K305R, Q312P, Q312K, T313I, T313V, Containing at least one selected from the group consisting of S322K, G325C, M331F, L332V, D337R, I341K, P347K, K358R, S359N, S359H, L361A, F369N, F369D, F369G, F369W, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435L, I440L, N442H, N442M, Q448V, Q448N, N451S, R457K, E459K, K479M, H508C, H508Y, K530R, and E555R Engineered Cas polypeptides.
5. In paragraph 1, wherein said one or more amino acid substitutions comprise an amino acid substitution at a residue corresponding to one or more residues selected from the group consisting of (i) I365, A368, H399, A402, A406, C433, I435, A436 and D437; Engineered Cas polypeptides.
6. In paragraph 5, wherein said one or more amino acid substitutions comprise (i) one or more selected from the group consisting of I365V, A368S, H399L, A402S, A406S, C433K, I435V, I435L, A436T and D437N; Engineered Cas polypeptides.
7. In paragraph 5, wherein said one or more amino acid substitutions comprise amino acid substitutions at residues corresponding to any one selected from I365 / A368, H399 / A402 / A406, C433 / I435 / A436 / D437 and any combination thereof. Engineered Cas polypeptides.
8. In paragraph 5, wherein said one or more amino acid substitutions comprises an amino acid substitution selected from I365V / A368S; H399L / A402S / A406S; C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N; and any combination thereof. Engineered Cas polypeptides.
9. In paragraph 8, wherein one or more of the amino acid substitutions comprises I365V / A368S, H399L / A402S / A406S, C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N. Engineered Cas polypeptides.
10. In paragraph 8, The one or more amino acid substitutions include I365V / A368S; H399L / A402S / A406S; and C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N. Engineered Cas polypeptides.
11. In paragraph 5, wherein said one or more amino acid substitutions further comprise one or more other amino acid substitutions in one or more domains selected from the group consisting of WED, ZF, REC, TNB, Linker and RuvC domains. Engineered Cas polypeptides.
12. In paragraph 5, One or more of the amino acid substitutions mentioned above (iii) K31, N32, T33, K39, N47, P221, L222, K224, Q225, Q229, Y230, T231, E234, N237, K245, I246, R250, Q252, K254, K255, E256, I257, K259, K265, F266, D267, E269, Q270, Q272, I278, L282, K288, K295, D296, E297, K304, K305, N308, D310, Q312, T313, K319, G321, S322, K323, G325, M331, L332, N333 and D337; (iv) S48, E50, K53, D57, E58, N61, I65, L67, N70, D72, T89 and A120; (v) Q124, E125, I126, L133, K135, I140, Y141, I154, K155, A162, S163, S164, E166, H167, D171, T175, R176, A178, E179, L180, I186, L190, L200, N205, M206, G209, S215, D216 and F218; (vi) I341, D342, K343, D346 and P347; (vii) T502, K505, H508, L509, N511, E516, K519, K527, K530 and E535; and (viii) further comprising a substitution at a residue corresponding to one or more residues selected from the group consisting of K358, S359, L361, I365, A368, F369, K385, H399, A402, A406, K407, T414, E418, S420, E421, C433, I435, A436, D437, I440, N442, T446, Q448, N451, S454, R457, E459, I465, K479, K501, L543, L550, K551, S552, E555 and E556. Engineered Cas polypeptides.
13. In paragraph 12, One or more of the amino acid substitutions mentioned above (iii) K31R, N32Q, T33S, K39R, N47Q, P221T, P221S, P221C, P221G, P221A, P221V, P221I, P221M, P221F, L222F, L222I, K224R, Q225R, Q225H, Q225D, Q225E, Q225N, Q225Y, Q225K, Q225F, Q225C, Q225V, Q225P, Q225S, Q225T, Q225A, Q229G, Q229K, Q229N, Q229R, Q229H, Q229C, Q229P, Y230E, Y230R, Y230H, Y230K, Y230D, Y230N, Y230Q, Y230F, Y230C, Y230W, T231E, T231K, T231P, E234D, N237Q, K245R, I246L, R250K, Q252N, K254R, K255R, E256D, I257L, K259R, K265R, F266C, D267E, E269D, Q270R, Q272N, Q272R, I278L, L282I, K288R, K295R, D296E, E297D, K304R, K305R, N308Q, D310E, Q312K, Q312P, Q312N, T313I, T313V, K319R, G321K, S322K, K323R, G325C, M331F, L332V, N333Q and D337R; (iv) S48T, E50D, K53R, D57E, E58D, N61Q, I65E, I65L, I65S, L67I, N70Q, D72E, T89S and A120T; (v) Q124N, E125M, E125N, E125S, I126M, L133C, L133M, L133S, L133V, K135R, I140M, I140V, Y141F, I154L, K155R, K155Y, A162D, A162Q, A162P, A162V, S163R, S163K, S163N, S163M, S163F, S163Y, S163G, S163P, S163A, S163L, S164D, E166D, H167K, H167R, D171N, D171R, D171E, T175R, R176K, A178S, E179A, L180N, I186L, L190I, L200F, N205H, M206H, G209R, S215H, D216E and F218H; (vi) I341K, D342E, K343R, D346N and P347K; (vii) T502S, K505R, H508C, H508Y, L509I, N511Q, E516D, K519R, K527R, K530R and E535D; and (viii) K358R, S359H, S359N, L361A, I365V, A368S, F369D, F369G, F369N, F369W, K385R, H399L, A402S, A406S, K407E, T414E, E418R, S420A, S420D, E421A, E421K, C433K, I435V, I435L, A436T, D437N, I440L, N442H, N442M, T446S, Q448N, Q448V, N451S, S454G, R457K, E459K, I465R, K479M, K501R, Additionally comprising at least one selected from the group consisting of L543I, L550I, K551R, S552T, E555D, E555R and E556D. Engineered Cas polypeptides.
14. In paragraph 5, wherein said one or more amino acid substitutions further comprise a substitution at a residue corresponding to one or more residues selected from the group consisting of K135, Y141, K305 and M331. Engineered Cas polypeptides.
15. In paragraph 14, The above one or more amino acid substitutions further include K135R / Y141F / K305R / M331F Engineered Cas polypeptides.
16. In paragraph 14, wherein said one or more substitutions further comprise a substitution at a residue corresponding to one or more residues selected from the group consisting of L67, Q124, K155 and E556. Engineered Cas polypeptides.
17. In paragraph 16, One or more of the above substitutions additionally include L67I / Q124N / K155R / E556D Engineered Cas polypeptides.
18. In paragraph 16, wherein said one or more substitutions further comprise a substitution at a residue corresponding to one or more residues selected from the group consisting of T175, G209, K358 and S454. Engineered Cas polypeptides.
19. In paragraph 18, One or more of the above substitutions further comprise T175R / G209R / K358R / S454G Engineered Cas polypeptides.
20. In paragraph 1, wherein said one or more amino acid substitutions comprise any one selected from the group consisting of amino acid substitutions listed in Table 9. Engineered Cas polypeptides.
21. In paragraph 1, The engineered Cas polypeptide has an amino acid sequence selected from the group consisting of SEQ ID NOs: 201 to 384 and 410 to 606. Engineered Cas polypeptides.
22. In paragraph 1, wherein said one or more amino acid substitutions comprise an engineered Cas polypeptide comprising the following amino acid substitutions: L67I / Q124N / K135R / Y141F / K155R / K305R / M331F / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K530R / E556D (M1010); L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1120); L67I / Q124N / K135R / Y141F / K155R / N205H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1223); L67I / Q124N / K135R / Y141F / K155R / M206H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1224); L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1225); Q272R / N205H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1231); Q272R / M206H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1232); Q272R / Q270R / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1233); Q225P / Q229C / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM622); Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM625); Q229R / Y230W / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM628); K230H / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM638); K230C / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM646); S163R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM674); S163K / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM676); T175R / G209R / K358R / S454G / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM714); D171R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM715); T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM716); I465R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM722); G209R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM724); D171R / T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM727); T175R / G209R / S454G / K385R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM728); D171R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM729); T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM730); I465R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM731); G209R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM732); D171R / T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM733); K358R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM734); or T175R / G209R / K358R / S454G / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM735).
23. In paragraph 5, wherein said one or more substitutions comprise 2 to 40 amino acid substitutions. Engineered Cas polypeptides.
24. In paragraph 5, The engineered Cas polypeptide comprises an amino acid sequence having at least 80% identity to any one amino acid sequence selected from the group consisting of SEQ ID NOs: 201 to 207, 284 to 384, and 410 to 606. Engineered Cas polypeptides.
25. In paragraph 24, The engineered Cas polypeptide comprises an amino acid sequence having at least 80% identity to any one amino acid sequence selected from the group consisting of SEQ ID NOs: 380 to 384 and 470 to 472. Engineered Cas polypeptides.
26. In paragraph 1, The above engineered Cas polypeptide is part of a fusion protein. Engineered Cas polypeptides.
27. In paragraph 1, The engineered Cas polypeptides exhibit improved indel efficiency, PAM sequence recognition range and / or editing window. Engineered Cas polypeptides.
28. In paragraph 1, The above engineered Cas polypeptide is based on SEQ ID NO:
1. (i) I365 / A368, (ii) H399 / A402 / A406, (iii) C433 / I435 / A436 / D437, (iv) K135 / Y141 / K305 / M331, (v) L67 / Q124 / K155 / E556, (vi) T175 / G209 / K358 / S454, or (vii) amino acid substitutions at residues corresponding to any combination of these; Engineered Cas polypeptides.
29. In paragraph 28, The above engineered Cas polypeptide is based on SEQ ID NO:
1. (i) I365V / A368S, (ii) H399L / A402S / A406S, (iii) C433K / I435V / A436T / D437N or C433K / I435L / A436T / D437N, (iv) K135R / Y141F / K305R / M331F, (v) L67I / Q124N / K155R / E556D, (vi) T175R / G209R / K358R / S454G, or (vii) containing amino acid substitutions corresponding to any combination of these; Engineered Cas polypeptides.
30. An isolated nucleic acid encoding an engineered Cas polypeptide of any one of claims 1 to 29.
31. A vector comprising a nucleic acid sequence encoding an engineered Cas polypeptide of any one of claims 1 to 29.
32. In paragraph 31, The above vector is a non-viral vector or a viral vector. vector.
33. In paragraph 32, The above non-viral vector is plasmid DNA, mRNA (transcript) or PCR amplicon. vector.
34. In paragraph 32, The above viral vector is selected from the group consisting of retrovirus, lentivirus, adenovirus, adeno-associated virus, vaccinia virus, poxvirus and herpes simplex virus. vector.
35. Contains one or more vectors, The one or more vectors comprise (a) a first nucleic acid construct operably linked to a nucleotide encoding an engineered Cas polypeptide of any one of claims 1 to 29; and (b) a second nucleic acid construct operably linked to a nucleotide encoding a first guide nucleic acid comprising a guide sequence complementary to a target sequence. The nucleic acid structures of (a) and (b) above are contained in the same or different vectors. Vector system.
36. In paragraph 35, The above vector system further comprises a third nucleic acid construct operably linked to a nucleotide encoding a second guide nucleic acid comprising a guide sequence complementary to the target sequence, The second guide nucleic acid targets a target sequence that is identical or different from the first guide nucleic acid, The third nucleic acid structure is contained in a vector that is identical or different from the first nucleic acid structure or the second nucleic acid structure. Vector system.
37. In paragraph 35 or 36, The first and / or second guide nucleic acid comprises a scaffold region sequence of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 151 to 153. Vector system.
38. In paragraph 35 or 36, The first and / or second guide nucleic acid further comprises a U-rich tail at the 3' end of the guide sequence, and the sequence of the U-rich tail is represented by 5'-(UmV)nUo-3', wherein V is each independently A, C or G, m and o are integers ranging from 1 to 20, and n is an integer ranging from 0 to 5. Vector system.
39. In paragraph 37, The first and / or second guide nucleic acid comprises a nucleic acid sequence selected from the group consisting of SEQ ID NO: 151 to SEQ ID NO:
153. Vector system.
40. A system comprising an engineered Cas polypeptide of any one of claims 1 to 29 or a nucleic acid encoding the polypeptide; and one or more guide nucleic acids or nucleic acids encoding the guide nucleic acids.
41. In paragraph 40, The guide nucleic acid comprises i) a first sequence comprising the nucleotide sequence 5'-GGAAUGCAAC-3' and ii) a second sequence that hybridizes to a target sequence of a target nucleic acid, wherein the first sequence is located at the 5' side of the second sequence. System.
42. In paragraph 41, The above guide nucleic acid is an engineered guide nucleic acid, and the engineered guide nucleic acid comprises, compared to an unmodified guide RNA, first, second, third and fourth stem-loop regions, a tracrRNA-crRNA complementary region and a guide sequence in the 5'-to-3'-end direction, (1) deletion of part or all of the first stem-loop region; (2) deletion of part or all of the second stem-loop region; (b) deletion of part or all of the tracrRNA-crRNA complementary region; (c) substitution of one or more Uracils (U) with A, G or C when three or more, four or more or five or more consecutive Us are present within the tracrRNA-crRNA complementary region; and (d) addition of a U-rich tail to the 3'-end of the crRNA sequence (wherein the sequence of the U-rich tail is represented as 5'-(UmV)nUo-3', wherein V is each independently A, C or G, m and o are integers ranging from 1 to 20, and n is an integer ranging from 0 to 5). System.
43. In paragraph 42, The above guide nucleic acid comprises a scaffold region sequence of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 13, SEQ ID NO: 149 to SEQ ID NO: 186 and SEQ ID NO:
196. System.
44. In paragraph 43, The above guide nucleic acid comprises a scaffold region sequence of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 151 to 153. System.
45. In paragraph 43, The above guide nucleic acid comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 156, 183 and 196. System.
46. In paragraph 40, The above engineered Cas polypeptide is L67I / Q124N / K135R / Y141F / K155R / K305R / M331F / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K530R / E556D (M1010); L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1120); L67I / Q124N / K135R / Y141F / K155R / N205H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1223); L67I / Q124N / K135R / Y141F / K155R / M206H / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1224); L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1225); Q272R / N205H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1231); Q272R / M206H / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1232); Q272R / Q270R / S322K / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / I278L / K288R / K304R / K305R / N308Q / D310E / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (M1233); Q225P / Q229C / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM622); Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM625); Q229R / Y230W / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM628); K230H / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM638); K230C / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM646); S163R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM674); S163K / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM676); T175R / G209R / K358R / S454G / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM714); D171R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM715); T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM716); I465R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM722); G209R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM724); D171R / T175R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM727); T175R / G209R / S454G / K385R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM728); D171R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM729); T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM730); I465R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM731); G209R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM732); D171R / T175R / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM733); K358R / Q229H / Y230K / A178S / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM734); or Contains amino acid substitutions of T175R / G209R / K358R / S454G / Q225H / Q229R / L67I / Q124N / K135R / Y141F / K155R / L222I / E234D / Q270R / I278L / K288R / K304R / K305R / N308Q / D310E / S322K / M331F / D342E / K343R / I365V / A368S / H399L / A402S / A406S / C433K / I435V / A436T / D437N / K501R / L509I / E556D (PM735), The above guide nucleic acid comprises a scaffold region sequence of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 151 to 153. System.
47. In paragraph 40, The system comprises two or more guide nucleic acids, each of the two or more guide nucleic acids independently comprising a spacer sequence complementary to the same or different target sequence. System.
48. A composition comprising an engineered Cas polypeptide of any one of claims 1 to 29 or a nucleic acid encoding the polypeptide, and one or more guide nucleic acids comprising a guide sequence complementary to a target sequence or a nucleic acid encoding the guide nucleic acid.
49. In paragraph 48, The above engineered Cas polypeptide and guide nucleic acid form a ribonucleoprotein complex. Composition.
50. An isolated cell comprising an engineered Cas polypeptide of any one of claims 1 to 29, a vector of any one of claims 30 to 34, a vector system of any one of claims 35 to 39, a system of any one of claims 40 to 47, or a composition of claim 48 or 49.
51. In paragraph 50, The above cells are eukaryotic cells Isolated cells.
52. A method for regulating or editing a target nucleic acid in a cell, comprising contacting the cell with an engineered Cas polypeptide of any one of claims 1 to 29, a vector of any one of claims 30 to 34, a vector system of any one of claims 35 to 39, a system of any one of claims 40 to 47, or a composition of claim 48 or 49.
53. In paragraph 52, The above cells are eukaryotic cells method.
54. In paragraph 52, The above method is performed in vitro or ex vivo. method.
55. A fusion protein comprising an engineered Cas polypeptide and an effector domain of any one of claims 1 to 29.
Citation Information
Patent Citations
Guided editor variants, constructs and methods for enhanced guided editing efficiency and precision
CN117321201A
Ultrasonic medium supply jig of mult-pass type and ultrasonic scanning method using the same
KR1020240175114A
An engineered guide RNA for the optimized CRISPR / Cas12f1 system and use thereof
KR102455623B1
Optimized crispr / spcas12f1 system, engineered guide RNA and use thereof
WO2023206871A1
Compositions and methods for treating broad chemoresistance through chemoresistance-specific regulatory components
WO2024130151A1