Iscb polypeptides and uses thereof
Enhanced IscB.m16* proteins with expanded target-adjacent motifs and fused deaminase domains provide efficient gene editing in mammalian cells, addressing the limitations of large Cas9 and Cas12 systems by achieving robust base editing through AAV delivery.
Patent Information
- Application Number
- PCT/CN2025/072127
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-13
- Publication Date
- 2025-07-17
AI Technical Summary
Existing genome editing tools, such as Cas9 and Cas12, are hindered by their large size and limited activity, particularly in applications using adeno-associated virus (AAV) vectors, and compact alternatives like IscB proteins offer improved RNA-guided DNA endonuclease activity for base editing in mammalian cells.
Engineering IscB proteins, specifically IscB.m16*, to enhance RNA-guided DNA endonuclease activity and target-adjacent motif scope, and fusing them with deaminase domains to create robust base editors for efficient gene editing in mammalian cells.
The engineered IscB.m16* system achieves enhanced base editing efficiency in mammalian cells, effectively restoring DMD proteins in disease models via single AAV delivery, offering compact and efficient gene editing tools for therapeutic applications.
Smart Images

Figure PCTCN2025072127-FTAPPB-I100001 
Figure PCTCN2025072127-FTAPPB-I100002 
Figure PCTCN2025072127-FTAPPB-I100003
Abstract
Description
ISCB POLYPEPTIDES AND USES THEREOF
[0001] REFERENCE TO RELATED APPLICATIONS
[0002] The instant application claims the priority to and the benefit of the filing date of PCT / CN2024 / 071744, filed on January 11, 2024, the entire contents of which, including any drawings and sequence listing, are incorporated herein by reference.
[0003] REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0004] The disclosure contains a Sequence Listing XML file which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on January 12, 2025, by software “WIPO Sequence” according to WIPO Standard ST. 26, is named HEP007PCT. xml, and is 515, 689 bytes in size.
[0005] According to WIPO Standard ST. 26, symbol “t” is used to denote both T in DNA and U in RNA. Thus, in the instant sequence listing prepared according to ST. 26, wherever a sequence is an RNA, the T in the sequence shall be deemed as U.BACKGROUND
[0006] CRIPSR Cas systems, such as type II Cas9 and type V Cas12 systems, serving as the prokaryotic adaptive immunity system against viruses, have been developed into genome editing tools in basic research and gene therapy1-3. Engineered Cas9 nickase (nCas9) or deactivated Cas9 (dCas9) versions fused with various domains have established base editing, prime editing, and epigenome editing technologies4-6. However, the large size of Cas9 and Cas12, particularly nCas9-based gene editing tool, hinders the application of gene editing based on adeno-associated virus (AAV) vectors. Recently compact Cas9 (7-9) , Cas12f homologs10-14 (400-700 aa) , and TnpB15, 16 (~400 aa) the ancestral branch of Cas12, have been reported. However, due to poor editing activity or lack of HNH domain, these proteins remain limited activity of base editing.
[0007] Citation or identification of any document in the disclosure is not an admission that such a document is available as prior art to the disclosure. Each of the references mentioned or cited in the disclosure is incorporated by reference in its entirety.SUMMARY
[0008] As the evolutionary ancestor of Cas9 nuclease, IscB proteins serve as compact RNA-guided DNA endonucleases, making it a strong candidate for base editing. Here, the inventor identified 10 out of 19 uncharacterized IscB proteins from uncultured microbes showing RNA-guided DNA endonuclease activity in mammalian cells. Through protein and ωRNA engineering, the inventor further enhanced the RNA-guided DNA endonuclease activity of IscB ortholog IscB. m16 and expanded its target-adjacent motif (TAM) scope from MRNRAA to NNNGNA, resulting in an enhanced IscB system named as IscB. m16*. By fusing the deaminase domains with IscB. m16*nickase, the inventor generated IscB. m16*-derived base editors that exhibited robust base editing efficiency in mammalian cells, and effectively restored DMD proteins in disease mice via single adeno-associated virus delivery. This study thus establishes a set of compact base editing tools for basic research and therapeutic applications.
[0009] The invention of the disclosure is not and shall not be used to edit any human germ cell (i.e., an embryonic cell, an egg cell, a sperm cell) containing any genetic material in any jurisdiction unless it is allowed by applicable laws and regulations in the jurisdiction.
[0010] In an aspect, provided in the disclosure is an IscB polypeptide comprising an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 1-19.
[0011] In some embodiments, the IscB polypeptide is a mutant of SEQ ID NO: 16 and comprises an amino acid mutation (e.g., substitution) relative to (compared to) SEQ ID NO: 16 at a position selected from the group consisting of K30, K37, H38, N39, P50, V53, E74, E77, V79, H83, E85, K93, A96, K99, Q103, A104, H107, V133, E142, E159, P160, L172, T179, K180, E182, K196, E201, D204, H205, H214, L217, F218, E221, S222, D224, D225, Y228, A229, E232, G233, K234, G239, I250, H253, E254, A261, S268, G269, D272, L273, A278, A280, D282, K283, A285, V287, K292, K307, T310, A313, D318, E326, S328, F329, I330, S333, A335, P337, S350, V351, G354, S356, H357, G360, Q367, M369, H380, Q381, V384, K387, N391, G392, K393, H394, H400, K401, Q405, K406, G407, E411, L414, Q415, K416, N417, P418, G419, E423, M424, A426, E429, H430, K431, V433, K435, N438, L445, K449, N450, D451, V452, K456, T459, P460, I461, T462, N463, T465, F467, Y468, E473, G474, Q475, R476, H477, K478, L481, K483, P484, L486, H487, L494, G495, N496, G498, G499, Y500, P501, P502, Q503, I504, L505, G506, T507, H508, D509, K510, K511, and E513 of SEQ ID NO: 16.
[0012] In some embodiments, the amino acid mutation is a substitution with R, S, H, L, V, or E.
[0013] In some embodiments, the IscB polypeptide comprises an amino acid mutation (e.g., substitution) relative to (compared to) SEQ ID NO: 16 at a position selected from the group consisting of E326, H380, Q381, M424, V433, T459, P460, I461, T462, N463, T465, F467, Y468, Q475, R476, K478, L481, and I504 of SEQ ID NO: 16.
[0014] In some embodiments, the amino acid mutation is a substitution with R, S, H, L, V, or E.
[0015] In some embodiments, the IscB polypeptide comprises an amino acid substitution selected from the group consisting of E326R, T459E, P460S, and T462H.
[0016] In some embodiments, the IscB polypeptide comprises an amino acid combination substitution of E326R + T459E + P460S + T462H.
[0017] In some embodiments, the IscB polypeptide comprises an amino acid mutation (e.g., substitution) at a position selected from the group consisting of D61, E193, and H248 of SEQ ID NO: 16.
[0018] In some embodiments, the amino acid mutation is a substitution with A.
[0019] In some embodiments, the IscB polypeptide comprises an amino acid combination substitution of D61A + E326R + T459E + P460S + T462H; or an amino acid combination substitution of D61A + H248A + E326R+ T459E + P460S + T462H.
[0020] In some embodiments, the IscB polypeptide (1) has endonuclease activity; (2) has nickase activity; or (3) is endonuclease deficient.
[0021] In some embodiments, the IscB polypeptide comprises, consists essentially of, or consists the amino acid sequence of SEQ ID NO: 239, 240, or 241, or an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 239, 240, or 241.
[0022] In another aspect, provided in the disclosure is a fusion protein comprising the IscB polypeptide of the disclosure fused to a functional domain.
[0023] In some embodiments, the functional domain is fused at the N-terminal or C-terminal of the IscB polypeptide, or fused internally with respect to the IscB polypeptide.
[0024] In some embodiments, the functional domain is selected from the group consisting of a nuclear localization signal (NLS) , a nuclear export signal (NES) , a base editing domain, a deaminase or a catalytic domain thereof, a glycosylase or a catalytic domain thereof, an uracil glycosylase inhibitor (UGI) , an uracil glycosylase (UNG) (e.g., UNG1, UNG2) , a methylpurine glycosylase (MPG) , a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease or a catalytic domain thereof (e.g., T5 exonuclease of SEQ ID NO:404) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having ssRNA cleavage activity, a moiety having dsRNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA, selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, and a catalytic domain thereof, and a functional fragment thereof, and any combination thereof.
[0025] In some embodiments, the base editing domain is a deaminase (e.g., adenine deaminase, cytidine deaminase) or a catalytic domain thereof or a glycosylase (e.g., MPG, UNG2) or a catalytic domain thereof.
[0026] In some embodiments, the fusion protein comprises, from N-to C-terminus, an adenine deaminase domain, an optional linker, the IscB polypeptide, an optional linker, and an adenine deaminase domain.
[0027] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 259.
[0028] In some embodiments, the fusion protein comprises, from N-to C-terminus, a cytidine deaminase domain, an optional linker, the IscB polypeptide, an optional linker, and a UGI domain (e.g., one UGI domain) .
[0029] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 260.
[0030] In another aspect, provided in the disclosure is a polynucleotide encoding the IscB polypeptide of the disclosure or the fusion protein of the disclosure (e.g., SEQ ID NOs: 39-57) .
[0031] In another aspect, provided in the disclosure is a system comprising:
[0032] (1) the IscB polypeptide of the disclosure or the fusion protein of the disclosure, or a polynucleotide encoding the IscB polypeptide or the fusion protein, and
[0033] (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0034] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide or the fusion protein; and
[0035] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA, wherein the guide sequence is 5’ to the scaffold sequence.
[0036] In some embodiments, the scaffold sequence has substantially the same secondary structure as the secondary structure of any one of SEQ ID NOs: 20-38, 58-238, and 242-252; or wherein the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 20-38, 58-238, and 242-252.
[0037] In some embodiments, the guide sequence is about or at least about 14 nucleotides in length.
[0038] In another aspect, provided in the disclosure is a vector comprising the polynucleotide the disclosure. In some embodiments, the vector further comprises a guide nucleic acid in the disclosure or a polynucleotide encoding the guide nucleic acid. In some embodiments, the vector is a plasmid vector, a recombinant AAV (rAAV) vector, a recombinant lentivirus vector, a RNP, or an LNP.
[0039] In another aspect, provided in the disclosure is a cell comprising the IscB polypeptide of the disclosure, the fusion protein of the disclosure, the polynucleotide of the disclosure, the system of the disclosure, or the vector the disclosure.
[0040] In another aspect, provided in the disclosure is a method for modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, whereby the target DNA is modified.
[0041] In an aspect, the disclosure provides an IscB polypeptide comprising an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of any one of SEQ ID NOs: 1-19.
[0042] In another aspect, the disclosure provides an IscB polypeptide, wherein the IscB polypeptide comprises an amino acid mutation relative to (compared to) a wild type IscB polypeptide.
[0043] In yet another aspect, the disclosure provides a method of increasing guide sequence-specific binding ability (e.g., represented by the guide sequence-specific endonuclease activity of the IscB system or guide sequence-specific base editing efficiency of the IscB system) of an IscB polypeptide for use in an IscB system comprising (1) the IscB polypeptide, or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide, and (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0044] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide; and
[0045] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA;
[0046] wherein the scaffold sequence is 3’ to the guide sequence,
[0047] said method comprising introducing an amino acid mutation into the IscB polypeptide.
[0048] In yet another aspect, the disclosure provides a method of widening target adjacent motif (TAM) recognition (e.g., represented by the increased guide sequence-specific endonuclease activity of the IscB system or increased guide sequence-specific base editing efficiency of the IscB system for a broader TAM, e.g., a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’, wherein N is A, T, G, or C) of an IscB polypeptide for use in an IscB system comprising (1) the IscB polypeptide, or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide, and (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0049] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide; and
[0050] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA;
[0051] wherein the scaffold sequence is 3’ to the guide sequence,
[0052] said method comprising introducing an amino acid mutation into the IscB polypeptide.
[0053] In some embodiments, the IscB polypeptide comprises an amino acid mutation relative to (compared to) a reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19.
[0054] In some embodiments, the IscB polypeptide comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%) and less than 100%to the amino acid sequence of any one of SEQ ID NOs: 1-19.
[0055] In some embodiments, the amino acid mutation is within a domain of the reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19 selected from the group consisting of PLMP domain, RuvC-I domain, Bridge Helix domain, Linker domain, RuvC-II domain, HNH domain, RuvC-III domain, P1D domain, and TID domain.
[0056] In some embodiments, the PLMP domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 1-54 of SEQ ID NO: 16.
[0057] In some embodiments, the RuvC-I domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 55-85 of SEQ ID NO: 16.
[0058] In some embodiments, the Bridge Helix domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 86-122 of SEQ ID NO: 16.
[0059] In some embodiments, the Linker domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 123-160 of SEQ ID NO: 16.
[0060] In some embodiments, the RuvC-II domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 161-196 of SEQ ID NO: 16.
[0061] In some embodiments, the HNH domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 197-297 of SEQ ID NO: 16.
[0062] In some embodiments, the RuvC-III domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 298-374 of SEQ ID NO: 16.
[0063] In some embodiments, the P1D domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 375-429 of SEQ ID NO: 16.
[0064] In some embodiments, the TID domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 430-513 of SEQ ID NO: 16.
[0065] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from the group consisting of M1, A2, N3, V4, I5, Y6, V7, I8, N9, K10, D11, G12, K13, P14, L15, M16, P17, T18, T19, R20, R21, G22, H23, V24, G25, Y26, L27, L28, R29, K30, K31, Q32, A33, R34, V35, V36, K37, H38, N39, P40, F41, T42, V43, Q44, L45, S46, Y47, E48, T49, P50, D51, K52, V53, Q54, E55, L56, T57, L58, G59, I60, D61, P62, G63, R64, T65, N66, I67, G68, I69, A70, V71, V72, D73, E74, T75, G76, E77, C78, V79, F80, S81, A82, H83, V84, E85, T86, R87, N88, K89, D90, V91, P92, K93, L94, M95, A96, K97, R98, K99, V100, H101, R102, Q103, A104, R105, R106, H107, Y108, G109, R110, R111, V112, K113, R114, Q115, R116, R117, A118, K119, A120, N121, G122, T123, V124, N125, E126, N127, G128, I129, I130, T131, R132, V133, L134, P135, Q136, T137, E138, T139, P140, I141, E142, C143, K144, L145, I146, K147, N148, K149, E150, A151, R152, F153, C154, N155, R156, E157, R158, E159, P160, G161, W162, L163, T164, P165, T166, A167, N168, Q169, L170, L171, L172, T173, H174, L175, N176, L177, V178, T179, K180, I181, E182, Q183, I184, L185, P186, I187, S188, K189, I190, A191, L192, E193, I194, N195, K196, F197, A198, F199, M200, E201, L202, D203, D204, H205, N206, I207, R208, P209, W210, E211, Y212, Q213, H214, G215, P216, L217, F218, G219, F220, E221, S222, R223, D224, D225, A226, V227, Y228, A229, L230, Q231, E232, G233, K234, C235, L236, L237, C238, G239, K240, P241, L242, I243, E244, H245, Y246, H247, H248, V249, I250, P251, K252, H253, E254, H255, G256, S257, D258, T259, I260, A261, N262, I263, V264, G265, L266, C267, S268, G269, C270, H271, D272, L273, V274, H275, R276, D277, A278, R279, A280, K281, D282, K283, L284, A285, K286, V287, H288, A289, G290, A291, K292, K293, K294, Y295, A296, G297, T298, S299, V300, L301, N302, Q303, I304, M305, P306, K307, L308, I309, T310, R311, L312, A313, S314, K315, D316, E317, D318, F319, T320, L321, V322, S323, A324, K325, E326, I327, S328, F329, I330, R331, R332, S333, S334, A335, L336, P337, K338, D339, H340, H341, I342, D343, A344, Y345, C346, I347, A348, M349, S350, V351, V352, D353, G354, E355, S356, H357, M358, N359, G360, M361, L362, R363, K364, P365, Y366, Q367, V368, M369, Q370, F371, R372, R373, H374, D375, R376, Q377, A378, R379, H380, Q381, A382, M383, V384, D385, R386, K387, Y388, Y389, L390, N391, G392, K393, H394, V395, A396, T397, N398, R399, H400, K401, R402, F403, E404, Q405, K406, G407, D408, S409, L410, E411, E412, F413, L414, Q415, K416, N417, P418, G419, V420, R421, P422, E423, M424, L425, A426, V427, R428, E429, H430, K431, P432, V433, Y434, K435, R436, M437, N438, R439, I440, A441, P442, G443, T444, L445, M446, R447, C448, K449, N450, D451, V452, F453, V454, Y455, K456, T457, G458, T459, P460, I461, T462, N463, G464, T465, P466, F467, Y468, A469, V470, D471, T472, E473, G474, Q475, R476, H477, K478, Y479, R480, L481, S482, K483, P484, V485, L486, H487, N488, T489, G490, I491, V492, V493, L494, G495, N496, R497, G498, G499, Y500, P501, P502, Q503, I504, L505, G506, T507, H508, D509, K510, K511, R512, and E513 of SEQ ID NO: 16.
[0066] In some embodiments, the amino acid mutation leads to an increased guide sequence-specific endonuclease activity, or wherein the IscB polypeptide comprising said amino acid mutation has an increased guide sequence-specific endonuclease activity compared to the reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0067] In some embodiments, the IscB polypeptide comprising said amino acid mutation leads to increased guide sequence-specific base editing efficiency compared to the reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0068] In some embodiments, the IscB polypeptide has decreased guide sequence-independent (off-target) endonuclease activity or substantially lacks guide sequence-independent (off-target) endonuclease activity.
[0069] In some embodiments, the amino acid mutation leads to a decreased guide sequence-independent (off-target) endonuclease activity, or wherein the IscB polypeptide comprising said amino acid mutation has a decreased guide sequence-independent (off-target) endonuclease activity compared to the reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19, e.g., a decrease by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%.
[0070] In some embodiments, the IscB polypeptide is capable of recognizing a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’ immediately 3’ adjacent to a protospacer sequence of a target DNA, wherein N is A, T, G, or C.
[0071] In some embodiments, the IscB polypeptide has an increased guide-sequence specific endonuclease activity compared to that of SEQ ID NO: 16 for a protospacer sequence of a target DNA immediately 5’ adjacent to a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’, wherein N is A, T, G, or C.
[0072] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from the group consisting of K30, K37, H38, N39, P50, V53, E74, E77, V79, H83, E85, K93, A96, K99, Q103, A104, H107, V133, E142, E159, P160, L172, T179, K180, E182, K196, E201, D204, H205, H214, L217, F218, E221, S222, D224, D225, Y228, A229, E232, G233, K234, G239, I250, H253, E254, A261, S268, G269, D272, L273, A278, A280, D282, K283, A285, V287, K292, K307, T310, A313, D318, E326, S328, F329, I330, S333, A335, P337, S350, V351, G354, S356, H357, G360, Q367, M369, H380, Q381, V384, K387, N391, G392, K393, H394, H400, K401, Q405, K406, G407, E411, L414, Q415, K416, N417, P418, G419, E423, M424, A426, E429, H430, K431, V433, K435, N438, L445, K449, N450, D451, V452, K456, T459, P460, I461, T462, N463, T465, F467, Y468, E473, G474, Q475, R476, H477, K478, L481, K483, P484, L486, H487, L494, G495, N496, G498, G499, Y500, P501, P502, Q503, I504, L505, G506, T507, H508, D509, K510, K511, and E513 of SEQ ID NO: 16.
[0073] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from the group consisting of E326, H380, Q381, M424, V433, T459, P460, I461, T462, N463, T465, F467, Y468, Q475, R476, K478, L481, and I504 of SEQ ID NO: 16.
[0074] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from or that is a position the group consisting of D61, E193, and H248 of SEQ ID NO: 16.
[0075] In some embodiments, the amino acid substitution is a conservative amino acid substitution or a non-conservative amino acid substitution.
[0076] In some embodiments, the amino acid substitution is an amino acid substitution with an amino acid residue that is different from the amino acid residue at the position of SEQ ID NO: 16.
[0077] In some embodiments, the amino acid substitution is an amino acid substitution with
[0078] (1) a non-polar amino acid residue (such as, Glycine (Gly / G) , Alanine (Ala / A) , Valine (Val / V) , Cysteine (Cys / C) , Proline (Pro / P) , Leucine (Leu / L) , Isoleucine (Ile / I) , Methionine (Met / M) , Tryptophan (Trp / W) , Phenylalanine (Phe / F) ,
[0079] (2) a polar amino acid residue (such as, Serine (Ser / S) , Threonine (Thr / T) , Tyrosine (Tyr / Y) , Asparagine (Asn / N) , Glutamine (Gln / Q) ) ,
[0080] (3) a positively charged amino acid residue (such as, Lysine (Lys / K) , Arginine (Arg / R) , Histidine (His / H) ) , or
[0081] (4) a negatively charged amino acid residue (such as, Aspartic Acid (Asp / D) , Glutamic Acid (Glue / E) ) .
[0082] In some embodiments, the amino acid substitution is an amino acid substitution with a positively charged amino acid residue, such as, Arginine (R) .
[0083] In some embodiments, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) .
[0084] In some embodiments, the amino acid mutation comprises a substitution that is corresponding to a substitution or that is a substitution selected from the group consisting of K30R, K37R, H38R, N39R, P50R, V53R, E74R, E77R, V79R, H83R, E85R, K93R, A96R, K99R, Q103R, A104R, H107R, V133R, E142R, E159R, P160R, L172R, T179R, K180R, E182R, K196R, E201R, D204R, H205R, H214R, L217R, F218R, E221R, S222R, D224R, D225R, Y228R, A229R, E232R, G233R, K234R, G239R, I250R, H253R, E254R, A261R, S268R, G269R, D272R, L273R, A278R, A280R, D282R, K283R, A285R, V287R, K292R, K307R, T310R, A313R, D318R, E326R, S328R, F329R, I330R, S333R, A335R, P337R, S350R, V351R, G354R, S356R, H357R, G360R, Q367R, M369R, Q381R, V384R, K387R, N391R, G392R, K393R, H394R, H400R, K401R, Q405R, K406R, G407R, E411R, L414R, Q415R, K416R, N417R, P418R, G419R, E423R, M424R, A426R, E429R, H430R, K431R, V433R, K435R, N438R, L445R, K449R, N450R, D451R, V452R, K456R, T459R, T462R, N463R, T465R, E473R, G474R, Q475R, H477R, K478R, K483R, P484R, L486R, H487R, L494R, G495R, N496R, G498R, G499R, Y500R, P501R, P502R, Q503R, I504R, L505R, G506R, T507R, H508R, D509R, K510R, K511R, E513R, and a combination of any two or more residues thereof, wherein the position is numbered according to SEQ ID NO: 16.
[0085] In some embodiments, the amino acid mutation comprises a substitution that is corresponding to a substitution or that is a substitution selected from the group consisting of T459E, P460S, T462H, T462L, T465V, and a combination of any two or more residues thereof, wherein the position is numbered according to SEQ ID NO: 16.
[0086] In some embodiments, the amino acid mutation comprises a substitution that is corresponding to a substitution or that is a substitution selected from the group consisting of E326R, T459E, P460S, T462H, and a combination of any two or more residues thereof, wherein the position is numbered according to SEQ ID NO: 16.
[0087] In some embodiments, the IscB polypeptide comprising said amino acid mutation comprises a substitution that is corresponding to a substitution or that is a substitution selected from the group consisting of D61A, E193A, and H248A, wherein the position is numbered according to SEQ ID NO: 16.
[0088] In some embodiments, the amino acid mutation comprises a combination substitution corresponding to a combination substitution of E326R, P460S, T462H, and T459E, wherein the position is numbered according to SEQ ID NO: 16.
[0089] In some embodiments, the amino acid mutation comprises a combination substitution corresponding to a combination substitution of D61A, E326R, P460S, T462H, and T459E, wherein the position is numbered according to SEQ ID NO: 16.
[0090] In some embodiments, the amino acid mutation comprises a combination substitution corresponding to a combination substitution of D61A, H248A, E326R, P460S, T462H, and T459E, wherein the position is numbered according to SEQ ID NO: 16.
[0091] In some embodiments, the IscB polypeptide comprising said amino acid mutation comprises, consists essentially of, or consists an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the amino acid sequence of SEQ ID NO: 16.
[0092] In some embodiments, the IscB polypeptide comprising said amino acid mutation comprises, consists essentially of, or consists an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the amino acid sequence of SEQ ID NO: 239 or an N-terminal truncation of the amino acid sequence of SEQ ID NO: 239 lacking the most N-terminal Methionine (M) (coded by start codon ATG) .
[0093] In some embodiments, the IscB polypeptide comprising said amino acid mutation comprises, consists essentially of, or consists an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the amino acid sequence of SEQ ID NO: 240 or 241 or an N-terminal truncation of the amino acid sequence of SEQ ID NO: 240 or 241 lacking the most N-terminal Methionine (M) (coded by start codon ATG) .
[0094] In some embodiments, the IscB polypeptide is an endonuclease or has guide sequence-specific endonuclease activity.
[0095] In some embodiments, the IscB polypeptide is a nickase or has guide sequence-specific nickase activity.
[0096] In some embodiments, the IscB polypeptide is endonuclease deficient.
[0097] In some embodiments, the IscB polypeptide is catalytically inactive.
[0098] In some embodiments, the IscB polypeptide is fused to a functional domain to form a fusion protein.
[0099] In some embodiments, the functional domain has transposase activity, methylase activity, demethylase activity, translation activation activity, translation repression activity, transcription activation activity, transcription repression activity, transcription release factor activity, chromatin modifying or remodeling activity, histone modification activity, nuclease activity, single-strand RNA cleavage activity, double-strand RNA cleavage activity, single-strand DNA cleavage activity, double-strand DNA cleavage activity, nucleic acid binding activity, detectable activity, or any combination thereof.
[0100] In some embodiments, the functional domain is fused N-terminally, C-terminally, or internally with respect to the IscB polypeptide.
[0101] In some embodiments, the functional domain is fused to the IscB polypeptide via a linker, e.g., a XTEN linker, a GS linker containing multiple glycine and serine residues, a GS linker containing multiple glycine and serine residues and a XTEN linker, a GS linker containing multiple glycine and serine residues and a BP NLS.
[0102] In some embodiments, the functional domain is selected from the group consisting of a nuclear localization signal (NLS) , a nuclear export signal (NES) , a deaminase or a catalytic domain thereof, an uracil glycosylase inhibitor (UGI) , an uracil glycosylase (UNG) , a methylpurine glycosylase (MPG) , a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease or a catalytic domain thereof (e.g., T5 exonuclease of SEQ ID NO: 404) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having ssRNA cleavage activity, a moiety having dsRNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA, selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, and a catalytic domain thereof, and a functional fragment thereof, and any combination thereof.
[0103] In some embodiments, the fusion protein comprises a NLS at the N-terminal and / or the C-terminal of the IscB polypeptide.
[0104] In some embodiments, the fusion protein comprises one or two NLS at the N-terminal and / or the C-terminal of the IscB polypeptide.
[0105] In some embodiments, the fusion protein comprises a NLS at the N-terminal and / or the C-terminal of the functional domain.
[0106] In some embodiments, the fusion protein comprises one or two NLS at the N-terminal and / or the C-terminal of the functional domain.
[0107] In some embodiments, the NLS comprises or is SV40 NLS (SEQ ID NO: 258) , bpSV40 NLS (BP NLS, bpNLS, SEQ ID NO: 256) , or NP NLS (Xenopus laevis Nucleoplasmin NLS, nucleoplasmin NLS, SEQ ID NO: 257) .
[0108] In some embodiments, the functional domain comprises a deaminase or a catalytic domain thereof.
[0109] In some embodiments, the deaminase or catalytic domain thereof is an adenine deaminase (e.g., TadA, such as, TadA8e, TadA8.17, TadA8.20, TadA9) or a catalytic domain thereof.
[0110] In some embodiments, the adenine deaminase or a catalytic domain thereof comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 253.
[0111] In some embodiments, the adenine deaminase or a catalytic domain thereof is TadA8EV106W (SEQ ID NO: 253) .
[0112] In some embodiments, the deaminase or catalytic domain thereof is a cytidine deaminase (e.g., APOBEC, such as, APOBEC3, for example, APOBEC3A, APOBEC3B, APOBEC3C; DddA) or a catalytic domain thereof.
[0113] In some embodiments, the cytidine deaminase or a catalytic domain thereof comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 254.
[0114] In some embodiments, the cytidine deaminase or a catalytic domain thereof is APOBEC3AW104A (SEQ ID NO: 254) .
[0115] In some embodiments, the functional domain comprises an uracil glycosylase inhibitor (UGI) domain.
[0116] In some embodiments, the UGI domain comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 255.
[0117] In some embodiments, the fusion protein comprises one, two, or three UGI domains, and optionally, one UGI domain.
[0118] In some embodiments, the functional domain comprises an uracil glycosylase (UNG) .
[0119] In some embodiments, the functional domain comprises a methylpurine glycosylase (MPG) .
[0120] In some embodiments, the functional domain comprises a reverse transcriptase or a catalytic domain thereof.
[0121] In some embodiments, the functional domain comprises a methylase or a catalytic domain thereof.
[0122] In some embodiments, the functional domain comprises a transcription activating domain.
[0123] In some embodiments, the fusion protein comprises, from N-to C-terminus, an adenine deaminase domain, an optional linker, the IscB polypeptide, an optional linker, and an adenine deaminase domain.
[0124] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 259.
[0125] In some embodiments, the fusion protein comprises, from N-to C-terminus, a cytidine deaminase domain, an optional linker, the IscB polypeptide, an optional linker, and a UGI domain (e.g., one UGI domain) .
[0126] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 260.
[0127] In yet another aspect, the disclosure provides a system comprising:
[0128] (1) the IscB polypeptide or method of the disclosure, or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide, and
[0129] (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0130] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide; and
[0131] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA;
[0132] wherein the scaffold sequence is 3’ to the guide sequence.
[0133] In some embodiments, the guide nucleic acid is a guide RNA (gRNA) (interchangeably used with omega RNA (ωRNA) ) .
[0134] In some embodiments, the guide nucleic acid is capable of directing guide sequence specific binding of the complex to the target sequence of the target DNA.
[0135] In some embodiments, the scaffold sequence comprises a nucleotide mutation relative to a reference or wild type scaffold sequence compatible to the IscB polypeptide.
[0136] In some embodiments, the scaffold sequence comprises a nucleotide mutation relative to a reference or wild type scaffold sequence of any one of SEQ ID NOs: 20-38.
[0137] In some embodiments, the nucleotide mutation comprises a deletion of about, at least about, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in a stem-loop region of the reference scaffold sequence.
[0138] In some embodiments, the nucleotide mutation comprises a substitution of about, at least about, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more thermodynamically unstable base pairs in a stem-loop region of the reference scaffold sequence with a G-C or C-G base pair.
[0139] In some embodiments, the thermodynamically unstable base pair is a A-U or U-Abase pair, a A-G or G-Abase pair, or a U-G or G-U base pair.
[0140] In some embodiments, the stem-loop region is selected from the first 5’ stem loop region, the second 5’ stem loop region, the third 5’ stem loop region, the fourth 5’ stem loop region, the fifth 5’ stem loop region, or the sixth 5’ stem loop region of the reference scaffold sequence, wherein the first, the second, the third, the fourth, the fifth, and the sixth 5’ stem loop region are counted from the 5’ end of the reference scaffold sequence.
[0141] In some embodiments, the stem-loop region is selected from the first 5’ stem loop region and the first 3’ stem loop region, wherein the first 5’ stem loop region is counted from the 5’ end of the reference scaffold sequence, and wherein the first 3’ stem loop region is counted from the 3’ end of the reference scaffold sequence.
[0142] In some embodiments, the nucleotide mutation leads to an increased guide sequence-specific endonuclease activity, or wherein the system comprising the guide nucleic acid comprising said nucleotide mutation has an increased guide sequence-specific endonuclease activity compared to an otherwise identical control system comprising a guide nucleic acid without said nucleotide mutation, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0143] In some embodiments, the nucleotide mutation leads to increased guide sequence-specific base editing efficiency, or wherein the system comprising the guide nucleic acid comprising said nucleotide mutation has increased guide sequence-specific base editing efficiency compared to an otherwise identical control system comprising a guide nucleic acid without said nucleotide mutation, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0144] In some embodiments, the system is capable of recognizing a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’ immediately 3’ adjacent to a protospacer sequence of a target DNA, wherein N is A, T, G, or C.
[0145] In some embodiments, the system comprising the guide nucleic acid comprising said nucleotide mutation has an increased guide-sequence specific endonuclease activity compared to that of an otherwise identical control system comprising a guide nucleic acid without said nucleotide mutation for a protospacer sequence of a target DNA immediately 5’ adjacent to a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’, wherein N is A, T, G, or C.
[0146] In some embodiments, the scaffold sequence has substantially the same secondary structure as the secondary structure of any one of SEQ ID NOs: 20-38, 58-238, and 242-252.
[0147] In some embodiments, the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 20-38, 58-238, and 242-252.
[0148] In some embodiments, the target sequence comprises about or at least about 14 contiguous nucleotides of the target DNA, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more contiguous nucleotides of the target DNA, or in a numerical range between any two of the preceding values, e.g., from about 14 to about 20 contiguous nucleotides of the target DNA, from about 14 to about 50 contiguous nucleotides of the target DNA; optionally, wherein the target sequence comprises about 14 contiguous nucleotides of the target DNA.
[0149] In some embodiments, the target sequence is immediately 3’ to a target adjacent motif (TAM) , or wherein the reversely complementary sequence of the target sequence (i.e., the protospacer sequence) is immediately 5’ to a target adjacent motif (TAM) .
[0150] In some embodiments, the TAM is 5’-NNNGNA-3’, wherein N is A, T, G, or C; and optionally, wherein the TAM is 5’-NNNGAN-3’, wherein N is A, T, G, or C.
[0151] In some embodiments, the guide sequence is about or at least about 14 nucleotides in length, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more nucleotides in length, or in a length of a numerical range between any two of the preceding values, e.g., in a length of from about 14 to about 20 nucleotides, in a length of from about 14 to about 50 nucleotides; optionally, wherein the guide sequence is about 14 nucleotides in length.
[0152] In some embodiments, (1) the guide sequence is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% (fully) , optionally about 100% (fully) , reversely complementary to the target sequence; (2) the guide sequence contains no more than 5, 4, 3, 2, or 1 mismatch or contains no mismatch with the target sequence; or (3) the guide sequence comprises no mismatch with the target sequence in the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nucleotides at the 5’ end of the guide sequence.
[0153] In some embodiments, the guide sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 261-403 or comprises a polynucleotide sequence having no more than 1, 2, 3, 4, 5, 6, 7, or 8 nucleotide difference from any one of SEQ ID NOs: 261-403.
[0154] In some embodiments, the system comprises two or more guide nuclei acids comprising two or more guide sequences capable of hybridizing to two or more target sequences of the same target DNA or different target DNAs, wherein the two or more guide sequences are the same or different, and wherein the two or more target sequences are the same or different.
[0155] In some embodiments, the target DNA is a target dsDNA, such as, a eukaryotic dsDNA, e.g., a gene in a eukaryotic cell.
[0156] In some embodiments, the target DNA is a target dsDNA, and wherein the target dsDNA comprises a protospacer sequence on a nontarget strand of the target dsDNA, wherein the dsDNA comprises a target deoxyribonucleotide (e.g., dA, dT, dC, dG) at a position of the protospacer sequence selected from the group consisting of position 1, position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9, position 10, and a combination thereof; or wherein the target deoxyribonucleotide is at a position of the protospacer sequence between position 1 and position 1 or between position 2 and position 5, both inclusive.
[0157] In yet another aspect, the disclosure provides a guide nucleic acid as defined in the disclosure.
[0158] In yet another aspect, the disclosure provides a polynucleotide encoding the IscB polypeptide or method of the disclosure (e.g., SEQ ID NOs: 39-57) .
[0159] In yet another aspect, the disclosure provides a polynucleotide encoding the guide nucleic acid of the disclosure.
[0160] In yet another aspect, the disclosure provides a polynucleotide encoding the IscB polypeptide or method of the disclosure (e.g., SEQ ID NOs: 39-57) and the guide nucleic acid of the disclosure.
[0161] In yet another aspect, the disclosure provides a delivery system comprising (1) the IscB polypeptide or method of the disclosure, the polynucleotide of the disclosure, or the system of the disclosure; and (2) a delivery vehicle.
[0162] In yet another aspect, the disclosure provides a vector comprising the polynucleotide of the disclosure; optionally, wherein the vector encodes a guide nucleic acid as defined in the disclosure; optionally, wherein the vector is a plasmid vector, a recombinant AAV (rAAV) vector, or a recombinant lentivirus vector.
[0163] In yet another aspect, the disclosure provides a recombinant AAV (rAAV) particle comprising the rAAV vector of the disclosure; optionally, wherein the rAAV vector is an RNA.
[0164] In yet another aspect, the disclosure provides a ribonucleoprotein (RNP) comprising the IscB polypeptide or method of the disclosure and a guide nucleic acid optionally as defined in the disclosure.
[0165] In yet another aspect, the disclosure provides a lipid nanoparticle (LNP) comprising an RNA (e.g., mRNA) encoding the IscB polypeptide or method of the disclosure and a guide nucleic acid optionally as defined in the disclosure.
[0166] In yet another aspect, the disclosure provides a cell comprising the IscB polypeptide or method of the disclosure, the system of the disclosure, the polynucleotide of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the RNP of the disclosure, or the LNP of the disclosure.
[0167] In some embodiments, the cell is not a human germ cell (i.e., an embryonic cell, an egg cell, a sperm cell) .
[0168] In some embodiments, the cell is not a human embryonic stem cell.
[0169] In yet another aspect, the disclosure provides a method for modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, or the lipid nanoparticle of the disclosure, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex.
[0170] In some embodiments, the method is ex vivo, in vivo, or in vitro.
[0171] In some embodiments, the method is non-therapeutical.
[0172] In some embodiments, the target DNA is in a cell;
[0173] optionally, wherein the cell is a eukaryotic cell (e.g., an animal cell, a vertebrate cell, a mammalian cell, a non-human mammalian cell, a non-human primate cell, a rodent (e.g., mouse or rat) cell, a human cell, a plant cell, or a yeast cell) or a prokaryotic cell (e.g., a bacteria cell) ;
[0174] optionally, wherein the cell is from a plant or an animal;
[0175] optionally, wherein the plant is a dicotyledon; optionally selected from the group consisting of soybean, cabbage (e.g., Chinese cabbage) , rapeseed, brassica, watermelon, melon, potato, tomato, tobacco, eggplant, pepper, cucumber, cotton, alfalfa, eggplant, grape;
[0176] optionally, wherein the plant is a monocotyledon; optionally selected from the group consisting of rice, corn, wheat, barley, oat, sorghum, millet, grasses, Poaceae, Zizania, Avena, Coix, Hordeum, Oryza, Panicum (e.g., Panicum miliaceum) , Secale, Setaria (e.g., Setaria italica) , Sorghum, Triticum, Zea, Cymbopogon, Saccharum (e.g., Saccharum officinarum) , Phyllostachys, Dendrocalamus, Bambusa, Yushania; and / or
[0177] optionally, wherein the animal is selected from the group consisting of pig, ox, sheep, goat, mouse, rat, alpaca, monkey, rabbit, chicken, duck, goose, fish (e.g., zebra fish) .
[0178] In yet another aspect, the disclosure provides a cell modified by the method of the disclosure.
[0179] In some embodiments, the cell is not a human germ cell (i.e., an embryonic cell, an egg cell, a sperm cell) .
[0180] In some embodiments, the cell is not a human embryonic stem cell.
[0181] In yet another aspect, the disclosure provides a pharmaceutical composition comprising (1) the system of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, or the cell of the disclosure; and (2) a pharmaceutically acceptable excipient.
[0182] In yet another aspect, the disclosure provides a method of detecting a target DNA, comprising contacting the target DNA with In some embodiments, the target DNA is modified by the complex, and wherein the modification detects the target DNA; optionally, wherein the modification generates a detectable signal, e.g., a fluorescent signal.
[0183] In yet another aspect, the disclosure provides a method of increasing guide sequence-specific binding ability (e.g., represented by the guide sequence-specific endonuclease activity of the IscB system or guide sequence-specific base editing efficiency of the IscB system) of a guide nucleic acid for use in an IscB system comprising (1) an IscB polypeptide (e.g., the IscB polypeptide or method of the disclosure) , or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide, and (2) the guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0184] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide; and
[0185] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA;
[0186] wherein the scaffold sequence is 3’ to the guide sequence.
[0187] In some embodiments, the guide nucleic acid is a guide RNA (gRNA) (interchangeably used with omega RNA (ωRNA) ) .
[0188] In some embodiments, the guide nucleic acid is capable of directing guide sequence specific binding of the complex to the target sequence of the target DNA.
[0189] In some embodiments, the method comprises introducing a nucleotide mutation into the scaffold sequence relative to a reference or wild type scaffold sequence compatible to the IscB polypeptide.
[0190] In some embodiments, the method comprises introducing a nucleotide mutation into the scaffold sequence relative to a reference or wild type scaffold sequence of any one of SEQ ID NOs: 20-38.
[0191] In some embodiments, the nucleotide mutation comprises a deletion of about, at least about, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in a stem-loop region of the reference scaffold sequence.
[0192] In some embodiments, the nucleotide mutation comprises a substitution of about, at least about, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more thermodynamically unstable base pairs in a stem-loop region of the reference scaffold sequence with a G-C or C-G base pair.
[0193] In some embodiments, the thermodynamically unstable base pair is a A-U or U-Abase pair, a A-G or G-Abase pair, or a U-G or G-U base pair.
[0194] In some embodiments, the stem-loop region is selected from the first 5’ stem loop region, the second 5’ stem loop region, the third 5’ stem loop region, the fourth 5’ stem loop region, the fifth 5’ stem loop region, or the sixth 5’ stem loop region of the reference scaffold sequence, wherein the first, the second, the third, the fourth, the fifth, and the sixth 5’ stem loop region are counted from the 5’ end of the reference scaffold sequence.
[0195] In some embodiments, the stem-loop region is selected from the first 5’ stem loop region and the first 3’ stem loop region, wherein the first 5’ stem loop region is counted from the 5’ end of the reference scaffold sequence, and wherein the first 3’ stem loop region is counted from the 3’ end of the reference scaffold sequence.
[0196] In some other aspects, the disclosure provides IscB-based base editors capable of A-to-G or C-to-T transition.
[0197] 1. Cytidine Base Editing System
[0198] In an aspect, provided in the disclosure is an IscB-based dsDNA base editing system, comprising:
[0199] (1) a base editor, or a polynucleotide encoding the base editor, the base editor comprising:
[0200] (i) an IscB domain;
[0201] (ii) a cytidine deamination domain; and
[0202] (iii) an uracil glycosylase inhibitor (UGI) domain capable of inhibiting a uracil-DNA glycosylase (UDG) ; and
[0203] (2) a guide RNA (or “ωRNA” throughout the disclosure) , or a polynucleotide encoding the guide RNA, the guide RNA comprising:
[0204] (i) a scaffold sequence capable of forming a complex with the IscB domain; and
[0205] (ii) a guide sequence capable of hybridizing to a target sequence on a target strand of a target dsDNA, thereby guiding the complex to the target dsDNA;
[0206] wherein the cytidine deamination domain is capable of deaminating a cytosine base of a target nucleotide of a protospacer sequence on the nontarget strand of the target dsDNA, wherein the protospacer sequence is complementary to the target sequence.
[0207] IscB domain
[0208] In some embodiments, the IscB domain is a IscB nickase (nIscB) domain capable of nicking the target strand.
[0209] In some embodiments, the IscB domain is a IscB nickase of the disclosure.
[0210] In some embodiments, the IscB domain is a IscB nickase (nIscB) domain that comprises a D-to-Amutation corresponding to the D61A mutation in the amino acid sequence of SEQ ID NO: 444.
[0211] In some embodiments, the IscB domain comprise an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of SEQ ID NO: 444.
[0212] Cytidine deaminase domain
[0213] In some embodiments, the cytidine deamination domain is a deaminase from an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase.
[0214] In some embodiments, the APOBEC family deaminase is selected from the group consisting of APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, and APOBEC3H deaminase.
[0215] In some embodiments, the cytidine deamination domain comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of SEQ ID NO: 254.
[0216] In some embodiments, the cytidine deaminase domain is an activation-induced deaminase (AID) .
[0217] In some embodiments, the cytidine deaminase domain is a cytidine deaminase 1 from Petromyzon marinus (pmCDA1) .
[0218] UGI domain
[0219] In some embodiments, the base editor comprises one UGI domain, two UGI domains, or three UGI domains; and optionally one UGI domain.
[0220] In some embodiments, the UGI domain comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of SEQ ID NO: 255.
[0221] Fusion protein
[0222] In some embodiments, the base editor comprises one IscB domain and one cytidine deamination domain.
[0223] In some embodiments, the base editor is a fusion protein comprising the IscB domain, the cytidine deamination domain, and the UGI domain, wherein any adjacent two of the IscB domain, the cytidine deamination domain, and the UGI domain are connected to each other with or without a linker.
[0224] In some embodiments, the IscB domain is at the N-or C-terminal, optionally C-terminal, of the cytidine deamination domain.
[0225] In some embodiments, the UGI domain is or is not between the IscB domain and the cytidine deamination domain.
[0226] In some embodiments, the UGI domain is at the N-or C-terminal of both the IscB domain and the cytidine deamination domain.
[0227] In some embodiments, the fusion protein comprises the structure:
[0228] NH2- [cytidine deaminase domain] - [IscB domain] - [UGI domain] -COOH, and wherein each instance of “-” comprises an optional linker.
[0229] NLS
[0230] In some embodiments, the base editor comprises a nuclear localization sequence (NLS) , e.g., one, two, three, or four NLSs.
[0231] In some embodiments, the base editor comprises a NLS N-or C-terminally fused to the cytidine deaminase domain.
[0232] In some embodiments, the base editor comprises a NLS N-or C-terminally fused to the IscB domain.
[0233] In some embodiments, the base editor comprises a NLS N-or C-terminally fused to the UGI domain.
[0234] In some embodiments, the NLS comprises the amino acid sequence PKKKRKV, MDSLLMNRRKFLYQFKNVRWAKGRRETYLC, KRTADGSEFESPKKKRKV (bpNLS 1, SEQ ID NO: 445) , KRTADGSESEPKKKRKV (bpNLS 1, SEQ ID NO: 446) , or KRPAATKKAGQAKKKK (npNLS, SEQ ID NO: 447) .
[0235] Linker
[0236] In some embodiments, the cytidine deaminase domain and the IscB domain are linked to each other via a linker comprising the amino acid sequence (GGGS) n, (GGGGS) n, (G) n, (EAAAK) n, (GGS) n, (SGGS) n, SGSETPGTSESATPES (XTEN linker, SEQ ID NO: 448) , GGGGGSGGGGSGGGGSGGGGS (SEQ ID NO: 449) , a GS linker containing a XTEN linker (SEQ ID NO: 448) (e.g., SGGSSGGSSGSETPGTSESATPESSGGSSGGS (GS-XTEN-GS linker, SEQ ID NO: 459) ) , a GS linker containing a NLS (e.g., bpNLS, such as SEQ ID NO: 445 or 446) (e.g., SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS (GS-bpNLS-GS linker, SEQ ID NO: 450) ) , or (XP) n motif, or a combination thereof, wherein n is independently an integer between 1 and 30, inclusive, and wherein X is any amino acid.
[0237] In some embodiments, the cytidine deaminase domain and the IscB domain are linked to each other via a linker comprising a XTEN linker (e.g., SEQ ID NO: 448) or a GS linker (e.g., SEQ ID NO: 449) .
[0238] In some embodiments, the UGI domain and the NLS are linked via a linker comprising the amino acid sequence of SGGSGGSGGS or SGGSSGGS .
[0239] Specific CBE
[0240] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of any one of SEQ ID NOs: 406-415.
[0241] 2. Adenine Base Editing System
[0242] In an aspect, provided in the disclosure is an IscB-based dsDNA base editing system, comprising:
[0243] (1) a base editor, or a polynucleotide encoding the base editor, the base editor comprising:
[0244] (i) an IscB domain;
[0245] (ii) an adenine deamination domain; and
[0246] (2) a guide RNA, or a polynucleotide encoding the guide RNA, the guide RNA comprising:
[0247] (i) a scaffold sequence capable of forming a complex with the IscB domain; and
[0248] (ii) a guide sequence capable of hybridizing to a target sequence on a target strand of a target dsDNA, thereby guiding the complex to the target dsDNA;
[0249] wherein the adenine deamination domain is capable of deaminating an adenine base of a target nucleotide of a protospacer sequence on the nontarget strand of the target dsDNA, wherein the protospacer sequence is complementary to the target sequence.
[0250] IscB domain
[0251] In some embodiments, the IscB domain is a IscB nickase (nIscB) domain capable of nicking the target strand.
[0252] In some embodiments, the IscB domain is a IscB nickase of the disclosure.
[0253] In some embodiments, the IscB domain is a IscB nickase (nIscB) domain that comprises a D-to-Amutation corresponding to the D61A mutation in the amino acid sequence of SEQ ID NO: 444.
[0254] In some embodiments, the IscB domain comprise an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of SEQ ID NO: 444.
[0255] Adenine deaminase domain
[0256] In some embodiments, the adenine deamination domain is a deaminase from a Bacterial tRNA adenosine deaminase (TadA) family deaminase.
[0257] In some embodiments, the APOBEC family deaminase is selected from the group consisting of E. coli TadA (ecTadA) , any TadA mutant in the patent family of WO2018027078A1 incorporated herein by reference, TadA8e, TadA8e-V106W, TadA9, any TadA mutant in Re-engineering the adenine deaminase TadA-8e for efficient and specific CRISPR-based cytosine base editing incorporated herein by reference.
[0258] In some embodiments, the adenine deamination domain comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of SEQ ID NO: 253.
[0259] Fusion protein
[0260] In some embodiments, the base editor comprises one IscB domain and one adenine deamination domain.
[0261] In some embodiments, the IscB domain is at the N-terminal or the C-terminal, optionally N-terminal, of the adenine deamination domain.
[0262] In some embodiments, the base editor comprises one IscB domain and two adenine deamination domains.
[0263] In some embodiments, the two adenine deamination domains are separated by the IscB domain.
[0264] In some embodiments, the two adenine deamination domains are in tandem and at the N-terminal or the C-terminal, optionally C-terminal, of the IscB domain.
[0265] In some embodiments, the base editor is a fusion protein comprising the IscB domain and the adenine deamination domain that are connected to each other with or without a linker.
[0266] In some embodiments, the fusion protein comprises the structure:
[0267] NH2- [IscB domain] - [adenine deaminase domain] -COOH, or
[0268] NH2- [IscB domain] - [adenine deaminase domain] - [adenine deaminase domain] -COOH,
[0269] and wherein each instance of “-” comprises an optional linker.
[0270] NLS
[0271] In some embodiments, the base editor comprises a nuclear localization sequence (NLS) , e.g., one, two, three, or four NLSs.
[0272] In some embodiments, the base editor comprises a NLS N-or C-terminally fused to the adenine deaminase domain.
[0273] In some embodiments, the base editor comprises a NLS N-or C-terminally fused to the IscB domain.
[0274] In some embodiments, the NLS comprises the amino acid sequence PKKKRKV, MDSLLMNRRKFLYQFKNVRWAKGRRETYLC, KRTADGSEFESPKKKRKV (bpNLS 1, SEQ ID NO: 445) , KRTADGSESEPKKKRKV (bpNLS 1, SEQ ID NO: 446) , or KRPAATKKAGQAKKKK (npNLS, SEQ ID NO: 447) .
[0275] Linker
[0276] In some embodiments, the adenine deaminase domain and the IscB domain are linked to each other via a linker comprising the amino acid sequence (GGGS) n, (GGGGS) n, (G) n, (EAAAK) n, (GGS) n, (SGGS) n, SGSETPGTSESATPES (XTEN linker, SEQ ID NO: 448) , GGGGGSGGGGSGGGGSGGGGS (SEQ ID NO: 449) , a GS linker containing a XTEN linker (SEQ ID NO: 448) (e.g., SGGSSGGSSGSETPGTSESATPESSGGSSGGS (GS-XTEN-GS linker, SEQ ID NO: 459) ) , a GS linker containing a NLS (e.g., bpNLS, such as SEQ ID NO: 445 or 446) (e.g., SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS (GS-bpNLS-GS linker, SEQ ID NO: 450) ) , or (XP) n motif, or a combination thereof, wherein n is independently an integer between 1 and 30, inclusive, and wherein X is any amino acid.
[0277] In some embodiments, the adenine deaminase domain and the IscB domain are linked to each other via a linker comprising a XTEN linker (e.g., SEQ ID NO: 448) or a bpNLS (e.g., SEQ ID NO: 445 or 446) , for example, a GS-XTEN-GS linker (SEQ ID NO: 459) , a GS-bpNLS-GS linker (SEQ ID NO: 450) .
[0278] In some embodiments, the two adenine deaminase domains are linked to each other via a linker comprising a XTEN linker (e.g., SEQ ID NO: 448) or a bpNLS (e.g., SEQ ID NO: 445 or 446) , for example, a GS-XTEN-GS linker (SEQ ID NO: 459) , a GS-bpNLS-GS linker (SEQ ID NO: 450) .
[0279] In some embodiments, the two adenine deaminase domains are in tandem and linked to each other via a linker comprising a bpNLS (e.g., SEQ ID NO: 445 or 446) , for example, a GS-bpNLS-GS linker (SEQ ID NO: 450) , and the two adenine deaminase domains in tandem are linked to the IscB domain via a linker comprising a bpNLS (e.g., SEQ ID NO: 445 or 446) , for example, a GS-bpNLS-GS linker (SEQ ID NO: 450) .
[0280] Specific ABE
[0281] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of any one of SEQ ID NOs: 416-427.
[0282] Guide RNA
[0283] In some embodiments, the guide sequence is in a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 or more nucleotides.
[0284] In some embodiments, the scaffold sequence is 3’ to the guide sequence.
[0285] In some embodiments, the scaffold sequence has substantially the same secondary structure as the secondary structure of SEQ ID NO: 442.
[0286] In some embodiments, the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the polynucleotide sequence of SEQ ID NO: 442.
[0287] The details of one or more embodiments of the disclosure are set forth in the description below. Other features or advantages of the disclosure will be apparent from the following drawings and detailed description of several embodiments, and also from the appended claims. It is understood that any aspect or embodiment of the disclosure can be combined with any other one or more aspects or embodiments of the disclosure, including aspects or embodiments only described in one sub-section, only in the examples, or only in the claims, to constitute another embodiment explicitly or implicitly disclosed herein unless otherwise indicated.BRIEF DESCRIPTION OF THE DRAWINGS
[0288] An understanding of the features and advantages of the disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure may be utilized, and the accompanying drawings of which:
[0289] FIG. 1 shows identification and characterization of functional IscB orthologs. a, Phylogenetic tree of 19 uncharacterized IscB orthologs. b, TAMs of 19 IscB proteins and two OgeuIscB proteins19 (wtOeguIscB and enOgeuIscB) determined by bacterial depletion assay. c, Schematics describing detection of editing activity based on the fluorescence signal of GFxxFP reporter activation in HEK293T cells. d, Fluorescence signal of EGFP activated by IscB-mediated double-strand breaks (DSB) quantified by flow cytometry. Non-target means non-targeting spacer with random sequence. Asterisks means >2-fold ratio of on-target and non-target, with on-target >1.0%, representing variants with activity in HEK293T cells. Values represent mean of three independent biological replicates.
[0290] FIG. 2 shows engineering of various IscB ωRNAs to improve editing efficiency in mammalian cells. a, Secondary structure of IscB. m16 ωRNA predicted by RNAfold. Five regions were indicated as R1, R2, R3, R4 and R5. b, Increased EGFP signal induced by IscB. m16 caused by truncation of stem loop in R1 (R1-Δ13b) and R5 (R5-Δ10) of ωRNA. c, Substitutions of A-U to G-C based on trimmed ωRNA (R1-Δ13b and R5-Δ10) enhanced the EGFP fluorescence signal mediated by IscB. m16. v2.27 represents IscB. m16-ωRNA variant with R1-Δ13, R5-Δ10, 24-G, 25-C, 57-G, 79-C, 117-C. d-e, ωRNA engineering for IscB. m17. The truncation of stem loop and / or the substitutions of A-U to G-C in ωRNA improved the editing efficiency of IscB. m17. f, Fluorescence EGFP intensity of various IscB orthologs along with their corresponding WT or optimized ωRNA, quantified by flow cytometry. Non-target means spacer with random sequence. Values represent mean of three independent biological replicates. The red dashed lines represent the value of WT. The red arrows represent the current optimal variants for each IscB ωRNA. Values and error bars represent mean and s.d., n = 3 independent biological replicates.
[0291] FIG. 3 shows protein engineering of IscB. m16 to improve endonuclease activity and expand TAM range in mammalian cells. a, Screening for highly active mutants by single substitutions of amino acid residues in RuvC domain of IscB. m16 with arginine (R) . Each dot represents the endonuclease activity of a single mutant. The dashed line indicates the endonuclease activity of IscB. m16 ( “WT” ) . b, Screening for highly active mutants by saturation mutagenesis at selected sites using GFxxFP reporters with TAM pool 1. Each dot represents the endonuclease activity of a single mutant. The dashed line indicates the endonuclease activity of IscB. m16 ( “WT” ) . c, Comparison of endonuclease activity of IscB. m16 and its mutants for 16 GFxxFP reporters with different TAMs. IscB. m16-Srepresents mutant IscB. IscB. m16-RS represents mutant IscB. IscB. m16-RSH represents mutant IscB. and IscB. m16-RSV represents mutant IscB. Colored dots reflect the mean of three independent biological replicates. d, Screening for mutants with improved endonuclease activity based on GFxxFP reporters containing three different TAM pools along with the same ωRNA-v2.27. Orange bar represents mutant IscB. m16RESH, which is the combination of IscB. m16RSH and single substitution T495E. e, The second round of ωRNA engineering by the substitution of C-G base pair on ωRNA-v2.27 (R1-Δ13, R5-Δ10, 24-G, 25-C, 57-G, 79-C, 117-C) based on IscB. m16RESH. f, Comparison of indel frequency of WT IscB. m16 and its mutants at five endogenous sites in HEK293T cells. Values and error bars represent mean and s.d., n = 3 independent biological replicates. g, TAM logs of IscB. m16 and IscB. m16*system.
[0292] FIG. 4 shows characterization of IscB-and SpG-derived base editors in mammalian cells. a, Schematic of IscB, OgueIscB and SpG with different sizes. b, Overview of TAM / PAM-matched sites used to compare IscB. m16-derived ABE to enOgueIscB-ABE and SpG-ABE. c, Editing window and base editing activity of IscB. m16-ABE, IscB. m16*-ABE, enOgueIscB-ABE and SpG-ABE at all protospacer positions. Value and error bars are presented as means ± s.e.m..d, Comparison of the A-to-G conversion efficiency of IscB. m16-ABE, IscB. m16*-ABE, enOgueIscB-ABE, and SpG-ABE at 33 endogenous loci. Each dot represents the average highest base editing activity at each endogenous target site of three independent biological replicates. e, Indel efficiency of IscB-and SpG-derived adenine base editors at 33 endogenous loci. f, Comparison of the A-to-G conversion efficiency of IscB. m16-ABE, IscB. m16*-ABE, enOgueIscB-ABE, and SpG-ABE grouped by TAM at 33 target sites. g, Comparison of the C-to-T conversion efficiency of IscB. m16-CBE, IscB. m16*-CBE, enOgueIscB-CBE and SpG-CBE at 8 target sites. Each dot represents the average highest base editing activity at each endogenous target site of three independent biological replicates. h, Indel efficiency of IscB-and SpG-derived cytosine base editors at 8 endogenous loci. P values determined by Tukey’ s multiple comparisons test following ordinary one-way analysis of variance. *P < 0.05, **P < 0.01, ***P < 0.001 and ****P < 0.0001. NS, not significant. Value and error bars are presented as means ± s.d..
[0293] FIG. 5 shows IscB. m16*based cytosine base editor mediates effective base editing and restores dystrophin expression in humanized DMDE51del mice. a, Schematics of the strategy of IscB. m16*-CBE treating DMD. IscB. m16*-CBE disrupts the conserved guanine within splice acceptor site for programmable exon 50 skipping, leading to the restoration of dystrophin expression. b, The C-to-T conversion efficiency of IscB. m16*-CBE, enOgeuIscB-CBE and SpG-CBE at the splice acceptor site of DMD intron between exon 49 and exon 50 in HEK293T cells. c, Schematics of single AAV9 carrying two versions of IscB. m16*-CBE delivered to the muscles in mice via TA muscle injection. Saline was injected in the left leg, while AAV9 cargo IscB. m16*-CBE in the right. d-e, In vivo G-to-H (C-to-D, including C-to-T, C-to-A, and C-to-G) editing efficiencies (d) and RNA level of exon 50 skipping (e) of AAV9-IscB. m16*-CBE were detected by targeted deep sequencing. f, Dystrophin immunohistochemistry showed the restoration of dystrophin expression following 3 weeks after TA injection of IscB. m16*-CBE. Dystrophin is shown in green. Scale bar, 100 μm. g Quantification of Dys+ fibers and dystrophin in cross sections of TA muscles from (f) . h. Western blot analysis of dystrophin and vinculin expression in TA muscles 3 weeks after injection with AAV9-IscB. m16*-CBE or saline. TA means tibialis anterior. Value and error bars are presented as means ± s.d., n = 3 independent biological replicates.
[0294] FIG. 6 shows maximum-likelihood tree of identified IscB orthologs and previously reported IscBs.
[0295] FIG. 7 shows functional properties of IscB systems. Distribution of conserved residues in 19 newly identified IscB proteins and 2 reported IscB proteins, OgeuIscB and AwaIscB. The conserved residues in RuvC, HNH, P1D and TID domains were marked as asterisks and blue lines in the bottom bar.
[0296] FIG. 8 shows identification of TAM profiling. a, Schematic illustrating the plasmid depletion experiment for the detection of the TAM. b, Depletion data distribution plot of TAM profiling. Diagonal lines in coordinates represent that the ratio of experiment and control normalized TAM abundance is 1. σ means standard deviation (STD) . Pink dots represent TAMs with over minus 3-fold STD, while blue dots denoting TAMs with less than minus 3-fold STD. c, The protein and TAM divergence plot of 19 identified IscB proteins and OgeuIscB. The x-axis represents pairwise protein sequence divergence, and y-axis represents pairwise TAM divergence. A value close to 1 indicates high consistency, and close to 0 indicates high diversity.
[0297] FIG. 9 shows secondary structures of ωRNA of IscB. m16 (a) , IscB. m17 (b) , IscB. m1 (c) , IscB. m8 (d) , IscB. m15 (e) , and IscB. m18 (f) predicted by RNAfold.
[0298] FIG. 10 shows engineering various IscB ωRNAs to improve their editing efficiency in mammalian cells. a-d, IscB. m1 (a) and IscB. m8 (b) , improved EGFP signal by truncation of stem loop of ωRNA, and IscB. m15 (c) and IscB. m18 (d) improved EGFP signal by truncation of stem loop in R1 and R5 of ωRNA. The red dotted lines represent the value of WT. The red arrows represent the current optimal variants for each IscB ωRNA. Data are shown as mean and s.d., n = 3 independent biological replicates.
[0299] FIG. 11 shows screening for IscB. m16 mutants with enhanced endonuclease activity using six GFxxFP reporters containing different TAMs. Substitutions of amino acid residues in P1D (a) and TID (b) domain of IscB. m16 protein with arginine (R) . The shade of the color represents the ratio of the endonuclease activity between WT IscB. m16 and its mutants.
[0300] FIG. 12 shows TAM profiling of IscB. m16 and its mutants using GFxxFP fluorescence reporter system in HEK293T cells. a-b, TAM profiling of IscB. m16 at 19 reporters containing different TAMs. c, Screening for highly active mutants from saturation mutants with a single substitution at H380, Q381 and M424 of IscB. m16 based on GFxxFP reporters containing a TAM pool 1 target. Each dot represents the endonuclease activity for a single mutant. The dashed line indicates the endonuclease activity of IscB. m16. d-e, Comparison of endonuclease activity among IscB. m16 and its mutants for 19 GFxxFP reporters, which contain different TAMs. Colored dots reflect the mean of three independent biological replicates. f, TAM profiling of IscB. m16RSH for 64 reporters containing the same spacer and 5’-NNNGAA-3’ TAM sequences. Values and error bars represent mean and s.d., n = 3 independent biological replicates.
[0301] FIG. 13 shows TAM profiling of IscB. m16 and IscB. m16RESH in HEK293T cells using fluorescence reporter system. a-b, TAM profiling of IscB. m16 and IscB. m16RESH for 72 reporters containing the same spacer and 5’-NNNGAA-3’ (a) or 5’-AAAGNN-3’ (b) TAM sequences, along with ωRNA-v2.27. Data shown as the mean of three independent biological replicates.
[0302] FIG. 14 shows characterization of IscB. m16 and IscB. m16*editing activities in HEK293T cells. a, Effect of the spacer length on cleavage activity generated by IscB. m16 WT and IscB. m16*at 2 sites on GFxxFP reporters. b, The cleavage patterns of indels generated by IscB. m16 WT and variants at five endogenous sites. Values and error bars were shown as mean and s.d., n = 3 independent biological replicates.
[0303] FIG. 15 shows introduction of D61A and / or H248A mutations results in nickase or abolished activity of IscB. m16 variants. D10A and H840A mutations in SpCas9 serve as positive controls. Values and error bars were displayed as mean and s.d., n = 3 independent biological replicates.
[0304] FIG. 16 shows A-to-G conversion efficiency of IscB-and SpG-based adenine base editors at endogenous loci in HEK293T cells. Comparisons of A-to-G conversion efficiency of IscB. m16-ABE, IscB. m16*-ABE, enOgueIscB-ABE, SpG-ABE at 33 target sites containing various TAMs in HEK293T cells. Value and error bars are presented as means ± s.d., n = 3 independent biological replicates.
[0305] FIG. 17 shows base editing and indel efficiency of IscB-and SpG-based adenine and cytosine base editors at endogenous loci in HEK293T cells. a, Indel frequency of different adenine base editors at endogenous sites from FIG. 10. b, Comparisons of C-to-T conversion efficiency of IscB. m16-CBE, IscB. m16*-CBE, enOgueIscB-CBE, SpG-CBE at 8 target sites containing various TAMs in HEK293T cells. c, Indel frequency of different cytosine base editors at endogenous sites from (b) . Value and error bars are presented as means ± s.d., n = 3 independent biological replicates.
[0306] FIG. 18 shows the gRNA-dependent off-target levels of ABEs at in-silico predicted off-target sites, determined by targeted deep sequencing. The left, middle and right panels are A-to-G conversion rate of IscB. m16*-ABE, enOgeuIscB-ABE and SpG-ABE targeting ALDH1A3-S2, VEGFA-S3 and VEGFA-S6 genes respectively. Value and error bars are presented as means ± s.d., n = 3 independent biological replicates.
[0307] FIG. 19 shows the gRNA-independent off-target levels of five R loops formed by dSaCas9 in IscB-and SpG-derived adenine base editors in HEK293T cells. a-c, The gRNA-independent off-target levels of base editors at ALDH1A3-S1 (a) , EMX1-S2 (b) , and VEGFA-S1 (c) . d, Statistical of gRNA-independent off-target levels at 3 endogenous sites. Value and error bars are presented as means ± s.d., n = 3 independent biological replicates. P values determined by Tukey’ s multiple comparisons test following ordinary one-way analysis of variance. NS, not significant.
[0308] Fig. 20 shows the secondary structure of the scaffold sequence of SEQ ID NO: 73.
[0309] Fig. 21 shows the secondary structure of the scaffold sequence of SEQ ID NO: 96.
[0310] Fig. 22 shows the secondary structure of the scaffold sequence of SEQ ID NO: 107.
[0311] Fig. 23 shows an exemplified target DNA and an exemplified IscB system comprising a guide nucleic acid and an IscB polypeptide.
[0312] FIG. 24 shows the constructs of the IscB based cytidine base editors of the disclosure.
[0313] FIG. 25 shows the constructs of the IscB based adenine base editors of the disclosure.
[0314] FIG. 26 shows schematic exemplary constructs of a target dsDNA and a corresponding ωRNA.
[0315] FIG. 27 shows the domain organization of IscB. P1D, P1 interaction domain; TID, TAM-interaction domain.
[0316] RuvC domain is separated into three segments: RuvC I, II, and III.
[0317] The figures herein are for illustrative purposes only and are not necessarily drawn to scale.DETAILED DESCRIPTION
[0318] The disclosure will be described with respect to particular embodiments, but the disclosure is not limited thereto in any respect. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which this disclosure belongs. Terms as set forth hereinafter are generally to be understood in their plain and ordinary meaning or common sense unless indicated otherwise.
[0319] Definition
[0320] The disclosure will be described with respect to particular embodiments, but the disclosure is not limited thereto in any respect. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which this disclosure belongs. Terms as set forth hereinafter are generally to be understood in their plain and ordinary meaning or common sense unless indicated otherwise. The definitions and explanations of terms in WO2022087494A1 are incorporated herein by reference in their entireties except for the extent that a different definition or explanation of a term is specifically provided herein.
[0321] IscB, as a nucleic acid programmable DNA endonuclease similar to Cas9 and Cas12, is capable of binding to a DNA (e.g., a dsDNA) as guided by a guide nucleic acid (e.g., a guide RNA) comprising a guide sequence targeting the DNA. IscB may be associated with the guide nucleic acid (e.g., a guide RNA) , which localizes / targets the IscB to a DNA that comprises a DNA strand (i.e., a target strand) that is reversely complementary to the guide nucleic acid, or a portion thereof (e.g., the guide sequence of a guide RNA) . In other words, the guide nucleic acid “programs” the IscB to localize and bind to the DNA. Binding of the IscB to the DNA enables the IscB or a construct comprising the IscB to access to and function on the DNA. IscB is discussed in Altae-Tran H, Kannan S, Demircioglu FE et al.. The widespread IS200 / IS605 transposon family encodes diverse programmable RNA-guided endonucleases. SCIENCE 2021; 374 (6563) : 57-65 (and all its supplementary materials) ; Kato, K., Okazaki, S., Kannan, S. et al. Structure of the IscB–ωRNA ribonucleoprotein complex, the likely ancestor of CRISPR-Cas9. Nat Commun 13, 6719 (2022) (and all its supplementary materials) ; and WO2022087494A1, the entire contents of which, including any drawings and sequence listing, are incorporated herein by reference.
[0322] Without wishing to be bound by theory, in some embodiments, the guide nucleic acid comprises a scaffold sequence responsible for forming a complex with the IscB, and a guide sequence that is intentionally designed to be responsible for hybridizing to a target sequence of the DNA, thereby guiding the complex comprising the IscB and the guide nucleic acid to the DNA such that the IscB is indirectly bound to the DNA.
[0323] The ability of IscB to be bound to a target DNA by being guided by such a guide nucleic acid makes IscB a nucleic acid programmable DNA binding protein (napDNAbp) similar to Cas9 and Cas12.
[0324] Referring to FIG. 23, an exemplary dsDNA is depicted to comprise a 5’ to 3’ s ingle DNA strand and a 3’ to 5’ s ingle DNA strand, the 5’ to 3’ s ingle DNA strand comprises an exemplary first deoxyribonucleotide dA, and the 3’ to 5’ s ingle DNA strand comprises an exemplary second deoxyribonucleotide dC that base pairs with the dT.
[0325] An exemplary guide nucleic acid is depicted to comprise a guide sequence and a scaffold sequence. The guide sequence is designed to hybridize to a part of the 3’ to 5’ s ingle DNA strand, and so the guide sequence “targets” that part. And thus, the 3’ to 5’ s ingle DNA strand is referred to as a “target strand (TS) ” of the dsDNA, while the opposite 5’ to 3’ s ingle DNA strand is referred to as a “nontarget strand (NTS) ” of the dsDNA. That part of the target strand based on which the guide sequence is designed and to which the guide sequence may hybridize is referred to as a “target sequence” , while the opposite part on the nontarget strand corresponding to that part is referred to as the “protospacer sequence” , which is typically 100%(fully) reversely complementary to the target sequence, if there is no intentional or unintentional mismatch.
[0326] Generally, as is conventional in the art, a nucleic acid sequence (e.g., a DNA sequence) is written in 5’ to 3’ direction / orientation unless explicitly indicated otherwise.
[0327] For example, for a DNA sequence of ATGC, it is usually understood as 5’-ATGC-3’ unless otherwise indicated. Its reverse sequence is 5’-CGTA-3’. Its fully complementary sequence is 5’-TACG-3’. Its fully reverse complementary sequence is 5’-GCAT-3’. Note that the fully complementary sequence usually does not have the ability to base-pair / hybridize with the original sequence.
[0328] Generally, the double-strand sequence of a dsDNA may be represented with the sequence of its 5’ to 3’ s ingle DNA strand conventionally written in 5’ to 3’ direction / orientation unless otherwise indicated.
[0329] For example, for a dsDNA having a 5’ to 3’ s ingle DNA strand of 5’-ATGC-3’ and a 3’ to 5’ single DNA strand of 3’-TACG-5’ as shown below, the dsDNA may be simply represented as 5’-ATGC-3’.
[0330] 5’-----ATGC -----3’
[0331] 3’-----TACG -----5’
[0332] It should be noted that either the 5’ to 3’ s ingle DNA strand or the 3’ to 5’ s ingle DNA strand of a dsDNA can be a nontarget strand from which a protospacer sequence is selected.
[0333] In the sense of base editing, the strand on which the target nucleotide (e.g., deoxyribonucleotide dA) to be edited is located is termed as an edited strand, and the opposite strand is termed as a non-edited strand. As used herein, the nontarget strand is the edited strand, and the target strand is the non-edited strand.
[0334] Typically for a gene, the 5’ to 3’ s ingle DNA strand of the gene is the sense strand, and the 3’ to 5’ s ingle DNA strand of the gene is the antisense strand. Either the sense strand or the antisense strand can be a nontarget strand from which a protospacer sequence is selected.
[0335] To hybridize to a dsDNA, such as, a dsDNA 5’-ATGC-3’, the guide sequence of a guide nucleic acid, in one embodiment, is designed to have a sequence of 5’-AUGC-3’ that is fully reversely complementary to the 3’ to 5’ strand of the dsRNA (3’-TACG-5’) , which would be set forth in ATGC in the electric sequence listing and marked as an RNA sequence according to WIPO standard ST. 26; and in another embodiment, the guide sequence of a guide nucleic acid is designed to have a sequence of 5’-GCAU-3’ that is fully reversely complementary to the 5’ to 3’ strand of the dsDNA (5’-ATGC-3’) , which would be set forth in GCAT in the electric sequence listing and marked as an RNA sequence according to WIPO standard ST. 26.
[0336] In the case that the guide sequence of a guide nucleic acid is fully reversely complementary to the target sequence and the target sequence is fully reversely complementary to the protospacer sequence, the guide sequence is identical to the protospacer sequence except for the difference between the U in the guide sequence due to its RNA nature and the corresponding T in the protospacer sequence due to its DNA nature. According to WIPO standard ST. 26, symbol “t” is used to denote both T in DNA and U in RNA (See “Table 1: List of nucleotides symbols” , the definition of symbol “t” is “thymine in DNA / uracil in RNA (t / u) ” ) . Thus, in the electronic sequence listing of the disclosure prepared according to WIPO standard ST. 26, such a guide sequence could be set forth in the same sequence as a corresponding protospacer sequence. For convenience, a single SEQ ID NO in the electronic sequence listing can be used to denote both such guide sequence and protospacer sequence, although such a single SEQ ID NO may be marked as either DNA or RNA in the electronic sequence listing. When a reference is made to such a SEQ ID NO that sets forth a protospacer / guide sequence, it refers to either a protospacer sequence that is a DNA sequence or a guide sequence that is an RNA sequence depending on the context, no matter whether it is marked as a DNA or an RNA in the electronic sequence listing.
[0337] Without wishing to be bound by theory, it is believed that IscB contains a RuvC nuclease domain separated into three segments: RuvC I, II, and III domains, which is responsible for single-strand cleavage at the nontarget strand of a dsDNA, and a HNH nuclease domain, which is responsible for single-strand cleavage at the target strand of a dsDNA, together leading to double-strand cleavage of the dsDNA.
[0338] As used herein, the terms “nucleic acid” , “nucleic acid molecule” , or “polynucleotide” are used interchangeably. They refer to a polymer of deoxyribonucleotides or ribonucleotides or their mixtures of any length in either single-or double-stranded form, and, unless otherwise stated, encompass known analogs of natural nucleotides that can function in a similar manner as naturally occurring nucleotides. The terms encompass nucleic acid-like structures with synthetic backbones, as well as amplification products. DNAs and RNAs are both polynucleotides. The polymer may include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine) , nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6) -methylguanine, and 2-thiocytidine) , chemically modified bases, biologically modified bases (e.g., methylated bases) , intercalated bases, modified sugars (e.g., 2′-fluororibose, ribose, 2′-deoxyribose, arabinose, and hexose) , or modified phosphate groups (e.g., phosphorothioates and 5′-N-phosphoramidite linkages) .
[0339] As used herein, the phrase “polynucleotide encodes / encoding polypeptide X” or a similar phrase refers to a polynucleotide that is translated to express polypeptide Y comprising polypeptide X, meaning that polypeptide X is all or part of polypeptide Y. For example, polynucleotide encoding a Cas protein can refer to (i) a polynucleotide that is translated to express a fusion protein comprising the Cas protein and one or more additional amino acids, wherein the Cas protein is part of the fusion protein; or alternatively, (ii) a polynucleotide that is translated to express the Cas protein per se without any additional amino acid.
[0340] As used herein, the phrase “polynucleotide encodes / encoding RNA X” or a similar phrase refers to a polynucleotide that is transcribed to RNA Y comprising RNA X, meaning that RNA X is all or part of RNA Y. For example, polynucleotide encoding a gRNA can refer to (i) a polynucleotide that is transcribed to an RNA comprising the gRNA and one or more additional nucleotides, wherein the gRNA is part of the transcribed RNA; or alternatively, (ii) a polynucleotide that is transcribed to the gRNA per se without any additional nucleotide.
[0341] As used herein, the term “polypeptide” and “protein” are used interchangeably to refer to a polymer of amino acids of any length. The polymer may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified; for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component.
[0342] As used herein, a “fusion protein” refers to a protein created through the joining of two or more originally separate proteins, or portions thereof. In some embodiments, a linker may be present between each protein.
[0343] As used herein, the term “heterologous, ” in reference to polypeptide domains, refers to the fact that the polypeptide domains do not naturally occur together (e.g., in the same polypeptide) . For example, in fusion proteins generated by the hand of man, a polypeptide domain from one polypeptide may be fused to a polypeptide domain from a different polypeptide. The two polypeptide domains would be considered “heterologous” with respect to each other, as they do not naturally occur together.
[0344] As used herein, the term “heterologous, ” in reference to nucleotide sequences, refers to the fact that the nucleotide sequences do not naturally occur together (e.g., in the same polynucleotide) . For example, in a guide nucleic acid generated by the hand of man, a guide sequence intentionally designed to target a human gene locus may be fused to a scaffold sequence from a microorganism. The two nucleotide sequences would be considered “heterologous” with respect to each other, as they do not naturally occur together.
[0345] As used herein, the term “guide nucleic acid” refers to a nucleic acid-based molecule capable of forming a complex with a nucleic acid programmable protein, for example, an IscB polypeptide of the disclosure (e.g., via a scaffold sequence of the guide nucleic acid) , and comprises a sequence (e.g., a guide sequences) that is sufficient to hybridize to a target nucleic acid and guides the complex to the target nucleic acid, which includes but is not limited to RNA-based molecules, e.g., a guide RNA. As used herein, the terms “guide RNA (gRNA) ” , “omega RNA” , “ωRNA” , and “RNA guide” are used interchangeably. As used in the disclosure, the term “guide sequence” is used interchangeably with the term “spacer sequence” .
[0346] As used herein, the term “complex” refers to a grouping of two or more molecules. In some embodiments, the complex comprises a nucleic acid and a polypeptide interacting with (e.g., binding to, coming into contact with, adhering to) one another. As used herein, the term “complex” can refer to a grouping of a guide nucleic acid and a polypeptide (e.g., an IscB polypeptide) . As used herein, the term “complex” can refer to a grouping of a guide nucleic acid, a polypeptide (e.g., an IscB polypeptide) , and a target nucleic acid (e.g., a target DNA) .
[0347] As used herein, if a DNA sequence, for example, 5’-ATGC-3’ is transcribed to an RNA sequence, with each dT (deoxythymidine, or “T” for short) in the primary sequence of the DNA sequence replaced with a U (uridine) and each dA (deoxyadenosine, or “A” for short) , dG (deoxyguanosine, or “G” for short) , and dC (deoxycytidine, or “C” for short) replaced with A (adenosine) , G (guanosine) , and C (cytidine) , respectively, for example, 5’-AUGC-3’, it is said in the disclosure that the DNA sequence “encodes” the RNA sequence.
[0348] As used herein, the terms “protospacer adjacent motif (PAM) ” and “target adjacent motif (TAM) ” are used interchangeably and refer to a short sequence (or a motif) adjacent to a protospacer sequence on the nontarget strand of a dsDNA recognizable by an IscB polypeptide or a complex comprising an IscB polypeptide and a guide nucleic acid. In some embodiments, the PAM or TAM is immediately 3’ to a protospacer sequence.
[0349] As used herein, the term “adjacent” includes instances wherein there is no nucleotide between the protospacer sequence and the PAM and also instances wherein there are a small number (e.g., 1, 2, 3, 4, or 5) of nucleotides between the protospacer sequence and the PAM. As used herein, A “immediately adjacent (to) ” B, A “immediately 5’ to” B, and A “immediately 3’ to” B mean that there is no nucleotide between A and B.
[0350] As described herein, the guide sequence is so designed to be substantially capable of hybridizing to a target sequence. As used herein, the term “hybridize” , “hybridizing” , or “hybridization” refers to a reaction in which one or more polynucleotide sequences react to form a complex that is stabilized via hydrogen bonding between the bases of the one or more polynucleotide sequences. The hydrogen bonding may occur by Watson Crick base pairing, Hoogstein binding, or in any other sequence specific manner. A polynucleotide sequence capable of hybridizing to a given polynucleotide sequence is referred to as the “complement” of the given polynucleotide sequence. As used herein, the hybridization of a guide sequence and a target sequence is so stabilized to permit an IscB polypeptide that is complexed with a guide nucleic acid comprising the guide sequence or a function domain (e.g., a deaminase domain) associated (e.g., fused) with the IscB polypeptide to act (e.g., cleave, deaminize) at or near the target sequence or its complement (e.g., a sequence of a target DNA or its complement) .
[0351] For the purpose of hybridization, in some embodiments, the guide sequence is reversely complementary to a target sequence. As used herein, the term “reverse complementary” refers to the ability of nucleobases of a first polynucleotide sequence, such as a guide sequence, to base pair with nucleobases of a second polynucleotide sequence, such as a target sequence, by traditional Watson-Crick base-pairing. Two reverse complementary polynucleotide sequences are able to non-covalently bind under appropriate temperature and solution ionic strength conditions. In some embodiments, a first polynucleotide sequence (e.g., a guide sequence) comprises 100% (fully) reverse complementarity to a second nucleic acid (e.g., a target sequence) . In some embodiments, a first polynucleotide sequence (e.g., a guide sequence) is reverse complementary to a second polynucleotide sequence (e.g., a target sequence) if the first polynucleotide sequence comprises at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%complementarity to the second nucleic acid (i.e., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%of the nucleotides of the first polynucleotide sequence can base-pair with the nucleotides of the second polynucleotide sequence) . As used herein, the term “substantially complementary” refers to a polynucleotide sequence (e.g., a guide sequence) that has a certain level of complementarity to a second polynucleotide sequence (e.g., a target sequence) (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%of the guide sequence can base-pair with the polynucleotide sequence of the target sequence, or at most 1, 2, 3, 4, or 5 contiguous or non-contiguous nucleotides of the guide sequence mismatch the nucleotides of the target sequence) . In some embodiments, the level of complementarity is such that the first polynucleotide sequence (e.g., a guide sequence) can hybridize to the second polynucleotide sequence (e.g., a target sequence) with sufficient affinity to permit an IscB polypeptide that is complexed with the first polynucleotide sequence or a nucleic acid comprising the first polynucleotide sequence or a function domain (e.g., a deaminase domain) associated (e.g., fused) with the IscB polypeptide to act (e.g., cleave, deaminize) on the target sequence or its complement (e.g., a sequence of a target DNA or its complement) . In some embodiments, a guide sequence that is substantially complementary to a target sequence has 100%or less than 100%complementarity to the target sequence. In some embodiments, a guide sequence that is substantially complementary to a target sequence has at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementarity to the target sequence, and / or has at most 1, 2, 3, 4, or 5 contiguous or non-contiguous nucleotide mismatches from the target sequence.
[0352] As used herein, the term “identity” refers to the overall relatedness between polymeric molecules, e.g., between nucleic acids (e.g., DNA and / or RNA) and / or between polypeptides. In some embodiments, polymeric molecules are considered to be “substantially identical” to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%identical. Calculation of the percent identity of two nucleic acids or polypeptides, for example, can be performed by aligning the two sequences for optimal comparison purpose (e.g., gaps can be introduced in one or both of a first and a second sequences for optimal alignment and non-identical sequences can be disregarded for comparison purposes) . In certain embodiments, the length of a sequence aligned for comparison purpose is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or substantially 100%of the length of a reference sequence. The nucleotides at corresponding positions are then compared. The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. As is well known in the art, nucleic acids or polypeptides may be compared using any of a variety of algorithms, including those available in commercial computer programs such as BLASTN for nucleotide sequences and BLASTP, gapped BLAST, and PSI-BLAST for amino acid sequences. In some embodiments, the sequence identity is calculated by global alignment, for example, using the Needleman-Wunsch algorithm and an online tool at ebi. ac. uk / Tools / psa / emboss_needle / . In some embodiments, the sequence identity is calculated by local alignment, for example, using the Smith-Waterman algorithm and an online tool at ebi. ac. uk / Tools / psa / emboss_water / .
[0353] As used herein, the term “variant” refers to an entity that shows significant structural identity with a reference entity (e.g., a wild-type sequence) but differs structurally from the reference entity in the presence or level of one or more chemical moieties as compared with the reference entity. In many embodiments, a variant also differs functionally from its reference entity. In general, whether a particular entity is properly considered to be a “variant” of a reference entity is based on its degree of structural identity with the reference entity. As will be appreciated by those skilled in the art, any biological or chemical reference entity has certain characteristic structural elements. A variant, by definition, is a distinct chemical entity that shares one or more such characteristic structural elements. To give but a few examples, a polypeptide may have a characteristic sequence element comprising a plurality of amino acids having designated positions relative to one another in linear or three-dimensional space and / or contributing to a particular biological function; a nucleic acid may have a characteristic sequence element comprising a plurality of nucleotide residues having designated positions relative to one another in linear or three-dimensional space. For example, a variant polypeptide may differ from a reference polypeptide as a result of one or more differences in amino acid sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, etc. ) covalently attached to the polypeptide backbone. In some embodiments, a variant polypeptide shows an overall sequence identity with a reference polypeptide (e.g., a nuclease described herein) that is at least about 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%or 99%. Alternatively or additionally, in some embodiments, a variant polypeptide does not share at least one characteristic sequence element with a reference polypeptide, for example, an IscB nickase as a variant of a reference IscB polypeptide does not share an active RuvC domain or an active HNH domain with the reference IscB polypeptide. In some embodiments, the reference polypeptide has one or more biological activities. In some embodiments, a variant polypeptide shares one or more of the biological activities of the reference polypeptide, e.g., nuclease activity. In some embodiments, a variant polypeptide lacks one or more of the biological activities of the reference polypeptide, for example, an IscB nickase as a variant of a reference IscB polypeptide does not share the endonuclease activity of the reference IscB polypeptide. In some embodiments, a variant polypeptide shows a reduced level of one or more biological activities (e.g., nuclease activity, e.g., off-target nuclease activity) as compared with the reference polypeptide. In some embodiments, a polypeptide of interest is considered to be a “variant” of a reference polypeptide if the polypeptide of interest has an amino acid sequence that is identical to that of the reference polypeptide but for a small number of sequence alterations at particular positions. Typically, fewer than about 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%of the residues in the variant are substituted as compared with the reference polypeptide. In some embodiments, a variant has about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 substituted residue as compared with a reference polypeptide. Often, a variant has a very small number (e.g., fewer than about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1) of substituted functional residues (i.e., residues that participate in a particular biological activity) . In some embodiments, a variant has not more than about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 additions or deletions, and often has no additions or deletions, as compared with the reference polypeptide. Moreover, any additions or deletions are typically fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 12, about 11, about 10, about 9, about 8, about 7, about 6, and commonly are fewer than about 5, about 4, about 3, or about 2 residues. In some embodiments, the reference polypeptide is a wild type polypeptide. A variant of a polynucleotide may be naturally occurring such as an allelic variant, or it may be a variant that is not known to occur naturally. Non-naturally occurring variants of a polynucleotide may be made by mutagenesis techniques, by direct synthesis, and by other recombinant methods known to skilled artisans.
[0354] As used herein, the terms “non-naturally occurring” and “engineered” are used interchangeably and refer to artificial participation. When these terms are used to describe a nucleic acid or a polypeptide, it is meant that the nucleic acid or polypeptide is at least substantially freed from at least one other component of its association in nature or as found in nature.
[0355] In some embodiments, a “conservative substitution” refers to a substitution of an amino acid made among amino acids within one of the following four groups:
[0356] (1) non-polar amino acids, including Glycine (Gly / G) , Alanine (Ala / A) , Valine (Val / V) , Cysteine (Cys / C) , Proline (Pro / P) , Leucine (Leu / L) , Isoleucine (Ile / I) , Methionine (Met / M) , Tryptophan (Trp / W) , and Phenylalanine (Phe / F) ;
[0357] (2) negatively charged amino acids, including Aspartic Acid (Asp / D) and Glutamic Acid (Glu / E) ;
[0358] (3) polar amino acids, including Serine (Ser / S) , Threonine (Thr / T) , Tyrosine (Tyr / Y) , Asparagine (Asn / N) , and Glutamine (Gln / Q) ; and
[0359] (4) positively charged amino acids, including Lysine (Lys / K) , Arginine (Arg / R) , and Histidine (His / H) .
[0360] As used herein, the term “wild type” has the meaning commonly understood by those skilled in the art to mean a typical form of an organism, a strain, a gene, or a feature that distinguishes it from a mutant or variant when it exists in nature. It can be isolated from sources in nature and not intentionally modified.
[0361] As used herein, the description of a mutant / engineered polypeptide (e.g., of WT OgeuIscB) “comprising an amino acid mutation (e.g., substitution) at a position corresponding to a given position (e.g., D61) of a given polypeptide (e.g., WT OgeuIscB of SEQ ID NO: 1) ” or similar description means that the given polypeptide serves as a parent or reference polypeptide that does not comprises an amino acid mutation at the given position, and the mutant is a mutant of the parent or reference polypeptide and comprises an amino acid mutation at a position of the amino acid sequence of the mutant corresponding to the given position of the amino acid sequence of the given polypeptide. The position of the amino acid mutation in the amino acid sequence of the mutant may be the same as the given position of the given polypeptide, for example, when the mutant has exactly the same length as the given polypeptide. The position of the amino acid mutation in the amino acid sequence of the mutant may be different from the given position of the given polypeptide, for example, when the mutant does not have exactly the same length as the given polypeptide, for example, when the mutant comprises a N-terminal truncation as compared with the given polypeptide and thus the first N-terminal amino acid of the mutant is not corresponding to the first N-terminal amino acid of the given polypeptide but to an internal amino acid within the given polypeptide, but the position of the amino acid mutation in the mutant can be determined by alignment of the mutant and the given polypeptide to identify the corresponding amino acids in the two sequences as understood by a skilled in the art. For example, if the mutant has a N-terminal truncation of 20 amino acids as compared with the given polypeptide, then the mutant comprising an amino acid mutation at a position corresponding to D61 of a given polypeptide means that the mutant comprises an amino acid mutation at position 41 of the mutant since position 41 in the mutant is corresponding to D61 in the given polypeptide as determined by alignment of the mutant and the given polypeptide.
[0362] As used herein, the description of a mutant / engineered polypeptide (e.g., of WT OgeuIscB) “comprising an amino acid substitution corresponding a given amino acid substitution (e.g., D61A) relative to a given polypeptide (e.g., WT OgeuIscB of SEQ ID NO: 1) ” means that the given polypeptide serves as a parent or reference polypeptide that does not comprise the given amino acid substitution, and the mutant is a mutant of the parent or reference polypeptide and comprises the same type of amino acid substitution (e.g., D-to-Asubstitution) as the given amino acid substitution at a position in the mutant corresponding to the position (e.g., D61) of the given amino acid substitution (e.g., D61A) numbered according to the given polypeptide. For example, an engineered IscB polypeptide comprising an amino acid substitution corresponding to D61A relative to SEQ ID NO: 1 (wild type OgeuIscB) refers to the fact that the parent or reference polypeptide of SEQ ID NO: 1 comprises amino acid D (Asp) at position 61, and the engineered IscB polypeptide comprises amino acid A (Ala) at a position corresponding to D61 of SEQ ID NO: 1. The corresponding relationship of positions in the two amino acid sequences as determined by alignment is explained in the previous paragraph.
[0363] As used herein, the terms “upstream” and “downstream” refer to relative positions within a single nucleic acid (e.g., DNA) sequence in a nucleic acid. “Upstream” and “downstream” relate to the 5’ to 3’ direction, respectively, in which transcription occurs. For a first sequence and a second sequence present on the same strand of a single nucleic acid written in 5’ to 3’ direction, the first sequence is upstream of the second sequence when the 3’ end of the first sequence is on the left side of the 5’ end of the second sequence, and the first sequence is downstream of the second sequence when the 5’ end of the first sequence is on the right side of the 3’ end of the second sequence. For example, a promoter is usually at the upstream of a sequence under the regulation of the promoter; and on the other hand, a sequence under the regulation of a promoter is usually at the downstream of the promoter.
[0364] As used herein, the term “regulatory element” refers to a DNA sequence that controls or impacts one or more aspects of transcription and / or expression and is intended to include promoters, enhancers, silencers, termination signals, internal ribosome entry sites (IRES) , and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences) . Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences) . Regulatory elements may also direct expression in a time-dependent manner, e.g., in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue or cell type specific.
[0365] As used herein, the term “operably linked” refers to a juxtaposition wherein the components described are in a relationship permitting them to function in their intended manner. A regulatory element “operably linked” to a functional element is associated in such a way that transcription, expression, and / or activity of the functional element is achieved under conditions compatible with the regulatory element. In some embodiments, “operably linked” regulatory elements are contiguous (e.g., covalently linked) with the functional elements of interest; in some embodiments, regulatory elements act in trans to or otherwise at a distance from the functional elements of interest.
[0366] As used herein, the term “cell” is understood to refer not only to a particular individual cell, but to the progeny or potential progeny of the cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term.
[0367] As used herein, the term “in vivo” means inside the body of an organism, and the terms “ex vivo” or “in vitro” means outside the body of an organism.
[0368] As used herein, the term “treat” , “treatment” , or “treating” is an approach for obtaining beneficial or desired results including clinical results. For purposes of the disclosure, the beneficial or desired clinical results include, but are not limited to, one or more of the following: alleviating one or more symptoms resulting from a disease, diminishing the extent of a disease, stabilizing a disease (e.g., delaying the worsening of a disease) , delaying the spread (e.g., metastasis) of a disease, delaying the recurrence of a disease, reducing recurrence rate of a disease, delay or slowing the progression of a disease, ameliorating a disease state, providing a remission (partial or total) of a disease, decreasing the dose of one or more other medications required to treat a disease, delaying the progression of a disease, increasing the quality of life, and prolonging survival. Also encompassed by the term is a reduction of pathological consequence of a disease (such as cancer) . The methods of the disclosure contemplate any one or more of these aspects of treatment.
[0369] As used herein, the term “disease” includes the terms “disorder” and “condition” and is not limited to those specific diseases that have been medically or clinically defined.
[0370] As used herein, reference to “not” a value or parameter generally means and describes “other than” a value or parameter. For example, the method is not used to treat cancer of type X means the method may be used to treat cancer of types other than X.
[0371] As used herein, the singular forms “a” , “an” , and “the” include plural referents unless the context clearly dictates otherwise. That is, articles “a / an” and “the” are used herein to refer to one or more than one (i.e., at least one) grammatical object of the article. For example, “an element” means one element or more than one element, e.g., two elements.
[0372] As used herein, the term “and / or” in a phrase such as “A and / or B” is intended to mean either or both of the alternatives, including both A and B, A or B, A (alone) , and B (alone) . Likewise, the term “and / or” in a phrase such as “A, B, and / or C” is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone) ; B (alone) ; and C (alone) .
[0373] As used herein, when the term “about” is ahead of a serious of numbers (for example, about 1, 2, 3) , it is understood that each of the serious of numbers is modified by the term “about” (that is, about 1, about 2, about 3) . The term “about X-Y” used herein has the same meaning as “about X to about Y. ”
[0374] As used herein, a numerical range includes the end values of the range, and each specific value within the range, for example, “16 to 100 nucleotides” includes 16 nucleotides and 100 nucleotides, and each specific value between 16 and 100, e.g., 17, 23, 34, 52, 78.
[0375] It is understood that embodiments of the disclosure described herein include “consisting” and / or “consisting essentially of” embodiments.
[0376] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely” , “only” , and the like in connection with the recitation of claim elements, or use of a “negative” limitation.
[0377] As used herein, the term “reference IscB polypeptide” is used in the context of designing and developing a new IscB polypeptide based on an original IscB polypeptide (e.g., a wild-type IscB polypeptide) , for example, the original IscB polypeptide is mutated to generate the new IscB polypeptide. In that case, the original IscB polypeptide is a reference of the new IscB polypeptide and termed as a reference IscB polypeptide. The properties of a new IscB polypeptide can be evaluated with a reference IscB polypeptide as a reference from which the new IscB polypeptide is derived. For example, one or more of the properties (e.g., endonuclease activity, nickase activity) of the new IscB polypeptide can be compared with the reference IscB polypeptide from which the new IscB polypeptide is derived. As used herein, the term “engineered IscB polypeptide” refers to an IscB polypeptide artificially designed and developed based on a reference IscB polypeptide (e.g., a wild-type IscB polypeptide) , for example, by introducing an amino acid mutation.
[0378] As used herein, the term “endonuclease activity” is used interchangeably with “dsDNA cleavage activity” herein, and the term “nickase activity” is used interchangeably with “ssDNA cleavage activity” herein. As used herein, the term “nick” is used interchangeably with “ssDNA cleavage” . Unless otherwise indicated, the term “endonuclease activity” in reference to IscB refers to guide sequence specific (on-target) endonuclease activity. Unless otherwise indicated, the term “nickase activity” in reference to IscB refers to guide sequence specific (on-target) nickase activity.
[0379] Although in the disclosure sometime reference is made to reduced nickase activity of a (e.g., engineered) IscB polypeptide compared to the nickase activity of a reference IscB polypeptide, it does not mean to acknowledge that the reference IscB polypeptide is a nickase. A reference IscB polypeptide that has endonuclease activity may also show positive results in a nickase activity evaluation assay due to its capability of cleaving one strand of a dsDNA, which however does not make it a nickase.
[0380] Overview
[0381] IscB proteins are thought to be the ancestor of Cas9 and contain HNH and RuvC domains like Cas9 (17, 18. A recent report has shown that engineered OgeuIscB based base editors (enOgeuIscB-BE) exhibit high base editing efficiency in mammalian cells19. It would be desired to develop higher-efficiency miniature base editors with broader TAM range. Here, the inventor identified 19 natural IscB-ωRNA systems with various TAM scopes from metagenome datasets. By engineering both the ωRNA scaffold and protein of wild type IscB. m16 system, the inventor generated the IscB. m16*system (IscB. m16 containing E326R / T459E / P460S / T462H substitutions (IscB. m16RESH) and enωRNA) with the robust editing activity and expanded TAM range to NNNGNA in mammalian cells. We further developed IscB. m16*-based adenine and cytosine base editors, demonstrating robust base editing efficiency and broad target recognition in mammalian cells and mouse models. Moreover, the inventor provide a comprehensive dataset of IscB-ωRNA systems with diverse TAM scopes and the strategy to widen TAM range. In addition, the inventor design optimal configurations of ABE and CBE based on IscB with high base editing efficiency.
[0382] The disclosure provides, in part, engineered IscB polypeptides that exhibit improved specificity for targeting a DNA, e.g., relative to a wild type IscB polypeptide. Improved specificity can be, e.g., (i) increased on-target binding, cleavage, and / or editing of DNA and / or (ii) decreased off-target binding, cleavage, and / or editing of DNA, e.g., relative to a wild type IscB polypeptide and / or to another engineered IscB polypeptide. The disclosure also provides, in part, optimal configurations of IscB-based base editors.
[0383] Thus, one of the objects of the disclosure is to provide engineered IscB polypeptide having enhanced endonuclease activity. Another one of the objects of the disclosure is to provide engineered IscB polypeptide substantially lacking endonuclease activity, e.g., IscB nickase, dead IscB, for use in, e.g., base editing. Another one of the objects of the disclosure is to provide optimal IscB-based base editors.
[0384] Representative IscB polypeptide
[0385] In an aspect, the disclosure provides an IscB polypeptide comprising an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of any one of SEQ ID NOs: 1-19.
[0386] In another aspect, the disclosure provides an IscB polypeptide, wherein the IscB polypeptide comprises an amino acid mutation relative to (compared to) a wild type IscB polypeptide.
[0387] In yet another aspect, the disclosure provides a method of increasing guide sequence-specific binding ability (e.g., represented by the guide sequence-specific endonuclease activity of the IscB system or guide sequence-specific base editing efficiency of the IscB system) of an IscB polypeptide for use in an IscB system comprising (1) the IscB polypeptide, or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide, and (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0388] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide; and
[0389] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA;
[0390] wherein the scaffold sequence is 3’ to the guide sequence,
[0391] said method comprising introducing an amino acid mutation into the IscB polypeptide.
[0392] In yet another aspect, the disclosure provides a method of widening target adjacent motif (TAM) recognition (e.g., represented by the increased guide sequence-specific endonuclease activity of the IscB system or increased guide sequence-specific base editing efficiency of the IscB system for a broader TAM, e.g., a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’, wherein N is A, T, G, or C) of an IscB polypeptide for use in an IscB system comprising (1) the IscB polypeptide, or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide, and (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0393] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide; and
[0394] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA;
[0395] wherein the scaffold sequence is 3’ to the guide sequence,
[0396] said method comprising introducing an amino acid mutation into the IscB polypeptide.
[0397] In some embodiments, the IscB polypeptide comprises an amino acid mutation relative to (compared to) a reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19.
[0398] In some embodiments, the IscB polypeptide comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%) and less than 100%to the amino acid sequence of any one of SEQ ID NOs: 1-19.
[0399] In some embodiments, the amino acid mutation is within a domain of the reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19 selected from the group consisting of PLMP domain, RuvC-I domain, Bridge Helix domain, Linker domain, RuvC-II domain, HNH domain, RuvC-III domain, P1D domain, and TID domain.
[0400] In some embodiments, the PLMP domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 1-54 of SEQ ID NO: 16.
[0401] In some embodiments, the RuvC-I domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 55-85 of SEQ ID NO: 16.
[0402] In some embodiments, the Bridge Helix domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 86-122 of SEQ ID NO: 16.
[0403] In some embodiments, the Linker domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 123-160 of SEQ ID NO: 16.
[0404] In some embodiments, the RuvC-II domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 161-196 of SEQ ID NO: 16.
[0405] In some embodiments, the HNH domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 197-297 of SEQ ID NO: 16.
[0406] In some embodiments, the RuvC-III domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 298-374 of SEQ ID NO: 16.
[0407] In some embodiments, the P1D domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 375-429 of SEQ ID NO: 16.
[0408] In some embodiments, the TID domain comprises, consists essentially of, or consists of amino acid residues of the reference or wild type IscB polypeptide at positions corresponding to positions 430-513 of SEQ ID NO: 16.
[0409] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from the group consisting of M1, A2, N3, V4, I5, Y6, V7, I8, N9, K10, D11, G12, K13, P14, L15, M16, P17, T18, T19, R20, R21, G22, H23, V24, G25, Y26, L27, L28, R29, K30, K31, Q32, A33, R34, V35, V36, K37, H38, N39, P40, F41, T42, V43, Q44, L45, S46, Y47, E48, T49, P50, D51, K52, V53, Q54, E55, L56, T57, L58, G59, I60, D61, P62, G63, R64, T65, N66, I67, G68, I69, A70, V71, V72, D73, E74, T75, G76, E77, C78, V79, F80, S81, A82, H83, V84, E85, T86, R87, N88, K89, D90, V91, P92, K93, L94, M95, A96, K97, R98, K99, V100, H101, R102, Q103, A104, R105, R106, H107, Y108, G109, R110, R111, V112, K113, R114, Q115, R116, R117, A118, K119, A120, N121, G122, T123, V124, N125, E126, N127, G128, I129, I130, T131, R132, V133, L134, P135, Q136, T137, E138, T139, P140, I141, E142, C143, K144, L145, I146, K147, N148, K149, E150, A151, R152, F153, C154, N155, R156, E157, R158, E159, P160, G161, W162, L163, T164, P165, T166, A167, N168, Q169, L170, L171, L172, T173, H174, L175, N176, L177, V178, T179, K180, I181, E182, Q183, I184, L185, P186, I187, S188, K189, I190, A191, L192, E193, I194, N195, K196, F197, A198, F199, M200, E201, L202, D203, D204, H205, N206, I207, R208, P209, W210, E211, Y212, Q213, H214, G215, P216, L217, F218, G219, F220, E221, S222, R223, D224, D225, A226, V227, Y228, A229, L230, Q231, E232, G233, K234, C235, L236, L237, C238, G239, K240, P241, L242, I243, E244, H245, Y246, H247, H248, V249, I250, P251, K252, H253, E254, H255, G256, S257, D258, T259, I260, A261, N262, I263, V264, G265, L266, C267, S268, G269, C270, H271, D272, L273, V274, H275, R276, D277, A278, R279, A280, K281, D282, K283, L284, A285, K286, V287, H288, A289, G290, A291, K292, K293, K294, Y295, A296, G297, T298, S299, V300, L301, N302, Q303, I304, M305, P306, K307, L308, I309, T310, R311, L312, A313, S314, K315, D316, E317, D318, F319, T320, L321, V322, S323, A324, K325, E326, I327, S328, F329, I330, R331, R332, S333, S334, A335, L336, P337, K338, D339, H340, H341, I342, D343, A344, Y345, C346, I347, A348, M349, S350, V351, V352, D353, G354, E355, S356, H357, M358, N359, G360, M361, L362, R363, K364, P365, Y366, Q367, V368, M369, Q370, F371, R372, R373, H374, D375, R376, Q377, A378, R379, H380, Q381, A382, M383, V384, D385, R386, K387, Y388, Y389, L390, N391, G392, K393, H394, V395, A396, T397, N398, R399, H400, K401, R402, F403, E404, Q405, K406, G407, D408, S409, L410, E411, E412, F413, L414, Q415, K416, N417, P418, G419, V420, R421, P422, E423, M424, L425, A426, V427, R428, E429, H430, K431, P432, V433, Y434, K435, R436, M437, N438, R439, I440, A441, P442, G443, T444, L445, M446, R447, C448, K449, N450, D451, V452, F453, V454, Y455, K456, T457, G458, T459, P460, I461, T462, N463, G464, T465, P466, F467, Y468, A469, V470, D471, T472, E473, G474, Q475, R476, H477, K478, Y479, R480, L481, S482, K483, P484, V485, L486, H487, N488, T489, G490, I491, V492, V493, L494, G495, N496, R497, G498, G499, Y500, P501, P502, Q503, I504, L505, G506, T507, H508, D509, K510, K511, R512, and E513 of SEQ ID NO: 16.
[0410] In some embodiments, the amino acid mutation leads to an increased guide sequence-specific endonuclease activity, or wherein the IscB polypeptide comprising said amino acid mutation has an increased guide sequence-specific endonuclease activity compared to the reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0411] In some embodiments, the IscB polypeptide comprising said amino acid mutation leads to increased guide sequence-specific base editing efficiency compared to the reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0412] In some embodiments, the IscB polypeptide has decreased guide sequence-independent (off-target) endonuclease activity or substantially lacks guide sequence-independent (off-target) endonuclease activity.
[0413] In some embodiments, the amino acid mutation leads to a decreased guide sequence-independent (off-target) endonuclease activity, or wherein the IscB polypeptide comprising said amino acid mutation has a decreased guide sequence-independent (off-target) endonuclease activity compared to the reference or wild type IscB polypeptide of any one of SEQ ID NOs: 1-19, e.g., a decrease by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%.
[0414] In some embodiments, the IscB polypeptide is capable of recognizing a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’ immediately 3’ adjacent to a protospacer sequence of a target DNA, wherein N is A, T, G, or C.
[0415] In some embodiments, the IscB polypeptide has an increased guide-sequence specific endonuclease activity compared to that of SEQ ID NO: 16 for a protospacer sequence of a target DNA immediately 5’ adjacent to a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’, wherein N is A, T, G, or C.
[0416] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from the group consisting of K30, K37, H38, N39, P50, V53, E74, E77, V79, H83, E85, K93, A96, K99, Q103, A104, H107, V133, E142, E159, P160, L172, T179, K180, E182, K196, E201, D204, H205, H214, L217, F218, E221, S222, D224, D225, Y228, A229, E232, G233, K234, G239, I250, H253, E254, A261, S268, G269, D272, L273, A278, A280, D282, K283, A285, V287, K292, K307, T310, A313, D318, E326, S328, F329, I330, S333, A335, P337, S350, V351, G354, S356, H357, G360, Q367, M369, H380, Q381, V384, K387, N391, G392, K393, H394, H400, K401, Q405, K406, G407, E411, L414, Q415, K416, N417, P418, G419, E423, M424, A426, E429, H430, K431, V433, K435, N438, L445, K449, N450, D451, V452, K456, T459, P460, I461, T462, N463, T465, F467, Y468, E473, G474, Q475, R476, H477, K478, L481, K483, P484, L486, H487, L494, G495, N496, G498, G499, Y500, P501, P502, Q503, I504, L505, G506, T507, H508, D509, K510, K511, and E513 of SEQ ID NO: 16.
[0417] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from the group consisting of E326, H380, Q381, M424, V433, T459, P460, I461, T462, N463, T465, F467, Y468, Q475, R476, K478, L481, and I504 of SEQ ID NO: 16.
[0418] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position that is corresponding to a position or that is a position selected from or that is a position the group consisting of D61, E193, and H248 of SEQ ID NO: 16.
[0419] In some embodiments, the amino acid substitution is a conservative amino acid substitution or a non-conservative amino acid substitution.
[0420] In some embodiments, the amino acid substitution is an amino acid substitution with an amino acid residue that is different from the amino acid residue at the position of SEQ ID NO: 16.
[0421] In some embodiments, the amino acid substitution is an amino acid substitution with
[0422] (1) a non-polar amino acid residue (such as, Glycine (Gly / G) , Alanine (Ala / A) , Valine (Val / V) , Cysteine (Cys / C) , Proline (Pro / P) , Leucine (Leu / L) , Isoleucine (Ile / I) , Methionine (Met / M) , Tryptophan (Trp / W) , Phenylalanine (Phe / F) ,
[0423] (2) a polar amino acid residue (such as, Serine (Ser / S) , Threonine (Thr / T) , Tyrosine (Tyr / Y) , Asparagine (Asn / N) , Glutamine (Gln / Q) ) ,
[0424] (3) a positively charged amino acid residue (such as, Lysine (Lys / K) , Arginine (Arg / R) , Histidine (His / H) ) , or
[0425] (4) a negatively charged amino acid residue (such as, Aspartic Acid (Asp / D) , Glutamic Acid (Glue / E) ) .
[0426] In some embodiments, the amino acid substitution is an amino acid substitution with a positively charged amino acid residue, such as, Arginine (R) .
[0427] In some embodiments, the amino acid substitution is an amino acid substitution with a non-polar amino acid residue, such as, Alanine (A) .
[0428] In some embodiments, the amino acid mutation comprises a substitution that is corresponding to a substitution or that is a substitution selected from the group consisting of K30R, K37R, H38R, N39R, P50R, V53R, E74R, E77R, V79R, H83R, E85R, K93R, A96R, K99R, Q103R, A104R, H107R, V133R, E142R, E159R, P160R, L172R, T179R, K180R, E182R, K196R, E201R, D204R, H205R, H214R, L217R, F218R, E221R, S222R, D224R, D225R, Y228R, A229R, E232R, G233R, K234R, G239R, I250R, H253R, E254R, A261R, S268R, G269R, D272R, L273R, A278R, A280R, D282R, K283R, A285R, V287R, K292R, K307R, T310R, A313R, D318R, E326R, S328R, F329R, I330R, S333R, A335R, P337R, S350R, V351R, G354R, S356R, H357R, G360R, Q367R, M369R, Q381R, V384R, K387R, N391R, G392R, K393R, H394R, H400R, K401R, Q405R, K406R, G407R, E411R, L414R, Q415R, K416R, N417R, P418R, G419R, E423R, M424R, A426R, E429R, H430R, K431R, V433R, K435R, N438R, L445R, K449R, N450R, D451R, V452R, K456R, T459R, T462R, N463R, T465R, E473R, G474R, Q475R, H477R, K478R, K483R, P484R, L486R, H487R, L494R, G495R, N496R, G498R, G499R, Y500R, P501R, P502R, Q503R, I504R, L505R, G506R, T507R, H508R, D509R, K510R, K511R, E513R, and a combination of any two or more residues thereof, wherein the position is numbered according to SEQ ID NO: 16.
[0429] In some embodiments, the amino acid mutation comprises a substitution that is corresponding to a substitution or that is a substitution selected from the group consisting of T459E, P460S, T462H, T462L, T465V, and a combination of any two or more residues thereof, wherein the position is numbered according to SEQ ID NO: 16.
[0430] In some embodiments, the amino acid mutation comprises a substitution that is corresponding to a substitution or that is a substitution selected from the group consisting of E326R, T459E, P460S, T462H, and a combination of any two or more residues thereof, wherein the position is numbered according to SEQ ID NO: 16.
[0431] In some embodiments, the IscB polypeptide comprising said amino acid mutation comprises a substitution that is corresponding to a substitution or that is a substitution selected from the group consisting of D61A, E193A, and H248A, wherein the position is numbered according to SEQ ID NO: 16.
[0432] In some embodiments, the amino acid mutation comprises a combination substitution corresponding to a combination substitution of E326R, P460S, T462H, and T459E, wherein the position is numbered according to SEQ ID NO: 16.
[0433] In some embodiments, the amino acid mutation comprises a combination substitution corresponding to a combination substitution of D61A, E326R, P460S, T462H, and T459E, wherein the position is numbered according to SEQ ID NO: 16.
[0434] In some embodiments, the amino acid mutation comprises a combination substitution corresponding to a combination substitution of D61A, H248A, E326R, P460S, T462H, and T459E, wherein the position is numbered according to SEQ ID NO: 16.
[0435] In some embodiments, the IscB polypeptide comprising said amino acid mutation comprises, consists essentially of, or consists an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the amino acid sequence of SEQ ID NO: 16.
[0436] In some embodiments, the IscB polypeptide comprising said amino acid mutation comprises, consists essentially of, or consists an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the amino acid sequence of SEQ ID NO: 239 or an N-terminal truncation of the amino acid sequence of SEQ ID NO: 239 lacking the most N-terminal Methionine (M) (coded by start codon ATG) .
[0437] In some embodiments, the IscB polypeptide comprising said amino acid mutation comprises, consists essentially of, or consists an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100%to the amino acid sequence of SEQ ID NO: 240 or 241 or an N-terminal truncation of the amino acid sequence of SEQ ID NO: 240 or 241 lacking the most N-terminal Methionine (M) (coded by start codon ATG) .
[0438] In some embodiments, the IscB polypeptide is an endonuclease or has guide sequence-specific endonuclease activity.
[0439] In some embodiments, the IscB polypeptide is a nickase or has guide sequence-specific nickase activity.
[0440] In some embodiments, the IscB polypeptide is endonuclease deficient.
[0441] In some embodiments, the IscB polypeptide is catalytically inactive.
[0442] In some embodiments, the IscB polypeptide is fused to a functional domain to form a fusion protein.
[0443] In some embodiments, the functional domain has transposase activity, methylase activity, demethylase activity, translation activation activity, translation repression activity, transcription activation activity, transcription repression activity, transcription release factor activity, chromatin modifying or remodeling activity, histone modification activity, nuclease activity, single-strand RNA cleavage activity, double-strand RNA cleavage activity, single-strand DNA cleavage activity, double-strand DNA cleavage activity, nucleic acid binding activity, detectable activity, or any combination thereof.
[0444] In some embodiments, the functional domain is fused N-terminally, C-terminally, or internally with respect to the IscB polypeptide.
[0445] In some embodiments, the functional domain is fused to the IscB polypeptide via a linker, e.g., a XTEN linker, a GS linker containing multiple glycine and serine residues, a GS linker containing multiple glycine and serine residues and a XTEN linker, a GS linker containing multiple glycine and serine residues and a BP NLS.
[0446] In some embodiments, the functional domain is selected from the group consisting of a nuclear localization signal (NLS) , a nuclear export signal (NES) , a deaminase or a catalytic domain thereof, an uracil glycosylase inhibitor (UGI) , an uracil glycosylase (UNG) , a methylpurine glycosylase (MPG) , a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease or a catalytic domain thereof (e.g., T5 exonuclease of SEQ ID NO: 404) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having ssRNA cleavage activity, a moiety having dsRNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA, selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, and a catalytic domain thereof, and a functional fragment thereof, and any combination thereof.
[0447] In some embodiments, the fusion protein comprises a NLS at the N-terminal and / or the C-terminal of the IscB polypeptide.
[0448] In some embodiments, the fusion protein comprises one or two NLS at the N-terminal and / or the C-terminal of the IscB polypeptide.
[0449] In some embodiments, the fusion protein comprises a NLS at the N-terminal and / or the C-terminal of the functional domain.
[0450] In some embodiments, the fusion protein comprises one or two NLS at the N-terminal and / or the C-terminal of the functional domain.
[0451] In some embodiments, the NLS comprises or is SV40 NLS (SEQ ID NO: 258) , bpSV40 NLS (BP NLS, bpNLS, SEQ ID NO: 256) , or NP NLS (Xenopus laevis Nucleoplasmin NLS, nucleoplasmin NLS, SEQ ID NO: 257) .
[0452] In some embodiments, the functional domain comprises a deaminase or a catalytic domain thereof.
[0453] In some embodiments, the deaminase or catalytic domain thereof is an adenine deaminase (e.g., TadA, such as, TadA8e, TadA8.17, TadA8.20, TadA9) or a catalytic domain thereof.
[0454] In some embodiments, the adenine deaminase or a catalytic domain thereof comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 253.
[0455] In some embodiments, the adenine deaminase or a catalytic domain thereof is TadA8EV106W (SEQ ID NO: 253) .
[0456] In some embodiments, the deaminase or catalytic domain thereof is a cytidine deaminase (e.g., APOBEC, such as, APOBEC3, for example, APOBEC3A, APOBEC3B, APOBEC3C; DddA) or a catalytic domain thereof.
[0457] In some embodiments, the cytidine deaminase or a catalytic domain thereof comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 254.
[0458] In some embodiments, the cytidine deaminase or a catalytic domain thereof is APOBEC3AW104A (SEQ ID NO: 254) .
[0459] In some embodiments, the functional domain comprises an uracil glycosylase inhibitor (UGI) domain.
[0460] In some embodiments, the UGI domain comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 255.
[0461] In some embodiments, the fusion protein comprises one, two, or three UGI domains, and optionally, one UGI domain.
[0462] In some embodiments, the functional domain comprises an uracil glycosylase (UNG) .
[0463] In some embodiments, the functional domain comprises a methylpurine glycosylase (MPG) .
[0464] In some embodiments, the functional domain comprises a reverse transcriptase or a catalytic domain thereof.
[0465] In some embodiments, the functional domain comprises a methylase or a catalytic domain thereof.
[0466] In some embodiments, the functional domain comprises a transcription activating domain.
[0467] In some embodiments, the fusion protein comprises, from N-to C-terminus, an adenine deaminase domain, an optional linker, the IscB polypeptide, an optional linker, and an adenine deaminase domain.
[0468] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 259.
[0469] In some embodiments, the fusion protein comprises, from N-to C-terminus, a cytidine deaminase domain, an optional linker, the IscB polypeptide, an optional linker, and a UGI domain (e.g., one UGI domain) .
[0470] In some embodiments, the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 260.
[0471] In yet another aspect, the disclosure provides a system comprising:
[0472] (1) the IscB polypeptide or method of the disclosure, or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide, and
[0473] (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0474] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide; and
[0475] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA;
[0476] wherein the scaffold sequence is 3’ to the guide sequence.
[0477] In some embodiments, the guide nucleic acid is a guide RNA (gRNA) (interchangeably used with omega RNA (ωRNA) ) .
[0478] In some embodiments, the guide nucleic acid is capable of directing guide sequence specific binding of the complex to the target sequence of the target DNA.
[0479] In some embodiments, the scaffold sequence comprises a nucleotide mutation relative to a reference or wild type scaffold sequence compatible to the IscB polypeptide.
[0480] In some embodiments, the scaffold sequence comprises a nucleotide mutation relative to a reference or wild type scaffold sequence of any one of SEQ ID NOs: 20-38.
[0481] In some embodiments, the nucleotide mutation comprises a deletion of about, at least about, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in a stem-loop region of the reference scaffold sequence.
[0482] In some embodiments, the nucleotide mutation comprises a substitution of about, at least about, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more thermodynamically unstable base pairs in a stem-loop region of the reference scaffold sequence with a G-C or C-G base pair.
[0483] In some embodiments, the thermodynamically unstable base pair is a A-U or U-Abase pair, a A-G or G-Abase pair, or a U-G or G-U base pair.
[0484] In some embodiments, the stem-loop region is selected from the first 5’ stem loop region, the second 5’ stem loop region, the third 5’ stem loop region, the fourth 5’ stem loop region, the fifth 5’ stem loop region, or the sixth 5’ stem loop region of the reference scaffold sequence, wherein the first, the second, the third, the fourth, the fifth, and the sixth 5’ stem loop region are counted from the 5’ end of the reference scaffold sequence.
[0485] In some embodiments, the stem-loop region is selected from the first 5’ stem loop region and the first 3’ stem loop region, wherein the first 5’ stem loop region is counted from the 5’ end of the reference scaffold sequence, and wherein the first 3’ stem loop region is counted from the 3’ end of the reference scaffold sequence.
[0486] In some embodiments, the nucleotide mutation leads to an increased guide sequence-specific endonuclease activity, or wherein the system comprising the guide nucleic acid comprising said nucleotide mutation has an increased guide sequence-specific endonuclease activity compared to an otherwise identical control system comprising a guide nucleic acid without said nucleotide mutation, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0487] In some embodiments, the nucleotide mutation leads to increased guide sequence-specific base editing efficiency, or wherein the system comprising the guide nucleic acid comprising said nucleotide mutation has increased guide sequence-specific base editing efficiency compared to an otherwise identical control system comprising a guide nucleic acid without said nucleotide mutation, e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, or more.
[0488] In some embodiments, the system is capable of recognizing a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’ immediately 3’ adjacent to a protospacer sequence of a target DNA, wherein N is A, T, G, or C.
[0489] In some embodiments, the system comprising the guide nucleic acid comprising said nucleotide mutation has an increased guide-sequence specific endonuclease activity compared to that of an otherwise identical control system comprising a guide nucleic acid without said nucleotide mutation for a protospacer sequence of a target DNA immediately 5’ adjacent to a target adjacent motif (TAM) comprising, consisting essentially of, or consisting of 5’-NNNGNA-3’, wherein N is A, T, G, or C.
[0490] In some embodiments, the scaffold sequence has substantially the same secondary structure as the secondary structure of any one of SEQ ID NOs: 20-38, 58-238, and 242-252.
[0491] In some embodiments, the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 20-38, 58-238, and 242-252.
[0492] In some embodiments, the target sequence comprises about or at least about 14 contiguous nucleotides of the target DNA, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more contiguous nucleotides of the target DNA, or in a numerical range between any two of the preceding values, e.g., from about 14 to about 20 contiguous nucleotides of the target DNA, from about 14 to about 50 contiguous nucleotides of the target DNA; optionally, wherein the target sequence comprises about 14 contiguous nucleotides of the target DNA.
[0493] In some embodiments, the target sequence is immediately 3’ to a target adjacent motif (TAM) , or wherein the reversely complementary sequence of the target sequence (i.e., the protospacer sequence) is immediately 5’ to a target adjacent motif (TAM) .
[0494] In some embodiments, the TAM is 5’-NNNGNA-3’, wherein N is A, T, G, or C; and optionally, wherein the TAM is 5’-NNNGAN-3’, wherein N is A, T, G, or C.
[0495] In some embodiments, the guide sequence is about or at least about 14 nucleotides in length, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more nucleotides in length, or in a length of a numerical range between any two of the preceding values, e.g., in a length of from about 14 to about 20 nucleotides, in a length of from about 14 to about 50 nucleotides; optionally, wherein the guide sequence is about 14 nucleotides in length.
[0496] In some embodiments, (1) the guide sequence is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% (fully) , optionally about 100% (fully) , reversely complementary to the target sequence; (2) the guide sequence contains no more than 5, 4, 3, 2, or 1 mismatch or contains no mismatch with the target sequence; or (3) the guide sequence comprises no mismatch with the target sequence in the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nucleotides at the 5’ end of the guide sequence.
[0497] In some embodiments, the guide sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 261-403 or comprises a polynucleotide sequence having no more than 1, 2, 3, 4, 5, 6, 7, or 8 nucleotide difference from any one of SEQ ID NOs: 261-403.
[0498] In some embodiments, the system comprises two or more guide nuclei acids comprising two or more guide sequences capable of hybridizing to two or more target sequences of the same target DNA or different target DNAs, wherein the two or more guide sequences are the same or different, and wherein the two or more target sequences are the same or different.
[0499] In some embodiments, the target DNA is a target dsDNA, such as, a eukaryotic dsDNA, e.g., a gene in a eukaryotic cell.
[0500] In some embodiments, the target DNA is a target dsDNA, and wherein the target dsDNA comprises a protospacer sequence on a nontarget strand of the target dsDNA, wherein the dsDNA comprises a target deoxyribonucleotide (e.g., dA, dT, dC, dG) at a position of the protospacer sequence selected from the group consisting of position 1, position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9, position 10, and a combination thereof; or wherein the target deoxyribonucleotide is at a position of the protospacer sequence between position 1 and position 1 or between position 2 and position 5, both inclusive.
[0501] In yet another aspect, the disclosure provides a guide nucleic acid as defined in the disclosure.
[0502] In yet another aspect, the disclosure provides a polynucleotide encoding the IscB polypeptide or method of the disclosure (e.g., SEQ ID NOs: 39-57) .
[0503] In yet another aspect, the disclosure provides a polynucleotide encoding the guide nucleic acid of the disclosure.
[0504] In yet another aspect, the disclosure provides a polynucleotide encoding the IscB polypeptide or method of the disclosure (e.g., SEQ ID NOs: 39-57) and the guide nucleic acid of the disclosure.
[0505] In yet another aspect, the disclosure provides a delivery system comprising (1) the IscB polypeptide or method of the disclosure, the polynucleotide of the disclosure, or the system of the disclosure; and (2) a delivery vehicle.
[0506] In yet another aspect, the disclosure provides a vector comprising the polynucleotide of the disclosure; optionally, wherein the vector encodes a guide nucleic acid as defined in the disclosure; optionally, wherein the vector is a plasmid vector, a recombinant AAV (rAAV) vector, or a recombinant lentivirus vector.
[0507] In yet another aspect, the disclosure provides a recombinant AAV (rAAV) particle comprising the rAAV vector of the disclosure; optionally, wherein the rAAV vector is an RNA.
[0508] In yet another aspect, the disclosure provides a ribonucleoprotein (RNP) comprising the IscB polypeptide or method of the disclosure and a guide nucleic acid optionally as defined in the disclosure.
[0509] In yet another aspect, the disclosure provides a lipid nanoparticle (LNP) comprising an RNA (e.g., mRNA) encoding the IscB polypeptide or method of the disclosure and a guide nucleic acid optionally as defined in the disclosure.
[0510] In yet another aspect, the disclosure provides a cell comprising the IscB polypeptide or method of the disclosure, the system of the disclosure, the polynucleotide of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the RNP of the disclosure, or the LNP of the disclosure.
[0511] In some embodiments, the cell is not a human germ cell (i.e., an embryonic cell, an egg cell, a sperm cell) .
[0512] In some embodiments, the cell is not a human embryonic stem cell.
[0513] In yet another aspect, the disclosure provides a method for modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, or the lipid nanoparticle of the disclosure, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex.
[0514] In some embodiments, the method is ex vivo, in vivo, or in vitro.
[0515] In some embodiments, the method is non-therapeutical.
[0516] In some embodiments, the target DNA is in a cell;
[0517] optionally, wherein the cell is a eukaryotic cell (e.g., an animal cell, a vertebrate cell, a mammalian cell, a non-human mammalian cell, a non-human primate cell, a rodent (e.g., mouse or rat) cell, a human cell, a plant cell, or a yeast cell) or a prokaryotic cell (e.g., a bacteria cell) ;
[0518] optionally, wherein the cell is from a plant or an animal;
[0519] optionally, wherein the plant is a dicotyledon; optionally selected from the group consisting of soybean, cabbage (e.g., Chinese cabbage) , rapeseed, brassica, watermelon, melon, potato, tomato, tobacco, eggplant, pepper, cucumber, cotton, alfalfa, eggplant, grape;
[0520] optionally, wherein the plant is a monocotyledon; optionally selected from the group consisting of rice, corn, wheat, barley, oat, sorghum, millet, grasses, Poaceae, Zizania, Avena, Coix, Hordeum, Oryza, Panicum (e.g., Panicum miliaceum) , Secale, Setaria (e.g., Setaria italica) , Sorghum, Triticum, Zea, Cymbopogon, Saccharum (e.g., Saccharum officinarum) , Phyllostachys, Dendrocalamus, Bambusa, Yushania; and / or
[0521] optionally, wherein the animal is selected from the group consisting of pig, ox, sheep, goat, mouse, rat, alpaca, monkey, rabbit, chicken, duck, goose, fish (e.g., zebra fish) .
[0522] In yet another aspect, the disclosure provides a cell modified by the method of the disclosure.
[0523] In some embodiments, the cell is not a human germ cell (i.e., an embryonic cell, an egg cell, a sperm cell) .
[0524] In some embodiments, the cell is not a human embryonic stem cell.
[0525] In yet another aspect, the disclosure provides a pharmaceutical composition comprising (1) the system of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, or the cell of the disclosure; and (2) a pharmaceutically acceptable excipient.
[0526] In yet another aspect, the disclosure provides a method of detecting a target DNA, comprising contacting the target DNA with In some embodiments, the target DNA is modified by the complex, and wherein the modification detects the target DNA; optionally, wherein the modification generates a detectable signal, e.g., a fluorescent signal.
[0527] In yet another aspect, the disclosure provides a method of increasing guide sequence-specific binding ability (e.g., represented by the guide sequence-specific endonuclease activity of the IscB system or guide sequence-specific base editing efficiency of the IscB system) of a guide nucleic acid for use in an IscB system comprising (1) an IscB polypeptide (e.g., the IscB polypeptide or method of the disclosure) , or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide, and (2) the guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0528] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide; and
[0529] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA;
[0530] wherein the scaffold sequence is 3’ to the guide sequence.
[0531] In some embodiments, the guide nucleic acid is a guide RNA (gRNA) (interchangeably used with omega RNA (ωRNA) ) .
[0532] In some embodiments, the guide nucleic acid is capable of directing guide sequence specific binding of the complex to the target sequence of the target DNA.
[0533] In some embodiments, the method comprises introducing a nucleotide mutation into the scaffold sequence relative to a reference or wild type scaffold sequence compatible to the IscB polypeptide.
[0534] In some embodiments, the method comprises introducing a nucleotide mutation into the scaffold sequence relative to a reference or wild type scaffold sequence of any one of SEQ ID NOs: 20-38.
[0535] In some embodiments, the nucleotide mutation comprises a deletion of about, at least about, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in a stem-loop region of the reference scaffold sequence.
[0536] In some embodiments, the nucleotide mutation comprises a substitution of about, at least about, or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more thermodynamically unstable base pairs in a stem-loop region of the reference scaffold sequence with a G-C or C-G base pair.
[0537] In some embodiments, the thermodynamically unstable base pair is a A-U or U-Abase pair, a A-G or G-Abase pair, or a U-G or G-U base pair.
[0538] In some embodiments, the stem-loop region is selected from the first 5’ stem loop region, the second 5’ stem loop region, the third 5’ stem loop region, the fourth 5’ stem loop region, the fifth 5’ stem loop region, or the sixth 5’ stem loop region of the reference scaffold sequence, wherein the first, the second, the third, the fourth, the fifth, and the sixth 5’ stem loop region are counted from the 5’ end of the reference scaffold sequence.
[0539] In some embodiments, the stem-loop region is selected from the first 5’ stem loop region and the first 3’ stem loop region, wherein the first 5’ stem loop region is counted from the 5’ end of the reference scaffold sequence, and wherein the first 3’ stem loop region is counted from the 3’ end of the reference scaffold sequence.
[0540] Fusion protein and functional domain
[0541] The IscB polypeptide of the disclosure may be combined / associated with one or more functional domains for additional function other than cleavage, e.g., deamination, base editing, prime editing, enabling various modifications of target DNA.
[0542] In some embodiments, the IscB polypeptide further comprises a functional domain fused to the IscB polypeptide with or without a linker to form a fusion protein. In some embodiments, the linker is a GS linker containing multiple glycine (GS) and serine (S) residues, XTEN linker (SEQ ID NO: 100) , XTEN&GS linker (SEQ ID NO: 99) containing XTEN linker (SEQ ID NO: 100) , or bpSV40 NLS&GS linker (SEQ ID NO: 111) containing bpSV40 NLS (SEQ ID NO: 110) .
[0543] In some embodiments, the functional domain is selected from the group consisting of a nuclear localization signal (NLS) , a nuclear export signal (NES) , a base editing domain, for example, a deaminase or a catalytic domain thereof, a base excising domain, an uracil glycosylase inhibitor (UGI) or a catalytic domain thereof, a glycosylase or a catalytic domain thereof, for example, an uracil glycosylase (UNG) or a catalytic domain thereof, a methylpurine glycosylase (MPG) or a catalytic domain thereof, a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease (e.g., T5 exonuclease (T5E) ) or a catalytic domain thereof, a non-LTR retrotransposon or a catalytic domain thereof, a destabilized domain (e.g., destabilized domains (DD) of E. coli dihydrofolate reductase (ecDHFR) ) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having RNA cleavage activity, a moiety having nickase activity, a moiety having endonuclease activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA, selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, and a catalytic domain thereof, and a functional fragment (e.g., a functional truncation) thereof, and any combination thereof.
[0544] In some embodiments, the NLS is SV40 NLS (such as, SEQ ID NO: 11) , bpSV40 NLS (BP NLS; bpNLS; such as, SEQ ID NO: 110) , or NP NLS (Xenopus laevis Nucleoplasmin NLS; nucleoplasmin NLS) (such as, SEQ ID NO: 12) .
[0545] In some embodiments, the exonuclease is T5 exonuclease (T5E) (SEQ ID NO: 404) .
[0546] In some embodiments, the deaminase or catalytic domain thereof is an adenine deaminase or a catalytic domain thereof (e.g., tRNA adenosine deaminase (TadA) , such as, TadA8e, TadA8.17, TadA8.20, TadA9, TadA8e-V106W (SEQ ID NO: 109) , TadA8EV106W+D108Q TadA-CDa, TadA-CDb, TadA-CDc, TadA-CDd, TadA-CDe, TadA-dual, TADAC-1.2, TADAC-1.14, TADAC-1.17, TADAC-1.19, TADAC-2.5, TADAC-2.6, TADAC-2.9, TADAC-2.19, TADAC-2.23, TadA8e-N46L, TadA8e-N46P) .
[0547] In some embodiments, the deaminase or catalytic domain thereof is a cytosine deaminase or a catalytic domain thereof (e.g., an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase, an activation induced deaminase (AID) , a cytidine deaminase 1 from Petromyzon marinus (pmCDA1) , or a functional variant thereof, e.g., APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, hAPOBEC3-W104A (SEQ ID NO: 113) ) .
[0548] In some embodiments, the UGI is human UGI domain (SEQ ID NO: 114) .
[0549] In some embodiments, the fusion protein comprises, from N-terminal to C-terminal, the engineered IscB polypeptide and an exonuclease, e.g., T5 exonuclease (T5E) (SEQ ID NO: 404) , with or without a linker between the engineered IscB polypeptide and the exonuclease.
[0550] In some embodiments, the engineered IscB polypeptide comprises the amino acid sequence of SEQ ID NO: 48, or an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of SEQ ID NO: 48.
[0551] In some embodiments, the fusion protein comprises the amino acid sequence of SEQ ID NO: 98, or an amino acid sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the amino acid sequence of SEQ ID NO: 98.
[0552] Representative systems
[0553] The IscB polypeptide or fusion protein of the disclosure may be used in combination with a guide nucleic acid as described herein to constitute a system comprising the IscB polypeptide or fusion protein and the guide nucleic acid, i.e., an IscB system.
[0554] In another aspect, the disclosure provides a system comprising:
[0555] (1) the IscB polypeptide or fusion protein of the disclosure, or a polynucleotide (e.g., a DNA, an RNA) encoding the IscB polypeptide or fusion protein, and
[0556] (2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:
[0557] (i) a scaffold sequence capable of forming a complex with the IscB polypeptide or fusion protein; and
[0558] (ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.
[0559] The components of the system are described more specifically elsewhere.
[0560] In some embodiments, the system is a complex comprising the IscB polypeptide or fusion protein complexed with the guide nucleic acid. In some embodiments, the complex further comprises the target DNA hybridized with the guide sequence.
[0561] In some embodiments, the system is a composition comprising the component (1) and the component (2) .
[0562] In some embodiments, the scaffold sequence is 3’ to the guide sequence.
[0563] In some embodiments, the guide nucleic acid is a guide RNA (gRNA) .
[0564] In some embodiments, the system further comprises a donor polynucleotide for integration or insertion into the target DNA.
[0565] In yet another aspect, the disclosure provides a guide nucleic acid described herein.
[0566] Scaffold sequence
[0567] For the purpose of the disclosure, the scaffold sequence is compatible with the IscB of the disclosure and is capable of complexing with the IscB. The scaffold sequence may be a naturally occurring scaffold sequence identified along with the IscB (e.g., WT OgeuIscB scaffold sequence of SEQ ID NO: 2) , or a variant thereof maintaining the ability to complex with the IscB. Generally, the ability to complex with the IscB is maintained as long as the secondary structure of the variant is substantially identical to the secondary structure of the naturally occurring scaffold sequence. A nucleotide deletion, insertion, or substitution in the primary sequence of the scaffold sequence may not necessarily change the secondary structure of the scaffold sequence (e.g., the relative locations and / or sizes of the stems, bulges, and loops of the scaffold sequence do not significantly deviate from that of the original stems, bulges, and loops) . For example, the nucleotide deletion, insertion, or substitution may be in a bulge or loop region of the scaffold sequence so that the overall symmetry of the bulge and hence the secondary structure remains largely the same. The nucleotide deletion, insertion, or substitution may also be in the stems of the scaffold sequence so that the lengths of the stems do not significantly deviate from that of the original stems (e.g., adding or deleting one base pair in each of two stems correspond to 4 total base changes) . On the other hand, engineering of the scaffold sequence may be applied to improve the activity of IscB system. The disclosure provides engineered scaffold sequence leading to improved endonuclease activity of IscB when used together.
[0568] In some embodiments, the scaffold sequence has substantially the same secondary structure as the secondary structure of any one of SEQ ID NOs: 2, 13-15, 17-47, and 353-361.
[0569] In some embodiments, the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 2, 13-15, 17-47, and 353-361.
[0570] In some embodiments, the scaffold sequence leads to an increased guide sequence-specific (on-target) endonuclease activity compared to that led by SEQ ID NO: 2 when both are used in otherwise identical guide nucleic acid in combination with a same IscB polypeptide (e.g., the engineered IscB polypeptide of the disclosure) , e.g., an increase by at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, or more.
[0571] In some embodiments, the scaffold sequence comprises a deletion in a stem-loop region of the scaffold sequence of SEQ ID NO: 2. In some embodiments, the deletion is in the R1 stem-loop region (positions 1-43 of SEQ ID NO: 2) of the scaffold sequence of SEQ ID NO: 2.
[0572] In some embodiments, the scaffold sequence comprises a deletion of about or at least about or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 nucleotides, or in a numerical range between any two of the preceding values, e.g., from about 5 to about 18 nucleotides. In some embodiments, the scaffold sequence comprises a deletion of about 15 nucleotides.
[0573] In some embodiments, the scaffold sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 13-15 and 17-19, or a polynucleotide sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the polynucleotide sequence of any one of SEQ ID NOs: 13-15 and 17-19.
[0574] In some embodiments, the scaffold sequence comprises a base pair substitution of a thermodynamically unstable base pair (e.g., a A-T base pair or a mismatched base pair (e.g., a A-G base pair, a T-G base pair) ) with a G-C base pair relative to the scaffold sequence of SEQ ID NO: 2. In some embodiments, the base pair substitution is in a stem-loop region (e.g., R1 stem-loop region (positions 1-43 of SEQ ID NO: 2) , R2 stem-loop region (positions 61-112 of SEQ ID NO: 2) , R3 stem-loop region (positions 126-146 of SEQ ID NO: 2) , and R4 stem-loop region (positions 168-184 of SEQ ID NO: 2) ) of the scaffold sequence of SEQ ID NO: 2.
[0575] In some embodiments, the scaffold sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 21-26, 30-31, 33, 36-38, 40, and 43-45, or a polynucleotide sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to the polynucleotide sequence of any one of SEQ ID NOs: 21-26, 30-31, 33, 36-38, 40, and 43-45.
[0576] In some embodiments, the scaffold sequence comprises the polynucleotide sequence of any one of SEQ ID NOs: 47 and 353-361. In some embodiments, the scaffold sequence comprises the polynucleotide sequence of SEQ ID NO: 47.
[0577] In yet another aspect, the disclosure provides a guide nucleic acid comprising the scaffold sequence of the disclosure.
[0578] Protospacer / target sequence
[0579] In some embodiments, the protospacer sequence comprises about or at least about 14 contiguous nucleotides of the target DNA, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more contiguous nucleotides of the target DNA, or in a numerical range between any two of the preceding values, e.g., from about 14 to about 50, or from about 17 to about 22 contiguous nucleotides of the target DNA. In some embodiments, the protospacer sequence comprises about 16 contiguous nucleotides of the target DNA. As used herein, in the context of a target dsDNA, the protospacer sequence is on the nontarget strand of the target dsDNA.
[0580] In some embodiments, the protospacer sequence is immediately 5’ to a target adjacent motif (TAM) . In some embodiments, the TAM is 5’-NNNNNN-3’, wherein N is A, T, G, or C. In some embodiments, the TAM is 5’-NNNGAN-3’, wherein N is A, T, G, or C.
[0581] In some embodiments, the target sequence comprises about or at least about 14 contiguous nucleotides of the target DNA, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more contiguous nucleotides of the target DNA, or in a numerical range between any two of the preceding values, e.g., from about 14 to about 50, or from about 17 to about 22 contiguous nucleotides of the target DNA. In some embodiments, the target sequence comprises about 16 contiguous nucleotides of the target DNA. As used herein, in the context of a target dsDNA, the target sequence is on the target strand of the target dsDNA.
[0582] In some embodiments, the target sequence is immediately 3’ to a target adjacent motif (TAM) , or the reversely complementary sequence of the target sequence (i.e., the protospacer sequence) is immediately 5’ to a target adjacent motif (TAM) . In some embodiments, the TAM is 5’-NNNNNN-3’, wherein N is A, T, G, or C. In some embodiments, the TAM is 5’-NNNGAN-3’, wherein N is A, T, G, or C.
[0583] Guide sequence
[0584] In some embodiments, the guide sequence is about or at least about 14 nucleotides in length, e.g., about or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more nucleotides in length, or in a length of a numerical range between any two of the preceding values, e.g., in a length of from about 14 to about 50 nucleotides, or from about 17 to about 22 nucleotides. In some embodiments, the guide sequence is about 16 nucleotides in length.
[0585] In some embodiments, (1) the guide sequence is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% (fully) reverse complementary to the target sequence; (2) the guide sequence contains no more than 5, 4, 3, 2, or 1 mismatch or contains no mismatch with the target sequence; or (3) the guide sequence comprises no mismatch with the target sequence in the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 nucleotides at the 3’ end of the guide sequence. In some embodiments, (1) the guide sequence is about 100% (fully) reverse complementary to the target sequence.
[0586] In some embodiments, the guide sequence comprises a sequence having a sequence identity of at least about 80% (e.g., at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) to the sequence of any one of SEQ ID NOs: 50-72, 116-125, 136-151, 168-212, and 258-302; or a sequence having at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotide differences, whether consecutive or not, compared to the sequence of any one of SEQ ID NOs: 50-72, 116-125, 136-151, 168-212, and 258-302.
[0587] In some embodiments, the guide sequence comprises any one of SEQ ID NOs: 50-72, 116-125, 136-151, 168-212, and 258-302.
[0588] In some embodiments, the system comprises two or more guide nuclei acids comprising two or more guide sequences capable of hybridizing to two or more target sequences of the same target DNA or different target DNAs, wherein the two or more guide sequences are the same or different, and wherein the two or more target sequences are the same or different.
[0589] In yet another aspect, the disclosure provides a guide nucleic acid comprising the guide sequence of the disclosure.
[0590] Target DNA
[0591] In some embodiments, the target DNA is a target dsDNA, such as, a eukaryotic dsDNA, e.g., a gene in a eukaryotic cell.
[0592] In some embodiments, the target dsDNA comprises a protospacer sequence on a nontarget strand of the target dsDNA, wherein the dsDNA comprises a target deoxyribonucleotide (e.g., dA, dT, dC, dG) at a position of the protospacer sequence selected from the group consisting of position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9, position 10, position 11, position 12, position 13, position 14, and a combination thereof; or wherein the target deoxyribonucleotide is at a position of the protospacer sequence between position 2 and position 12 or between position 3 and position 14, both inclusive.
[0593] In some embodiments, the target deoxyribonucleotide is the N2 nucleotide in a motif of N1N2N3, wherein N1, N2, or N3 is A, T, G, or C.
[0594] Polynucleotide
[0595] In yet another aspect, the disclosure provides a polynucleotide encoding the IscB polypeptide or fusion protein of the disclosure. In some embodiments, the polynucleotide encodes a guide nucleic acid as described herein.
[0596] In yet another aspect, the disclosure provides a polynucleotide comprising or encoding the guide nucleic acid of the disclosure.
[0597] Regulation of guide nucleic acid
[0598] In some embodiments, the polynucleotide encoding the guide nucleic acid is a DNA, a RNA, or a DNA / RNA mixture. By “DNA / RNA mixture” it refers to a nucleic acid comprising both one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, whether consecutive or not. However, by “DNA” or “RNA” it may also refer to a DNA containing one or more modified or unmodified ribonucleotides, whether consecutive or not, or an RNA containing one or more modified or unmodified deoxyribonucleotides, whether consecutive or not.
[0599] In some embodiments, the guide nucleic acid is operably linked to or under the regulation of a promoter.
[0600] In some embodiments, the promoter is a ubiquitous, tissue-specific, cell-type specific, constitutive, or inducible promoter.
[0601] Suitable promoters are known in the art and include, for example, a Cbh promoter, a Cba promoter, a pol I promoter, a pol II promoter, a pol III promoter, a T7 promoter, a U6 promoter, a H1 promoter, a retroviral Rous sarcoma virus LTR promoter, a cytomegalovirus (CMV) promoter, a SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, an elongation factor 1α short (EFS) promoter, a β glucuronidase (GUSB) promoter, a cytomegalovirus (CMV) immediate-early (Ie) enhancer and / or promoter, a chicken β-actin (CBA) promoter or derivative thereof such as a CAG promoter, CB promoter, a (human) elongation factor 1α-subunit (EF1α) promoter, a ubiquitin C (UBC) promoter, a prion promoter, a neuron-specific enolase (NSE) , a neurofilament light (NFL) promoter, a neurofilament heavy (NFH) promoter, a platelet-derived growth factor (PDGF) promoter, a platelet-derived growth factor B-chain (PDGF-β) promoter, a synapsin (Syn) promoter, a synapsin 1 (Syn1) promoter, a methyl-CpG binding protein 2 (MeCP2) promoter, a Ca2+ / calmodulin-dependent protein kinase II (CaMKII) promoter, a metabotropic glutamate receptor 2 (mGluR2) promoter, a neurofilament light (NFL) promoter, a neurofilament heavy (NFH) promoter, a β-globin minigene nβ2 promoter, a preproenkephalin (PPE) promoter, an enkephalin (Enk) promoter, an excitatory amino acid transporter 2 (EAAT2) promoter, a glial fibrillary acidic protein (GFAP) promoter, and a myelin basic protein (MBP) promoter.
[0602] Regulation of polypeptides
[0603] In some embodiments, the polynucleotide encoding the polypeptide (e.g., the IscB polypeptide, the fusion protein) of the disclosure is a DNA, a RNA, or a DNA / RNA mixture.
[0604] In some embodiments, the polynucleotide encoding the polypeptide of the disclosure is operably linked to or under the regulation of a promoter.
[0605] In some embodiments, the promoter is a ubiquitous, tissue-specific, cell-type specific, constitutive, or inducible promoter.
[0606] Suitable promoters are known in the art and include, for example, a Cbh promoter, a Cba promoter, a pol I promoter, a pol II promoter, a pol III promoter, a T7 promoter, a U6 promoter, a H1 promoter, a retroviral Rous sarcoma virus LTR promoter, a cytomegalovirus (CMV) promoter, a SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, an elongation factor 1α short (EFS) promoter, a β glucuronidase (GUSB) promoter, a cytomegalovirus (CMV) immediate-early (Ie) enhancer and / or promoter, a chicken β-actin (CBA) promoter or derivative thereof such as a CAG promoter, CB promoter, a (human) elongation factor 1α-subunit (EF1α) promoter, a ubiquitin C (UBC) promoter, a prion promoter, a neuron-specific enolase (NSE) , a neurofilament light (NFL) promoter, a neurofilament heavy (NFH) promoter, a platelet-derived growth factor (PDGF) promoter, a platelet-derived growth factor B-chain (PDGF-β) promoter, a synapsin (Syn) promoter, a human synapsin (hSyn) promoter, a synapsin 1 (Syn1) promoter, a methyl-CpG binding protein 2 (MeCP2) promoter, a Ca2+ / calmodulin-dependent protein kinase II (CaMKII) promoter, a metabotropic glutamate receptor 2 (mGluR2) promoter, a neurofilament light (NFL) promoter, a neurofilament heavy (NFH) promoter, a β-globin minigene nβ2 promoter, a preproenkephalin (PPE) promoter, an enkephalin (Enk) promoter, an excitatory amino acid transporter 2 (EAAT2) promoter, a glial fibrillary acidic protein (GFAP) promoter, a myelin basic protein (MBP) promoter, a OTOF promoter, a GRK1 promoter, a CRX promoter, a NRL promoter, a MECP2 promoter, a mMECP2 promoter, a hMECP2 promoter, an APP promoter, and a RCVRN promoter.
[0607] Delivery
[0608] Various ways of delivery can be applied to the IscB polypeptide or fusion protein of the disclosure or the system of the disclosure as needed in practices.
[0609] In yet another aspect, the disclosure provides a delivery system comprising (1) the engineered IscB polypeptide of the disclosure, the polynucleotide of the disclosure, or the system of the disclosure; and (2) a delivery vehicle.
[0610] In yet another aspect, the disclosure provides a vector comprising the polynucleotide of the disclosure. In some embodiments, the vector encodes a guide nucleic acid as described herein. In some embodiments, the vector is a plasmid vector, a recombinant AAV (rAAV) vector, or a recombinant lentivirus vector.
[0611] In yet another aspect, the disclosure provides a recombinant AAV (rAAV) particle comprising the rAAV vector of the disclosure. In some embodiments, the rAAV vector is an RNA. A simple introduction of AAV for delivery may refer to “Adeno-associated Virus (AAV) Guide” (addgene. org / guides / aav / ) .
[0612] Adeno-associated virus (AAV) , when engineered to delivery, e.g., a protein-encoding sequence of interest, may be termed as a (r) AAV vector, a (r) AAV vector particle, or a (r) AAV particle, where “r” stands for “recombinant” . And the nucleic acid packaged in AAV vectors for delivery may be termed as a (r) AAV vector genome, vector genome, or vg for short, while viral genome may refer to the original viral genome of natural AAVs.
[0613] The serotypes of the capsids of rAAV particles can be matched to the types of target cells. For example, Table 2 of WO2018002719A1 lists exemplary cell types that can be transduced by the indicated AAV serotypes (incorporated herein by reference) .
[0614] In some embodiments, the rAAV particle comprising a capsid with a serotype suitable for delivery into ear cells (e.g., inner hair cells) . In some embodiments, the rAAV particle comprising a capsid with a serotype of AAV1, AAV2, AAV3A, AAV3B, AAV4, AAV5, AAV6, AAV7, AAVrh74, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV-DJ, or AAV. PHP. eB, a member of the Clade to which any of the AAV1-AAV13 belong, or a functional variant (e.g., a functional truncation) thereof, encapsidating the rAAV vector genome. In some embodiments, the serotype of the capsid is AAV9 or a functional variant thereof.
[0615] General principles of rAAV particle production are known in the art. In some embodiments, rAAV particles may be produced using the triple transfection method (described in detail in U.S. Pat. No. 6,001,650) .
[0616] The vector titers are usually expressed as vector genomes per ml (vg / ml) . In some embodiments, the vector titer is above 1×109, above 5×1010, above 1×1011, above 5×1011, above 1×1012, above 5×1012, or above 1×1013 vg / ml.
[0617] Instead of packaging a single strand (ss) DNA as a vector genome of a rAAV particle, systems and methods of packaging an RNA as a vector genome into a rAAV particle is recently developed and applicable herein. See PCT / CN2022 / 075366, which is incorporated herein by reference in its entirety.
[0618] When the vector genome is RNA as in, for example, PCT / CN2022 / 075366, for simplicity of description and claiming, sequence elements described herein for DNA vector genomes, when present in RNA vector genomes, should generally be considered to be applicable for the RNA vector genomes except that the deoxyribonucleotides in the DNA sequence are the corresponding ribonucleotides in the RNA sequence (e.g., dT is equivalent to U, and dA is equivalent to A) and / or the element in the DNA sequence is replaced with the corresponding element with a corresponding function in the RNA sequence or omitted because its function is unnecessary in the RNA sequence and / or an additional element necessary for the RNA vector genome is introduced.
[0619] As used herein, a coding sequence, e.g., as a sequence element of rAAV vector genomes herein, is construed, understood, and considered as covering and covers both a DNA coding sequence and an RNA coding sequence. When it is a DNA coding sequence, an RNA sequence can be transcribed from the DNA coding sequence, and optionally further a protein can be translated from the transcribed RNA sequence as necessary. When it is an RNA coding sequence, the RNA coding sequence per se can be a functional RNA sequence for use, or an RNA sequence can be produced from the RNA coding sequence, e.g., by RNA processing, or a protein can be translated from the RNA coding sequence.
[0620] For example, an IscB polypeptide coding sequence encoding an IscB polypeptide covers either an IscB polypeptide DNA coding sequence from which an IscB polypeptide is expressed (indirectly via transcription and translation) or an IscB polypeptide RNA coding sequence from which an IscB polypeptide is translated (directly) .
[0621] For example, a gRNA coding sequence encoding a gRNA covers either a gRNA DNA coding sequence from which a gRNA is transcribed or a gRNA RNA coding sequence (1) which per se is the functional gRNA for use, or (2) from which a gRNA is produced, e.g., by RNA processing.
[0622] In some embodiments for rAAV RNA vector genomes, 5’-ITR and / or 3’-ITR as DNA packaging signals may be unnecessary and can be omitted at least partly, while RNA packaging signals can be introduced. In some embodiments for rAAV RNA vector genomes, a promoter to drive transcription of DNA sequences may be unnecessary and can be omitted at least partly. In some embodiments for rAAV RNA vector genomes, a sequence encoding a polyA signal may be unnecessary and can be omitted at least partly, while a polyA tail can be introduced. Similarly, other DNA elements of rAAV DNA vector genomes can be either omitted or replaced with corresponding RNA elements and / or additional RNA elements can be introduced, in order to adapt to the strategy of delivering an RNA vector genome by rAAV particles.
[0623] In yet another aspect, the disclosure provides a ribonucleoprotein (RNP) comprising the engineered IscB polypeptide of the disclosure and a guide nucleic acid. In some embodiments, the guide nucleic acid is as described herein.
[0624] In yet another aspect, the disclosure provides a lipid nanoparticle (LNP) comprising an RNA (e.g., mRNA) encoding the engineered IscB polypeptide of the disclosure and a guide nucleic acid. In some embodiments, the guide nucleic acid is as described herein.
[0625] In yet another aspect, the disclosure provides a cell comprising the engineered IscB polypeptide of the disclosure, the system of the disclosure, the polynucleotide of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the RNP of the disclosure, or the LNP of the disclosure.
[0626] Method of modifying
[0627] The system of the disclosure comprising the IscB polypeptide or fusion protein of the disclosure has a wide variety of utilities, including modifying (e.g., cleaving, deleting, inserting, base editing, translocating, inactivating, or activating) a target DNA in a multiplicity of cell types. The system has a broad spectrum of applications requiring high activity / efficiency and small sizes, e.g., drug screening, disease diagnosis and prognosis, and treating various genetic disorders.
[0628] The method and / or the system of the disclosure can be used to modify a target DNA, for example, to modify the translation and / or transcription of one or more genes of the cells. For example, the modification may lead to increased transcription / translation / expression of a gene. In other embodiments, the modification may lead to decreased transcription / translation / expression of a gene.
[0629] In yet another aspect, the disclosure provides a method for modifying a target DNA, comprising contacting the target DNA with the system of the disclosure, the vector of the disclosure, the ribonucleoprotein of the disclosure, or the lipid nanoparticle of the disclosure, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified by the complex.
[0630] In some embodiments, the modification includes indel event, double-strand cleavage (double stranded break, DSB) , single-strand cleavage (nick) (e.g., on either target strand or nontarget stand of a dsDNA) , base editing (e.g., single base editing) , prime editing, and integration or insertion of exogenous donor (e.g., by homologous recombination) .
[0631] In some embodiments, the method is in vitro, in vivo, or ex vivo.
[0632] In some embodiments, the target DNA is in a cell.
[0633] In yet another aspect, the disclosure provides a cell comprising the system of the disclosure.
[0634] In yet another aspect, the disclosure provides a cell modified by the system of the disclosure or the method of the disclosure. In some embodiments, the cell is modified in vitro, in vivo, or ex vivo.
[0635] In some embodiments, the cell is a eukaryotic cell (e.g., an animal cell, a vertebrate cell, a mammalian cell, a non-human mammalian cell, a non-human primate cell, a rodent (e.g., mouse or rat) cell, a human cell, a plant cell, or a yeast cell) or a prokaryotic cell (e.g., a bacteria cell) .
[0636] In some embodiments, the cell is from a plant or an animal. In some embodiments, the cell is not from a plant.
[0637] In some embodiments, the cell is a non-human mammalian cell, such as a cell from a non-human primate (e.g., monkey) , an ox / cow / bull / cattle, sheep, goat, pig, horse, dog, cat, rodent (such as rabbit, mouse, rat, hamster, etc. ) , alpaca. In some embodiments, the cell is from fish (such as salmon, zebra fish) , bird (such as poultry bird, including chick, duck, goose) , reptile, shellfish (e.g., oyster, clam, lobster, shrimp) , insect, worm, yeast, etc.
[0638] In some embodiments, the plant is a dicotyledon. In some embodiments, the dicotyledon is selected from the group consisting of soybean, cabbage (e.g., Chinese cabbage) , rapeseed, brassica, watermelon, melon, potato, tomato, tobacco, eggplant, pepper, cucumber, cotton, alfalfa, eggplant, grape. In some embodiments, the plant is a monocotyledon. In some embodiments, the monocotyledon is selected from the group consisting of rice, corn, wheat, barley, oat, sorghum, millet, grasses, Poaceae, Zizania, Avena, Coix, Hordeum, Oryza, Panicum (e.g., Panicum miliaceum) , Secale, Setaria (e.g., Setaria italica) , Sorghum, Triticum, Zea, Cymbopogon, Saccharum (e.g., Saccharum officinarum) , Phyllostachys, Dendrocalamus, Bambusa, Yushania.
[0639] In some embodiments, the cell is from a plant, such as monocot or dicot. In certain embodiment, the plant is a food crop such as barley, cassava, cotton, groundnuts or peanuts, maize, millet, oil palm fruit, potatoes, pulses, rapeseed or canola, rice, rye, sorghum, soybeans, sugar cane, sugar beets, sunflower, and wheat. In certain embodiment, the plant is a cereal (barley, maize, millet, rice, rye, sorghum, and wheat) . In certain embodiment, the plant is a tuber (cassava and potatoes) . In certain embodiment, the plant is a sugar crop (sugar beets and sugar cane) . In certain embodiment, the plant is an oil-bearing crop (soybeans, groundnuts or peanuts, rapeseed or canola, sunflower, and oil palm fruit) . In certain embodiment, the plant is a fiber crop (cotton) . In certain embodiment, the plant is a tree (such as a peach or a nectarine tree, an apple or pear tree, a nut tree such as almond or walnut or pistachio tree, or a citrus tree, e.g., orange, grapefruit or lemon tree) , a grass, a vegetable, a fruit, or an algae. In certain embodiment, the plant is a nightshade plant; a plant of the genus Brassica; a plant of the genus Lactuca; a plant of the genus Spinacia; a plant of the genus Capsicum; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.
[0640] In some embodiments, the cell is a stem cell. In some embodiments, the cell is an embryonic stem cell. In some embodiments, the cell is a primary human cell or an established human cell line.
[0641] In some embodiments, the cell is not a human or animal embryonic stem cell. In some embodiments, the cell is not a human or animal germ cell. In some embodiments, the cell is not a plant cell.
[0642] Pharmaceutical composition
[0643] In yet another aspect, the disclosure provides a pharmaceutical composition comprising (1) the system of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, or the cell of the disclosure; and (2) a pharmaceutically acceptable excipient.
[0644] In some embodiments, the pharmaceutical composition comprises the rAAV particle in a concentration selected from the group consisting of about 1×1010 vg / mL, 2×1010 vg / mL, 3×1010 vg / mL, 4×1010 vg / mL, 5×1010 vg / mL, 6×1010 vg / mL, 7×1010 vg / mL, 8×1010 vg / mL, 9×1010 vg / mL, 1×1011 vg / mL, 2×1011 vg / mL, 3×1011 vg / mL, 4×1011 vg / mL, 5×1011 vg / mL, 6×1011 vg / mL, 7×1011 vg / mL, 8×1011 vg / mL, 9×1011 vg / mL, 1×1012 vg / mL, 2×1012 vg / mL, 3×1012 vg / mL, 4×1012 vg / mL, 5×1012 vg / mL, 6×1012 vg / mL, 7×1012 vg / mL, 8×1012 vg / mL, 9×1012 vg / mL, 1×1013 vg / mL, or in a concentration of a numerical range between any of two preceding values, e.g., in a concentration of from about 9×1010 vg / mL to about 8×1011 vg / mL.
[0645] In some embodiments, the pharmaceutical composition is an injection.
[0646] In some embodiments, the volume of the injection is selected from the group consisting of about 1 microliter, 10 microliters, 50 microliters, 100 microliters, 150 microliters, 200 microliters, 250 microliters, 300 microliters, 350 microliters, 400 microliters, 450 microliters, 500 microliters, 550 microliters, 600 microliters, 650 microliters, 700 microliters, 750 microliters, 800 microliters, 850 microliters, 900 microliters, 950 microliters, 1000 microliters, and a volume of a numerical range between any of two preceding values, e.g., in a concentration of from about 10 microliters to about 750 microliters.
[0647] Method of diagnosing, preventing, or treating
[0648] In yet another aspect, the disclosure provides a method for diagnosing, preventing, or treating a disease in a subject in need thereof, comprising administering to the subject the system of the disclosure, the vector of the disclosure, the rAAV particle of the disclosure, the ribonucleoprotein of the disclosure, the lipid nanoparticle of the disclosure, the cell of the disclosure, or the pharmaceutical composition of the disclosure, wherein the disease is associated with a target DNA, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, wherein the target DNA is modified, and wherein the modification of the target DNA diagnose, prevents, or treats the disease.
[0649] In some embodiments, the disease is selected from the group consisting of Angelman syndrome (AS) , Alzheimer's disease (AD) , transthyretin amyloidosis (ATTR) , transthyretin amyloid cardiomyopathy (ATTR-CM) , cystic fibrosis (CF) , hereditary angioedema, diabetes, progressive pseudohypertrophic muscular dystrophy, Duchenne muscular dystrophy (DMD) , Becker muscular dystrophy (BMD) , spinal muscular atrophy (SMA) , alpha-1-antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington’s disease (HTT) , fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis (ALS) , frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA) , sickle cell disease, thalassemia (e.g., β-thalassemia) , Parkinson's disease (PD) , myelodysplastic syndrome (MDS) , retinitis pigmentosa (RP) , age-related macular degeneration (AMD) , Hepatitis B, nonalcoholic fatty liver disease (NAFLD) , Acquired Immune Deficiency Syndrome, corneal dystrophy (CD) , hypercholesterolemia, familial hypercholesterolemia (FH) , heart disease (e.g., hypertrophic cardiomyopathy (HCM) ) , and cancer.
[0650] In some embodiments, the target DNA encodes a mRNA, a tRNA, a ribosomal RNA (rRNA) , a microRNA (miRNA) , a non-coding RNA, a long non-coding (lnc) RNA, a nuclear RNA, an interfering RNA (iRNA) , a small interfering RNA (siRNA) , a ribozyme, a riboswitch, a satellite RNA, a microswitch, a microzyme, or a viral RNA.
[0651] In some embodiments, the target DNA is a eukaryotic DNA.
[0652] In some embodiments, the eukaryotic DNA is a mammal DNA, such as a non-human mammalian DNA, a non-human primate DNA, a human DNA, a plant DNA, an insect DNA, a bird DNA, a reptile DNA, a rodent (e.g., mouse, rat) DNA, a fish DNA, a nematode DNA, or a yeast DNA.
[0653] In some embodiments, the target DNA is in a eukaryotic cell, for example, a human cell, a non-human primate cell, or a mouse cell.
[0654] In some embodiments, the administrating comprises local administration or systemic administration.
[0655] In some embodiments, the administrating comprises intrathecal administration, intramuscular administration, intravenous administration, transdermal administration, intranasal administration, oral administration, mucosal administration, intraperitoneal administration, intracranial administration, intracerebroventricular administration, or stereotaxic administration.
[0656] In some embodiments, the administration is injection or infusion.
[0657] In some embodiments, the subject is a human, a non-human primate, or a mouse.
[0658] In some embodiments, the level of the transcript (e.g., mRNA) of the target DNA is decreased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the transcript (e.g., mRNA) of the target DNA in the subject prior to the administration.
[0659] In some embodiments, the level of the transcript (e.g., mRNA) of the target DNA is increased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the transcript (e.g., mRNA) of the target DNA in the subject prior to the administration.
[0660] In some embodiments, the level of the expression product (e.g., protein) of the target DNA is decreased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the expression product (e.g., protein) of the target DNA in the subject prior to the administration.
[0661] In some embodiments, the level of the expression product (e.g., protein) of the target DNA is increased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the expression product (e.g., protein) of the target DNA in the subject prior to the administration. In some embodiments, the expression product is a functional mutant of the expression product of the target DNA.
[0662] In some embodiments, the median survival of the subject suffering from the disease but receiving the administration is 5 days, 10 days, 20 days, 30 days, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 1.5 year, 2 years, 2.5 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years or more longer than that of a subject or a population of subjects suffering from the disease and not receiving the administration.
[0663] The therapeutically effective dose may be either via a single dose, or multiple doses. One skilled in the art understands that the actual dose may vary greatly depending upon a variety of factors, such as the vector choices, the target cells, organisms, tissues, the general conditions of the subject to be treated, the degrees of transformation / modification sought, the administration routes, the administration modes, the types of transformation / modification sought, etc.
[0664] For example, the therapeutically effective dose of the rAAV particle may be about 1.0E+8, 2.0E+8, 3.0E+8, 4.0E+8, 6.0E+8, 8.0E+8, 1.0E+9, 2.0E+9, 3.0E+9, 4.0E+9, 6.0E+9, 8.0E+9, 1.0E+10, 2.0E+10, 3.0E+10, 4.0E+10, 6.0E+10, 8.0E+10, 1.0E+11, 2.0E+11, 3.0E+11, 4.0E+11, 6.0E+11, 8.0E+11, 1.0E+12, 2.0E+12, 3.0E+12, 4.0E+12, 6.0E+12, 8.0E+12, 1.0E+13, 2.0E+13, 3.0E+13, 4.0E+13, 6.0E+13, 8.0E+13, 1.0E+14, 2.0E+14, 3.0E+14, 4.0E+14, 6.0E+14, 8.0E+14, 1.0E+15, 2.0E+15, 3.0E+15, 4.0E+15, 6.0E+15, 8.0E+15, 1.0E+16, 2.0E+16, 3.0E+16, 4.0E+16, 6.0E+16, 8.0E+16, or 1.0E+17 vg, or within a range of any two of the those point values. vg stands for vector genomes of rAAV particles for administration.
[0665] Method of detecting
[0666] In yet another aspect, the disclosure provides a method of detecting a target DNA, comprising contacting the target DNA with the system of the disclosure, wherein the target DNA is modified, and wherein the modification detects the target DNA. In some embodiments, the modification generates a detectable signal, e.g., a fluorescent signal.
[0667] Kits
[0668] In yet another aspect, the disclosure provides a kit comprising the IscB polypeptide or fusion protein of the disclosure, the system of the disclosure, the polynucleotide of the disclosure, the vector of the disclosure, the RNP of the disclosure, the LNP of the disclosure, the delivery system of the disclosure, the cell of the disclosure, or the pharmaceutical composition of the disclosure, or any one, two, or all components of the same.
[0669] In some embodiments, the kit further comprises an instruction to use the component (s) contained therein, and / or instructions for combining with additional component (s) that may be available or necessary elsewhere.
[0670] In some embodiments, the kit further comprises one or more buffers that may be used to dissolve any of the component (s) contained therein, and / or to provide suitable reaction conditions for one or more of the component (s) . Such buffers may include one or more of PBS, HEPES, Tris, MOPS, Na2CO3, NaHCO3, NaB, or combinations thereof. In some embodiments, the reaction condition includes a proper pH, such as a basic pH. In some embodiments, the pH is between 7-10.
[0671] In some embodiments, any one or more of the kit components may be stored in a suitable container or at a suitable temperature, e.g., 4 Celsius degree.
[0672] Further embodiments are illustrated in the following Examples which are given for illustrative purposes only and are not intended to limit the scope of the disclosure.
[0673] EXAMPLES
[0674] The following examples are provided to further illustrate some embodiments of the disclosure but are not intended to limit the scope of the invention; it will be understood by their exemplary nature that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.
[0675] Methods
[0676] Computational analysis of IscB systems
[0677] More than 200 Gb metagenome assemblies were downloaded from ENA database with accession PRJEB31266. Firstly, the inventor used TBLASTN and OgeuIscB protein to identify IscB-containing sequences of metagenomes with E value < 1e-50 (ref. 20) . Then, prodigal was used to annotate the proteins of IscB-containing sequence32. The inventor further used search and previously trained ωRNA model to annotate the ωRNA sequences with E value < 1e-10. RNAfold was used to predict the secondary structure of the ωRNA33, 34. MEGAX was used to construct the phylogenetic tree35.
[0678] Plasmids constructions
[0679] All E. coli codon-optimized IscB coding genes and their associated ωRNA scaffolds were synthesized by Shanghai Huagene Biotechnology Co., Ltd. and assembled into a pUC19-derived vector (EcoNI + XbaI) under the lac and J23119 promoters by 2× -Basic Seamless Cloning and Assembly Kit (TransGen Biotech Co., Ltd. ) , respectively. All human codon-optimized IscB coding sequences were synthesized by GenScript Co., Ltd. and incorporated into a mammalian expression vector under CBH promoter. For endogenous genome editing experiments in HEK293T cells, the guide RNA oligos were synthesized and cloned into a BpiI-digested backbone of U6 promoter using T4 ligase (Thermo Fisher Scientific) .
[0680] Generation of the TAM library and TAM depletion assay
[0681] A randomized TAM library containing a target sequence followed by 8 randomized bases downstream was constructed. The synthesized single-stranded DNA (ssDNA) (HuaGene Co., Ltd. ) was converted into double stranded by annealing with a short ssDNA and second strand synthesis using Large (Klenow) fragment (NEB) . The resulting dsDNA was then assembled into pACYC184 vectors using Gibson assembly (NEB) . The products were purified using Isopropanol, electroporated into TransforMax EC100 Electrocompetent E. coli according to the manufacturer’s instructions and plated on chloramphenicol plates. After 13 hours of growth at 37 ℃, cells were scraped from the plates and extracted using NucleoBond Xtra Midiprep kit (Machery Nagel) .
[0682] For bacterial TAM depletion assay, the inventor co-transformed 200 ng of TAM library plasmids and 300 ng of plasmids expressing E. coli codon-optimized IscB and ωRNA into TransforMax EC100 Electrocompetent E. coli cells by electroporation. Then the transformed cells were recovered for 1 hour, and plated on 250 mm × 250 mm carbenicillin and chloramphenicol plates. After 13 hours of growth, cells were harvested and plasmid DNA was extracted using NucleoBond Xtra Midiprep kit (Machery Nagel) . TAM-containing region was amplified by Phanta Max Super-Fidelity DNA Polymerase (Vazyme Biotech) for 12 cycles, and Illumina adaptors and unique barcodes were added by a second round PCR for 18 cycles. The resulting PCR products were purified with Gel extraction kit (Omega) and sequenced by Illumina NovaSeq 6000 platform with 150-bp paired-end reads (Genewiz Co. Ltd) .
[0683] TAM regions were extracted, counted, and then normalized to the total TAM counts for each sample. For each specific TAM, TAMs that appeared more than once were filtered, and the log fold change (logFC) of its frequency was measured as the log ratio compared to non-targeting control. Depletions with a logFC < -3σ (standard deviation) were considered statistically significant. A position weight matrix (PWM) was built from all significantly depleted sequences, with -logFC values serving as the corresponding weight. A sequence logo was generated based on this PWM using WebLogo version 3.7.12 (ref. 17) .
[0684] Cell culture and transfection
[0685] HEK293T cells were cultivated in DMEM (Sigma) supplemented with 10%FBS (Gibco) , 1%Pen-Stre-Glutamine (Gibco) , and 1%minimum essential medium nonessential amino acid (Gibco) in a humidified incubator at 37 ℃ with 5%CO2. For the detection of IscB nuclease activities and screening its variants, HEK293T cells were seeded in 24-well plates with 70%-80%confluence. After a 12-hour incubation, 1.6 μg of plasmids were co-transfected into HEK293T cells using polyetherimide (PEI) following the manufacturer’s manual. The plasmids included those encoding BFP-T2A-GFxxFP and IscB systems, with a molar ratio of 1: 1 mixture. For genome or base editing at endogenous loci, 1.6 μg of all-in-one plasmids were transfected to express guide RNA and nuclease or base editor system. After 48 hours, cells were sorted by FACS analysis.
[0686] Fluorescence-activated cell sorting (FACS) analysis
[0687] Before FACS analysis, cells were subjected to treatment with 0.25%Trypsin-EDTA (Gibco) for dissociation and suspended in FBS-containing DMEM. For the assessment of IscB nuclease activity and screening of variants using fluorescence reporter system, cells were analyzed for EGFP, mCherry and BFP fluorescence. A total of 25,000 single cells were recorded to analyze efficiency by Beckman CytoFlex flow cytometer 48 hours after transfection. Data analysis was performed by FlowJo X (v. 10.0.7) . For genome editing analysis, approximately 15,000 transfection-positive cells (defined as those with a fluorescence intensity ≥ 103 among fluorescence-positive cells) were sorted 48 hours after transfection using BD FACS Aria III flow cytometer. Following FACS sorting, genomic DNA from the collected cells was extracted by cell lysis with 25 μl of proteinase K-added lysis buffer (Vazyme Biotech) per sample, as described previously. The cell lysates were stored at -20 ℃ until further use.
[0688] Targeted deep sequencing and analysis
[0689] To detect the editing efficiency at endogenous loci, the target genome regions of interest were amplified from cell lysates by PCR using Phanta Max Super-Fidelity DNA Polymerase (Vazyme Biotech) . For targeted deep sequencing analysis, PCR reactions were performed using primers with unique barcodes. The amplified products were purified using Gel extraction kit (Omega) and sequenced by Illumina NovaSeq 600 platform with 150-bp paired-end reads (Genewiz Co., Ltd. ) . The deep-sequencing data were first demultiplexed by a custom script based on sample barcodes. The demultiplexed reads were then analyzed by CRISPResso2 (ref. 36) for the quantification of editing efficiency, including indels and base conversions at each target locus.
[0690] Guide RNA (gRNA) -dependent off-target analysis
[0691] To examine the gRNA-dependent off-target effects of IscB. m16*-ABE, enOgeuIscB-ABE and SpG-ABE, the CRISPR RGEN Tools (Cas-OFFinder, http: / / www. rgenome. net / cas-offinder / ) was used to predict potential off-target sites as described previously26. For the adenosine base editor based on IscB. m16*, the search queries covered both the 14-nt on-target spacer sequences and “NNNGNA” . The PAM of search was set as “NNN” and the number of mismatches was set to 3. The search queries of enOgeuIscB-ABE were set similar to IscB. m16*-ABE, but with a 16-nt spacer sequence and 6-nt TAM sequence containing “NWRRNA” . For adenosine base editor based on SpG, search queries covered 20-nt on-target spacer sequences, PAM type was set to ‘NG’ and the number of mismatches were set to 4. All other parameters were left as default. Off-target sites for each gRNA in each group were manually selected in order of the number of mismatches from low to high. Sites with a 5’-NNNGNA-3’ TAM were retained for IscB. m16*-ABE, and sites with a 5’-NWRRNA-3’ TAM for enOgueIscB-ABE.
[0692] Orthogonal R-loop assay
[0693] Orthogonal R-loop assay was performed to detect the gRNA-independent off-target editing as described previously27.0.8 μg of plasmids that encode IscB. m16*-ABE, enOgueIscB-ABE or SpG-ABE with their respective ωRNA or single-guide RNA (sgRNA) , and 0.8 μg of dSaCas9 plasmids with corresponding sgRNA targeting five previously reported R-loop sites were co-transfected into HEK293T cells using PEI. After a 48-hour cultivation, transfected cells were analyzed by FACS followed by genomic DNA extraction with 25 μl of freshly prepared lysis buffer (Vazyme) containing proteinase K. Amplification and targeted deep sequencing were performed at ABE on-target sites and dSaCas9 R-loop off-target sites.
[0694] Animals
[0695] All animal experiments in this study were performed following approved protocols and guidelines set by the Animal Care and Use Committee of Huidagene Therapeutics Co., Ltd, located in Shanghai, China. Mice were housed in a controlled barrier facility with a 12-hour light / dark cycle and 18-23 ℃ with 40-60%humidity. Diet and water were accessible at all times. DMDΔmE5051, KIhE50 / Y mice were generated in the C57BL / 6 J background using the CRISPR-Cas9 system. Duchenne muscular dystrophy (DMD) is the most common sex-linked lethal disease in human, and thus male mice were selected for this study.
[0696] Production and delivery of AAV9 to DMDΔmE5051, KIhE50 / Y mice
[0697] AAVs were manufactured by HuidaGene Therapeutics Inc (Shanghai, China) . Briefly, Cells were grown in culture until they reached a confluency of 70 -90%. Before transfection, the growth media was replaced with pre-warmed growth media. For each 15-cm dish, a mixture of 20 μg of pHelper, 10 μg of pRepCap, and 10 μg of GOI plasmid was prepared and added dropwise to the cell media. After a three-day incubation period, AAVs were harvested and purified using iodixanol density gradient centrifugation. For intramuscular injection, three-week-old DMDΔmE5051, KIhE50 / Y mice were anesthetized, and their tibialis anterior (TA) muscle was injected with either 30 μL of AAV9 (2.5 × 1011 vg) preparations or an equivalent volume of saline solution. Tissue samples were collected for genomic DNA, RNA, immunoblotting, and immunofluorescence analyses four weeks post-treatment.
[0698] Western Blot Analysis
[0699] Tissues samples were homogenized using RIPA buffer supplemented with a protease inhibitor cocktail. The supernatants of the lysates were quantified using a Pierce BCA protein assay kit (Thermo Fisher Scientific, 23225) and adjusted to a uniform concentration using H2O. Equal volumes of the samples were mixed with NuPAGE LDS sample buffer (Invitrogen, NP0007) and 10%β-mercaptoethanol, then subjected to boiling at 70 ℃ for 10 min. A total of 10 μg of protein per lane was loaded into 3%to 8%tris-acetate gels (Invitrogen, EA03752BOX) and underwent electrophoresis for 1 hour at 200 V. Proteins were then transferred onto a PVDF membrane under wet conditions at 350 mA for 3.5 hours. The membrane was then blocked in 5%non-fat milk in TBST buffer and incubated with the primary antibody to mark the target protein. After three-times washes with TBST, the membrane was incubated with an HRP-conjugated secondary antibody specific to the IgG of the species from which the primary antibody against dystrophin (Sigma, D8168) or vinculin (CST, 13901S) was derived. The target proteins were visualized using Chemiluminescent Substrates (Invitrogen, WP20005) .
[0700] Immunofluorescence
[0701] Tissues were encased in optimal cutting temperature (OCT) compound and rapidly frozen in liquid nitrogen. Serial frozen cryosections, each measuring 10 μm in thickness, were fixed for two hours at 37 ℃, followed by permeabilization with PBS containing 0.4%Triton-X for 30 min. After washing with PBS, samples were blocked with 10%goat serum for 1 hour at room temperature. Following this, the slides were incubated overnight at 4 ℃ with primary antibodies against dystrophin (Abcam, ab15277) and spectrin (Millipore, MAB1622) . The following day, samples were thoroughly washed with PBS and incubated with compatible secondary antibodies (Alexa 488 AffiniPure donkey anti-rabbit IgG (Jackson ImmunoResearch labs, 711-545-152) or Alexa Fluor 647 AffiniPure donkey anti-mouse IgG (Jackson ImmunoResearch labs, 715-605-151) ) and DAPI for 3 h at room temperature. After a 15-minute PBS wash, slides were sealed with fluoromount-G mounting medium. All images were captured using Nikon C2. The number of Dys+ muscle fibers is represented as a percentage of the total spectrin-positive muscle fibers.
[0702] Statistical analysis
[0703] All values are shown as mean ± standard deviation (s.d. ) , except values of editing window from base editors as mean ± standard error of the mean (s.e.m. ) . One-wayay ANOVA was used for statistical comparisons and P-value < 0.05 was considered to be statistically significant. Details of statistical values are provided in Supplementary Tables. The experiments were not randomized and the investigators were not blinded to allocation during experiments and outcome assessment. GraphPad Prism (v 8.2.1) was used for statistics (www. graphpad. com / ) .
[0704] Example 1: Functional identification of uncharacterized IscB orthologs from uncultured microbes
[0705] The inventor discovered 19 uncharacterized IscB systems as shown in Table A1 below. The amino acid sequences of the wild type IscB proteins (named as IscB. m1 to IscB. m19, respectively) of these IscB systems are set forth in SEQ ID NOs: 1-19, respectively. The human codon-optimized coding sequences of these IscB proteins are set forth in SEQ ID NO: 39-57, respectively. The scaffold sequences of the ωRNA in the IscB systems corresponding to these IscB proteins one by one are set forth in SEQ ID NOs: 20-38, respectively.
[0706] Table A1
[0707] These IscB systems were phylogenetically clustered into three subgroups based on sequence alignment of the IscB proteins (FIG. 1a, FIG. 6) . Through the protein sequence alignment encompassing 500 amino acids, the inventor identified the conserved residues within the RuvC, HNH, P1D and TID domains of the IscB proteins, suggesting the possibility of endonuclease and nickase activity of the identified IscB proteins (FIG. 7) .
[0708] To detect whether the identified IscB-ωRNA systems comprising these IscB proteins and the corresponding ωRNAs were capable of cleaving DNA and to characterize their TAM recognition ability, the inventor performed a bacterial depletion assay. The inventor co-transformed E. coli cells with plasmids carrying the IscB and ωRNA with a spacer (a. k. a., a spacer sequence, a guide sequence) and a TAM library plasmid carrying target sequences complementary to the spacer and an 8-base pair (bp) randomized sequences (as potential TAM) (FIG. 8a) . Through this assay, a series of specific depleted TAM sequences were enriched associated with each IscB-ωRNA system, indicating that these natural IscB orthologs have RNA-guided DNA endonuclease activity (FIG. 1b, FIG. 8b) . Subsequently, the inventor analyzed the relationship between the divergence of the IscB proteins and the difference in TAMs and observed that most of the IscB proteins have significant distinctions in both terms of IscB protein sequences and TAM recognition ability (FIG. 8c) .
[0709] To further assess the endonuclease activity of these IscB orthologs in human cells, the inventor employed a fluorescence reporter system. This system involved co-transfecting a reporter plasmid encoding unactivated GFxxFP and an expression plasmid expressing one of the IscB proteins and its corresponding GFxxFP-targeting ωRNA into cultured HEK293T cells. The inventor then measured the EGFP signal intensity of the GFxxFP reporter which was activated by IscB-mediated double-strand breaks (DSB) 21 (FIG. 1c) . Using the GFxxFP reporter with the experimentally determined TAM for each IscB, 10 (marked with “*” in FIG. 1d) out of the 19 IscBs were observed to have a significant increase (>2-fold ratio of on-target relative to non-target, with on-target >1.0%) in EGFP signal intensity relative to non-target control (using a non-targeting spacer) , indicating their RNA-guided DNA endonuclease activity in mammalian cells. Notably, IscB. m16 (SEQ ID NO: 16; coded by SEQ ID NO: 54; with ωRNA scaffold sequence of SEQ ID NO: 35) exhibited the highest signal intensity, indicating the highest endonuclease activity (FIG. 1d) .
[0710] Example 2: Engineering scaffold sequence of ωRNA to improve editing efficiency
[0711] To enhance the editing efficiency of natural IscB. m16 system, the inventor engineered the wild type scaffold sequence (SEQ ID NO: 35) of ωRNA identified along with IscB. m16 by truncation or mutagenesis, generating multiple scaffold sequence variants (SEQ ID NOs: 58-97; FIG. 2b) with a change in one of five stem-loop regions namely R1, R2, R3, R4, and R5 (FIG. 2a, FIG. 9a) .
[0712] SEQ ID NO: 35, IscB. m16, scaffold sequence of ωRNA.
[0713] R1: position 2 to position 47 of the scaffold sequence of SEQ ID NO: 35.
[0714] R2: position 51 to position 121 of the scaffold sequence of SEQ ID NO: 35.
[0715] R3: position 122 to position 133 of the scaffold sequence of SEQ ID NO: 35.
[0716] R4: position 136 to position 154 of the scaffold sequence of SEQ ID NO: 35.
[0717] R5: position 157 to position 180 of the scaffold sequence of SEQ ID NO: 35 (including the stem-loop in color and the subsequent 3’ end shown in FIG. 2a; collectively, “R5” ) .
[0718] As shown in FIG. 2a and FIG. 9a, the five stem-loop regions are parts of the secondary structure of the scaffold sequence of ωRNA. The secondary structure of the scaffold sequence of an IscB ωRNA can be depicted, like any other RNA, by well-known methods, e.g., online tool RNAfold (http: / / rna. tbi. univie. ac. at / cgi-bin / RNAWebSuite / RNAfold. cgi) .
[0719] The inventor then tested the endonuclease activity of the IscB system comprising the wild type IscB. m16 (SEQ ID NO: 16) and the ωRNA comprising the same GFxxFP-targeting guide sequence and a distinct scaffold sequence variant (one of SEQ ID NOs: 58-97) using the GFxxFP reporter in Example 1. The results are shown in FIG. 2b. It was observed that several scaffold sequence variants led to increased endonuclease activity of the IscB system compared to an otherwise identical control IscB system ( “WT” in FIG. 2b) with the wild type scaffold sequence of SEQ ID NO: 35. In particular, the scaffold sequence variants R1-Δ13b (SEQ ID NO: 73) and R5-Δ10 (SEQ ID NO: 96) achieved remarkable improvement.
[0720] SEQ ID NO: 73, IscB. m16-scaffold-R1Δ13-b, where and in R1 of the WT scaffold sequence (SEQ ID NO: 35) were deleted.
[0721] As shown from the secondary structure of the WT scaffold sequence (SEQ ID NO: 35) , several base pairs in R1 (red box in FIG. 20) were truncated to generate IscB. m16-scaffold-R1Δ13-b (SEQ ID NO: 73) .
[0722] SEQ ID NO: 96, IscB. m16-scaffold-R5Δ10, where in R5 of WT scaffold sequence (SEQ ID NO: 35) was deleted.
[0723] As shown from the secondary structure of the WT scaffold sequence (SEQ ID NO: 35) , the 3’ end in R5 (red box in FIG. 21) was truncated to generate IscB. m16-scaffold-R5Δ10 (SEQ ID NO: 96) .
[0724] Furthermore, combination of R1Δ13-b truncation and R5Δ10 truncation was made to the WT scaffold sequence (SEQ ID NO: 35) , generating the scaffold sequence variant (R1Δ13-R5Δ10) (SEQ ID NO: 98) that led to even higher endonuclease activity of the IscB system (FIG. 2c) .
[0725] SEQ ID NO: 98, IscB. m16-scaffold-R1Δ13-R5Δ10, where and in R1 and in R5 of WT scaffold sequence (SEQ ID NO: 35) were deleted.
[0726] To increase the stability of ωRNA, the inventor replaced the A-U or mismatched base pairs in one or more of stem regions to thermodynamically stable G-C base-pairs based on the IscB. m16-scaffold-R1Δ13-R5Δ10 (SEQ ID NO: 98) , generating multiple scaffold sequence variants (SEQ ID NOs: 99-117) . Among them, several scaffold sequence variants achieved higher endonuclease activity than IscB. m16-scaffold-R1Δ13-R5Δ10 (SEQ ID NO: 98) , and in particular, scaffold sequence v2.27 (R1-Δ13, R5-Δ10, T24G, G25C, T57G, T79C, A117C) (SEQ ID NO: 107; IscB. m16-scaffold-R1Δ13-R5Δ10-M10) achieved the highest endonuclease activity for AAAGCA TAM reporter (FIG. 2c) .
[0727] SEQ ID NO: 98, IscB. m16-scaffold-R1Δ13-R5Δ10
[0728] SEQ ID NO: 107, IscB. m16-scaffold-R1-Δ13, R5-Δ10, T24G, G25C, T57G, T79C, A117C
[0729] As shown in FIG. 22 and the comparison of IscB. m16-scaffold-R1Δ13-R5Δ10 (SEQ ID NO: 98) and IscB. m16-scaffold-R1-Δ13, R5-Δ10, T24G, G25C, T57G, T79C, A117C (SEQ ID NO: 107) , a U-G mismatch in R1 of IscB. m16-scaffold-R1Δ13-R5Δ10 at positions 24 and 25 (numbered according to WT IscB. m16 scaffold sequence of SEQ ID NO: 35) was replaced with a g-c base pair, a U-Abase pair in R2 of IscB. m16-scaffold-R1Δ13-R5Δ10 at positions 57 and 117 (numbered according to WT IscB. m16 scaffold sequence of SEQ ID NO: 35) was replaced with a g-c base pair, and the uracil (U) of a U-G mismatch in R2 of IscB. m16-scaffold-R1Δ13-R5Δ10 at position 79 (numbered according to WT IscB. m16 scaffold sequence of SEQ ID NO: 35) was replaced with cytidine (c) .
[0730] Similarly, to improve the endonuclease activity of IscB. m17 system, the inventor introduced truncation into six stem loop (R1, R2, R3, R4, R5, R6) of the secondary structure (FIG. 9b) of the wild type IscB. m17 scaffold sequence (SEQ ID NO: 36) , generating multiple scaffold sequence variants (SEQ ID NOs: 118-151) . It was observed that multiple variants achieved improved endonuclease activity than wild type IscB. m17 scaffold sequence (SEQ ID NO: 36) , and in particular, variant R1-Δ59 (SEQ ID NO: 136) achieved the highest endonuclease activity.
[0731] SEQ ID NO: 136, R1-Δ59, where 59 nucleotides at positions 18-76 of WT IscB. m17 scaffold sequence (SEQ ID NO: 36) was deleted.
[0732] The inventor then generated scaffold sequence variants based on R1-Δ59 (SEQ ID NO: 136) by the replacement of A-U to C-G base-pairs in the R1 stem loop and identified a variant (R1-Δ59-M2; SEQ ID NO: 153) with even higher endonuclease activity than R1-Δ59 (FIG. 2e) .
[0733] SEQ ID NO: 153, R1-Δ59-M2, wherein a A-U base pair in stem loop R1 at positions 4 and 89 of R1-Δ59 (SEQ ID NO: 136) (numbered according to wild type IscB. m17 scaffold sequence (SEQ ID NO: 36) ) was replaced with a c-g base pair.
[0734] In view of the various truncations and base pair substitutions tested above, an IscB scaffold sequence engineering strategy was established, comprising modifying the scaffold sequence of an IscB ωRNA by (1) deleting one or more base pair (mismatched or not) in the first (e.g., R1) and / or the last (e.g., R5) stem loop of the scaffold sequence, and / or (2) substituting a A-U base pair or mismatched base pair in a stem region of a stem loop of the scaffold sequence with a G-C base-pair, the endonuclease activity of the IscB system comprising the modified IscB ωRNA may be increased in comparison to an otherwise identical control IscB system comprising a ωRNA without such a modification.
[0735] Wild type IscB. m18 scaffold sequence (SEQ ID NO: 37) , wild type IscB. m15 scaffold sequence (SEQ ID NO: 34) , wild type IscB. m1 scaffold sequence (SEQ ID NO: 20) , and wild type IscB. m8 scaffold sequence (SEQ ID NO: 27) (FIG. 9c-f) were selected to validate this deduction. Scaffold sequence variants SEQ ID NOs: 168-195 based on SEQ ID NO: 37, scaffold sequence variants SEQ ID NOs: 196-222 based on SEQ ID NO:34, scaffold sequence variants SEQ ID NOs: 223-232 based on SEQ ID NO: 20, and scaffold sequence variants SEQ ID NOs: 233-238 based on SEQ ID NO: 27 were generated and tested with respective IscB. m18, IscB. 15, IscB. m1, and IscB. m8 for endonuclease activity.
[0736] It was observed that significantly increased endonuclease activity was successfully achieved for all the four distinct IscB, for example, H1M4 (SEQ ID NO: 226) for IscB. m1 (FIG. 10a) , H8M6 (SEQ ID NO: 238) for IscB. m8 (FIG. 10b) , H15M10 (SEQ ID NO: 232) for IscB. m15 (FIG. 10c) , H18M2 (SEQ ID NO: 169) for IscB. m18 (FIG. 10d) , confirming the feasibility of the engineering strategy above.
[0737] Example 3: Engineering IscB protein to expand the targeting scope and enhance activity
[0738] Unless otherwise indicated, the scaffold sequence of the ωRNA used in combination with IscB. m16 or IscB. m16 mutant in Example 3 is SEQ ID NO: 54.
[0739] IscB. m16 (SEQ ID NO: 16) is composed of, from N-to C-terminus, PLMP domain (positions 1-54) , RuvC-I domain (positions 55-85) , Bridge Helix (positions 86-122) , Linker (positions 123-160) , RuvC-II domain (positions 161-196) , HNH domain (positions 197-297) , RuvC-III domain (positions 298-374) , P1D domain (positions 375-429) , and TID domain (positions 430-513) . The domain architecture of other IscB can be determined by sequence alignment with IscB. m16.
[0740] According to the predictive structural analysis of IscB. m16 (SEQ ID NO: 16) , the inventor performed an arginine scanning mutagenesis, where a single amino acid residue arginine (R) was introduced to substitute the original amino acid residue at one indicated position of IscB. m16 (SEQ ID NO: 16) (Table A2) . For example, A2R denotes IscB. m16-A2R mutant containing a single substitution of A2R relative to IscB. m16 (SEQ ID NO: 16) at position A2 of IscB. m16 (SEQ ID NO: 16) . Based on the activated EGFP fluorescence intensity of cells with an AAAGAA TAM reporter using the reporter system in Example 1, 141 of all 467 IscB. m16 mutants with single substitution with R exhibited improved endonuclease activity compared to IscB. m16 (SEQ ID NO: 16) using the same ωRNA scaffold sequence. Mutant IscB. m16-E326R showed the highest endonuclease activity among those IscB. m16 mutants containing a single substitution with R in RuvC I, RuvC II, or RuvC III domain (FIG. 3a) .
[0741] Table A2. Endonuclease activities (%EGFP+ / mCherry+ BFP+) (n=3 or 1) of IscB. m16 mutants
[0742] Table 2A (Con’)
[0743] Table 2A (Con’)
[0744] The inventor next screened 124 variants with single substitution with R in P1D and TID domain of IscB. m16 using six GFxxFP reporters with different TAMs to broaden TAM recognition (FIG. 11) . These reporters had the same target sequences but different 6-base TAMs (AAAGAA, CAAGAA, ACAGAA, AACGAA, AAAGCA, AAAGAC) . Compared to IscB. m16 (SEQ ID NO: 16) , 7 (seven) variants (each with single substitution M424R, T462R, N463R, T465R, Q475R, K478R, or I504R) (indicated by red arrows in FIG. 11) showed improved endonuclease activity and broader TAM recognition, as evidence by the increase (>1.05) of EGFP fluorescence intensity in all the six reporters relative to the IscB. m16 (FIG. 11) . Meanwhile, through predictive structural analysis of IscB. m16, the inventor has identified 11 potential sites associated with TAM recognition, which are H380, Q381, V433, T459, P460, I461, F467, Y468, R476, K478, and L481.
[0745] In order to broaden the TAM (in other words, achieving improved endonuclease activity for a given TAM for which the endonuclease activity used to be low) , the inventor conducted saturation mutagenesis at these 18 sites, including the 7 sites from P1D and TID screening (M424, T462, N463, T465, Q475, K478, and I504) and the 11 predicted sites (H380, Q381, V433, T459, P460, I461, F467, Y468, R476, K478, and L481) .
[0746] The inventor tested these saturation mutants using a TAM pool characterized by low activity. TAM pool 1 consisted of ACAGAA, AATGAA, AAACAA, AAAGCA, and AAAGAC, which were TAMs for which IscB. m16 showed relative low endonuclease activity (FIG. 12a-b) . Through similar fluorescence reporter system, the inventor found that various IscB. m16 mutants with a single substitution, especially P460S, T462H, T462L or T465V mutant, greatly enhanced the endonuclease activity for TAM pool 1 (FIG. 3b, FIG. 12c) .
[0747] To validate the enhanced endonuclease activity of these four IscB. m16 mutants, P460S, T462H, T462L and T465V, the inventor evaluated their TAM recognition ability with 16 reporters including TAMs NAAGAA, ANAGAA, AANGAA, AAAGNA and AAAGAN. The results demonstrated that the four mutants exhibited higher EGFP fluorescence intensity relative to IscB. m16, suggesting their superior endonuclease activity (FIG. 12d) .
[0748] Based on the results above, the inventor combined the single substitutions of E326R, P460S, T462H, T462L, and T465V in various ways, and obtained the best performing combination mutant of E326R + P460S + T462H, named as IscB. m16RSH (SEQ ID NO: 405) (FIG. 3c, FIG. 12e) . To test the TAM preference of IscB. m16RSH, the inventor detected EGFP activation using 64 TAM reporters with 5’-NNNGAA-3’ TAM, and IscB. m16RSH showed high endonuclease activity (FIG. 12f) .
[0749] Considering the characteristics of TAM recognition, the inventor designed three additional TAM pools, pool 2 (TTTGAA, TTGGAA, TCAGAA, CTAGAA and CTGGAA) and pool 3 (GTAGAA, GTTGAA, GTCGAA, GTGGAA and GCAGAA) , as well as pool 4 (ATAGAA, TGTGAA, CTCGAA, and GAGGAA) as a positive pool (FIG. 12f) . To further improve the endonuclease activity at more TAM sites, the inventor selected sites (T459, N463, Q475, L481 and I504) that had shown improved endonuclease activity for TAM pool 1. The inventor separately combined a single substitution at a site of T459, N463, Q475, L481 or I504 with IscB. m16RSH, and evaluated the resulting mutants using reporters with TAM pool 2, pool 3 or pool 4 (FIG. 3d; “3M” refers to IscB. m16RSH) . The inventor found that the combination of T459E with IscB. m16RSH, named as IscB. m16RESH (SEQ ID NO: 239) , exhibited increased editing efficiency at pool2 and pool3 reporters relative to IscB. m16RSH (FIG. 3d) . To assess the target range (TAM range) and endonuclease activity of IscB. m16RESH, the inventor used reporters with 64 NNNGAA and 16 AAAGNN TAMs and found that IscB. m16RESH (with ωRNA v2.27) exhibited significant improved endonuclease activity for these reporters, compared with that of IscB. m16 (with ωRNA v2.27) (FIG. 13) .
[0750] The inventor further optimized IscB. m16 scaffold sequence variant v2.27 (SEQ ID NO: 107) based on IscB. m16RESH, and found that variant v2.27-M21 (R1-Δ13, R5-Δ10, T24G, G25C, T57C, T79C, A117G) (SEQ ID NO: 242) showed significantly enhanced endonuclease activity compared to v2.27 (SEQ ID NO: 107) , hereafter named as “enωRNA” (FIG. 3e) . Then the IscB system composed of IscB. m16RESH and an ωRNA composed of a guide sequence and the scaffold sequence variant enωRNA was designated as “IscB. m16*” . Unless otherwise indicated, enωRNA was used in all the subsequent experiments.
[0751] SEQ ID NO: 242, IscB. m16 scaffold sequence variant v2.27-M21 (R1-Δ13, R5-Δ10, T24G, G25C, T57C, T79C, A117G) (hereafter named as “enωRNA” ) , wherein the g-c base pair at positions 57 and 117 (numbered according to WT IscB. m16 scaffold sequence of SEQ ID NO: 35) of IscB. m16 scaffold sequence variant v2.27 was further replaced with a c-g base pair (FIG. 22) .
[0752] The inventor explored the optimal guide sequence length (spacer length) using the GFxxFP fluorescence reporter, and found that IscB. m16*achieved maximum endonuclease activity with a 14-nt spacer length at two different targets (FIG. 14a) .
[0753] The inventor examined the indel efficiency of IscB. m16*at five endogenous loci in cultured HEK293T cells and found that IscB. m16*showed the highest endonuclease activity and the broadest range of deletion (FIG. 3f, FIG. 14b) compared to the other combinations of wild type IscB. m16 scaffold sequence or scaffold sequence variant enωRNA and wild type IscB. m16 or IscB. m16RESH mutant. Furthermore, TAM identification of IscB. m16*using bacterial depletion indicated that IscB. m16*recognized a TAM of 5’-NNNGNA-3’, while IscB. m16 recognized a TAM of 5’-MRNRAA-3’ (R denotes A or G) (FIG. 3g) .
[0754] Together, these results demonstrate that IscB. m16*exhibits both high endonuclease activity and highly flexible 5’-NNNGNA-3’ TAM recognition.
[0755] Example 4: IscB. m16*-mediated base editing in mammalian cells
[0756] The inventor constructed endonuclease-deficient IscB. m16 mutant D61A in RuvC-I, H248A in HNH domain, and D61A+H248A on the basis of IscB. m16 and IscB. m16*, respectively. The inventor tested their nickase activity using the dual target reporter according to previous study19. Consistently, IscB. m16*D61A system showed the highest nickase activity (IscB nickase) , and IscB. m16*D61A / H428A system showed substantially no activity (dead IscB) (FIG. 15) . The amino acid sequence of IscB. m16RESH-D61A (nickase) is set forth in SEQ ID NO: 240, and the amino acid sequence of IscB. m16RESH-D61A+H248A (dead) is set forth in SEQ ID NO: 241.
[0757] The inventor next fused IscB. m16RESH-D61A with TadA8eV106W (SEQ ID NO: 253) and NLS and linkers to generate IscB. m16*-ABE (SEQ ID NO: 259) , or with human APOBEC3AW104A (SEQ ID NO: 254) and 1xUGI domain (SEQ ID NO: 255) and NLS and linkers to generate IscB. m16*-CBE (SEQ ID NO: 260) 24, 25.
[0758] To comprehensively evaluate the base editing performance of IscB. m16*-ABE, the inventor designed dozens of TAM / PAM-matched endogenous loci for testing IscB. m16-ABE, IscB. m16*-ABE, enOgeuIscB-ABE19, and SpG-ABE22 (FIG. 4a-b) . The inventor found that the base editing window of IscB. m16*-ABE ranged from positions 1 to 10 (counting the TAM as positions 15-20) , while the optimal base editing occurred within positions 2-5 (FIG. 4c) . At these matched G-containing TAM / PAM sites in HEK293T cells, IscB. m16*-ABE showed significantly higher A-to-G base editing efficiency (46.15 ± 4.08%) than that of IscB. m16-ABE (9.19 ± 2.34%) and enOgeuIscB-ABE (31.34 ± 4.90%) , and comparable base editing efficiency to SpG-ABE (50.77 ± 4.13%) but with much smaller base editor size (FIG. 4a and 4d, FIG. 16) . In addition, the indel activity of IscB. m16*-ABE was similar to that of enOgeuIscB-ABE but lower than that of SpG-ABE (FIG. 4e, FIG. 17a) .
[0759] To characterize the TAM compatibility of IscB. m16*-ABE, the inventor further analyzed the base editing results and found that it showed A-to-G base editing at all different TAM sites, while enOgeuIscB-ABE showed substantially no base editing at some TAM sites of N3GCA, N3GGA, and N3GTA (FIG. 4f, FIG. 16) .
[0760] To further evaluate the specificity of IscB. m16*-ABE in HEK293T cells, the inventor conducted gRNA-dependent off-target DNA editing at predictive sites using CasOFFinder26, and gRNA-independent off-target DNA editing using the orthogonal R-loop assay27 at ALDH1A3-S1, VEGFA-S1, and EMX1-S2 target sites, respectively. Targeted deep sequencing analysis revealed that IscB. m16*-ABE exhibited similar gRNA-dependent off-target effects as enOgeuIscB-ABE and SpG-ABE at predicted off-target sites (FIG. 18) .
[0761] Using five previously reported SaCas9 target sites, the inventor observed that IscB. m16*-ABE showed comparably low gRNA-independent off-target events to enOgeuIscB-ABE and SpG-ABE (FIG. 4g-h, FIG. 19) .
[0762] On the other hand, IscB. m16*-CBE exhibited comparable base editing efficiency and indel efficiency with enOgeuIscB-CBE and SpG-CBE, with base editing efficiencies of 60.01 ± 8.08, 63.72 ± 5.33, and 68.06 ± 5.88, respectively (FIG. 4g-h, FIG. 17b-c) .
[0763] Collectively, these results indicate that the IscB. m16*-based base editors exhibit highly active base editing, broad target range, and low off-target effects in mammalian cells.
[0764] Example 5: IscB-mediated base editing restores the expression of dystrophin in mice
[0765] Taking advantage of its small size, the IscB. m16*-based base editors can be packaged with ωRNA into a single rAAV vector, making it a promising candidate for the treatment of certain genetic diseases, such as Duchenne muscular dystrophy (DMD) 28, 29. Previous study has shown that exon 50 skipping of the dystrophin gene can restore the dystrophin expression in a mouse model with an exon 51 deletion, a mutation occurring in nearly 8%of DMD patients30, 31. To access the therapeutical potential of IscB. m16* based base editing in DMD, the inventor devised a strategy to disrupt the splicing signal with IscB. m16*-CBE by converting the G within the splicing acceptor site ( ‘AG’ ) to other bases (A / C / T) , resulting in exon skipping (FIG. 5a) . The inventor first tested the IscB. m16*-CBE with ωRNA targeting the AG site adjacent to exon 50 in HEK293T cells. The inventor observed that IscB. m16*-CBE displayed approximate 25%conversion rate at position 10, which is the splicing acceptor site, while enOgeuIscB-CBE and SpG-CBE showed nearly no base editing activity at that position (FIG. 5b) . To conveniently package the IscB. m16*-based base editor into a single AAV, the inventor removed the UGI domain from the IscB. m16*-CBE and packaged it into AAV9 capsid and then detected base editing efficiency in mice. The IscB. m16*-CBE carried two versions of nuclear localization signal (NLS) patterns and was delivered to the muscle of mice with humanized exon 51 deletion (FIG. 5c) . IscB. m16*-CBE-v1 carried one copy of BpNLS (bpSV40 NLS) (SEQ ID NO: 256) at the N-terminal of the CBE and one copy of NpNLS (SEQ ID NO: 257) at the C-terminal of the CBE, and IscB. m16*-CBE-v2 carried two copies of BpNLS at the N-terminal of the CBE and one copy of NpNLS followed by one copy of BpNLS at the C-terminal of the CBE.
[0766] 3-weeks after injection, the inventor performed the base editing efficiency evaluation, western blot analysis, and histological staining for dystrophin expression. Targeted deep sequencing analysis showed that IscB. m16*-CBE-v2 achieved an approximate 7%of G-to-H (G-to-A, G-to-T, and G-to-C) conversion and up to 30%level of exon 51 skipping (FIG. 5d-e) . Western blotting and histological staining quantitative analysis of tibialis anterior (TA) muscle and immunostaining results indicated that IscB. m16*-CBE-v2 restored the dystrophin protein levels in myofibers to 40%of WT control (FIG. 5f-h) .
[0767] Together, these results indicate that IscB. m16*-based base editor, as a highly effective and broad-TAM miniature base editing tool, provides a promising approach for basic research and therapeutic applications.
[0768] Discussion
[0769] In summary, through computational mining of metagenomic sequence datasets, the inventor identified 19 natural IscB orthologs with various TAM recognition, and 10 of those IscBs showed endonuclease activity in mammalian cells, highlighting the diversity of IscB family. By examining the results of engineered ωRNAs, the inventor found that the truncation of the first (R1) and the last (e.g., R5 / R6) stem loops of ωRNA scaffold sequence usually enhanced the endonuclease activity of IscBs. By structure-guided design and protein engineering of P1D, TID, and RuvC domains of IscB, the inventor developed IscB. m16* system that exhibited remarkably high endonuclease activity and extended TAM scope of 5’-NNNGNA-3’, which is significantly broader than previously reported enIscB with 5’-NWRRNA-3’ TAM19. Furthermore, IscB. m16*-based base editors showed base editing efficiency comparable to SpG-BE but with much smaller size, and even higher than SpG-BE and enOgeuIscB-BE at some disease-related loci, such as DMD. Therefore, considering their compact size and extended editing scope, IscB. m16*-based base editors have high potential to be alternatives to enOgeuIscB-and Cas9-based base editors, especially for AAV based therapeutics applications.
[0770] Example 6: Evaluation of C-to-T base editing efficiency of mini IscB cytidine base editor (miCBE)
[0771] This Example demonstrates the C-to-T base editing efficiency of the miCBE of the disclosure to design optimal IscB-based CBE configuration.
[0772] Designs and constructions:
[0773] A miCBE expression plasmid and a guide RNA expression plasmid were constructed for the detection of the C-to-T base editing efficiency of the miCBE of the disclosure.
[0774] The miCBE expression plasmid comprised, from 5’ to 3’, CMV enhancer 1 (SEQ ID NO: 428) , CAG promoter (SEQ ID NO: 429) , hybrid intron (SEQ ID NO: 430) , a sequence encoding a miCBE based on OgeuIscB, a bGH polyA signal coding sequence (SEQ ID NO: 431) , CMV enhancer 2 (SEQ ID NO: 432) , CMV promoter (SEQ ID NO: 433) , and a mCherry (SEQ ID NO: 434) coding sequence indicative of successful transfection and expression of the miCBE expression plasmid.
[0775] The guide RNA expression plasmid comprised, from 5’ to 3’, U6 promoter (SEQ ID NO: 435) , a sequence encoding a EXM1-S1-targeting-or PCSK9-S4-targeting-guide RNA (SEQ ID NO: 436 or 439) composed of a EXM1-S1-targeting-or PCSK9-S4-targeting-guide sequence (SEQ ID NO: 437 or 440) and a scaffold sequence (SEQ ID NO: 442) 3’ to the guide sequence, CMV enhancer 2 (SEQ ID NO: 432) , CMV promoter (SEQ ID NO: 433) , and a EGFP (SEQ ID NO: 443) coding sequence indicative of successful transfection and expression of the guide RNA expression plasmid.
[0776] The structures of the miCBEs (miCBE-v1 to v10) tested in this Example are shown in FIG. 24. A3AW104A = APOBEC3A-W104A mutant. OgeuIscBD61A = OgeuIscB-D61A+E85R+H369R+S387R+S457R (IscB nickase) . The amino acid sequences of miABE-v1 to v10 fusions are set forth in SEQ ID NOs: 406-415, respectively.
[0777] Transfection and Detection:
[0778] HEK293T cells were cultured in 24-well tissue culture plates according to standard methods for 12 hours, before the miCBE expression plasmid and the guide RNA expression plasmid (synthesized by GenScript Co., Ltd. ) were co-transfected into the cells using standard polyethyleneimine (PEI) transfection. The transfected cells were then cultured at 37℃ under CO2 for 48 hours. mCherry and EGFP dual-positive cells were sorted from the cultured cells by flow cytometry.
[0779] About ten thousand sorted cells were lysed in 20 μL of lysis buffer with proteinase K (Vazyme Biotech) following the manufacturer’s manual. The genome sequence regions of interests were amplified with nested PCR by Phanta Max Super-Fidelity DNA Polymerase (Vazyme Biotech) . For targeted deep sequence analysis, PCR reactions were performed using primers with barcodes. The DNA products were purified with Gel extraction kit (Omega) and sequenced by 150-bp paired-end reads Illumina NovaSeq 6000 platform (Genewiz Co. Ltd. ) . The deep sequencing data were first de-multiplexed by Cutadapt (v. 2.8) based on sample barcodes. The de-multiplexed reads were then processed by CRISPResso2 for the quantification of editing efficiency, including C-to-T conversions at each target site.
[0780] Results:
[0781] Table 1 shows the base editing efficiency (%) of each tested miCBE at each indicated site for target EXM1-S1. Sites C9, C11, C12, and C13 refer to the cytidines at positions 9, 11, 12, and 13 of the EXM1-S1 protospacer sequence (SEQ ID NO: 438) , respectively. The data in Table 1 shows that all the tested miCBEs achieved significant C-to-T conversions at the indicated sites of EXM1 gene (n=3) . Untreated = blank HEK293 cells with no transfection of the two expression plasmids.
[0782] Table 1
[0783] Table 2 shows the base editing efficiency (%) of each tested miCBE at each indicated site for target PCSK9-S4. Sites C4, C5, C6, C8, C9, C10, and C11 refer to the cytidines at positions 4, 5, 6, 8, 9, 10, and 11 of the PCSK9-S4 protospacer sequence (SEQ ID NO: 441) , respectively. The data in Table 2 shows that all the tested miCBEs achieved significant C-to-T conversions at the indicated sites of PCSK9 gene (n=3) . Untreated = blank HEK293 cells with no transfection of the two expression plasmids.
[0784] Table 2
[0785] Overall, the above results demonstrate the C-to-T base editing activity of the miCBEs of the disclosure.
[0786] (1) Effect of number of UGI domain
[0787] The data of miCBE-v10, miCBE-v1, and miCBE-v9 in Tables 1 and 2 are rearranged into Tables 3 and 4, respectively, for direct comparison.
[0788] Table 3
[0789] Table 4
[0790] As noted, miCBE-v10, miCBE-v1, and miCBE-v9 are only difference in the number of UGI domain contained therein. miCBE-v10 contains one UGI domain, miCBE-v1 contains two UGI domains, and miCBE-v9 contains three UGI domains. As shown in Tables 3 and 4, for either EMX1 or PCSK9 targe gene, miCBE-v1 with two UGI domains achieved significantly improved base editing efficiency than that of miCBE-v9 with three UGI domains at all the indicated sites, and miCBE-v10 with just one UGI domain even achieved significantly improved base editing efficiency than that of miCBE-v1 with two UGI domains at all the indicated sites. It is thus concluded that one UGI domain is superior to two UGI domains for the miCBE of the disclosure, and two UGI domains are superior to three UGI domains for the miCBE of the disclosure.
[0791] Further and overall, it is observed that for EMX1 target gene, miCBE-v10 achieved a higher base editing efficiency than not only miCBE-v1 but all miCBE-v1 to v8 that contain two UGI domains at all the four sites. It is thus concluded that miCBE-v10 is superior to all the tested miCBEs with two UGI domains.
[0792] For PCSK9 target gene, miCBE-v10 achieved a higher base editing efficiency than not only miCBE-v1 but all miCBE-v1 to v8 that contain two UGI domains at sites C6, C10, and C11, and achieved a base editing efficiency comparable to the best two-UGI-containing miCBE at sites C4 (slightly lower than miCBE-v6 by about 3%) , C5 (slightly lower than miCBE-v5 by about 3%) , C8 (slightly lower than miCBE-v6 by about 1%) , and C9 (slightly lower than miCBE-v6 by about 1%) . Considering the benefit of saving one UGI domain (e.g., reduced tool size) , it is thus concluded that miCBE-v10 is superior to all the tested miCBEs with two UGI domains.
[0793] (2) Effect of cytidine deaminase domain -IscB nickase orientation
[0794] The data of miCBE-v1 and miCBE-v2 in Tables 1 and 2 are rearranged into Tables 5 and 6, respectively, for direct comparison. The position of UGI domains was constant across miCBE-v1 and miCBE-v2.
[0795] Table 5
[0796] Table 6
[0797] As noted, miCBE-v1 and miCBE-v3 are only difference in the orientation of cytidine deaminase domain and IscB nickase domain. miCBE-v1 has a N -cytidine deaminase domain -IscB nickase -C configuration, whereas miCBE-v3 has a N -IscB nickase -cytidine deaminase domain -C configuration. As shown in Tables 5 and 6, for EMX1 targe gene, miCBE-v1 achieved significantly improved base editing efficiency than that of miCBE-v3; and for PCSK9 targe gene, miCBE-v1 achieved comparable or improved base editing efficiency than that of miCBE-v3. It is thus concluded that N -cytidine deaminase domain -IscB nickase -C configuration is superior to N -IscB nickase -cytidine deaminase domain -C configuration for the miCBE of the disclosure.
[0798] Similarly, miCBE-v2 and miCBE-v4 are rearranged and directly compared in Tables 7 and 8, and miCBE-v5 and miCBE-v6 are rearranged and directly compared in Tables 9 and 10. All those results generally demonstrate that N -cytidine deaminase domain -IscB nickase -C configuration is superior to N -IscB nickase -cytidine deaminase domain -C configuration for the miCBE of the disclosure.
[0799] Table 7
[0800] Table 8
[0801] Table 9
[0802] Table 10
[0803] (3) Effect of position of UGI
[0804] The data of miCBE-v1, miCBE-v2, and miCBE-v6 in Tables 1 and 2 are rearranged into Tables 11 and 12, respectively, for direct comparison. The orientation of N -cytidine deaminase domain -IscB nickase -C was constant across miCBE-v1, miCBE-v2, and miCBE-v6, and the UGI domains were placed at the C-terminal, in the middle, and at the N-terminal of the base editor, respectively.
[0805] The comparison shows that for EMX1 target gene, the base editing efficiency with UGI domains at the C-terminal of miCBE is superior to the UGI domains either in the middle or at the N-terminal of miCBE; and for PCSK9 target gene, the base editing efficiency with UGI domains at the C-terminal of miCBE is superior to the UGI domains in the middle of miCBE, and the base editing efficiency with UGI domains at the N-terminal of miCBE is superior to the UGI domains at the C-terminal of miCBE. It is thus concluded that the position of UGI domains may be at either the N-terminal or C-terminal of miCBE to adapt to specific targets, and it is generally not desired to place UGI domains in the middle of miCBE between the cytidine deaminase domain and the IscB nickase.
[0806] Table 11
[0807] Table 12
[0808] (4) Effect of linker between cytidine deaminase domain and IscB nickase
[0809] The data of miCBE-v1, miCBE-v7, and miCBE-v8 in Tables 1 and 2 are rearranged into Tables 13 and 14, respectively, for direct comparison.
[0810] As noted, miCBE-v1, v7, and v8 are only difference in the linker between the cytidine deaminase domain and the IscB nickase, which is 16 aa XTEN linker (SEQ ID NO: 448) in v1, 21 aa GS linker (SEQ ID NO: 449) in v7, and 34 aa GS-bpNLS-GS linker (SEQ ID NO: 450) in v8.
[0811] The comparison shows that the XTEN linker achieved slightly higher base editing efficiency than the GS linker for EMX1 target gene at C9 and C11 and slightly lower base editing efficiency at C12 and C13; and achieved slightly higher base editing efficiency than the GS linker for PCSK9 target gene at all indicated sites. It is thus concluded that both the XTEN linker and the GS linker can be used for the miCBE of the disclosure, and generally speaking, the XTEN linker is preferred.
[0812] Furthermore, the comparison shows that the XTEN linker achieved higher base editing efficiency than the GS-bpNLS-GS linker for both EMX1 and PCSK9 target genes at all indicated sites. The comparison also shows that the GS linker achieved significantly higher base editing efficiency than the GS-bpNLS-GS linker for EMX1 target gene and slightly higher or lower base editing efficiency than the GS-bpNLS-GS linker for PCSK9 target gene. Overall speaking, the introduction of the additional bpNLS did not bring about a substantial advantage over the either the XTEN linker or the GS linker but undesirably increased the overall size of the miCBE.
[0813] Table 13
[0814] Table 14
[0815] Example 7: Evaluation of A-to-G base editing efficiency of mini IscB adenine base editor (miABE)
[0816] This Example demonstrates the A-to-G base editing efficiency of the miABE of the disclosure to design optimal IscB-based ABE configuration.
[0817] Designs and constructions:
[0818] A miABE expression plasmid and a guide RNA expression plasmid were constructed for the detection of the C-to-T base editing efficiency of the miABE of the disclosure.
[0819] The miABE expression plasmid comprised, from 5’ to 3’, CMV enhancer 1 (SEQ ID NO: 428) , CAG promoter (SEQ ID NO: 429) , hybrid intron (SEQ ID NO: 430) , a sequence encoding a miABE based on OgeuIscB, a bGH polyA signal coding sequence (SEQ ID NO: 431) , CMV enhancer 2 (SEQ ID NO: 432) , CMV promoter (SEQ ID NO: 433) , and a mCherry (SEQ ID NO: 434) coding sequence indicative of successful transfection and expression of the miABE expression plasmid.
[0820] The guide RNA expression plasmid comprised, from 5’ to 3’, U6 promoter (SEQ ID NO: 435) , a sequence encoding a VEGFa-S1-targeting-or VEGFa-S3-targeting-guide RNA (SEQ ID NO: 451 or 454) composed of a VEGFa-S1-targeting-or VEGFa-S3-targeting-guide sequence (SEQ ID NO: 452 or 455) and an OgeuIscB scaffold sequence (SEQ ID NO: 442) 3’ to the guide sequence, CMV enhancer 2 (SEQ ID NO: 432) , CMV promoter (SEQ ID NO: 433) , and a EGFP (SEQ ID NO: 443) coding sequence indicative of successful transfection and expression of the guide RNA expression plasmid.
[0821] The structures of the miABEs (miABE-v1 to v11) tested in this Example are shown in FIG. 25. OgeuIscBD61A = OgeuIscB-D61A+E85R+H369R+S387R+S457R (IscB nickase) . The amino acid sequences of miABE-v1 to v10 fusions are set forth in SEQ ID NOs: 416-426, respectively. The amino acid sequence of the SpCas9-based ABE, SpG-ABE, is set forth in SEQ ID NO: 427. The scaffold sequence of the guide RNA for the SpG-ABE is set forth in SEQ ID NO: 458. The amino acid sequence of the SpG Cas9-D10A nickase is set forth in SEQ ID NO: 457.
[0822] Transfection and Detection:
[0823] HEK293T cells were cultured in 24-well tissue culture plates according to standard methods for 12 hours, before the miABE expression plasmid and the guide RNA expression plasmid (synthesized by GenScript Co., Ltd. ) were co-transfected into the cells using standard polyethyleneimine (PEI) transfection. The transfected cells were then cultured at 37℃ under CO2 for 48 hours. mCherry and EGFP dual-positive cells were sorted from the cultured cells by flow cytometry.
[0824] About ten thousand sorted cells were lysed in 20 μL of lysis buffer with proteinase K (Vazyme Biotech) following the manufacturer’s manual. The genome sequence regions of interests were amplified with nested PCR by Phanta Max Super-Fidelity DNA Polymerase (Vazyme Biotech) . For targeted deep sequence analysis, PCR reactions were performed using primers with barcodes. The DNA products were purified with Gel extraction kit (Omega) and sequenced by 150-bp paired-end reads Illumina NovaSeq 6000 platform (Genewiz Co. Ltd. ) . The deep sequencing data were first de-multiplexed by Cutadapt (v. 2.8) based on sample barcodes. The de-multiplexed reads were then processed by CRISPResso2 for the quantification of editing efficiency, including C-to-T conversions at each target site.
[0825] Results:
[0826] Tables 15 and 16 shows the base editing efficiency (%) of each tested miABE at each indicated site for target VEGFA-S1. Sites A2, A3, A4, A6, A10, and A11 refer to the adenines at positions 2, 3, 4, 6, 10, and 11 of the VEGFA-S1 protospacer sequence (SEQ ID NO: 453) , respectively. The data in Tables 15 and 16 shows that all the tested miABEs achieved significant A-to-G conversions at the indicated sites of VEGFA gene (n=3) .
[0827] Table 15
[0828] Table 16
[0829] Tables 17 and 18 shows the base editing efficiency (%) of each tested miABE at each indicated site for target VEGFA-S3. Sites A3, A4, A5, A7, A9, and A11 refer to the adenines at positions 3, 4, 5, 7, 9, and 11 of the VEGFA-S3 protospacer sequence (SEQ ID NO: 456) , respectively. The data in Tables 17 and 18 shows that all the tested miABEs achieved significant A-to-G conversions at the indicated sites of VEGFA gene (n=3) .
[0830] Table 17
[0831] Table 18
[0832] Overall, the above results demonstrate the A-to-G base editing activity of the miABEs of the disclosure, and particularly, all of the miABEs of the disclosure achieved significantly improved A-to-G base editing than conventional SpCas9-based ABE (SpG-ABE) at adenines at or near the 3’ end of a protospacer sequence, for example, A10 and A11 of VEGFa-S1, and A9 and A11 of VEGFa-S3.
[0833] (1) Effect of adenine deaminase domain -IscB nickase orientation
[0834] The data of miABE-v1 and -v2 in Tables 15 and 17 are rearranged into Tables 19 and 20, respectively, for direct comparison.
[0835] Table 19. Base editing efficiency of miABE-v1 and -v2 and fold change of miABE-v2 relative to miABE-v1
[0836] Table 20. Base editing efficiency of miABE-v1 and -v2 and fold change of miABE-v2 relative to miABE-v1
[0837] As noted, miABE-v1 and -v2 are only different in the orientation of IscB nickase and the adenine deaminase domain (TadA8e-V106W) . As shown from the comparison, miABE-v1 achieved significantly improved base editing efficiency than miABE-v2 for both VEGFa-S1 and VEGFa-S3 targets at all indicated sites. It is thus concluded that the configuration of N-IscB nickase-adenine deaminase domain-C is superior to the configuration of N-adenine deaminase domain-IscB nickase-C for the miABEs of the disclosure.
[0838] (2) Effect of linker between IscB nickase and adenine deaminase domain
[0839] The data of miABE-v1, -v3, and -v4 in Tables 15 and 17 are rearranged into Tables 21 and 22, respectively, for direct comparison.
[0840] Table 21. Base editing efficiency of miABE-v1, v3, and -v4 and fold changes
[0841] Table 22. Base editing efficiency of miABE-v1, v3, and -v4 and fold changes
[0842] As noted, miABE-v1, v3, and -v4 are only different in the linker between the IscB nickase and the adenine deaminase domain (TadA8e-V106W) , which is 32 aa GS-XTEN-GS linker (SEQ ID NO: 459) in v1, 21 aa GS linker (SEQ ID NO: 449) in v3, and 34 aa GS-bpNLS-GS linker (SEQ ID NO: 450) in v4. As shown from the comparison, both miABE-v1 and -v4 achieved significantly improved base editing efficiency than miABE-v3 for both VEGFa-S1 and VEGFa-S3 targets. It is thus concluded that it would be advantageous to use either the GS-XTEN-GS linker (SEQ ID NO: 459) or the GS-bpNLS-GS linker (SEQ ID NO: 450) than the GS linker (SEQ ID NO: 449) for the miABEs of the disclosure. In addition, miABE-v1 achieved higher base editing efficiency than miABE-v4 for VEGFa-S3 target but lower base editing efficiency for VEGFa-S1 target, indicating that the selection of use of the GS-XTEN-GS linker (SEQ ID NO: 459) or the GS-bpNLS-GS linker (SEQ ID NO: 450) for the miABEs of the disclosure may depend on a target sequence.
[0843] Similarly, the data of miABE-v6 and v7 in Tables 16 and 18 are rearranged into Tables 23 and 24, respectively, for direct comparison. As noted, miABE-v6 and -v7 are only different in the linker between the N-terminal adenine deaminase domain (TadA8e-V106W) and the IscB nickase, which is 32 aa GS-XTEN-GS linker (SEQ ID NO: 459) in v6 and 34 aa GS-bpNLS-GS linker (SEQ ID NO: 450) in v7. As shown from the comparison, miABE-v6 achieved a higher or lower base editing efficiency than miABE-v7 depending on the target site of the target sequence. It was thus further confirmed that the selection of use of the GS-XTEN-GS linker (SEQ ID NO: 459) or the GS-bpNLS-GS linker (SEQ ID NO: 450) for the miABEs of the disclosure may depend on a target sequence.
[0844] Table 23. Base editing efficiency of miABE-v6 and -v7 and fold change of miABE-v7 relative to miABE-v6
[0845] Table 24. Base editing efficiency of miABE-v6 and -v7 and fold change of miABE-v7 relative to miABE-v6
[0846] (3) Effect of C-terminal NLS
[0847] The data of miABE-v4 and -v5 in Tables 15 and 17 are rearranged into Tables 25 and 26, respectively, for direct comparison.
[0848] Table 25. Base editing efficiency of miABE-v4 and -v5 and fold change of miABE-v5 relative to miABE-v4
[0849] Table 26. Base editing efficiency of miABE-v4 and -v5 and fold change of miABE-v5 relative to miABE-v4
[0850] As noted, miABE-v4 and -v5 are only different in the C-terminal NLS, which is npNLS (SEQ ID NO: 447) in v4 and bpNLS 1 (SEQ ID NO: 445) in v5. As shown from the comparison, miABE-v4 achieved higher base editing efficiency than miABE-v5 for VEGFa-S1 target but lower base editing efficiency for VEGFa-S3 target, indicating that the selection of use of the C-terminal npNLS (SEQ ID NO: 447) or the C-terminal bpNLS 1 (SEQ ID NO: 445) for the miABEs of the disclosure may depend on a target sequence.
[0851] (4) Effect of additional adenine deaminase domain
[0852] The data of miABE-v5 and -v6 in Tables 15 and 17 are rearranged into Tables 27 and 28, respectively, for direct comparison.
[0853] Table 27. Base editing efficiency of miABE-v5 and -v6 and fold change of miABE-v6 relative to miABE-v5
[0854] Table 28. Base editing efficiency of miABE-v5 and -v6 and fold change of miABE-v6 relative to miABE-v5
[0855] miABE-v6 was designed based on miABE-v5 by the introduction of an additional adenine deaminase domain (TadA8e-V106W, SEQ ID NO: 253) that was N-terminally fused to the IscB nickase via the GS-XTEN-GS linker (SEQ ID NO: 459) . As shown from the comparison, miABE-v6 with two adenine deaminase domain achieved significantly improved base editing efficiency than miABE-v5 with one adenine deaminase domain for both VEGFa-S1 and VEGFa-S3 targets at all indicated sites except for sites A2 and A3 of VEGFa-S1. Since the base editing efficiency of both miABE-v5 and -v6 at the two sites A2 and A3 of VEGFa-S1 was relatively lower than other sites, the two sites themselves might not be ideal target sites for base editing with the miABEs of the disclosure. It is thus still concluded that, for the purpose of high base editing efficiency for practical use, it would be advantageous to use two adenine deaminase domains than one adenine deaminase domain for the miABEs of the disclosure.
[0856] (5) Effect of position of adenine deaminase domains relative to IscB nickase
[0857] The data of miABE-v7 and -v10 in Tables 15 and 17 are rearranged into Tables 29 and 30, respectively, for direct comparison.
[0858] Table 29. Base editing efficiency of miABE-v7 and -v10 and fold change of miABE-v10 relative to miABE-v7
[0859] Table 30. Base editing efficiency of miABE-v7 and -v10 and fold change of miABE-v10 relative to miABE-v7
[0860] As noted, miABE-v7 and -v10 are only different in the position of the two adenine deaminase domains relative to the IscB nickase. In miABE-v7, the two adenine deaminase domains are separated by the IscB nickase, while in miABE-v10, the two adenine deaminase domains are in tandem on one side of the IscB nickase. As shown from the comparison, miABE-v10 achieved significantly improved base editing efficiency than miABE-v7 at A2, A3, A4, and A6 for VEGFa-S1 target and significantly improved base editing efficiency than miABE-v7 at A3, A4, A5, A7, and A9 for VEGFa-S3 target and comparable base editing efficiency at A11. It is thus generally concluded that it would be advantageous to use two adenine deaminase domains in tandem (not being separated by the IscB nickase) than two adenine deaminase domains separated by the IscB nickase for the miABEs of the disclosure.
[0861] (6) Effect of position of adenine deaminase domain in tandem
[0862] The data of miABE-v8 to v11 in Tables 15 and 17 are rearranged into Tables 31 and 32, respectively, for direct comparison.
[0863] Table 31. Base editing efficiency of miABE-v8 to v11 and fold changes
[0864] Table 32. Base editing efficiency of miABE-v8 to v11 and fold changes
[0865] As noted, miABE-v8 and -v9 are only different in the orientation of the IscB nickase and the two adenine deaminase domains (TadA8e-V106W) in tandem. As shown from the comparison, miABE-v8 achieved significantly improved base editing efficiency than miABE-v9 for both VEGFa-S1 and VEGFa-S3 targets at all indicated sites except for site A11 of VEGFa-S1 target.
[0866] Similarly, as noted, miABE-v10 and -v11 are only different in the orientation of the IscB nickase and the two adenine deaminase domains (TadA8e-V106W) in tandem. As shown from the comparison, miABE-v10 achieved significantly improved base editing efficiency than miABE-v11 for both VEGFa-S1 and VEGFa-S3 targets at all indicated sites.
[0867] It is thus concluded that the configuration of N-IscB nickase-adenine deaminase domains in tandem-C is superior to the configuration of N-adenine deaminase domains in tandem-IscB nickase-C for the miABEs of the disclosure, which is consistent with the conclusion in item (1) that the configuration of N-IscB nickase-adenine deaminase domain-C is superior to the configuration of N-adenine deaminase domain-IscB nickase-C.
[0868] (7) Effect of linker combination in the presence of two adenine deaminase domains
[0869] The data of miABE-v8 to v11 in Tables 15 and 17 are rearranged into Tables 33 and 34, respectively, for direct comparison.
[0870] Table 33. Base editing efficiency of miABE-v8 to v11 and fold changes
[0871] Table 34. Base editing efficiency of miABE-v8 to v11 and fold changes
[0872] As noted, miABE-v8 and -v10 are only different in the linker combination of the first linker between the IscB nickase and the two adenine deaminase domains (TadA8e-V106W) in tandem and the second linker between the two adenine deaminase domains (TadA8e-V106W) in tandem, which is two GS-XTEN-GS linkers (SEQ ID NO: 459) in miABE-v8 and two GS-bpNLS-GS linkers (SEQ ID NO: 450) in miABE-v10. As shown from the comparison, miABE-v10 achieved significantly improved base editing efficiency than miABE-v8 for both VEGFa-S1 and VEGFa-S3 targets at all indicated sites.
[0873] On the other hand, as noted, miABE-v9 and -v11 are only different in the linker combination of the first linker between the IscB nickase and the two adenine deaminase domains (TadA8e-V106W) in tandem and the second linker between the two adenine deaminase domains (TadA8e-V106W) in tandem, which is two GS-XTEN-GS linkers (SEQ ID NO: 459) in miABE-v9 and two the GS-bpNLS-GS linkers (SEQ ID NO: 35) in miABE-v11. As shown from the comparison, miABE-v9 achieved a higher or lower base editing efficiency than miABE-v11 depending on target site of the target sequence.
[0874] It is thus concluded that for the miABEs of the disclosure with two adenine deaminase domains in tandem at C-terminal (e.g., miABE-v8 and -v10) , it would be advantageous to use two GS-bpNLS-GS linkers (SEQ ID NO: 450) than two GS-XTEN-GS linkers (SEQ ID NO: 459) , while for the miABEs of the disclosure with two adenine deaminase domains in tandem at N-terminal (e.g., miABE-v9 and -v11) , either two GS-bpNLS-GS linkers (SEQ ID NO: 450) or two GS-XTEN-GS linkers (SEQ ID NO: 459) can be used.
[0875] References
[0876] 1. Mali, P. et al. RNA-Guided Human Genome Engineering via Cas9. Science 339, 823-826 (2013) .
[0877] 2. Zetsche, B. et al. Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System. Cell 163, 759-771 (2015) .
[0878] 3. Doudna, J.A. The promise and challenge of therapeutic genome editing. Nature 578, 229-236 (2020) .
[0879] 4. Anzalone, A.V., Koblan, L.W. & Liu, D.R. Genome editing with CRISPR–Cas nucleases, base editors, transposases and prime editors. Nature Biotechnology 38, 824-844 (2020) .
[0880] 5. Gaudelli, N.M. et al. Programmable base editing of A. T to G. C in genomic DNA without DNA cleavage. Nature 551, 464-+ (2017) .
[0881] 6. Komor, A.C., Kim, Y.B., Packer, M.S., Zuris, J.A. & Liu, D.R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-+ (2016) .
[0882] 7. Chen, S.Y. et al. Compact Cje3Cas9 for Efficient Genome Editing and Adenine Base Editing. Crispr J 5, 472-486 (2022) .
[0883] 8. Davis, J.R. et al. Efficient in vivo base editing via single adenoassociated viruses with size-optimized genomes encoding compact adenine base editors (Jul, 10.1038 / s41551-022-00911-4, 2022) . Nat Biomed Eng 6, 1317-1317 (2022) .
[0884] 9. Zhang, H. et al. Adenine Base Editing In Vivo with a Single Adeno-Associated Virus Vector. GEN Biotechnology 1, 285-299 (2022) .
[0885] 10. Wu, Z. et al. Programmed genome editing by a miniature CRISPR-Cas12f nuclease. Nature Chemical Biology 17, 1132-1138 (2021) .
[0886] 11. Kim, D.Y. et al. Efficient CRISPR editing with a hypercompact Cas12f1 and engineered guide RNAs delivered by adeno-associated virus. Nature Biotechnology 40, 94-102 (2021) .
[0887] 12. Kong, X. et al. Engineered CRISPR-OsCas12f1 and RhCas12f1 with robust activities and expanded target range for genome editing. Nature Communications 14 (2023) .
[0888] 13. Hino, T. et al. An AsCas12f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell 186, 4920-4935. e4923 (2023) .
[0889] 14. Xu, X. et al. Engineered miniature CRISPR-Cas system for mammalian genome regulation and editing. Molecular Cell 81, 4333-4345. e4334 (2021) .
[0890] 15. Karvelis, T. et al. Transposon-associated TnpB is a programmable RNA-guided DNA endonuclease. Nature 599, 692-696 (2021) .
[0891] 16. Kim, D.Y. et al. Hypercompact adenine base editors based on a Cas12f variant guided by engineered RNA. Nature Chemical Biology 18, 1005-1013 (2022) .
[0892] 17. Altae-Tran, H. et al. The widespread IS200 / IS605 transposon family encodes diverse programmable RNA-guided endonucleases. Science 374, 57-+ (2021) .
[0893] 18. Schuler, G., Hu, C.Y. & Ke, A.L. Structural basis for RNA-guided DNA cleavage by IscB-ωRNA and mechanistic comparison with Cas9. Science 376, 1476-+ (2022) .
[0894] 19. Han, D. et al. Development of miniature base editors using engineered IscB nickase. Nature Methods 20, 1029-1036 (2023) .
[0895] 20. Stewart, R.D. et al. Compendium of 4, 941 rumen metagenome-assembled genomes for rumen microbiome biology and enzyme discovery. Nature Biotechnology 37, 953-+ (2019) .
[0896] 21. Zhang, H. et al. An engineered xCas12i with high activity, high specificity and broad PAM range. Protein & Cell (2022) .
[0897] 22. Walton, R.T., Christie, K.A., Whittaker, M.N. & Kleinstiver, B.P. Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants. Science 368, 290-+ (2020) .
[0898] 23. Nishimasu, H. et al. Engineered CRISPR-Cas9 nuclease with expanded targeting space. Science 361, 1259-1262 (2018) .
[0899] 24. Richter, M.F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity (vol 15, pg 891, 2020) . Nature Biotechnology 38, 901-901 (2020) .
[0900] 25. Wang, X. et al. Cas12a Base Editors Induce Efficient and Specific Editing with Low DNA Damage Response. Cell Rep 31 (2020) .
[0901] 26. Bae, S., Park, J. & Kim, J.S. Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473-1475 (2014) .
[0902] 27. Doman, J.L., Raguram, A., Newby, G.A. & Liu, D.R. Evaluation and minimization of Cas9-independent off-target DNA editing by cytosine base editors. Nature Biotechnology 38, 620-+ (2020) .
[0903] 28. Min, Y.L., Bassel-Duby, R. & Olson, E.N. CRISPR Correction of Duchenne Muscular Dystrophy. Annu Rev Med 70, 239-255 (2019) .
[0904] 29. Olson, E.N. Toward the correction of muscular dystrophy by gene editing. P Natl Acad Sci USA 118 (2021) .
[0905] 30. Chemello, F. et al. Precise correction of Duchenne muscular dystrophy exon deletion mutations by base and prime editing. Sci Adv 7 (2021) .
[0906] 31. Yuan, J.J. et al. Genetic Modulation of RNA Splicing with a CRISPR-Guided Cytidine Deaminase. Molecular Cell 72, 380-+ (2018) .
[0907] 32. Hyatt, D. et al. Prodigal: prokaryotic gene recognition and translation initiation site identification. Bmc Bioinformatics 11 (2010) .
[0908] 33. Hofacker, I.L. Vienna RNA secondary structure server. Nucleic Acids Res 31, 3429-3431 (2003) .
[0909] 34. Lorenz, R. et al. ViennaRNA Package 2.0. Algorithm Mol Biol 6 (2011) .
[0910] 35. Kumar, S., Stecher, G., Li, M., Knyaz, C. & Tamura, K. MEGA X: Molecular Evolutionary Genetics Analysis across Computing Platforms. Mol Biol Evol 35, 1547-1549 (2018) .
[0911] 36. Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nature Biotechnology 37, 224-226 (2019) .
[0912] * * *
[0913] Various modifications and variations of the described products, methods, and uses of the disclosure will be apparent to those skilled in the art without departing from the scope and spirit of the disclosure. Although the disclosure has been described in connection with specific embodiments, it will be understood that it is capable of further modifications and that the disclosure as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the disclosure that are obvious to those skilled in the art are intended to be within the scope of the disclosure. This application is intended to cover any variations, uses, or adaptations of the disclosure following, in general, the principles of the disclosure and including such departures from the present disclosure come within known customary practice within the art to which the disclosure pertains and may be applied to the essential features herein before set forth.
Claims
1.An IscB polypeptide comprising an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 1-19.2.The IscB polypeptide of claim 1, wherein the IscB polypeptide is a mutant of SEQ ID NO: 16 and comprises an amino acid mutation (e.g., substitution) relative to (compared to) SEQ ID NO: 16 at a position selected from the group consisting of K30, K37, H38, N39, P50, V53, E74, E77, V79, H83, E85, K93, A96, K99, Q103, A104, H107, V133, E142, E159, P160, L172, T179, K180, E182, K196, E201, D204, H205, H214, L217, F218, E221, S222, D224, D225, Y228, A229, E232, G233, K234, G239, I250, H253, E254, A261, S268, G269, D272, L273, A278, A280, D282, K283, A285, V287, K292, K307, T310, A313, D318, E326, S328, F329, I330, S333, A335, P337, S350, V351, G354, S356, H357, G360, Q367, M369, H380, Q381, V384, K387, N391, G392, K393, H394, H400, K401, Q405, K406, G407, E411, L414, Q415, K416, N417, P418, G419, E423, M424, A426, E429, H430, K431, V433, K435, N438, L445, K449, N450, D451, V452, K456, T459, P460, I461, T462, N463, T465, F467, Y468, E473, G474, Q475, R476, H477, K478, L481, K483, P484, L486, H487, L494, G495, N496, G498, G499, Y500, P501, P502, Q503, I504, L505, G506, T507, H508, D509, K510, K511, and E513 of SEQ ID NO: 16; and optionally, wherein the amino acid mutation is a substitution with R, S, H, L, V, or E.3.The IscB polypeptide of claim 2, wherein the IscB polypeptide comprises an amino acid mutation (e.g., substitution) relative to (compared to) SEQ ID NO: 16 at a position selected from the group consisting of E326, H380, Q381, M424, V433, T459, P460, I461, T462, N463, T465, F467, Y468, Q475, R476, K478, L481, and I504 of SEQ ID NO: 16; and optionally, wherein the amino acid mutation is a substitution with R, S, H, L, V, or E.4.The IscB polypeptide of claim 3, wherein the IscB polypeptide comprises an amino acid substitution selected from the group consisting of E326R, T459E, P460S, and T462H, and optionally, an amino acid combination substitution of E326R + T459E + P460S + T462H.5.The IscB polypeptide of any preceding claim, wherein the IscB polypeptide comprises an amino acid mutation (e.g., substitution) at a position selected from the group consisting of D61, E193, and H248 of SEQ ID NO: 16; and optionally, wherein the amino acid mutation is a substitution with A.6.The IscB polypeptide of claim 5, wherein the IscB polypeptide comprises an amino acid combination substitution of D61A + E326R + T459E + P460S + T462H; or an amino acid combination substitution of D61A + H248A + E326R+ T459E + P460S + T462H.7.The IscB polypeptide of any preceding claim, wherein the IscB polypeptide (1) has endonuclease activity; (2) has nickase activity; or (3) is endonuclease deficient.8.The IscB polypeptide of any preceding claim, wherein the IscB polypeptide comprises, consists essentially of, or consists the amino acid sequence of SEQ ID NO: 239, 240, or 241, or an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 239, 240, or 241.9.A fusion protein comprising the IscB polypeptide of any preceding claim fused to a functional domain; optionally, wherein the functional domain is fused at the N-terminal or C-terminal of the IscB polypeptide, or fused internally with respect to the IscB polypeptide.10.The fusion protein of claim 9, wherein the functional domain is selected from the group consisting of a nuclear localization signal (NLS) , a nuclear export signal (NES) , a base editing domain, a deaminase or a catalytic domain thereof, a glycosylase or a catalytic domain thereof, an uracil glycosylase inhibitor (UGI) , an uracil glycosylase (UNG) (e.g., UNG1, UNG2) , a methylpurine glycosylase (MPG) , a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease or a catalytic domain thereof (e.g., T5 exonuclease of SEQ ID NO: 404) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having ssRNA cleavage activity, a moiety having dsRNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA, selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, and a catalytic domain thereof, and a functional fragment thereof, and any combination thereof.11.The fusion protein of claim 10, wherein the base editing domain is a deaminase (e.g., adenine deaminase, cytidine deaminase) or a catalytic domain thereof or a glycosylase (e.g., MPG, UNG2) or a catalytic domain thereof.12.The fusion protein of claim 9, wherein the fusion protein comprises, from N-to C-terminus, an adenine deaminase domain, an optional linker, the IscB polypeptide, an optional linker, and an adenine deaminase domain; and optionally, wherein the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 259.13.The fusion protein of claim 9, wherein the fusion protein comprises, from N-to C-terminus, a cytidine deaminase domain, an optional linker, the IscB polypeptide, an optional linker, and a UGI domain (e.g., one UGI domain) ; and optionally, wherein the fusion protein comprises an amino acid sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to SEQ ID NO: 260.14.A polynucleotide encoding the IscB polypeptide of any one of claims 1-8 or the fusion protein of any one of claims 9-13 (e.g., SEQ ID NOs: 39-57) .15.A system comprising:(1) the IscB polypeptide of any one of claims 1-8 or the fusion protein of any one of claims 9-13, or a polynucleotide encoding the IscB polypeptide or the fusion protein, and(2) a guide nucleic acid or a polynucleotide (e.g., a DNA, an RNA) encoding the guide nucleic acid, the guide nucleic acid comprising:(i) a scaffold sequence capable of forming a complex with the IscB polypeptide or the fusion protein; and(ii) a guide sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA, wherein the guide sequence is 5’ to the scaffold sequence.16.The system of claim 15, wherein the scaffold sequence has substantially the same secondary structure as the secondary structure of any one of SEQ ID NOs: 20-38, 58-238, and 242-252; or wherein the scaffold sequence comprises a polynucleotide sequence having a sequence identity of at least about 60% (e.g., at least about 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) to any one of SEQ ID NOs: 20-38, 58-238, and 242-252.17.The system of claim 15, wherein the guide sequence is about or at least about 14 nucleotides in length.18.A vector comprising the polynucleotide of claim 14; optionally, wherein the vector further comprises a guide nucleic acid as defined in claim 15 or 16 or a polynucleotide encoding the guide nucleic acid; optionally, wherein the vector is a plasmid vector, a recombinant AAV (rAAV) vector, a recombinant lentivirus vector, a RNP, or an LNP.19.A cell comprising the IscB polypeptide of any one of claims 1-8, the fusion protein of any one of claims 9-13, the polynucleotide of claim 14, the system of any one of claims 15-17, or the vector of claim 18.20.A method for modifying a target DNA, comprising contacting the target DNA with the system of any one of claims 15-17, wherein the guide sequence is capable of hybridizing to a target sequence of the target DNA, whereby the target DNA is modified.
Citation Information
Patent Citations
IS200 / IS60S transposon ISCB mutant protein and application thereof
CN116656649A
IscB fusion protein expression vector as well as construction method and application of IscB fusion protein expression vector in gene editing
CN117247938A
Reprogrammable ISCB nucleases and uses thereof
WO2022087494A1
Reprogrammable ISCB nucleases and uses thereof
WO2023097228A1
Use of ISCB in genome editing
WO2023215915A1