Deaminase variant, base editor comprising same, and use thereof
By mutation of deaminase variants at specific amino acid sites, the editing efficiency and specificity of the base editor are improved, and the problem of insufficient efficiency and specificity of existing base editors in genomic modification is solved.
Patent Information
- Application Number
- PCT/CN2025/079351
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-26
- Filing Date
- 2025-02-26
- Publication Date
- 2025-09-04
AI Technical Summary
When existing base editors modify the genome, there are limitations on editing efficiency and specificity, making it difficult to achieve efficient and stable single-base mutations.
A deaminase variant is provided that forms non-natural proteins by mutation at specific amino acid sites of wild-type deaminases, improving editing efficiency and specificity.
Enhance the editing efficiency and specificity of the base editor, and achieve more efficient genomic modification.
Smart Images

Figure PCTCN2025079351-FTAPPB-I100001 
Figure PCTCN2025079351-FTAPPB-I100002 
Figure PCTCN2025079351-FTAPPB-I100003
Abstract
Description
A deaminase variant, a base editor comprising the same, and applications thereof
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of and priority to patent application No. 202410212048.2, filed on February 26, 2024, entitled “A deaminase, a base editor comprising the same, and its application,” the entire contents of which, including any sequence listing and drawings, are incorporated herein by reference in their entirety.
[0003] References to electronic sequence listings
[0004] This disclosure contains an electronic sequence listing (created by the software "WIPOSequence" in accordance with WIPO Standard ST.26), which is incorporated herein by reference in its entirety. In accordance with WIPO Standard ST.26, the symbol "t" is used to represent both T in DNA and U in RNA. Therefore, in the sequence listing prepared in accordance with ST.26, in any case where the sequence is RNA, T in the sequence should be regarded as U. Technical Field
[0005] The present disclosure relates to the field of gene editing, and in particular, to a deaminase variant, a base editor comprising the same, and applications thereof. Background Art
[0006] How to accurately and efficiently modify the genome is a key research goal in the life sciences. Traditional CRISPR / Cas9 technology creates double-strand breaks (DSBs) at target sites, thereby inducing homologous recombination (HR) and non-homologous end joining (NHEJ) repair pathways within cells, thereby achieving site-specific modifications such as knockout, replacement, and insertion of genomic DNA. However, DSB-induced DNA repair struggles to achieve efficient and stable single-base mutations.
[0007] Currently available base editors include cytidine base editors that convert the target C·G base pair to T·A (e.g., BE4) and adenine base editors that convert A·T to G·C (e.g., ABE8e). For applications requiring higher editing efficiency, the editing efficiency of currently available base editors may be limited. This field requires base editors with higher specificity and editing efficiency, and it is necessary to improve base editors. Summary of the Invention
[0008] The main purpose of the present disclosure is to provide base editors with higher specificity and editing efficiency.
[0009] The first aspect of the present disclosure provides a deaminase variant, which is a non-natural protein and has a mutation at the core amino acid position of the wild-type deaminase corresponding to SEQ ID NO. 1 as shown below:
[0010] (a)A46;
[0011] (b)I47;
[0012] (c) T48;
[0013] (d) L49;
[0014] (e) V104;
[0015] (f)Q148;
[0016] (g) P150;
[0017] (h)E152;
[0018] (i) V153;
[0019] (j) F154; and
[0020] (k)N155.
[0021] In some embodiments, the deaminase variant has at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% sequence identity to SEQ ID NO. 1.
[0022] In some embodiments, the deaminase is adenine deaminase.
[0023] In some embodiments, the A at position 46 is mutated to C, V, L, I, S, T or G, preferably, to C or V, more preferably, to C.
[0024] In some embodiments, the I at position 47 is mutated to V, Y, L, M, A, F or G, preferably, to L, V or Y, more preferably, to V or Y.
[0025] In some embodiments, the T at position 48 is mutated to H, G, S, N, A or Q, preferably, to H or G.
[0026] In some embodiments, the L at position 49 is mutated to H, N, I, V, M, A, F or G, preferably, to H, N or I, more preferably, to H or N.
[0027] In some embodiments, the V at position 104 is mutated to M, I, L, F, A or G, preferably, to L or M, more preferably, to M.
[0028] In some embodiments, the Q at position 148 is mutated to R, N, S or T, preferably, mutated to R.
[0029] In some embodiments, the P at position 150 is mutated to L, R, A, C or M, preferably, to L or R.
[0030] In some embodiments, the E at position 152 is mutated to G, L, P, V or D, preferably, mutated to G, L, P or V.
[0031] In some embodiments, the V mutation at position 153 is A, C, F, G, I, L, or M, preferably selected from L, A, C, F, or G, and more preferably selected from A, C, F, or G.
[0032] In some embodiments, the F at position 154 is mutated to E, K, P, V, L, I, A, Y or W, preferably, to L, E, K, P or V, more preferably, to E, K, P or V.
[0033] In some embodiments, the N at position 155 is mutated to K, G, R, T, Q, H or S, preferably, to Q, K, G, R, T, more preferably, to K, G, R or T.
[0034] In some embodiments, the deaminase variant further comprises a mutation in a core amino acid position as shown below:
[0035] (i) D167;
[0036] (ii) R168;
[0037] (iii) A169; and
[0038] (iv)D170.
[0039] In some embodiments, the D at position 167 is mutated to A, K, G, I, Q, R, S, V or E, preferably, it is mutated to A, K, G, I, Q, R, S or V.
[0040] In some embodiments, the R at position 168 is mutated to D, F, K, G, V, Y, Q, N or H, preferably, to K, D, F, G, V or Y.
[0041] In some embodiments, the A at position 169 is mutated to C, K, G, R, S, V, L, I or T, preferably, to C, K, G, R, S or V.
[0042] In some embodiments, the D at position 170 is mutated to A, C, F, I, S, T, V or E, preferably, to A, C, F, I, S, T or V.
[0043] In some embodiments, the deaminase variant further comprises a mutation in a core amino acid position as shown below:
[0044] (i) Q66;
[0045] (ii) I67;
[0046] (iii) V68;
[0047] (iv) Q69;
[0048] (v) C139;
[0049] (vi) S140;
[0050] (vii) M142;
[0051] (viii) L163;
[0052] (ix) N164;
[0053] (x) Q165; and
[0054] (xi)P166.
[0055] In some embodiments, the Q at position 66 is mutated to K, N, R, S or T, preferably, to K, N or R.
[0056] In some embodiments, the I at position 67 is mutated to Q, R, L, V, M, A, F, G or I, preferably, to Q, R or L, more preferably, to Q or R.
[0057] In some embodiments, the V at position 68 is mutated to L, I, M, F, A, G or V, preferably, to L or I.
[0058] In some embodiments, the Q at position 69 is mutated to H, R, T, N or S, preferably, to H, R or T.
[0059] In some embodiments, the C at position 139 is mutated to A, Q, V, S, M or P, preferably, to A, Q or V.
[0060] In some embodiments, the S at position 140 is mutated to A, F, L, Y, T, S or N, preferably, to A, F, L or Y.
[0061] In some embodiments, the M at position 142 is mutated to L, F, I, C, P, M or V, preferably, mutated to L.
[0062] In some embodiments, the L at position 163 is mutated to E, I, V, M, A, F, G or L, preferably, to I or E, more preferably, to E.
[0063] In some embodiments, the N at position 164 is mutated to V, Q, H, K, R, S or T, preferably, to Q or V, more preferably, to V.
[0064] In some embodiments, the Q at position 165 is mutated to S, T or N, preferably, mutated to N.
[0065] In some embodiments, the P at position 166 is mutated to L, A, C or M, preferably, mutated to L.
[0066] In some embodiments, the deaminase variant further comprises a mutation in a core amino acid position as shown below:
[0067] (i) D167;
[0068] (ii) R168;
[0069] (iii) A169; and
[0070] (iv) D170; and mutations in one or more amino acid positions selected from the group consisting of:
[0071] (a)A156;
[0072] (b) E157;
[0073] (c) R158;
[0074] (d) S2;
[0075] (e) E3;
[0076] (f) L4;
[0077] (g)N5;
[0078] (h)A15;
[0079] (i) L16;
[0080] (j)Q18;
[0081] (k) K19;
[0082] (l)A20;
[0083] (m)R21;
[0084] (n)Q66;
[0085] (o)I67;
[0086] (p)V68;
[0087] (q)Q69;
[0088] (r)C139;
[0089] (s)S140;
[0090] (t)M142;
[0091] (u)E141;
[0092] (v)A111;
[0093] (w)A112;
[0094] (x)G113.
[0095] In some embodiments, the A at position 156 is mutated to H, P, V, L, I, G, S or T, preferably, to V, H, P, more preferably, to H or P.
[0096] In some embodiments, the E at position 157 is mutated to K, R or D, preferably, to K or R.
[0097] In some embodiments, the R at position 158 is mutated to M, N, K, Q or H, preferably, to K, M or N, preferably, to M or N.
[0098] In some embodiments, the S at position 2 is mutated to N, Q, K, G, R, T or V, preferably, to K, G, R, T or V.
[0099] In some embodiments, the E at position 3 is mutated to I, G, R, P, Q, V or D, preferably, to I, G, R, P, Q or V.
[0100] In some embodiments, the L at position 4 is mutated to A, G, P, T, V, Y, I, M or F, preferably, to I, A, G, P, T, V or Y, more preferably, to A, G, P, T, V, Y.
[0101] In some embodiments, the N at position 5 is mutated to A, D, M, G, R, P, Q, H or K, preferably, to Q, A, D, M, G, R or P, more preferably, to A, D, M, G, R or P.
[0102] In some embodiments, the A at position 15 is mutated to C, V, L, I, G, S or T, preferably, to C or V, more preferably, to C.
[0103] In some embodiments, the L at position 16 is mutated to A, I, V, M, F, G or L, preferably, to A or I, more preferably, to A.
[0104] In some embodiments, the Q at position 18 is mutated to C, M, H, R, N, S or T, preferably, to C, M, H or R.
[0105] In some embodiments, the K at position 19 is mutated to H, Q, R, Y or N, preferably, to H, Q, R or Y.
[0106] In some embodiments, the A at position 20 is mutated to S, T, V, L, I or G, preferably, to T, S, V, more preferably, to S or T.
[0107] In some embodiments, the R at position 21 is mutated to F, L, S, K, Q, N or H, preferably, to F, S, L or K, more preferably, to F, L or S.
[0108] In some embodiments, the Q at position 66 is mutated to S, T, K, N or R, preferably, mutated to K, N or R.
[0109] In some embodiments, the I at position 67 is mutated to G, Q, R, L, V, M, A or F, preferably, to Q, R or L, more preferably, to Q or R.
[0110] In some embodiments, the V at position 68 is mutated to L, I, M, F, A or G, preferably, to L or I.
[0111] In some embodiments, the Q at position 69 is mutated to S, H, R, T or N, preferably, to H, R or T.
[0112] In some embodiments, the C at position 139 is mutated to M, P, A, Q, V or S, preferably, to A, Q or V.
[0113] In some embodiments, the S at position 140 is mutated to A, F, L, Y or T, preferably, to A, F, L or Y.
[0114] In some embodiments, the M at position 142 is mutated to L, F, I, C or P, preferably, mutated to L.
[0115] In some embodiments, the E at position 141 is mutated to F, L, V or D, preferably, to F, L or V.
[0116] In some embodiments, the A at position 111 is mutated to G, V, L, I, S or T, preferably, to V or G, more preferably, to G.
[0117] In some embodiments, the A at position 112 is mutated to L, V, I, G, S or T, preferably, to L or V, and more preferably, to L.
[0118] In some embodiments, the G at position 113 is mutated to N, P, A, V, L, or I, preferably to N or A, and more preferably to N.
[0119] In some embodiments, the deaminase variant further comprises a mutation in the amino acid position shown below:
[0120] (a) C144;
[0121] (b) Q145; and
[0122] (c)Q149.
[0123] In some embodiments, the C at position 144 is mutated to W, S, M, or P, preferably, to W.
[0124] In some embodiments, the Q at position 145 is mutated to K, N, S, or T, preferably, to K.
[0125] In some embodiments, the Q at position 149 is mutated to R, N, S, or T, preferably, to R.
[0126] In some embodiments, the mutation is a mutation occurring in combination at the following positions of the amino acid sequence as shown in SEQ ID NO: 1: A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155;
[0127] Preferably, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0128] (1)A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0129] (4)S2+E3+L4+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0130] (7)A15+L16+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0131] (8)S2+E3+L4+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0132] (9)S2+E3+L4+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0133] (11)S2+E3+L4+N5+Q18+K19+A20+R21+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0134] (12) S2+E3+L4+N5+Q18+A20+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; or
[0135] (16)S2+E3+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0136] Preferably, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0137] (3)A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;
[0138] (5)A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170; or
[0139] (10)S2+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;
[0140] In some embodiments, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0141] A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;
[0142] Preferably, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0143] S2+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;
[0144] In some embodiments, the mutation is a mutation occurring in a combination of the following positions of the amino acid sequence as shown in SEQ ID NO: 1: A46+I47+T48+L49+Q66+Q69+V104+Q148+P150+E152+V153+F154+N155;
[0145] Preferably, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0146] (13)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148 +P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0147] <h2 style=";text-align:left;direction:ltr">(14)Q18+K19+R21+A46+I47+T48+L49+V104+C139+S140+E141+M142+Q148 +P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0148] <h2 style=";text-align:left;direction:ltr"> (15)A46+I47+T48+L49+Q66+V68+Q69+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0149] <h2 style=";text-align:left;direction:ltr"> (17)S2+E3+L4+N5+A46+I47+T48+L49+V104+C139+S140+E141+M142+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0150] <h2 style=";text-align:left;direction:ltr"> (18)A46+I47+T48+L49+Q66+I67+Q69+V104+C139+S140+E141+M142+Q148 +P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0151] <h2 style=";text-align:left;direction:ltr"> (19)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+L163+N164+Q165+P166;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0152] <h2 style=";text-align:left;direction:ltr"> (20)S2+E3+L4+N5+A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0153] <h2 style=";text-align:left;direction:ltr"> (21)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0154] (22)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;
[0155] (23)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;
[0156] (24)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+D167+R168+A16+D170;
[0157] (25)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;
[0158] (26)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+A111+A112+G113+C139+S140+M1 42+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0159] (27)Q18+K19+A20+R21+A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M 142+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170;
[0160] (28)Q18+K19+A20+R21+A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; or
[0161] (29)A46+I47+T48+L49+Q66+I67+V68+Q69+V104+C139+S140+M142+Q148+P150+E152+V153+F154+N155+D167+R168+A169+D170;
[0162] In some embodiments, the mutation is a mutation occurring in combination at the following positions of the amino acid sequence as shown in SEQ ID NO: 1: A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155;
[0163] Preferably, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0164] (2) S2+E3+L4+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+D167+R168+A169+D170; or
[0165] (6)A46+I47+T48+L49+V104+C144+Q145+Q148+Q149+P150+E152+V153+F154+N155;
[0166] In some specific embodiments, the mutation is a substitution occurring in combination at the following positions of the amino acid sequence as shown in SEQ ID NO: 1:
[0167] A46C+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F; or
[0168] A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0169] Furthermore, preferably, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0170] (8)A15C+L16A+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0171] (9)S2K+E3G+L4P+N5M+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0172] (10)S2G+E3I+L4Y+N5G+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0173] (12)S2R+E3R+L4G+N5A+Q18C+K19H+A20T+R21F+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0174] (13)S2R+E3R+L4G+N5A+Q18R+A20T+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0175] (18)S2V+E3P+N5D+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0176] (2)S2G+E3I+L4Y+N5G+A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152V+V153C+F154E+N155G+A156P+D167K+R168K+A169R+D170C;
[0177] (3)A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152V+V153G+F154P+N155G+D167K+R168K+A169S+D170C;
[0178] (4)S2G+E3G+L4A+N5R+A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0179] (5)A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152G+V153G+F154K+N155G+D167K+R168K+A169S+D170C;
[0180] (6)A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152L+V153A+F154P+N155K+A156H+E157K+R158N+D167G+R168G+A169K+D170V;
[0181] (7)A46C+I47Y+T48G+L49H+V104M+C144W+Q145K+Q148R+Q149R+P150R+E152P+V153F+F154V+N155T;
[0182] (11)S2R+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152V+V153G+F154P+N155G+D167V+R168G+A169R+D170F;
[0183] (14)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0184] (15)Q18H+K19Y+R21L+A46C+I47V+T48H+L49N+V104M+C139V+S140L+E141F+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0185] (16)A46C+I47Y+T48G+L49H+Q66N+V68I+Q69T+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0186] (17)A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0187] (19)S2T+E3V+L4T+N5P+A46C+I47V+T48H+L49N+V104M+C139Q+S140A+E141V+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0188] (20)A46C+I47V+T48H+L49N+Q66R+I67Q+Q69R+V104M+C139A+S140F+E141L+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0189] (21)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+L163E+N164V+Q165N+P166L;
[0190] (22)S2V+E3Q+L4V+N5R+A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F 154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0191] (23)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167A+R168G+A169V+D170S;
[0192] (24)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167S+R168G+A169S+D170T;
[0193] (25)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167I+R168Y+A169G+D170T;
[0194] (26)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167V+R168D+A169V+D170I;
[0195] (27)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M1 42L+Q148R+P150L+E152L+V153A+F154P+N155G+D167Q+R168V+A169C+D170A;
[0196] (28)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+A111G+A112L+G113N+C139A+S140Y+M1 42L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0197] (29)Q18M+K19R+A20S+R21S+A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M 142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0198] (30)Q18R+K19Q+A20S+R21F+A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F; or
[0199] (31)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M1 42L+Q148R+P150L+E152L+V153A+F154P+N155G+D167R+R168F+A169G+D170T.
[0200] In some embodiments, the mutation is selected from the group consisting of S2K, S2G, S2R, S2T, S2V, E3G, E3Q, E3P, E3R, E3I, E3V, L4A, L4G, L4T, L4P, L4V, L4Y, N5A, N5D, N5M, N5G, N5P, N5R, A15C, L16A, A16V, Q18C, Q18H, Q18M, Q18R, K19H, K19R, K19Q, K19Y, A20S, A20T, R21F, R21L, R21S, A46C, I47V, I47Y, T48H, T48G, L49H, L49N, Q66K, Q66N, Q66R, I67Q, I67R, V68L, V68I, Q69H, Q69R, Q69T, V104M, A111G, A112L, G113N, C139A, C139Q, C139V, S140A, S140F, S140L, S140Y, E141F, E141L, E141V, M1 42L, C144W, Q145K, Q148R, Q149R, P150L, P150R, E152G, E152L, E152P, E152V, V153A, V153C, V153F, V153G, F1 54E, F154K, F154P, F154V, N155G, N155K, N155R, N155T, A156H, A156P, E157K, E157R, R158M, R158N, L163E, N16 4V, Q165N, P166L, D167A, D167K, D167G, D167Q, D167I, D167R, D167S, D167V, R168D, R168F, R168K, R168G, R168V, R168Y, A169C, A169K, A169G, A169R, A169S, A169V, D170A, D170C, D170F, D170I, D170S, D170T, D170V, or a combination thereof.
[0201] In some embodiments, the amino acid sequence of the deaminase variant is identical or substantially identical to that of the wild-type deaminase except for the mutation site.
[0202] In certain embodiments, the substantial identity is that there are at most 50 (preferably 1-20, more preferably 1-10, and more preferably 1-5) amino acid differences, wherein the differences include amino acid substitutions, deletions, or additions, and the gene editing activity of the deaminase variant is improved.
[0203] In certain embodiments, the deaminase variant is at least 80% homologous to the wild-type deaminase, preferably at least 85% or 90%, more preferably at least 95%, and most preferably at least 98% or 99%.
[0204] In certain embodiments, the variant is formed by mutation of the wild-type deaminase.
[0205] In yet another aspect, the present disclosure provides a fusion protein comprising the variant described in the present disclosure; and a nucleic acid programmable nucleotide binding domain.
[0206] In some embodiments, the nucleic acid programmable nucleotide binding domain is a Cas protein or an Ago protein.
[0207] In some embodiments, the Cas protein includes at least one of a type II CRISPR-Cas polypeptide, a type I CRISPR-Cas polypeptide, a type III CRISPR-Cas polypeptide, a type IV CRISPR-Cas polypeptide, a type V CRISPR-Cas polypeptide, a type VI CRISPR-Cas polypeptide, a type VII CRISPR-Cas polypeptide, an IscB polypeptide, a TnpB polypeptide, and an IsrB polypeptide.
[0208] In some embodiments, type I CRISPR-Cas polypeptides include I-A, I-B, I-C, I-D, I-E, and I-F CRISPR-Cas proteins.
[0209] In some embodiments, the V-type CRISPR-Cas polypeptide includes a Cas12 protein.
[0210] In some embodiments, the Type VI CRISPR-Cas polypeptide comprises a Cas13 protein.
[0211] In some embodiments, the Cas protein is selected from Cas9, CasX, CasY, Cas12a (Cpf1), Cas12b (C2cl), Cas13a (C2c2), Cas12c (C2c3), Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, or a combination thereof.
[0212] In some embodiments, the Cas protein is selected from Cas9 protein (e.g., SpCas9, SaCas9, GeoCas9, CjCas9, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9, xCas9, SpCas9-NG); Cas12 protein (e.g., Cas12a, AsCas12a, LbCas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f (Cas14), Cas12g, Cas12h, Cas12i, xCas12i, Cas12Max, hfCas12Max, Cas12j, Cas12k, Cas12l, Cas12m, Cas12n, Cas12o, Cas12p, Cas12q, Cas12r, Cas12s, Cas12t, Cas12u, Cas12v, Cas12w, Cas12x, Cas12y, Cas12z); Cas13 proteins (e.g., Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, Cas13x, Cas13y); Csn2; and mutants thereof.
[0213] In some embodiments, the Cas protein comprises an amino acid sequence having at least 95%, 96%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs. 41-45.
[0214] In some embodiments, the Cas protein is a nickase (nCas) comprising amino acids having at least 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO. 54.
[0215] In some embodiments, the Cas protein is a dCas protein, which comprises an amino acid sequence having at least 95%, 96%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs. 46-53 and 57.
[0216] In some embodiments, the Cas protein is a dCas protein comprising an amino acid sequence of SEQ ID NO. 57 having at least 95%, 96%, 97%, 98%, or 99% sequence identity.
[0217] In some embodiments, the Cas protein is a dCas protein, which comprises the amino acid sequence shown in SEQ ID NO.57.
[0218] In some embodiments, the AGO protein is selected from the group consisting of pAgo, eAgo, Ago1, Ago2, Ago3, and Ago4.
[0219] In some embodiments, the fusion protein has the following structure from N-terminus to C-terminus:
[0220] Z1-Z2(I); or
[0221] Z2-Z1(II); or
[0222] Z3-Z1-Z4(III);
[0223] wherein Z1 is a deaminase variant described in the present disclosure;
[0224] Z2 is a nucleic acid programmable nucleotide binding domain;
[0225] Z3 is the N-terminal fragment of the nucleic acid programmable nucleotide binding domain;
[0226] Z4 is the C-terminal fragment of the nucleic acid programmable nucleotide binding domain;
[0227] Furthermore, each "-" is independently a bond or a linker.
[0228] In some embodiments, the variant is operably linked to one end of the nucleic acid programmable nucleotide binding domain or is embedded in the nucleic acid programmable nucleotide binding domain.
[0229] In some preferred embodiments, the operably linked is linked via a linker.
[0230] In some embodiments, the linker preferably comprises an amino acid sequence as shown in one or more of SEQ ID NOs: 3-12.
[0231] In some embodiments, the chimeric site is located in the carboxy-terminal domain of the nucleic acid programmable nucleotide binding domain.
[0232] In some embodiments, the nucleic acid programmable nucleotide binding domain retains partial or no nucleotide chain cleavage activity.
[0233] In some embodiments, the fusion protein further comprises a nuclear localization signal sequence; the nuclear localization signal sequence is linked to the N-terminus and / or C-terminus of the fusion protein, and / or is linked to the N-terminus and / or C-terminus of the variant.
[0234] In some embodiments, the nuclear localization signal sequence is linked to the N-terminus and the C-terminus of the fusion protein.
[0235] In some embodiments, the structure of the fusion protein from N-terminus to C-terminus is: optional nuclear localization signal sequence-variant-nucleic acid programmable nucleotide binding domain-optional nuclear localization signal sequence.
[0236] In some embodiments, when the nucleic acid programmable nucleotide binding domain is a Cas protein, such as a Cas9 protein (e.g., spCas9), the chimeric site is located between positions 1249-1250 of Cas9.
[0237] In some embodiments, the fusion protein comprises an amino acid sequence as shown in any one of SEQ ID NOs: 23-26.
[0238] In another aspect, the present disclosure provides a base editing system comprising:
[0239] (i) variants and nucleic acid programmable nucleotide binding domains as described herein;
[0240] or (ii) a fusion protein as described herein;
[0241] and guide RNA.
[0242] In some embodiments, the nucleic acid programmable nucleotide binding domain or the fusion protein forms a ribonucleoprotein complex with the guide RNA and binds to the target nucleic acid under the guidance of the guide RNA.
[0243] In another aspect, the present disclosure provides an isolated polynucleotide encoding the variant described in the present disclosure, the fusion protein described in the present disclosure, or the base editing system described in the present disclosure.
[0244] In certain embodiments, the isolated nucleotide comprises a humanization-optimized sequence.
[0245] In certain embodiments, the polynucleotide further comprises auxiliary elements flanking the ORF of the variant selected from the group consisting of a signal peptide, a secretory peptide, a tag sequence (such as 6His), or a combination thereof.
[0246] In certain embodiments, the polynucleotide is selected from the group consisting of a genomic sequence, a cDNA sequence, an RNA sequence, or a combination thereof.
[0247] In certain embodiments, the polynucleotide further comprises a promoter operably linked to the ORF sequence of the variant.
[0248] In certain embodiments, the promoter is selected from the group consisting of a constitutive promoter, a tissue-specific promoter, an inducible promoter, or a strong promoter.
[0249] In certain embodiments, the polynucleotide is codon-optimized according to the codon preference of the host cell.
[0250] In certain embodiments, the host cell comprises a prokaryotic cell or a eukaryotic cell.
[0251] In certain embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).
[0252] In certain embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.
[0253] In certain embodiments, the yeast cell is selected from one or more yeasts of the following sources: Pichia pastoris, Kluyveromyces, or a combination thereof; preferably, the yeast cell includes: Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.
[0254] In certain embodiments, the host cell is selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.
[0255] In some embodiments, the polynucleotide encoding the fusion protein comprises a nucleotide sequence as shown in any one of SEQ ID NOs: 27-30.
[0256] In yet another aspect, the present disclosure provides a vector comprising the polynucleotide described in the present disclosure.
[0257] In some embodiments, the polynucleotide is located on one or more vectors.
[0258] In certain embodiments, the vector comprises one or more promoters operably linked to the nucleic acid sequence, enhancer, transcription termination signal, polyadenylation sequence, origin of replication, selectable marker, nucleic acid restriction site, and / or homologous recombination site.
[0259] In some embodiments, the promoter is selected from one or more of a constitutive promoter, an inducible promoter, a ubiquitin promoter, a cell type-specific promoter, and a tissue-specific promoter.
[0260] In certain embodiments, the vector comprises a plasmid or a viral vector.
[0261] In certain embodiments, the viral vector is selected from the group consisting of adeno-associated virus (AAV), adenovirus, lentivirus, retrovirus, herpes virus, SV40, poxvirus, or a combination thereof.
[0262] In certain embodiments, the vector includes a cloning vector, a transformation vector, an expression vector, a shuttle vector, an integration vector, and a multifunctional vector.
[0263] In another aspect, the present disclosure provides a delivery composition comprising a delivery vector and one or more selected from the following: the variant described in the present disclosure, the fusion protein described in the present disclosure, the system described in the present disclosure, the polynucleotide described in the present disclosure, and the vector described in the present disclosure.
[0264] In certain embodiments, the delivery vehicle is a particle.
[0265] In certain embodiments, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, nanoparticles, cell-penetrating peptides, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).
[0266] In some embodiments, the composition further comprises a nucleic acid programmable nucleotide binding domain.
[0267] In some embodiments, the composition further comprises one or more lipid moieties selected from the group consisting of ionizable lipids, neutral lipids, PEG lipids, steroids or derivatives thereof, or combinations thereof.
[0268] In some embodiments, the ionizable lipid comprises a compound represented by Formula I (described in PCT application number PCT / CN2023 / 106421):
[0269] in,
[0270] R1 is selected from -OH and R a (R b )N-, where R a and R b are independently hydrogen, C1-C 10 Alkyl or C1-C 10 alkyl halide;
[0271] R2 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0272] R3 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl or R c -(CH2)n-, wherein n is a positive integer of 1-20, preferably a positive integer of 1-14, more preferably a positive integer of 1-10;
[0273] R c is selected from the following structures: in Represents the connection key;
[0274] L1, L2, L3, L4, L5 are none or are independently selected from the following groups:
[0275] X is none or -CH- or N;
[0276] R4 is none or C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0277] R5 is none or C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0278] R6 is C1-C 30 Alkyl, C2-C 30 Alkenyl, C2-C 30 Alkynyl;
[0279] R7 is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl;
[0280] R8 and R9 are none or independently C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl or -R h -C1-C 14 Alkyl, -R h -C2-C 14 Alkenyl, -R h -C2-C 14 Alkynyl, where R h O or S;
[0281] R 10 、R 11 Each independently is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl, or -R h -C1-C 14 Alkyl, -R h -C2-C 14 Alkenyl, -R h-C2-C 14 Alkynyl, where R h is O or S. In some embodiments, the neutral lipid includes any lipid molecule disclosed or undisclosed that exists in an uncharged form or a neutral zwitterionic form at a selected pH value or range. The selected useful pH value or range corresponds to the pH conditions of the environment in which the lipid is intended to be used, such as physiological pH.
[0282] In some embodiments, the neutral lipid comprises phosphatidylcholine (PC), phosphatidylethanolamine (PE), phosphatidylserine (PS), phosphatidic acid (PA), or phosphatidylglycerol (PG).
[0283] As non-limiting examples, neutral lipids that can be used in conjunction with the present disclosure include, but are not limited to, phosphatidylcholines, such as 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), 1,2-dipalmitoyl-sn-glycero-3-phosphocholine (DPPC), 1,2-dimyristoyl-sn-glycero-3-phosphocholine (DMPC), 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC), 1,2-dioleoyl-sn-glycero-3-phosphocholine (DOPC); phosphatidylethanolamines, such as 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 2-((2,3-bis(oleoyloxy)propyl))dimethylammonio)ethyl hydrogenphosphate (DOCP); sphingomyelin (SM); ceramides; steroids, such as sterols and their derivatives. The neutral lipids provided herein can be synthetic or derived from (isolated or modified from) natural sources or compounds.
[0284] In some embodiments, exemplary phospholipids that can form part of the nanoparticle compositions of the present disclosure include, but are not limited to, 1,2-dioleoyl-sn-glycero-3-phosphatidylethanolamine (DOPE), 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), 1,2-dipalmitoyl-sn-glycero-3-phosphatidylcholine (DPPC), 1,2-dioleoyl-sn-glycero-3-phosphatidylcholine (DOPC), dipalmitoylphosphatidylglycerol (DPPG), and 1,2-dioleoyl-sn-glycero-3-phosphatidylcholine (DOPC). ), oleoylphosphatidylcholine (POPC), 1-palmitoyl-2-oleoylphosphatidylethanolamine (POPE), 1,2-dipalmitoyl-sn-glycero-3-phosphoethanolamine (DPPE), 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine (DMPE), distearoylphosphatidylethanolamine (DSPE) and 1-stearoyl-2-oleoylphosphatidylethanolamine (SOPE), or lipids modified with anionic or cationic modifying groups.
[0285] In some embodiments, the PEG lipids include 1,2-dimyristoyl-sn-glyceromethoxy-polyethylene glycol (PEG-DMG), dimyristoylglycerol-polyethylene glycol (PEG-c-DMG), polyethylene glycol-dimyristoylglycerol (PEG-C14), PEG-1,2-dimyristoyloxypropyl-3-amine (PEG-c-DMA), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[amino(polyethylene glycol)] (PEG-DSPE), pegylated phosphatidylethanolamine (PEG-PE), PEG-modified ceramides, PEG-modified dialkylamines, PEG-modified diacylglycerols, Tween-20, Tween-80, 1,2-dipalmityl-sn-glycerol-methoxypolyethylene glycol PEG-DPG, 4-O-(2',3'-di(tetradecanoyloxy)propyl-1-O-(ω-methoxy(polyethoxy)ethyl)succinate (PEG-s-DMG), PEG-dialkoxypropyl (PEG-DAA), mPEG2000-1,2-di-O-alkyl-sn3-carbamoylglycerol ester (PEG-c-DOMG) and N-acetylgalactosamine ((R)-2,3-bis(octadecyloxy)propyl-1-(methoxypoly(ethylene glycol) 2000)propylcarbamate)) (GalNAc-PEG-DSG) or a combination of two or more thereof.
[0286] In some embodiments, the compositions may include one or more structural lipids. Without being bound by theory, it is expected that structural lipids can stabilize the amphiphilic structure of nanoparticles, such as, but not limited to, the lipid bilayer structure of nanoparticles. Exemplary structural lipids that can be used in conjunction with the present disclosure include, but are not limited to, cholesterol, coprosterol, sitosterol, ergosterol, campesterol, stigmasterol, brassicasterol, tomatine, tomatine, ursolic acid, alpha-tocopherol, and mixtures thereof. In certain embodiments, the structural lipid is cholesterol. In some embodiments, the structural lipid includes cholesterol and corticosteroids (such as prednisolone, dexamethasone, prednisone, and hydrocortisone) or a combination thereof.
[0287] In some embodiments, the composition may further include anionic lipids, including one or a combination of two or more of phosphatidylserine, phosphatidylinositol, phosphatidic acid, phosphatidylglycerol, dioleoylphosphatidylglycerol DOPG, 1,2-dioleoyl-sn-glycero-3-phosphatidylserine DOPS, and dimyristoylphosphatidylglycerol.
[0288] In some embodiments, the composition comprises ionizable lipids, anionic lipids, neutral lipids, structural lipids, and PEG lipids in a molar ratio of (20-65):(0-20):(5-25):(25-55):(0.3-15).
[0289] Exemplarily, the above molar ratio can be 20:20:5:50:5, 30:5:25:30:10, 20:5:5:55:15, 65:0:9.7:25:0.3, etc.; wherein, the molar ratio of the compound (compound shown in Formula I) or its pharmaceutically acceptable salt, stereoisomer, tautomer, solvate, chelate, non-covalent complex or prodrug in the ionizable lipid and other cationic or ionizable lipids is (1 to 10): (0 to 10); exemplarily, the molar ratio can be 1:1, 1:2, 1:5, 1:7.5, 1:10, 2:1, 5:1, 7.5:1, 10:1, etc.
[0290] In some embodiments, the molar ratio of ionizable lipids, anionic lipids, neutral lipids, structured lipids and polymer-bound lipids is (20-55):(0-13):(5-25):(25-51.5):(0.5-15); wherein, the molar ratio of the compound (compound shown in Formula I) or its pharmaceutically acceptable form such as salt, stereoisomer, tautomer, solvate, chelate, non-covalent complex or prodrug in the ionizable lipid and other cationic or ionizable lipids is (3-4):(0-5).
[0291] In some embodiments, the composition includes ionizable lipids, neutral lipids, structured lipids, and polymer-bound lipids in a molar ratio of 20-55:5-25:25-55:0.5-15.
[0292] In some embodiments, the substructures of the ionizable lipids are as shown in the following formulas I-1, I-2, I-3, and I-4, respectively:
[0293] In formula I-1, R1-R2, R6-R 11 ,L2-L5 are defined as above;
[0294] R3 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl.
[0295] In formula I-2, R1-R2, R6-R7, R 10 -R 11 , the definitions of L1-L5 are as described above;
[0296] R3 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, or R c -(CH2)n-, wherein n is a positive integer of 1-20, preferably a positive integer of 1-14, more preferably a positive integer of 1-10, R c The definition of is as above;
[0297] R5 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0298] R8 and R9 are each independently C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl.
[0299] In formula I-3, R1-R3, R6, R7, R 10 、R 11 , the definitions of L1-L5 are as described above;
[0300] R4 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0301] R5 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0302] R8 and R9 are each independently C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl or -R h -C1-C 14 Alkyl, -R h -C2-C 14 Alkenyl, -R h -C2-C 14 Alkynyl, where R h is O or S;
[0303] In formula I-4, R1-R7, R 10 -R 11 , L1-L3, and L5 are defined as above.
[0304] In a more preferred embodiment of the present disclosure, the ionizable lipid has a structure selected from the group consisting of those shown in Table 1.
[0305] Table 1
[0306] In some embodiments, the composition comprises the compound represented by Formula I (or its subformula), DSPC, cholesterol, and PEG-DMG in a molar ratio of 50:10:38.5:1.5.
[0307] In yet another aspect, the present disclosure provides a host cell comprising the variant, the fusion protein, the system, the polynucleotide, the vector or the delivery composition described herein.
[0308] In certain embodiments, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammals).
[0309] In certain embodiments, the host cell is a prokaryotic cell, such as Escherichia coli.
[0310] In certain embodiments, the yeast cell is selected from one or more yeasts of the following sources: Pichia pastoris, Kluyveromyces, or a combination thereof; preferably, the yeast cell includes: Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.
[0311] In certain embodiments, the host cell is selected from the group consisting of Escherichia coli, wheat germ cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, or a combination thereof.
[0312] In another aspect, the present disclosure provides a kit comprising one or more components selected from the following: the variant described in the present disclosure, the fusion protein described in the present disclosure, the system described in the present disclosure, the polynucleotide described in the present disclosure, the vector described in the present disclosure, the delivery composition described in the present disclosure, and the host cell described in the present disclosure.
[0313] In certain embodiments, the kit further comprises a label or instructions.
[0314] In certain embodiments, the kit is used for one or more of gene or genome editing, disease treatment, targeting a target gene, and cleaving a target gene or a non-target gene.
[0315] In another aspect, the present disclosure provides a composition comprising the variant described in the present disclosure, the fusion protein described in the present disclosure, the base editing system described in the present disclosure, the polynucleotide described in the present disclosure, the vector described in the present disclosure, the delivery composition described in the present disclosure, or the host cell described in the present disclosure.
[0316] In some embodiments, the composition further comprises a pharmaceutically acceptable carrier and / or excipient.
[0317] In certain embodiments, the composition comprises a pharmaceutical composition.
[0318] In certain embodiments, the dosage form of the composition is selected from the group consisting of a lyophilized formulation, a liquid formulation, or a combination thereof.
[0319] In certain embodiments, the composition is in the form of a liquid preparation.
[0320] In certain embodiments, the composition is in the form of an injection.
[0321] In certain embodiments, the composition is a cell preparation.
[0322] In another aspect, the present disclosure provides an enzyme preparation, comprising the variant described in the present disclosure, the fusion protein described in the present disclosure, the base editing system described in the present disclosure, the polynucleotide described in the present disclosure, the vector described in the present disclosure, and the delivery composition described in the present disclosure.
[0323] In certain embodiments, the enzyme preparation includes an injection and / or a lyophilized preparation.
[0324] In yet another aspect, the present disclosure provides a kit comprising:
[0325] A first container, and the system of the present disclosure, the polynucleotide of the present disclosure, the vector of the present disclosure, or the delivery composition of the present disclosure located in the first container, or a drug containing the system of the present disclosure, the polynucleotide of the present disclosure, the vector of the present disclosure, or the delivery composition of the present disclosure.
[0326] In certain embodiments, the drug in the first container is a single preparation containing the system described in the present disclosure, the polynucleotide described in the fourth aspect of the present disclosure, the vector described in the present disclosure, or the delivery composition described in the sixth aspect of the present disclosure.
[0327] In certain embodiments, the dosage form of the drug is selected from the group consisting of a lyophilized formulation, a liquid formulation, or a combination thereof.
[0328] In certain embodiments, the dosage form of the drug is an oral dosage form or an injectable dosage form.
[0329] In certain embodiments, the kit further comprises instructions.
[0330] In yet another aspect, the present disclosure provides a kit comprising:
[0331] (a1) a first container, and the variant described in the present disclosure, the fusion protein described in the present disclosure, or a gene encoding the same or an expression vector thereof, or a drug containing the variant described in the present disclosure, the fusion protein described in the present disclosure, or a gene encoding the same or an expression vector thereof, located in the first container;
[0332] (b1) an optional second container, and a guide RNA or its expression vector, or a drug containing the guide RNA or its expression vector, located in the second container.
[0333] In certain embodiments, the first container and the second container are different containers.
[0334] In certain embodiments, the medicine in the first container is a single-ingredient preparation containing the variant described in the first aspect of the present disclosure, the fusion protein described in the present disclosure, or the gene encoding it or its expression vector.
[0335] In certain embodiments, the drug in the second container is a single-ingredient preparation containing the guide RNA or its expression vector.
[0336] In certain embodiments, the dosage form of the drug is selected from the group consisting of a lyophilized formulation, a liquid formulation, or a combination thereof.
[0337] In certain embodiments, the dosage form of the drug is an oral dosage form or an injectable dosage form.
[0338] In certain embodiments, the kit further comprises instructions.
[0339] In another aspect, the present disclosure provides a lipid nanoparticle composition comprising an ionizable lipid and a nucleotide sequence encoding the variant described herein, a nucleotide sequence encoding the fusion protein described herein, a nucleotide sequence encoding the system described herein, a polynucleotide described herein, or a vector described herein.
[0340] In some embodiments, the ionizable lipid comprises a compound represented by Formula I or a subformula thereof, or a pharmaceutically acceptable salt, stereoisomer, tautomer, solvate, chelate, non-covalent complex, or precursor thereof.
[0341] In another aspect, the present disclosure provides a pharmaceutical formulation comprising an ionizable lipid, and a variant, a fusion protein, a system, a polynucleotide, or a vector as described herein, and a pharmaceutically acceptable excipient, carrier, or diluent; or the pharmaceutical formulation comprises the lipid nanoparticle composition as described herein, and a pharmaceutically acceptable excipient, carrier, or diluent.
[0342] In some embodiments, the particle size of the pharmaceutical preparation is 30 to 500 nm. For example, the particle size can be 30 nm, 50 nm, 100 nm, 150 nm, 250 nm, 350 nm, 500 nm, etc.
[0343] In some embodiments, the encapsulation efficiency of the bioactive substance in the pharmaceutical preparation is greater than 50%. Exemplarily, the encapsulation efficiency can be 55%, 60%, 65%, 70%, 75%, 79%, 80%, 85%, 89%, 90%, 93%, 95%, etc.
[0344] In some embodiments, the hydrated particle size of the drug is 50-200 nm, preferably 70-150 nm, and most preferably 75-110 nm.
[0345] In some embodiments, the pharmaceutical preparation can be used for the treatment and / or prevention of a disease.
[0346] In some embodiments, the disorder or disease is associated with one or more C>A point mutations or C>T point mutations; preferably, the disorder or disease includes the diseases shown in Table A1 below.
[0347] In some embodiments, the dosage form of the pharmaceutical preparation is selected from the group consisting of injection, lyophilized preparation, nebulized inhalation preparation, and smear-type preparation.
[0348] In some embodiments, the pharmaceutical formulation is administered by injection, ie, intravenously, intramuscularly, intradermally, subcutaneously, intrathecally, intraduodenally, or intraperitoneally.
[0349] In some embodiments, the pharmaceutical formulation is administered by inhalation, such as intranasally.
[0350] In some embodiments, the pharmaceutical preparation is administered transdermally, for example, by transdermal application or electrode introduction. In another aspect, the present disclosure provides a method for targeting and editing a target gene, comprising: contacting the variant described herein, the fusion protein described herein, the base editing system described herein, the polynucleotide described herein, the vector described herein, the delivery composition described herein, the host cell described herein, the composition described herein, the enzyme preparation described herein, the lipid nanoparticle composition described herein, the pharmaceutical composition described herein, or the pharmaceutical preparation described herein with the target gene, or delivering it to a cell comprising the target gene, wherein the target sequence is present in the target gene.
[0351] In certain embodiments, the target gene is present in a cell.
[0352] In certain embodiments, the cell is a prokaryotic cell.
[0353] In certain embodiments, the cell is a eukaryotic cell, such as a mammalian cell (eg, a human cell) or a plant cell.
[0354] In certain embodiments, the target gene is present in a nucleic acid molecule (eg, a plasmid) in vitro.
[0355] In certain embodiments, the editing of the target gene comprises a mutation of the target sequence.
[0356] In certain embodiments, the base editing system is capable of effectively generating a desired mutation in a nucleic acid (e.g., a nucleic acid within a subject's genome). In certain embodiments, the base editing system is capable of generating at least 0.01% of the desired mutation. In certain embodiments, the desired mutation is a mutation generated by a specific base editor bound to a gRNA, specifically designed to alter or correct a mutation in a target gene.
[0357] In certain embodiments, the expected mutation is a mutation produced by a specific base editor bound to a gRNA, specifically designed to change or correct the expected mutation. In certain embodiments, the expected mutation is a mutation that produces a stop codon, such as a premature stop codon in a gene coding region. In certain embodiments, the expected mutation is a mutation that eliminates a stop codon. In certain embodiments, the expected mutation is a mutation that changes gene splicing. In certain embodiments, the expected mutation is a mutation that changes the regulatory sequence of a gene (e.g., a gene promoter or a gene repressor).
[0358] In certain embodiments, the target gene comprises DNA.
[0359] In certain embodiments, the DNA comprises single-stranded DNA or double-stranded DNA.
[0360] In some embodiments, the method includes the steps of contacting the variant described herein, the fusion protein described herein, the base editing system described herein, the polynucleotide described herein, the vector described herein, the delivery composition described herein, or the host cell described herein, or the composition described herein, or the enzyme preparation described herein with the target gene and causing a deamination reaction.
[0361] In some embodiments, the method is in vivo or in vitro.
[0362] In another aspect, the present disclosure provides a method for inducing a change in a cell state, the method comprising contacting the variant described herein, the fusion protein described herein, the base editing system described herein, the polynucleotide described herein, the vector described herein, the delivery composition described herein, or the host cell described herein, or the composition described herein, or the enzyme preparation described herein, or the lipid nanoparticle composition described herein, or the pharmaceutical composition described herein, or the pharmaceutical preparation described herein with a target gene in a cell.
[0363] In yet another aspect, the present disclosure provides a cell or progeny thereof obtained by the method of any one of the present disclosures, wherein the cell comprises a modification that is not present in its wild type.
[0364] In yet another aspect, the present disclosure provides a cell product of the cell of the present disclosure or its progeny.
[0365] In another aspect, the present disclosure provides an in vitro, ex vivo or in vivo cell or cell line or their progeny, which comprises: the variant described in the present disclosure, the fusion protein described in the present disclosure, the base editing system described in the present disclosure, the polynucleotide described in the present disclosure, the vector described in the present disclosure, the delivery composition described in the present disclosure, or the composition described in the present disclosure, or the lipid nanoparticle composition described in the present disclosure.
[0366] In certain embodiments, the cell is a prokaryotic cell.
[0367] In certain embodiments, the cell is a eukaryotic cell, such as a mammalian cell (eg, a human cell) or a plant cell.
[0368] In certain embodiments, the cell is a stem cell or a stem cell line.
[0369] In another aspect, the present disclosure provides uses of the variants described herein, the fusion proteins described herein, the base editing systems described herein, the polynucleotides described herein, the vectors described herein, the delivery compositions described herein, or the host cells described herein, or the compositions described herein, or the enzyme preparations described herein, or the lipid nanoparticle compositions described herein, or the pharmaceutical compositions described herein, or the pharmaceutical preparations described herein for preparing a drug or preparation for treating a disease associated with or caused by a point mutation.
[0370] In some embodiments, the disorder or disease is associated with one or more C>A point mutations or C>T point mutations; preferably, the disorder or disease includes the diseases shown in Table A1.
[0371] In some embodiments, the disease or condition comprises one or more of hypercholesterolemia, transthyretin amyloidosis, and beta-hemoglobinopathies.
[0372] In another aspect, the present disclosure provides a method for treating a condition or disease, comprising administering to a subject in need thereof an effective amount of a variant as described herein, a fusion protein as described herein, a base editing system as described herein, a polynucleotide as described herein, a vector as described herein, a delivery composition as described herein, or a host cell as described herein, or a composition as described herein, or an enzyme preparation as described herein, or a lipid nanoparticle composition as described herein, or a pharmaceutical composition as described herein, or a pharmaceutical preparation as described herein.
[0373] In some embodiments, the disorder or disease is associated with one or more C>A point mutations or C>T point mutations; preferably, the disorder or disease includes the diseases shown in Table A1 below.
[0374] It should be understood that within the scope of the present disclosure, the above-mentioned technical features of the present disclosure and the technical features described in detail below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be listed here one by one. BRIEF DESCRIPTION OF THE DRAWINGS
[0375] Figure 1 shows the editing efficiency of LNP-delivered TA9999-nCas9 and ABE8e in HepG2 cells.
[0376] Figure 2 shows the editing efficiency of TA9999-nCas9 in HepG2 cells and the changes in PCSK9 protein content in cells under different doses.
[0377] Figure 3 shows the editing efficiency of TA9999-nCas9 and ABE8e in mice. Under the same dose (0.05mpk), the editing efficiency of TA9999-nCas9 is much higher than that of ABE8e. Compared with 0.05mpk, the editing efficiency of TA9999-nCas9 is significantly improved at a dose of 2mpk.
[0378] Figure 4 shows the changes in LDL-C and PCSK9 protein levels after TA9999-nCas9 editing in mice.
[0379] Figure 5 shows the editing efficiency of TA9999-nCas9 mRNA and mPCSK9-sgRNA delivered to mice using LNPs containing Yoltech Lipid1-5 and ALC-0315, respectively.
[0380] FIG6 shows the editing efficiency of TA9999-nCas9 targeting the PCSK9 gene target in cynomolgus monkeys.
[0381] Figure 7 shows the changes in PCSK9 protein content in cynomolgus monkey plasma after TA9999-nCas9 edited the PCSK9 gene target.
[0382] Figure 8 shows the editing activity of TA7904-dC05440 and dC05440-TA7904 in the hHao1 and hKLF4 genes; Figure 8A shows the editing activity of TA7904-dC05440 in the hHao1 gene target, and Figure 8B shows the editing activity of TA7904-dC05440 in the hKLF4 gene target, wherein N-ABE / N-terminus refers to the base editor obtained by fusing the deaminase TA7904 to the N-terminus of dC05440, and C-ABE / C-terminus refers to the base editor obtained by fusing the deaminase TA7904 to the C-terminus of dC05440. DETAILED DESCRIPTION
[0383] The deaminase variants disclosed herein have better gene editing activity than wild-type deaminase, can effectively edit target genes, and can effectively treat disorders or diseases in subjects in need.
[0384] the term
[0385] Unless otherwise indicated, the experiments and procedures described in the examples were performed essentially according to conventional methods well known in the art and described in various references.
[0386] In addition, where specific conditions are not specified in the examples, the experiments were performed under conventional conditions or the conditions recommended by the manufacturer. Reagents or instruments used where the manufacturer is not specified are conventional products that can be obtained commercially. It is understood by those skilled in the art that the terms used herein are only used to describe specific embodiments by way of example and are not intended to limit the subject matter protected by the claims. All publications and other references mentioned herein are incorporated herein by reference in their entirety.
[0387] In order to more easily understand the present disclosure, some terms are first defined. As used in this application, unless otherwise expressly provided herein, each of the following terms should have the meaning given below. Other definitions are set forth throughout the application.
[0388] The term "about" can refer to a value or composition that is within an acceptable error range for a particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined. For example, as used herein, the expression "about 100" includes all values between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).
[0389] As used herein, the terms "comprising" or "including" may be open, semi-closed, or closed. In other words, the terms also include "consisting essentially of" or "consisting of."
[0390] Sequence identity (or homology) is determined by comparing two aligned sequences along a predetermined comparison window (which may be 50%, 60%, 70%, 80%, 90%, 95% or 100% of the length of the reference nucleotide sequence or protein) and determining the number of positions at which identical residues occur. Typically, this is expressed as a percentage. The measurement of sequence identity of nucleotide sequences is a method well known to those skilled in the art. The term "mutant" refers to a protein that has been produced by mutation or recombinant DNA procedures.
[0391] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction. Deaminases herein are nucleobase deaminases, and the terms "deaminase" and "nucleobase deaminase" are used interchangeably herein. The deaminase can be a naturally occurring deaminase or an active fragment or variant thereof. The deaminase can be active on single-stranded nucleic acids such as ssDNA or ssRNA, or double-stranded nucleic acids such as dsDNA or dsRNA. In some embodiments, the deaminase can only deaminate ssDNA and has no effect on dsDNA.
[0392] The term "adenosine deaminase" or "adenosine deaminase protein" refers to a protein, a polypeptide, or one or more functional domains of a protein or polypeptide that can catalyze the hydrolytic deamination reaction of converting adenine (or the adenine portion of a molecule) into hypoxanthine (or the hypoxanthine portion of a molecule). In some embodiments, the adenine-containing molecule is adenosine (A), and the hypoxanthine-containing molecule is inosine (I). The adenine-containing molecule can be a deoxyribonucleic acid (DNA) or a ribonucleic acid (RNA). Adenosine deaminases include, but are not limited to, members of the enzyme family known as adenosine deaminases acting on RNA (ADARs), members of the enzyme family known as adenosine deaminases acting on tRNA (ADATs), and other family members containing adenosine deaminase domains (ADADs). According to the present disclosure, adenosine deaminases can target adenine in RNA / DNA and RNA duplexes. In specific embodiments, adenosine deaminases have been modified to increase their ability to edit DNA in RNA / DNA heteroduplexes of RNA duplexes.
[0393] The term "base editor" is a fusion protein including a nucleic acid programmable nucleotide binding protein (napDNAbp) (such as a nuclease) and a deaminase. "Base editor (BE)" or "nucleobase editor" refers to a reagent that binds to a polynucleotide and has a nuclear base modification activity. In various embodiments, the base editor includes a nuclear base modification polypeptide (e.g., a deaminase) and a nucleic acid programmable nucleotide binding domain (e.g., a nucleic acid programmable DNA binding protein) bound to a guide polynucleotide (e.g., a guide RNA). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2c3, and Argonaute protein (AGO). In various embodiments, the reagent is a biomolecular complex comprising a protein domain with base editing activity, i.e., capable of modifying bases (e.g., A, T, C, G, or U) in a nucleic acid molecule (e.g., DNA, RNA). In some embodiments, the polynucleotide programmable DNA binding domain is fused or connected to a deaminase domain. In one embodiment, the reagent is a fusion protein comprising a domain having base editing activity. In some embodiments, the domain having base editing activity is capable of deaminating bases within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases within a DNA molecule. In some embodiments, the base editor is an adenosine base editor (ABE).
[0394] The term "nuclease" refers to an enzyme that catalyzes the cleavage of phosphodiester bonds between nucleotides in a nucleic acid molecule. In some embodiments, the DNA binding polypeptide is an endonuclease that is capable of cleaving phosphodiester bonds between nucleotides in a nucleic acid molecule. In certain embodiments, the DNA binding polypeptide is an exonuclease that is capable of cleaving nucleotides at either end (5' or 3') of a nucleic acid molecule. In some embodiments, the nuclease is selected from a group consisting of a homing endonuclease (Meganuclease), a zinc finger nuclease (ZFN), a TAL effector DNA nuclease fusion protein (TALEN), and an RNA-guided nuclease or homologs or variants thereof, wherein the nuclease activity is reduced or inhibited.
[0395] The term "homing endonuclease" or "meganuclease" refers to an endonuclease that binds to a recognition site within dsDNA of 12 to 40 bp in length. Exemplary, non-limiting examples of homing endonucleases include the LAGLIDADG series. "Homing endonuclease" can refer to either a dimeric or single-chain meganuclease.
[0396] The term "zinc finger nuclease" or "Zinc Finger Nuclease, ZFN" refers to a chimeric protein comprising a zinc finger DNA binding domain and a nuclease domain.
[0397] The term "TAL effector DNA binding domain nuclease fusion protein" or "TALEN" refers to a chimeric protein comprising a TAL effector DNA binding domain and a nuclease domain.
[0398] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" can be used interchangeably with "polynucleotide programmable nucleotide binding domain" and "nucleic acid programmable nucleotide binding domain" to refer to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or a guide polynucleotide (e.g., gRNA), that guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the nucleic acid programmable nucleotide binding protein is an RNA-guided nucleic acid programmable nucleotide binding protein. In some embodiments, the RNA-guided nucleic acid programmable nucleotide binding protein is an RNA-guided nuclease.
[0399] In some embodiments, the RNA-guided nuclease is selected from type II CRISPR-Cas polypeptides, type I CRISPR-Cas polypeptides, type III CRISPR-Cas polypeptides, type IV CRISPR-Cas polypeptides, type V CRISPR-Cas polypeptides, type VI CRISPR-Cas polypeptides, type VII CRISPR-Cas polypeptides, IscB polypeptides, TnpB polypeptides, IsrB polypeptides. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein can be associated with a guide RNA that guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, napDNAbp is a Cas9 domain, such as a nuclease-active Cas9, a Cas9 nickase (nCas9), or a nuclease-inactivated Cas9 (dCas9). Non-limiting examples of nucleic acid-programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Cas13a (C2c2), Cas13b, Cas13c, Cas13d.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csn1, Csn2, Csn3, Csn4, Csn5e, Csn1, Csn2, Csn3, Csn4, Csn5e, Csn1 c1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector protein, type V Cas effector protein, type VI Cas effector protein, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically listed in the present disclosure. See, for example, Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Wherefrom Here." (CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr.2018.0033); Yan et al., "Functionally diverse type V CRISPR-Cas systems" (Science. 2019 Jan 4; 363(6422): 88-91. doi: 10.1126 / science.aav7271), which are incorporated herein by reference in their entirety.
[0400] As used herein, "base editing activity" refers to the activity of chemically altering a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is adenosine or adenine deaminase activity, for example, converting a target A·T to C·G.
[0401] In some embodiments, base editing activity is assessed by editing efficiency. Base editing efficiency can be measured by any suitable means, for example, by Sanger sequencing or next-generation sequencing. In some embodiments, base editing efficiency is measured by the percentage of total sequencing reads with nuclear base conversions affected by the base editor, for example, the percentage of total sequencing reads with target C·G base pairs converted to A·T base pairs. In some embodiments, when base editing is performed in a cell population, base editing efficiency is measured by the percentage of total cells with nuclear base conversions affected by the base editor.
[0402] As used herein, the term "base editor system" refers to a system for editing a nucleobase of a target nucleotide sequence. In various embodiments, the base editor system comprises: (1) a nucleic acid programmable nucleotide binding domain (e.g., Cas9); (2) a deaminase domain (e.g., adenosine deaminase) for deaminating the nucleobase; and (3) one or more guide polynucleotides (e.g., guide RNA).
[0403] "Guide polynucleotide," "guide RNA," or "gRNA" refers to a polynucleotide that can specifically target a target sequence and can form a complex with a nucleic acid programmable nucleotide binding domain protein (e.g., Cas9). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule may be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species includes two domains: (1) a domain that has homology to the target nucleic acid (e.g., guides the binding of the Cas9 complex to the target nucleic acid); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical to or homologous to the tracrRNA provided in Jinek et al., Science 337:816-821 (2012). Other examples of gRNAs may be those disclosed in U.S. Provisional Patent Application USSN 61 / 874,682, filed on September 6, 2013, entitled “Switchable Cas9 Nucleases and Uses Thereof” and U.S. Provisional Patent Application USSN 61 / 874,746, filed on September 6, 2013, entitled “Delivery System For Functional Nucleases.” In some embodiments, the gRNA includes two or more of domains (1) and (2) and may be referred to as an “extended gRNA.” The extended gRNA will bind to two or more Cas9 proteins and bind to the target nucleic acid at two or more different regions. The gRNA includes a nucleotide sequence complementary to the target site that mediates binding of the nuclease / RNA complex to the target site, providing sequence specificity for the nuclease:RNA complex.
[0404] The term "identity" is used to refer to the matching of sequences between two polypeptides or between two nucleic acids. "Identity" represents the percentage of the number of identical residues between the polypeptide or nucleic acid sequences to the total number of residues, and the calculation of the total number of residues is determined based on the mutation type. Mutation types include insertions (extensions) at either or both ends of the sequence, deletions (truncations) at either or both ends of the sequence, substitutions / alternations of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence. Taking a polypeptide sequence as an example, if the mutation type is one or more of the following: substitutions / alternations of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence, the total number of residues is calculated based on the larger of the molecules being compared. If the mutation type also includes insertions (extensions) at either or both ends of the sequence or deletions (truncations) at either or both ends of the sequence, the number of amino acids inserted or deleted at either or both ends (e.g., the number of insertions or deletions at both ends is less than 20) is not included in the total number of residues. When calculating the percentage of identity, the sequences being compared are aligned in a manner that produces the maximum match between the sequences, and gaps (if any) in the alignment are resolved by a specific algorithm. The same principle applies to the calculation of nucleotide identity.
[0405] The terms "sequence identity" and "sequence homology" are used interchangeably herein and, as used in conjunction with a polynucleotide or polypeptide, refer to the percentage of bases or amino acids that are identical and in the same relative position when comparing or aligning two sequences of a polypeptide or polynucleotide. Sequence identity can be determined in a variety of different ways. For example, sequences can be aligned using various methods and computer programs (e.g., BLAST, T-COFFEE, MUSCLE, MAFFT, etc.). See, for example, Altschul et al., (1990) J. Mol. Biol. 215:403-10.
[0406] The term "DNA sequence or DNA polynucleotide sequence encoding a specific RNA is a sequence of DNA that can be transcribed into RNA. A DNA polynucleotide can encode an RNA (mRNA) that is translated into a protein, or a DNA polynucleotide can encode an RNA that is not translated into a protein (e.g., tRNA, rRNA, or guide RNA; also referred to as "non-coding" RNA or "ncRNA"). A DNA sequence or DNA polynucleotide sequence can also "encode" a specific polypeptide or protein sequence, wherein, for example, DNA directly encodes an mRNA that can be translated into a polypeptide or protein sequence. A "protein coding sequence" or a sequence encoding a specific protein or polypeptide is a nucleic acid sequence that can be transcribed into mRNA (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vitro or in vivo when placed under the control of an appropriate regulatory sequence. The boundaries of the coding sequence can be determined by a translation termination nonsense codon at the 5' end (N-terminus) and a translation termination nonsense codon at the 3' end (C-terminus). The coding sequence can include, but is not limited to, cDNA from prokaryotic or eukaryotic mRNA, genomic DNA sequences from prokaryotic or eukaryotic DNA, and synthetic nucleic acids. A transcription termination sequence will usually be located 3' to the coding sequence.
[0407] The term "promoter" or "promoter sequence" is a DNA regulatory sequence that is capable of promoting transcription of an operably linked coding or non-coding sequence (e.g., a downstream (3' direction) coding or non-coding sequence) (e.g., capable of causing detectable levels of transcription and / or increasing detectable levels of transcription (relative to the level provided in the absence of the promoter)), such as by binding RNA polymerase. In some embodiments, the promoter sequence is bounded at its 3' terminus by the transcription start site and extends upstream (5' direction) to include the minimum number of bases or elements to initiate transcription at a detectable level above background. In some embodiments, the promoter sequence may include a transcription start site and a protein binding domain responsible for binding RNA polymerase. In addition to sequences sufficient to initiate transcription, a promoter may also include sequences of other regulatory elements involved in regulating transcription (e.g., enhancers, Kozak sequences, and introns). Various promoters, including inducible promoters and constitutive promoters, can be used to drive the vectors disclosed herein. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, will result in the production of the gene product in the cell under most or all physiological conditions of the cell. Inducible promoter refers to the presence of endogenous or exogenous stimulation, such as by chemical compound (chemical inducer) response, or to environment, hormone, chemical and / or development signal response, the promoter of selective expression coding sequence or function RNA.Inducible or regulated promoter includes for example by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormone, wound or chemical (such as ethanol, abscisic acid (ABA), jasmonate, salicylic acid or safener) induction or the promoter of regulation. The example of the promoter that can be used in certain embodiments (such as, in the viral vector disclosed herein) as known in the art includes CMV promoter, CBA promoter, smCBA promoter and the promoter derived from immunoglobulin gene, SV40 or other tissue-specific genes (such as: RLBP1, RPE, VMD2). In addition, the standard technology of producing functional promoter by mixing and matching known regulatory elements is known in the art. Fragments of promoters can also be used, such as those retaining at least a minimum number of bases or elements to start the fragment of transcription above the detectable level of background.
[0408] "Operably linked" refers to a juxtaposition in which the components are in a relationship permitting them to function in their intended manner. For example, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence.
[0409] The terms "heterologous promoter" and "heterologous control region" refer to promoters and other control regions not normally associated with a particular nucleic acid in nature. For example, a "transcriptional control region heterologous to a coding region" is a transcriptional control region not normally associated with a coding region in nature.
[0410] The term "vector" refers to a nucleic acid molecule that is capable of transporting another nucleic acid molecule to which it is attached. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other various polynucleotides known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection so that the genetic material elements it carries are expressed in the host cell. A vector can be introduced into a host cell to thereby produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. The vector may also contain a replication initiation site. Vectors include plasmids and viral vectors. The plasmid refers to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. Viral vector, wherein the DNA or RNA sequence derived from the virus is present in the vector for packaging the virus, including for example retrovirus, replication defective retrovirus, adenovirus, replication defective adenovirus and adeno-associated virus. Viral vector also comprises the polynucleotide carried by the virus for transfection into a host cell. Some vectors (for example, bacterial vectors and additional mammalian vectors with bacterial replication origin) can replicate autonomously in the host cell into which they are introduced. Other vectors (for example, non-additional mammalian vectors) are integrated into the genome of the host cell after introducing the host cell, and thus replicate together with the host genome. Moreover, some vectors can instruct the expression of the gene that they are operably connected. Such vectors are referred to as "expression vectors".
[0411] The term "wild type" has the meaning generally understood by those skilled in the art, which refers to the typical form of an organism, strain, gene, protein or the characteristics that distinguish it from mutant or variant forms when it exists in nature, which can be isolated from a source in nature and has not been intentionally modified by man.
[0412] The terms "variant," "derivative," and "analog" refer to polypeptides that substantially retain the function or activity of a protein. Generally, derivatization of a protein does not adversely affect the desired activity of the protein, i.e., the derivative of the protein has the same activity as the protein. Modified forms of "derivatives" include those in which one or more amino acids of the protein may be deleted, inserted, modified, and / or substituted.
[0413] The terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of human effort.
[0414] As used herein, a "functional fragment" or "active fragment" of a polynucleotide or polypeptide may refer to any subset of consecutive nucleotides or consecutive amino acids, respectively, that retains the original (e.g., wild-type) activity (or substantially similar activity) of the polynucleotide or polypeptide. In some embodiments, the "functional fragment" or "active fragment" comprises any portion or subsequence of the original (e.g., wild-type) or mutant polynucleotide or polypeptide. In some embodiments, the activity of the "functional fragment" or "active fragment" of the polynucleotide or polypeptide, for example, relative to the activity of the original (e.g., wild type), the "functional fragment" or "active fragment" can be about 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or less than 10% of the activity.
[0415] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As a further guide, an "orthologue" of a protein as described herein refers to a protein belonging to a different species that performs the same or similar function as the protein to which it is an orthologue.
[0416] The nucleic acid cleavage disclosed herein includes: DNA or RNA breakage in the target nucleic acid produced by the Cas protein (Cis cleavage), and DNA or RNA breakage in a side branch nucleic acid substrate (single-stranded nucleic acid substrate) caused by the Cas protein side cutting activity (i.e., non-specific or non-targeted, Trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.
[0417] The terms “clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system” or “CRISPR-Cas system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements associated with the expression of CRISPR-associated (“Cas”) genes, or transcripts or other elements capable of directing the activity of the Cas genes.
[0418] The terms "target nucleic acid" and "target sequence" are used interchangeably and refer to a specific nucleic acid that comprises a nucleic acid sequence that is fully or partially complementary to the guide sequence in the gRNA. A "target sequence" refers to a polynucleotide targeted by the guide sequence in the gRNA, such as a sequence complementary to the guide sequence, wherein the hybridization between the target sequence and the guide sequence will promote the formation of a CRISPR / Cas complex (including Cas protein and gRNA). Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex. In some embodiments, the target nucleic acid comprises a non-coding region (e.g., a promoter or terminator). In some embodiments, the target nucleic acid is single-stranded or double-stranded. The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the nucleus, cytoplasm, or organelles (e.g., mitochondria or chloroplasts) of the cell. The target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be related to a protospacer adjacent motif (PAM).
[0419] The term "reporter nucleic acid" refers to a molecule that can be cut or otherwise deactivated by an activated CRISPR system protein as described herein. The reporter nucleic acid comprises a nucleic acid element (e.g., a single-stranded nucleic acid detector) that can be cut by a CRISPR protein. The cutting of the nucleic acid element produces a detectable signal. Before cutting, or when the reporter nucleic acid is in an "active" state, the reporter nucleic acid prevents the generation or detection of a positive detectable signal. It will be understood that in certain example embodiments, a minimal background signal can be generated in the presence of active reporter nucleic acid. A positive detectable signal can be any signal that can be detected using optical, fluorescent, chemiluminescent, electrochemical, or other detection methods known in the art. For example, in certain embodiments, when a reporter nucleic acid is present, a first signal (i.e., a negative detectable signal) can be detected, which is then converted to a second signal (e.g., a positive detectable signal) after detecting the target molecule and cutting or deactivating the activated CRISPR protein. The reporter nucleic acid can be a single-stranded DNA molecule, a single-stranded RNA molecule, or a single-stranded DNA-RNA hybrid.
[0420] The detection method disclosed herein can be used for quantitative detection of target nucleic acids. The quantitative detection index can be quantified based on the signal strength of the reporter group, such as the luminescence intensity of the fluorescent group, or the width of the color band.
[0421] The term "regulatory element" includes promoters, enhancers, internal ribosome entry sites (IRES) and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals, poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). In some cases, regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct the nucleotide sequence to be expressed only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or special cell types (e.g., lymphocytes). In other cases, regulatory elements can also direct expression in a temporally dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell-type specific.
[0422] The term "host cell" refers to a eukaryotic cell (e.g., an animal cell, a plant cell, a fungal cell, etc.), a prokaryotic cell (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or a cell from a multicellular organism cultured as a unicellular entity (e.g., a cell line), which serves as a recipient of nucleic acid (e.g., an expression vector) and includes the descendants of the original cell that has been genetically modified by the nucleic acid.
[0423] It is understood that the progeny of a single cell may not necessarily have completely the same morphology, genome, etc. as the original parent cell due to natural, accidental, or deliberate mutation. A "recombinant host cell" (also called a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an expression vector, has been introduced.
[0424] Those skilled in the art will appreciate that the design of the expression vector may depend on factors such as the choice of the host cell to be transformed, the level of expression desired, and the like.
[0425] The term "NLS" refers to a "nuclear localization sequence" or "nuclear localization signal," which refers to an amino acid sequence that promotes protein entry into the cell nucleus. Nuclear localization sequences are known in the art (e.g., Plank et al., International PCT Application PCT / EP2000 / 011690, filed November 23, 2000, and described in WO / 2001 / 038547, published on May 31, 2001), which is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is optimized, for example, as described in Koblan et al., Nature Biotech.2018doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the following amino acid sequence:
[0426] KRTADGSEFESPKKKRKV(SEQ ID NO::15),AVKRPAATKKAGQAKKKKLD(SEQ ID NO::16),KRPAATKKAGQAKKKK(SEQ ID NO:17),KKTELQTTNA ENKTKKL(SEQ ID NO::18),KRGINDRNFWRGENGRKTR(SEQ ID NO:19),RKSGKIAAIVVKRPRK(SEQ ID NO:20), PKKKRKV (SEQ ID NO:21) or MDSLL MNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO:22).
[0427] "Operably linked" means that the nucleotide sequence of interest is linked to regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). Advantageous vectors include lentiviruses and adeno-associated viruses, and the type of these vectors can also be selected to target specific cell types.
[0428] The term "complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence using traditional Watson-Crick or other non-traditional types. Percent complementarity represents the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 5, 6, 7, 8, 9, or 10 out of 10 are complementary, resulting in percentages of 50%, 60%, 70%, 80%, 90%, and 100%). "Perfect complementarity" means that all consecutive residues of one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. "Substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.
[0429] The term "stringent conditions" in relation to hybridization refers to conditions under which a nucleic acid having complementarity with a target sequence predominantly hybridizes to the target sequence and substantially does not hybridize to non-target sequences. Stringent conditions are generally sequence-dependent and depend on many factors. Generally speaking, the longer the sequence, the higher the temperature at which it will specifically hybridize to its target sequence.
[0430] The term "hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding of the bases between the nucleotide residues. The complex may comprise two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute one step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing to a given sequence is referred to as the "complement" of that given sequence.
[0431] Hybridization of the target sequence with the gRNA indicates that at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or represents that at least 12, 15, 16, 17, 18, 19, 20 or more bases of the nucleic acid sequences of the target sequence and the gRNA can complement each other and hybridize to form a complex.
[0432] The term "nucleic acid expression" includes one or more of production of an RNA template from a DNA sequence (e.g., transcription), processing of the RNA transcript (e.g., by splicing, editing, 5' capping, and / or 3' end processing), translation of the RNA into a polypeptide or protein, or post-translational modification of the polypeptide or protein.
[0433] The term "delivery" refers to providing an entity (such as a drug) to a destination. For example, the components of the CRISPR-Cas system / composition of the present disclosure can be delivered in various forms, such as a combination of DNA / RNA or RNA / RNA or protein RNA. For example, the Cas protein can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein.
[0434] The term "linker" refers to a linear polypeptide formed by connecting multiple amino acid residues through peptide bonds. The linker can be an artificially synthesized amino acid sequence or a naturally occurring polypeptide sequence.
[0435] In some embodiments, the linker sequence is an amino acid sequence as shown in any one of SEQ ID NOs. 3-12.
[0436] The term "effective amount" or "therapeutically effective amount" refers to a dosage sufficient to achieve beneficial or desired results. A therapeutically effective amount may depend on the individual and disease condition being treated, the individual's weight and age, the severity of the disease condition, the mode of administration, etc., and can be readily determined by one skilled in the art.
[0437] The terms "treatment," "treating," and the like refer to obtaining a desired pharmacological and / or physiological effect, e.g., treating or curing a condition in a subject, delaying the onset of symptoms of a condition, and / or delaying the severity of a condition. The effect may be prophylactic, in terms of completely or partially preventing a disease or its symptoms, and / or therapeutic, in terms of partially or completely curing a disease and / or side effects attributable to the disease. As used herein, "treatment" encompasses any treatment of a disease in a mammal (e.g., a human), and includes: (a) preventing the occurrence of a disease in a subject who may be susceptible to the disease but has not yet been diagnosed with the disease; (b) inhibiting the disease, i.e., arresting its development; and (c) relieving the disease, i.e., causing the disease to regress.
[0438] The terms "individual," "subject," "host," and "patient" refer to individual organisms, including but not limited to various animals, plants, and microorganisms. Animals include mammals, including but not limited to bovines, equines, ovines, porcines, canines, felines, lagomorphs, rodents (e.g., mice or rats), apes, non-human primates (e.g., macaques or cynomolgus monkeys), humans, mammalian farm animals, mammalian sports animals, and mammalian pets. In certain embodiments, the subject (e.g., a human) suffers from a disorder (e.g., a disorder caused by a disease-related gene defect). "Plant" is any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development.
[0439] The term "each independently" means that at least two groups (or ring systems) with the same or similar numerical ranges present in a structure may have the same or different meanings in certain circumstances. For example, if substituent X and substituent Y are each independently hydrogen, halogen, hydroxy, cyano, alkyl, or aryl, then when substituent X is hydrogen, substituent Y may be hydrogen, halogen, hydroxy, cyano, alkyl, or aryl; similarly, when substituent Y is hydrogen, substituent X may be hydrogen, halogen, hydroxy, cyano, alkyl, or aryl.
[0440] The terms "comprising" and "including" are used in an open, non-limiting sense.
[0441] The term "alkyl" refers to a monovalent straight or branched chain alkyl group consisting solely of carbon and hydrogen atoms, containing no unsaturation, connected to other moieties by single bonds, including but not limited to methyl, ethyl, propyl, isopropyl, butyl, sec-butyl, isobutyl, and tert-butyl. For example, the term "C 1-30 "Alkyl" refers to a saturated monovalent straight or branched hydrocarbon group containing 1 to 30 carbon atoms.
[0442] The term "alkylene" refers to a divalent straight or branched chain alkyl group consisting solely of carbon and hydrogen atoms, containing no saturation, and connected to other moieties by two single bonds, including but not limited to methylene, 1,1-ethylene, and 1,2-ethylene. For example, "C 1-30 "Alkylene" refers to a saturated divalent straight or branched chain alkyl group containing 1 to 30 carbon atoms.
[0443] The term "cycloalkyl" refers to a saturated, monocyclic or polycyclic (e.g., bicyclic, tricyclic, or tetracyclic) non-aromatic hydrocarbon group consisting solely of carbon and hydrogen atoms. Cycloalkyl groups may include parallel, bridged, or spirocyclic ring systems. For example, the term "C 3-6 "Cycloalkyl" refers to a cycloalkyl group having 3 to 6 carbon atoms. For example, the cycloalkyl group may be cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl or bicyclo[2.2.1]heptyl, etc.
[0444] The term "cycloalkylene" refers to a divalent group obtained by removing a hydrogen atom from the cycloalkyl group defined above, including but not limited to cyclopropylene, cyclobutylene, cyclopentylene, cyclohexylene, cycloheptylene, etc. For example, "C 3-30 "Cycloalkylene" refers to a divalent group derived from a cycloalkyl group containing 3 to 30 carbon atoms by removing a hydrogen atom.
[0445] The term "branched alkyl" refers to an alkyl group attached to a parent molecule and forming at least two self-branched structures. For example:
[0446] The term "alkenyl" refers to a monovalent straight or branched chain alkane consisting solely of carbon and hydrogen atoms, containing at least one double bond and connected to other moieties by single bonds, including but not limited to ethenyl, propenyl, allyl, isopropenyl, butenyl, and isobutenyl. For example, the term "C 2-30 "Alkenyl" refers to a monovalent straight or branched hydrocarbon group containing 2 to 30 carbon atoms and having at least one carbon-carbon double bond (>C=C<).
[0447] The term "alkenylene" refers to a divalent straight or branched chain alkane consisting only of carbon atoms and hydrogen atoms, containing at least one double bond and connected to other fragments by two single bonds, including (but not limited to) vinylene. For example, the term "C 2-30 "Alkenylene" refers to a divalent straight or branched chain hydrocarbon group containing 2 to 30 carbon atoms and having at least one carbon-carbon double bond (>C=C<).
[0448] The term "alkynyl" refers to a monovalent straight or branched chain alkane consisting solely of carbon and hydrogen atoms, containing at least one carbon-carbon triple bond and connected to other moieties by single bonds, including but not limited to ethynyl, propynyl, butynyl, and pentynyl. For example, the term "C 2-30 "Alkynyl" refers to a monovalent straight or branched chain hydrocarbon radical containing 2 to 30 carbon atoms and having at least one carbon-carbon triple bond.
[0449] The term "alkynylene" refers to a divalent straight or branched alkane consisting only of carbon atoms and hydrogen atoms, containing at least one carbon-carbon triple bond and connected to other fragments by two single bonds, including (but not limited to) acetylene. For example, "C 2-30 "Alkyne" refers to a divalent straight or branched chain hydrocarbon group containing 2 to 30 carbon atoms and at least one carbon-carbon triple bond.
[0450] The term "cycloalkenyl" refers to an unsaturated, monocyclic or polycyclic (e.g., bicyclic, tricyclic, or tetracyclic) non-aromatic hydrocarbon group composed solely of carbon and hydrogen atoms. Cycloalkenyl groups may include fused, bridged, or spirocyclic ring systems. Examples include cyclopropenyl and cyclobutene.
[0451] The term "cycloalkenylene" refers to a divalent group derived from a cycloalkenyl group as defined above by removing a hydrogen atom, including but not limited to cyclopropenylene and cyclobutenylene. For example, the term "C 3-30 The "cycloalkenylene group" refers to a divalent group obtained by removing a hydrogen atom from a cycloalkenyl group containing 3 to 30 carbon atoms.
[0452] The term "branched alkenyl" refers to an alkenyl group attached to a parent molecule and forming at least two self-branched structures. For example
[0453] The term "heterocyclyl" refers to a saturated or partially saturated, monocyclic or polycyclic (e.g., bicyclic, parallel, bridged or spirocyclic) non-aromatic group, the ring atoms of which include one carbon atom and at least one heteroatom selected from N, O and S, wherein the S atom is optionally substituted to form S(=O), S(=O)2 or S(=O)(=NR x) , R x Independently selected from H or C 1-4 Alkyl. If the valence bond requirements are met, the heterocyclic group can be connected to the rest of the molecule through any one of the ring atoms. For example, the term "3-8 membered heterocyclic group" as used in this disclosure refers to a heterocyclic group having 3 to 8 ring atoms. For example, but not limited to, the heterocyclic group can be an oxiranyl, aziridine, azetidinyl, oxetanyl, tetrahydrofuranyl, dioxolyl, pyrrolidinyl, pyrrolidonyl, imidazolidinyl, pyrazolidinyl, tetrahydropyranyl, piperidinyl, piperazinyl, morpholinyl, thiomorpholinyl, dithianyl or trithianyl.
[0454] The term "aryl" refers to a monocyclic or dense polycyclic aromatic hydrocarbon group having a conjugated π electron system. For example, the term "C 6-10 The term "aryl" refers to an aromatic group having 6 to 10 carbon atoms. For example, the aromatic group may be phenyl, naphthyl, anthracenyl, phenanthrenyl, acenaphthenyl, azulenyl, fluorenyl, indenyl, pyrenyl, and the like.
[0455] The term "heteroaryl" refers to a monocyclic or fused polycyclic aromatic group having a conjugated π-electron system, wherein the ring atoms consist of one carbon atom and at least one heteroatom selected from N, O, and S. A heteroaromatic group may be attached to the rest of the molecule via any one of the ring atoms if valence requirements are met. For example, the term "5-10 membered heteroaryl" as used in this disclosure refers to a heteroaryl group having 5 to 10 ring atoms. Common heteroaryl groups include, but are not limited to, thienyl, furanyl, pyrrolyl, oxazolyl, thiazolyl, imidazolyl, pyrazolyl, isoxazolyl, isothiazolyl, oxadiazolyl, triazolyl, thiadiazolyl, pyridyl, pyridazinyl, pyrimidinyl, pyrazinyl, triazinyl and their benzo derivatives, pyrrolopyridinyl, pyrrolopyridinyl, pyrazolopyridinyl, imidazopyridinyl, pyrrolopyrimidinyl, pyrazolopyrimidinyl, purinyl, and the like.
[0456] The term "halogen" refers to fluorine (F), chlorine (Cl), bromine (Br) and iodine (I).
[0457] The term "hydroxy" refers to -OH.
[0458] The term "cyano" refers to -CN.
[0459] The term "amino" refers to -NH2.
[0460] The term "nitro" refers to -NO2.
[0461] The term "oxo" refers to (=0).
[0462] Deaminase variants and nucleic acids encoding same
[0463] As used herein, the terms "deaminase variant," "variant of the present disclosure," "mutant of the present disclosure," and "mutant" are used interchangeably to refer to a non-naturally occurring mutant deaminase, wherein the mutant is mutated at the core amino acid position corresponding to SEQ ID NO. 1 as shown below in the wild-type deaminase:
[0464] (a)A46;
[0465] (b)I47;
[0466] (c) T48;
[0467] (d) L49;
[0468] (e) V104;
[0469] (f)Q148;
[0470] (g) P150;
[0471] (h)E152;
[0472] (i) V153;
[0473] (j) F154; and
[0474] (k)N155.
[0475] The term "core amino acids" refers to sequences based on a wild-type deaminase that are at least 80%, such as 84%, 85%, 90%, 92%, 95%, 98% or 99% homologous to the wild-type deaminase, at the corresponding positions of the specific amino acids described herein. For example, based on the wild-type deaminase, the core amino acids are:
[0476] (a)A46;
[0477] (b)I47;
[0478] (c) T48;
[0479] (d) L49;
[0480] (e) V104;
[0481] (f)Q148;
[0482] (g) P150;
[0483] (h)E152;
[0484] (i) V153;
[0485] (j) F154; and
[0486] (k)N155.
[0487] Furthermore, the mutant protein obtained by mutating the above core amino acids has higher editing activity than the wild-type deaminase (SEQ ID NO.1).
[0488] In some embodiments, the mutein has at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% sequence identity to SEQ ID NO. 1.
[0489] In some embodiments, the mutation is a mutation occurring in combination at the following positions of the amino acid sequence as shown in SEQ ID NO: 1: A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155;
[0490] In some embodiments, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0491] (2) S2+E3+L4+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+D167+R168+A169+D170; or
[0492] (6)A46+I47+T48+L49+V104+C144+Q145+Q148+Q149+P150+E152+V153+F154+N155;
[0493] In some specific embodiments, the mutation is a substitution occurring in combination at the following positions of the amino acid sequence as shown in SEQ ID NO: 1:
[0494] A46C+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F; or
[0495] A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0496] In some embodiments, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1:
[0497] (8)A15C+L16A+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0498] (9)S2K+E3G+L4P+N5M+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152 L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0499] (10)S2G+E3I+L4Y+N5G+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152 L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0500] (12)S2R+E3R+L4G+N5A+Q18C+K19H+A20T+R21F+A46C+I47V+T48H+L49N+V104M+Q148R+ P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0501] (13)S2R+E3R+L4G+N5A+Q18R+A20T+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D 167V+R168G+A169R+D170F;
[0502] (18)S2V+E3P+N5D+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F
[0503] (2)S2G+E3I+L4Y+N5G+A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152V+V153C+F154E+N155G+A156P+D167K+R168K+A169R+D170C;
[0504] (3)A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152V+V153G+F154P+N155G+D167K+R168K+A169S+D170C;
[0505] (4)S2G+E3G+L4A+N5R+A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0506] (5)A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152G+V153G+F154K+N155G+D167K+R168K+A169S+D170C;
[0507] (6)A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152L+V153A+F154P+N155K+A156H+E157K+R158N+D167G+R168G+A169K+D170V;
[0508] (7)A46C+I47Y+T48G+L49H+V104M+C144W+Q145K+Q148R+Q149R+P150R+E152P+V153F+F154V+N155T;
[0509] (11)S2R+A46C+I47V+T48H+L49N+V104M+Q148R+P150L+E152V+V153G+F154P+N155G+D167V+R168G+A169R+D170F;
[0510] (14)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0511] (15)Q18H+K19Y+R21L+A46C+I47V+T48H+L49N+V104M+C139V+S140L+E141F+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0512] (16)A46C+I47Y+T48G+L49H+Q66N+V68I+Q69T+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0513] (17)A46C+I47Y+T48G+L49H+V104M+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0514] (19)S2T+E3V+L4T+N5P+A46C+I47V+T48H+L49N+V104M+C139Q+S140A+E141V+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0515] (20)A46C+I47V+T48H+L49N+Q66R+I67Q+Q69R+V104M+C139A+S140F+E141L+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0516] (21)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+L163E+N164V+Q165N+P166L;
[0517] (22)S2V+E3Q+L4V+N5R+A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0518] (23)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167A+R168G+A169V+D170S;
[0519] (24)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167S+R168G+A169S+D170T;
[0520] (25)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167I+R168Y+A169G+D170T;
[0521] (26)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167V+R168D+A16V+D170I;
[0522] (27)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155G+D167Q+R168V+A169C+D170A;
[0523] (28)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+A111G+A112L+G113N+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0524] (29)Q18M+K19R+A20S+R21S+A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F;
[0525] (30)Q18R+K19Q+A20S+R21F+A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M142L+Q148R+P150L+E152L+V153A+F154P+N155R+A156H+E157R+R158M+D167V+R168G+A169R+D170F; or
[0526] (31)A46C+I47V+T48H+L49N+Q66K+I67R+V68L+Q69H+V104M+C139A+S140Y+M1 42L+Q148R+P150L+E152L+V153A+F154P+N155G+D167R+R168F+A169G+D170T.
[0527] It should be understood that the amino acid numbering in the mutant proteins disclosed herein is based on the wild-type deaminase. When the sequence homology of a specific mutant protein to the wild-type deaminase reaches 80% or more, the amino acid numbering of the mutant protein may be misplaced relative to the amino acid numbering of the wild-type deaminase, such as misplacement of 1-100 positions toward the N-terminus or C-terminus of the amino acid. Using conventional sequence alignment techniques in the art, those skilled in the art can generally understand that such misplacement is within a reasonable range, and mutant proteins with a homology of 80% (such as 90%, 95%, 98%), having the same or similar gene editing activity as the wild-type deaminase (SEQ ID NO. 1) and having higher gene editing activity should not be excluded from the scope of the mutant proteins disclosed herein due to the misplacement of amino acid numbering.
[0528] The muteins of the present invention are synthetic or recombinant proteins, i.e., they can be the product of chemical synthesis or produced using recombinant technology from prokaryotic or eukaryotic hosts (e.g., bacteria, yeast, plants). Depending on the host used in the recombinant production protocol, the muteins of the present invention can be glycosylated or non-glycosylated. The muteins of the present invention may or may not include an initial methionine residue.
[0529] The present disclosure also includes fragments, derivatives and analogs of the mutant protein. As used herein, the terms "fragment", "derivative" and "analog" refer to proteins that substantially retain the same biological function or activity of the mutant protein.
[0530] The mutant protein fragments, derivatives or analogs disclosed herein can be (i) mutant proteins in which one or more conservative or non-conservative amino acid residues (preferably conservative amino acid residues) are substituted, and such substituted amino acid residues may or may not be encoded by the genetic code, or (ii) mutant proteins with substitution groups in one or more amino acid residues, or (iii) mutant proteins formed by fusion of a mature mutant protein with another compound (such as a compound that extends the half-life of the mutant protein, such as polyethylene glycol), or (iv) mutant proteins formed by fusion of an additional amino acid sequence to the mutant protein sequence (such as a leader sequence or secretory sequence or a sequence or protein prosequence used to purify the mutant protein, or a fusion protein formed with an antigen IgG fragment). According to the teachings of this document, these fragments, derivatives and analogs are within the scope known to those skilled in the art.
[0531] In certain embodiments, other selected groups of amino acids are considered to be conservative substitutions for each other.
[0532] Table A
[0533] In certain embodiments, other selected groups of amino acids are considered to be conservative substitutions for each other (see, e.g., Creighton, Proteins (1984)):
[0534] Table B
[0535] In certain embodiments, other selected groups of amino acids that are considered conservative substitutions for each other are:
[0536] Table C
[0537] The active mutant protein disclosed herein has higher gene editing activity than the wild-type deaminase (SEQ ID NO. 1).
[0538] In addition, the mutant proteins disclosed herein can also be modified. Modifications (usually without changing the primary structure) include: chemical derivatization of the mutant proteins in vivo or in vitro, such as acetylation or carboxylation. Modifications also include glycosylation, such as those mutant proteins produced by glycosylation modification during the synthesis and processing of the mutant protein or in further processing steps. Such modifications can be accomplished by exposing the mutant protein to a glycosylation enzyme (such as a mammalian glycosylase or deglycosylation enzyme). Modified forms also include sequences with phosphorylated amino acid residues (such as phosphotyrosine, phosphoserine, phosphothreonine). Also included are mutant proteins that have been modified to improve their resistance to proteolysis or optimize their solubility.
[0539] The term "polynucleotide encoding a mutant protein" may include a polynucleotide encoding a mutant protein disclosed herein, or may also include additional coding and / or non-coding sequences.
[0540] The present disclosure also relates to variants of the aforementioned polynucleotides, which encode fragments, analogs, and derivatives of polypeptides or muteins having the same amino acid sequence as those disclosed herein. These nucleotide variants include substitution variants, deletion variants, and insertion variants. As is known in the art, an allelic variant is an alternative form of a polynucleotide, which may contain one or more nucleotide substitutions, deletions, or insertions that do not substantially alter the function of the encoded mutein.
[0541] The present disclosure also relates to polynucleotides that hybridize to the above-mentioned sequences and have at least 50%, preferably at least 70%, and more preferably at least 80% identity between the two sequences. The present disclosure particularly relates to polynucleotides that can hybridize to the polynucleotides described in the present disclosure under stringent conditions (or stringent conditions). In the present disclosure, "stringent conditions" refer to: (1) hybridization and elution at relatively low ionic strength and relatively high temperature, such as 0.2×SSC, 0.1% SDS, 60°C; or (2) the addition of a denaturing agent during hybridization, such as 50% (v / v) formamide, 0.1% calf serum / 0.1% Ficoll, 42°C, etc.; or (3) hybridization occurs only when the identity between the two sequences is at least 90%, more preferably at least 95%.
[0542] The muteins and polynucleotides of the present disclosure are preferably provided in isolated form, and more preferably, purified to homogeneity.
[0543] The full-length sequences of the polynucleotides disclosed herein can generally be obtained by PCR amplification, recombinant methods, or synthetic methods. For PCR amplification, primers can be designed based on the nucleotide sequences disclosed herein, particularly the open reading frame sequences, and amplified using commercially available cDNA libraries or cDNA libraries prepared by conventional methods known to those skilled in the art as templates to obtain the relevant sequences. When the sequences are long, two or more PCR amplifications are often required, followed by splicing the fragments amplified in the correct order.
[0544] Once the relevant sequence is obtained, it can be obtained in large quantities by recombinant methods. This is usually done by cloning it into a vector, then transferring it into cells, and then isolating the relevant sequence from the propagated host cells by conventional methods.
[0545] In addition, the sequences can also be synthesized by artificial synthesis, especially when the fragment length is shorter. Usually, a long fragment can be obtained by synthesizing multiple small fragments and then connecting them.
[0546] At present, the DNA sequence encoding the protein of the present invention (or its fragment, or its derivative) can be obtained completely by chemical synthesis. The DNA sequence can then be introduced into various existing DNA molecules (or vectors) and cells known in the art. In addition, mutations can also be introduced into the protein sequence of the present invention by chemical synthesis. The method of amplifying DNA / RNA using PCR technology is preferably used to obtain the polynucleotides of the present invention. In particular, when it is difficult to obtain full-length cDNA from a library, the RACE method (RACE-rapid amplification of cDNA ends) can be preferably used. The primers used for PCR can be appropriately selected based on the sequence information of the present invention disclosed herein and can be synthesized by conventional methods. The amplified DNA / RNA fragments can be separated and purified by conventional methods such as by gel electrophoresis.
[0547] Fusion protein
[0548] In one aspect, the present disclosure provides a fusion protein comprising any of the aforementioned deaminase variants and a nucleic acid programmable nucleotide binding domain.
[0549] In some embodiments, the nucleic acid programmable nucleotide binding domain is a Cas protein or an Ago protein.
[0550] In some embodiments, the Cas protein is selected from at least one of a type II CRISPR-Cas polypeptide, a type I CRISPR-Cas polypeptide, a type III CRISPR-Cas polypeptide, a type IV CRISPR-Cas polypeptide, a type V CRISPR-Cas polypeptide, a type VI CRISPR-Cas polypeptide, a type VII CRISPR-Cas polypeptide, an IscB polypeptide, a TnpB polypeptide, and an IsrB polypeptide.
[0551] In some embodiments, type I Cas proteins include I-A, I-B, I-C, I-D, I-E, and I-F proteins.
[0552] In some embodiments, the V-type Cas protein includes a Cas12 protein.
[0553] In some embodiments, the type VI Cas protein includes a Cas13 protein.
[0554] In some embodiments, the Cas protein is selected from Cas9, CasX, CasY, Cpf1, C2c1, C2c2, and C2c3. Cas9, CasX, CasY, Cas12a (Cpf1), Cas12b (C2cl), Cas13a (C2c2), Cas12c (C2c3), Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, Argonaute (Ago).
[0555] In some embodiments, the Cas protein is selected from Cas9 protein (e.g., SpCas9, SaCas9, GeoCas9, CjCas9, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9, xCas9, SpCas9-NG); Cas12 protein (e.g., Cas12a, AsCas12a, LbCas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f (Cas14), Cas12g, Cas12h, Cas12i , xCas12i, Cas12Max, hfCas12Max, Cas12j, Cas12k, Cas12l, Cas12m, Cas12n, Cas12o, Cas12p, Cas12q, Cas12r, Cas12s, Cas12t, Cas12u, Cas12v, Cas12w, Cas12x, Cas12y, Cas12z); Cas13 protein (e.g., Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, Cas13x, Cas13y); Csn2 and its mutants.
[0556] In some embodiments, the Cas protein has an amino acid sequence that has at least 95%, 96%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs. 41-45.
[0557] In some embodiments, the Cas protein is a nickase (nCas) having amino acids with at least 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO. 54.
[0558] In some embodiments, the Cas protein is a dCas protein having an amino acid sequence with at least 95%, 96%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs. 46-53 and 57.
[0559] In some embodiments, the Cas protein is a dCas protein having an amino acid sequence shown in SEQ ID NO.57.
[0560] In some embodiments, the AGO protein is selected from the group consisting of pAgo, eAgo, Ago1, Ago2, Ago3, and Ago4.
[0561] polynucleotides
[0562] In one aspect, the present disclosure provides a polynucleotide, which is a polynucleotide sequence encoding the deaminase variant, or a polynucleotide sequence encoding the aforementioned fusion protein, or a polynucleotide sequence encoding a base editing system containing the deaminase variant and a nucleic acid programmable nucleotide binding domain, or a polynucleotide sequence encoding a base editing system containing a fusion protein.
[0563] In one embodiment, the polynucleotide is a DNA molecule that is codon-optimized according to the codon preference of the host cell;
[0564] The optimization described in the present disclosure may require mutations of the nucleotide sequence encoding the protein (e.g., the deaminase variants of the present disclosure) to simulate the codon preferences of the intended host organism or cell when encoding the same protein. Thus, the codons may be changed, but the encoded protein remains unchanged. For example, if the intended target cell is a human cell, a nucleotide sequence encoding the protein optimized by human codons may be used. As another non-limiting example, if the intended host cell is an animal cell (e.g., a mouse cell, an insect cell), a nucleotide sequence encoding the protein optimized by the animal codons may be generated. As another non-limiting example, if the intended host cell is a plant cell, a nucleotide sequence encoding the protein optimized by plant codons may be generated.
[0565] Lists of codon usage are readily available, for example, at the "Codon Usage Database" available at www.kazusa.or.jp / codon. In some cases, the nucleic acids of the present disclosure comprise nucleotide sequences encoding deaminase variants, or fusion proteins thereof, which are codon-optimized for expression in eukaryotic cells. In some cases, the nucleic acids of the present disclosure comprise nucleotide sequences encoding deaminase variants, or fusion proteins thereof, which are codon-optimized for expression in animal cells. In some cases, the nucleic acids of the present disclosure comprise nucleotide sequences encoding deaminase variants, or fusion proteins thereof, which are codon-optimized for expression in fungal cells. In some cases, the nucleic acids of the present disclosure comprise nucleotide sequences encoding deaminase variants, or fusion proteins thereof, which are codon-optimized for expression in plant cells.
[0566] In one embodiment, the host cell comprises a prokaryotic cell or a eukaryotic cell;
[0567] CRISPR system
[0568] The terms “clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements associated with the expression of CRISPR-associated (“Cas”) genes, or transcripts or other elements capable of directing the activity of the Cas genes.
[0569] guide RNA (gRNA)
[0570] The terms "guide RNA (gRNA)", "mature crRNA", "crRNA", "guide sequence", and "guide RNA" are used interchangeably and have meanings commonly understood by those skilled in the art. Generally speaking, a guide RNA can comprise a direct repeat (DR) sequence and a spacer sequence, or consist essentially of or consist of a direct repeat (DR) sequence and a spacer sequence.
[0571] In some cases, the spacer sequence is any polynucleotide sequence that has sufficient complementarity with the target sequence to hybridize with the target sequence and guide the CRISPR-Cas complex to specific binding to the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between the spacer sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. The guide sequence comprises a sequence (e.g., a direct repeat (DR) sequence) that has sufficient complementarity with the target nucleic acid sequence to hybridize with the target nucleic acid sequence and guide the sequence-specific binding of the complex to the target nucleic acid sequence.
[0572] It is known in the art that complete complementarity is not required on the basis of sufficient complementarity to function, and therefore, if necessary, the cleavage efficiency can be adjusted by introducing mismatches (e.g., one or more mismatches between the spacer sequence and the target nucleic acid, such as 1 or 2 nucleotide mismatches (including the position of the mismatch along the spacer sequence / target sequence). For example, if a cleavage rate of less than 100% of the target is desired (e.g., in a cell population), 1 or 2 mismatches between the spacer sequence and the target sequence can be introduced into the spacer sequence.
[0573] Identity
[0574] "Identity" refers to the matching of sequences between two polypeptides or between two nucleic acids. "Identity" represents the percentage of residues that are identical between the polypeptide or nucleic acid sequences, and is calculated based on the total number of residues determined by the type of mutation. Mutation types include insertions (extensions) at either or both ends of the sequence, deletions (truncations) at either or both ends of the sequence, substitutions / alternations of one or more amino acids / nucleotides, insertions within the sequence, and deletions within the sequence.
[0575] Taking polypeptide sequences as an example, if the mutation type is one or more of the following: replacement / substitution of one or more amino acids / nucleotides, insertion within the sequence, and deletion within the sequence, the total number of residues is calculated based on the larger of the compared molecules. If the mutation type also includes insertions (extensions) at either or both ends of the sequence or deletions (truncations) at either or both ends of the sequence, the number of amino acids inserted or deleted at either or both ends (for example, the number of insertions or deletions at both ends is less than 20) is not included in the total number of residues. When calculating the percent identity, the sequences being compared are aligned in a manner that produces the maximum match between the sequences, and gaps in the alignment (if any) are resolved using a specific algorithm. The same applies to nucleotide identity calculations.
[0576] carrier
[0577] A vector is a nucleic acid molecule capable of transporting another nucleic acid molecule to which it has been linked.
[0578] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, no free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and other various polynucleotides known in the art. Vectors can be introduced into host cells by transformation, transduction, or transfection so that the genetic material elements they carry are expressed in the host cells. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including protein variants, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. The vector may also contain a replication initiation site.
[0579] The term "vector" refers to a plasmid or viral vector, wherein the plasmid is a circular double-stranded DNA loop that can be inserted into another DNA fragment by, for example, standard molecular cloning techniques. Viral vectors include virally derived DNA or RNA sequences that are present in vectors for packaging viruses, and viruses include, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. Viral vectors also include polynucleotides carried by viruses for transfection into a host cell. Some vectors (for example, bacterial vectors and episomal mammalian vectors with bacterial origins of replication) can replicate autonomously in the host cells into which they are introduced.
[0580] Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell upon introduction into the host cell and are replicated together with the host genome. Furthermore, some vectors are capable of directing the expression of genes to which they are operably linked. Such vectors are referred to as "expression vectors."
[0581] In some embodiments, the vector (e.g., a viral vector or a non-viral vector, such as a lentiviral vector or a plasmid) can be delivered to the target tissue by, for example, intramuscular injection, intravenous administration, transdermal administration, intranasal administration, oral administration, or mucosal administration. The above-mentioned delivery can be carried out via a single dose or multiple doses. It will be understood by those skilled in the art that the actual dose to be delivered herein can vary to a large extent according to a variety of factors, including but not limited to vector selection, target cells, organisms, tissues, the general condition of the subject to be treated, the degree of transformation / modification sought, the route of administration, the mode of administration, and the type of transformation / modification sought.
[0582] host cells
[0583] As used herein, "host cell" refers to a eukaryotic cell (e.g., an animal cell, a plant cell, a fungal cell, etc.), a prokaryotic cell (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or a cell from a multicellular organism cultured as a unicellular entity (e.g., a cell line), which serves as a recipient of nucleic acid (e.g., an expression vector) and includes descendants of the original cell that has been genetically modified by the nucleic acid.
[0584] It is understood that the progeny of a single cell may not necessarily have completely the same morphology, genome, etc. as the original parent cell due to natural, accidental, or deliberate mutation. A "recombinant host cell" (also called a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an expression vector, has been introduced.
[0585] Those skilled in the art will appreciate that the design of the expression vector may depend on factors such as the choice of the host cell to be transformed, the level of expression desired, and the like.
[0586] In another aspect, the present disclosure also provides a host cell or its progeny, wherein the host cell comprises the pre-deaminase variant, or the aforementioned fusion protein, or the aforementioned polynucleotide, or the aforementioned vector system, or the aforementioned base editing system, or the aforementioned composition.
[0587] In one embodiment, the host cell comprises a non-human mammal, human, insect, bird, reptile, amphibian, rodent, fish, worm, nematode, or yeast cell.
[0588] In one aspect, the present disclosure also provides a multicellular organism comprising the aforementioned cell or its progeny.
[0589] In one embodiment, the multicellular organism is an animal model or a plant model for a relevant disease.
[0590] operably connected
[0591] "Operably linked" means that the nucleotide sequence of interest is linked to regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). Advantageous vectors include lentiviruses and adeno-associated viruses, and the type of these vectors can also be selected to target specific cell types.
[0592] complementary
[0593] "Complementarity" refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence using traditional Watson-Crick or other non-traditional types. Percent complementarity represents the percentage of residues in one nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 5, 6, 7, 8, 9, or 10 out of 10 are complementary, resulting in percentages of 50%, 60%, 70%, 80%, 90%, and 100%). "Complete complementarity" means that all consecutive residues of one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. "Substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.
[0594] The term "stringent conditions" in relation to hybridization refers to conditions under which a nucleic acid having complementarity with a target sequence predominantly hybridizes to the target sequence and substantially does not hybridize to non-target sequences. Stringent conditions are generally sequence-dependent and depend on many factors. Generally speaking, the longer the sequence, the higher the temperature at which it will specifically hybridize to its target sequence.
[0595] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding of the bases between the nucleotide residues. The complex may comprise two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute one step in a broader process, such as the initiation of PCR, or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing to a given sequence is referred to as the "complement" of that given sequence.
[0596] Hybridization of the target sequence with the gRNA indicates that at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or represents that at least 12, 15, 16, 17, 18, 19, 20 or more bases of the nucleic acid sequences of the target sequence and the gRNA can complement each other and hybridize to form a complex.
[0597] Express
[0598] Nucleic acid expression includes one or more of production of an RNA template from a DNA sequence (e.g., transcription), processing of the RNA transcript (e.g., by splicing, editing, 5' capping, and / or 3' end processing), translation of the RNA into a polypeptide or protein, or post-translational modification of the polypeptide or protein.
[0599] deliver
[0600] "Delivery" refers to providing an entity (such as a drug) to a destination. For example, the components of the systems / compositions of the present disclosure can be delivered in various forms, such as DNA / RNA or RNA / RNA or protein / RNA combinations. For example, the deaminase variant can be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein.
[0601] In one aspect, the present disclosure further provides a delivery system comprising the deaminase variant, the fusion protein, the polynucleotide, or the composition.
[0602] In one embodiment, the delivery system further comprises a delivery vehicle, and the delivery vehicle comprises nanoparticles, liposomes, exosomes, microbubbles, a gene gun, or an electroporation device.
[0603] In addition, when the delivery object is a plant cell, a method such as cell penetrating peptide (CPP) delivery can also be adopted. For example, in one embodiment, a nuclease, a deaminase variant and / or at least one guide RNA is coupled to one or more CPPs, thereby effectively transporting the CPP coupled with a nuclease, a deaminase variant and / or a guide RNA into a plant cell (e.g., in the protoplast). CPP has a short peptide of less than 35 amino acids, which is derived from a protein or derived from a chimeric sequence and can transport biomolecules across the cell membrane in a non-receptor-dependent manner. CPP can be a cationic peptide, a peptide with a hydrophobic sequence, an amphipathic peptide, a peptide rich in proline and an antimicrobial sequence, and a chimeric or dichotomous peptide. CPP can penetrate biological membranes and therefore trigger different biomolecules to move across the cell membrane into the cytoplasm and improve their intracellular pathways and therefore promote the interaction between biomolecules and targets.
[0604] Exemplary CPPs include Tat (a nuclear transcription activator protein required for viral replication by HIV type 1), penetratin, Kaposi fibroblast growth factor (FGF) signal peptide sequence, integrin β3 signal peptide sequence, polyarginine peptide Arg sequence, guanine-rich molecule transporter, sweet arrow peptide, etc.
[0605] Lipid nanoparticles (LNPs)
[0606] The term "lipid nanoparticle" or "LNP" refers to a particle with nanometer scale (nm) (e.g., 1nm to 1,000nm) comprising one or more types of lipid molecules. LNP provided herein can further comprise at least one non-lipid payload molecule (e.g., one or more nucleic acid molecules). In some embodiments, LNP comprises a non-lipid payload molecule partially or completely encapsulated in a lipid shell. In particular, in some embodiments, wherein payload is a negatively charged molecule (e.g., mRNA encoding a viral protein), and the lipid component of LNP comprises at least one cationic lipid. It is contemplated that cationic lipids can interact with negatively charged payload molecules and promote payload incorporation and / or encapsulation into LNP during LNP formation. As provided herein, other lipids that can form a part for LNP include but are not limited to neutral lipids and charged lipids, such as steroids, polymer-conjugated lipids and various zwitterionic lipids. Nucleic acids can be encapsulated in cationic lipid particles (e.g., liposomes) by LNP, and can be delivered to cells relatively easily. In some instances, lipid nanoparticles do not contain any viral components, which helps to minimize safety and immunogenicity issues. Lipid particles can be used for in vitro, ex vivo, and in vivo delivery.Lipid particles can be used for cell populations of various sizes.
[0607] In some instances, LNP can be used to deliver DNA molecules (e.g., comprising nuclease-deaminase fusion protein and / or gRNA) and / or RNA molecules (e.g., mRNA of nuclease-deaminase fusion protein, gRNA). In some cases, LNP can be used to deliver RNP complexes of nuclease-deaminase fusion protein / gRNA.
[0608] The components of LNPs may include cationic lipids 1,2-dilinoleoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxyketo-N,N-dimethyl-3-aminopropane (DLinK-DMA), 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA), (3-o-[2-(methoxypolyoxy)amino]-1,2-dimethyl ...
[0015] LNPs may be prepared from poly(ethylene glycol) succinyl]-1,2-dimyristoyl-sn-glycerol (PEG-S-DMG), R-3-[(p-methoxy-poly(ethylene glycol) 2000)carbamoyl]-1,2-dimyristyloxypropyl-3-amine (PEG-C-DOMG), and any combination thereof. The preparation and encapsulation of LNPs may be adapted from Rosin et al., Molecular Therapy, Vol. 19, No. 12, pp. 1286-2200, December 2011).
[0609] In one embodiment, the LNP comprises one or more lipids of formula (I) (and subformulae thereof) as described herein (described in PCT application number PCT / CN2023 / 106421).
[0610] in,
[0611] R1 is selected from -OH and R a (R b )N-, where R a and R b are independently hydrogen, C1-C 10 Alkyl or C1-C 10 alkyl halide;
[0612] R2 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0613] R3 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl or R c -(CH2)n-, wherein n is a positive integer of 1-20, preferably a positive integer of 1-14, more preferably a positive integer of 1-10;
[0614] R c is selected from the following structures: in Represents the connection key;
[0615] L1, L2, L3, L4, L5 are none or are independently selected from the following groups:
[0616] X is none or -CH- or N;
[0617] R4 is none or C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0618] R5 is none or C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0619] R6 is C1-C 30 Alkyl, C2-C 30 Alkenyl, C2-C 30 Alkynyl;
[0620] R7 is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl;
[0621] R8 and R9 are none or independently C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl or -R h -C1-C 14 Alkyl, -R h -C2-C 14 Alkenyl, -R h -C2-C 14 Alkynyl, where R h O or S;
[0622] R 10 、R 11 Each independently is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl, or -R h -C1-C 14 Alkyl, -R h -C2-C 14 Alkenyl, -R h -C2-C 14 Alkynyl, where R h O or S.
[0623] In another embodiment, R a and R b Each is independently hydrogen, C1-C6 alkyl or C1-C6 haloalkyl.
[0624] In another embodiment, R2 is selected from -(CH2)m-, wherein m is a positive integer of 1-20, preferably a positive integer of 1-14, more preferably a positive integer of 1-10, and more preferably a positive integer of 1-6.
[0625] In another embodiment, R6 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl.
[0626] In another embodiment, R2 is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl.
[0627] In another embodiment, R2 and R7 are each independently C1-C 10Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl.
[0628] In another embodiment, R4 and R5 are each independently -(CH2)q-CH3, wherein q is selected from a positive integer of 1-20, preferably a positive integer of 1-12, and more preferably a positive integer of 1-8.
[0629] In another embodiment, R4 and R5 each independently have -R d -R e -structure;
[0630] where R d is none or selected from the following functional groups: -(CH2)-(CH=CH)-;
[0631] R e It is -(CH2)p-CH3, wherein p is selected from a positive integer of 0-20, preferably p is selected from a positive integer of 0-12, and more preferably p is selected from a positive integer of 0-8.
[0632] In another embodiment, R8 and R9 are each independently -(CH2)q-CH3, wherein q is selected from a positive integer of 1-14, preferably a positive integer of 1-12, and more preferably a positive integer of 1-8.
[0633] In another embodiment, R7 is -(CH2)q-CH3, wherein q is selected from a positive integer of 1-14, preferably a positive integer of 1-12, and more preferably a positive integer of 1-8.
[0634] In another embodiment, R6 has -R f -R g - structure;
[0635] where R f Selected from the following functional groups: -(CH2)s1-(CH=CH)-CH2-(CH=CH)-(CH2)s2-CH3, s1 is a positive integer of 1-14, preferably, a positive integer of 1-10, s2 is a positive integer of 1-8, preferably, a positive integer of 1-6;
[0636] R g It is -(CH2)m-, wherein m is a positive integer of 1-14, preferably a positive integer of 1-10, and more preferably a positive integer of 1-6.
[0637] In another embodiment, R 10 、R 11 Each independently is C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl or -Rh -C1-C 10 Alkyl, -R h -C2-C 10 Alkenyl, -R h -C2-C 10 Alkynyl, where R h O or S.
[0638] In another embodiment, R 11 It is C 1- C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl.
[0639] In another embodiment, R8 and R9 are each independently C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl or -R h -C1-C 10 Alkyl, -R h -C2-C 10 Alkenyl, -R h -C2-C 10 Alkynyl, where R h O or S.
[0640] In another embodiment, R9 is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl.
[0641] In another embodiment, X is absent or is -CH-.
[0642] In some embodiments, the substructures of the ionizable lipids are as shown in the following formulas I-1, I-2, I-3, and I-4, respectively:
[0643] In formula I-1, R1-R2, R6-R 11 ,L2-L5 are defined as above;
[0644] R3 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl.
[0645] In another embodiment, R3 is C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl.
[0646] In another embodiment, R3 is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl.
[0647] In formula I-2, R1-R2, R6-R7, R 10 -R 11 , the definitions of L1-L5 are as described above;
[0648] R3 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, or R c -(CH2)n-, wherein n is a positive integer of 1-20, preferably a positive integer of 1-14, more preferably a positive integer of 1-10, R c The definition of is as above;
[0649] R5 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0650] R8 and R9 are each independently C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl.
[0651] In another embodiment, R5 is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl.
[0652] In another embodiment, R5 is C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl.
[0653] In another embodiment, R3 is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl.
[0654] In another embodiment, R3 is C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 In formula I-3, R1-R3, R6, R7, R 10 、R 11 , the definitions of L1-L5 are as described above;
[0655] R4 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0656] R5 is C1-C20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl;
[0657] R8 and R9 are each independently C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl or -R h -C1-C 14 Alkyl, -R h -C2-C 14 Alkenyl, -R h -C2-C 14 Alkynyl, where R h is O or S;
[0658] In another embodiment, R4 is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl.
[0659] In another embodiment, R4 is C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl.
[0660] In another embodiment, R5 is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl.
[0661] In another embodiment, R8 and R9 are each independently C1-C 10 Alkyl, C2-C 10 Alkenyl, C2-C 10 Alkynyl.
[0662] In another embodiment, R9 is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl.
[0663] In formula I-4, R1-R7, R 10 -R 11 , L1-L3, and L5 are defined as above.
[0664] In one embodiment, the ionizable lipid has a structure selected from those shown in Table 1.
[0665] Table 1
[0666] In one embodiment, the present disclosure further provides a lipid nanoparticle (LNP), wherein the lipid nanoparticle comprises an ionizable lipid represented by formula (I).
[0667] In one embodiment, the lipid nanoparticle (LNP) comprises an ionizable lipid represented by formula (I), DSPC (distearoylphosphatidylcholine), cholesterol and DMG-PEG (polyethylene glycol-dimyristyl glyceride).
[0668] connector
[0669] As used herein, "linker" refers to a chemical group or molecule that connects two molecules or moieties, such as two domains of a fusion protein, such as a nuclease and a deaminase variant. In some connection schemes, a linker is located between or flanking two groups, molecules, or other moieties and connects the two by a covalent bond.
[0670] In some embodiments, the linker is a linear polypeptide formed by amino acids or multiple amino acid residues connected by peptide bonds. In some embodiments, the linker is an organic molecule, group, polymer or chemical moiety. The length and type of the linker can be designed as needed. In some embodiments, the linker can be selected from artificially synthesized amino acid sequences or naturally occurring polypeptide sequences.
[0671] Reagent test kit
[0672] In one aspect, the present disclosure provides a kit comprising the aforementioned deaminase variant, the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned composition, and the use of the aforementioned host cell in preparing the kit, wherein the components of the kit are in the same or different containers.
[0673] In one aspect, the present disclosure also provides a container comprising the aforementioned kit.
[0674] In one embodiment, the container comprises a sterile container;
[0675] In one embodiment, the container comprises a syringe.
[0676] In some embodiments, the kit also includes instructions for using the kit, such as instructions in more than one language. The kit may also include one or more reagents for use in the process of utilizing one or more of the above-mentioned components. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The above-mentioned reagents may be provided in a form (for example, in concentrated or lyophilized form) that requires the addition of one or more other components before use; the buffer may be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer and combinations thereof. The buffer may have a suitable pH value, for example, it may be alkaline. In some embodiments, the pH of the buffer is between about 7-10.
[0677] treat
[0678] "Treatment" refers to treating or curing the subject's condition, delaying the onset of the symptoms of the condition, and / or delaying the severity of the condition. The term "subject" includes, but is not limited to, various animals, plants, and microorganisms. Animals include mammals, such as bovines, equines, ovines, porcines, canines, felines, lagomorphs, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., a human) suffers from a condition (e.g., a condition caused by a disease-related gene defect). "Plant" is any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development.
[0679] In one aspect, the present disclosure also provides use of the aforementioned deaminase variant, the aforementioned fusion protein, the aforementioned polynucleotide, the aforementioned composition, and the aforementioned host cell in the preparation of a medicament for treating a condition or disease in a subject in need thereof.
[0680] In one embodiment, the use comprises administering the deaminase variant, fusion protein, base editing system, delivery composition, or host cell or composition or enzyme preparation to the subject or to cells ex vivo of the subject.
[0681] In one embodiment, the condition or disease includes the condition or disease being associated with one or more C>A point mutations or C>T point mutations; preferably, the condition or disease includes the diseases shown in Table A1 below:
[0682] Table A1
[0683] In some embodiments, the disease or condition comprises one or more of hypercholesterolemia, transthyretin amyloidosis, and beta-hemoglobinopathies.
[0684] The main advantages of the present disclosure include:
[0685] (a) The present disclosure obtains new deaminase variants by mutating wild-type deaminase. The deaminase variants disclosed herein have better gene editing activity than the wild-type deaminase, can effectively edit the target gene, and can effectively treat the condition or disease of a subject in need.
[0686] (b) The deaminases provided herein, when used in base editors and base editing systems, significantly improve editing efficiency. The base editing systems constructed therefrom can be used for base editing of nucleic acids. For example, base editors can be used to target adenine (A) in nucleic acids (e.g., DNA) to guanine (G). Such changes can alter the amino acid sequence of a protein, thereby destroying or creating new start codons, creating stop codons, disrupting splice donors, disrupting splice acceptors, or editing regulatory sequences, thereby correcting pathogenic genes and achieving therapeutic purposes.
[0687] Where specific conditions are not specified in the examples, the experiments were carried out under conventional conditions or the conditions recommended by the manufacturer. Reagents or instruments used without specifying the manufacturer are all conventional products that can be obtained commercially. It is understood by those skilled in the art that the examples describe the present disclosure by way of example and are not intended to limit the scope of protection claimed in the present disclosure. All publications and other references mentioned herein are incorporated herein by reference in their entirety. For any embodiment of the present disclosure described herein, including embodiments described only in the examples or claims or only in one aspect / part below, it should be understood that unless expressly denied or inappropriate in combination, the embodiments can be combined with any other one or more embodiments of the present disclosure. The experimental methods for which specific conditions are not specified in the following examples are generally carried out under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and weight parts.
[0688] Example 1: Obtaining deaminase mutants
[0689] (1) Obtaining 005V1 deaminase mutant
[0690] In order to construct an adenosine deaminase with higher editing efficiency and specificity, the amino acid site function prediction was performed on the amino acid sequence of the known adenosine deaminase 005V1 (deaminase 005V1 of CN114634923A, whose amino acid sequence is shown in SEQ ID NO: 1 and the nucleotide coding sequence is shown in SEQ ID NO: 2), and multiple positions that may improve editing efficiency and specificity were found. By site-directed PCR mutagenesis of the base editor 005V1-nCas9 expression vector containing the wild-type deaminase 005V1 (the amino acid sequence of the base editor 005V1-nCas9 is shown in SEQ ID NO: 13 and the nucleotide coding sequence is shown in SEQ ID NO: 14), an expression vector containing the adenosine deaminase 005V1 variant-nCas9 was constructed.
[0691] The amino acid sequence of 005V1-nCas9 is as follows: Among them, the bold sequence represents the sequence derived from nCas9; the italic represents the linker sequence; the double-underlined sequence represents the nuclear localization sequence; the single-underlined sequence is the 005V1 deaminase sequence; the asterisk at the C-terminus represents the stop codon position.
[0692] Base editors with different deaminase variants were generated by PCR-based site-directed mutagenesis. The specific method was to amplify the DNA sequence encoding the base editor 005V1-nCas9 (SEQ ID NO: 14) centered on multiple amino acids near the mutation site, and at the same time introduce the sequence to be mutated on the primer. Different mutant base editors were obtained by homologous recombination and connection of the amplified fragments (Table 2).
[0693] Table 2 Mutation patterns of 005V1 deaminase mutants
[0694] The specific mutation methods are as follows:
[0695] The plasmid expressing the 005V1-nCas9 base editor was used as a template, and the plasmid of the 005V1-nCas9 base editor was amplified using amplification primers containing the mutant sequence using Vazyme's high-fidelity enzyme kit (Vazyme, P501-d2).
[0696] The amplification system is shown in Table 3:
[0697] Table 3 Plasmid mutation amplification system of 005V1-nCas9 base editor
[0698] The PCR amplification program is shown in Table 4 below:
[0699] Table 4 Plasmid mutation amplification PCR program for 005V1-nCas9 base editor
[0700] The amplified PCR product was recovered and purified using a kit (Tiangen, Universal DNA Purification and Recovery Kit, DP214). The purified PCR product was transformed into Escherichia coli DH5a competent cells (Weidi Biotechnology, DL1001) and cultured. Single colonies were picked and, after sequencing confirmation, positive clones were shaken and plasmids were extracted using an endotoxin-free plasmid extraction kit (TIANGEN: DP120-01) and stored at -20°C until use.
[0701] Example 2. Verification of editing activity of deaminase mutants
[0702] To test the editing activity of each deaminase mutant, the PCSK9 target sequence was determined for the human PCSK9 gene, and the hPCSK9-sgRNA spacer sequence was designed: cccgcaccttggcgcagcgg (SEQ ID NO: 35).
[0703] The construction process of sgRNA expression vector (sgRNA plasmid) is as follows:
[0704] sgRNAs were designed based on the target sequence and oligonucleotides were synthesized. The sgRNA spacer sequences used were shown in SEQ ID NO: 35. A CACC sequence was added to the 5' end of the upstream sequence of each sgRNA spacer, and an AAAC sequence was added to the 5' end of the downstream sequence. The upstream and downstream primer sequences used for synthesis were hPCSK9-sgRNA-F (SEQ ID NO: 36) and hPCSK9-sgRNA-R (SEQ ID NO: 37), respectively.
[0705] After synthesis, the upstream and downstream sequences were annealed using a preset PCR program (95°C, 5 min; 95°C–85°C at −2°C / s; 85°C–25°C at −0.1°C / s; maintained at 4°C), and the annealed products were ligated into the lenti U6-sgRNA / EF1a-mCherry vector (Addgene, Plasmid, #114199) linearized with BbsI (NEB, R3539S).
[0706] Among them, the system used in the construction of sgRNA plasmid is as follows:
[0707] The linearization system of lenti U6-sgRNA / EF1a-mCherry vector is as follows: 3 μg of vector; 6 μL of buffer (NEB: R0539L); 2 μL of BbsI; ddH2O is added to 60 μL, and enzyme digestion is carried out at 37°C overnight.
[0708] The sgRNA annealing product and linearized vector ligation system is as follows: T4 ligase buffer (NEB, M0202L) 1 μL, linearized vector 20 ng, annealed oligo fragment (10 μM) 5 μL, T4 ligase (NEB: M0202L) 0.5 μL, ddH2O is added to 10 μL, and ligation is carried out at 16°C overnight.
[0709] The ligated vector was transformed into Escherichia coli DH5α competent cells (Weidi Biotech, DL1001). The specific process is as follows: Remove the DH5α competent cells from -80°C and quickly place them on ice. After 5 minutes, allow the bacterial mass to thaw. Add the ligated product and gently mix by flicking the bottom of the centrifuge tube. Let it rest on ice for 25 minutes. Heat shock the cells in a 42°C water bath for 45 seconds, quickly return them to ice, and let them rest for 2 minutes. Add 700 μL of sterile LB medium without antibiotics to the centrifuge tube, mix thoroughly, and then recover the cells at 37°C, 200 rpm, for 60 minutes. Centrifuge at 5000 rpm for one minute to harvest the cells. Collect approximately 100 μL of the supernatant, gently pipette to resuspend the bacterial mass, and spread the supernatant onto LB medium containing the Amp antibiotic. Incubate the plate upside down in a 37°C incubator overnight. Single colonies were picked, and after sequencing confirmation, the positive clones were shaken and the sgRNA plasmid was extracted using an endotoxin-free plasmid extraction kit (TIANGEN, DP120-01) and the concentration was measured. The plasmids were then stored in a -20°C refrigerator for later use.
[0710] HEK293T cells (purchased from ATCC) were seeded in DMEM medium (Gibco, 11965092) supplemented with 10% FBS (v / v) and 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) and cultured in a 37°C cell culture incubator with 5% CO2. Cells for transfection were seeded in 24-well cell culture plates the day before and cultured. The cells were observed the next day and transfected when the cells grew to a cell density of approximately 80%. The amount of editor fusion protein plasmid transfected per well of the 24-well plate was 0.4 μg, and the amount of sgRNA plasmid transfected was 0.4 μg.
[0711] After mixing the plasmids, dilute them with 25 μL of reduced serum medium (Source Bio, L530KJ), add 2 μL of p3000 reagent, pipette and mix well as reagent A, and let it stand for 5 minutes. At the same time, dilute 2 μL of Lipofectamine 3000 transfection reagent (Thermo, 11668019) with 25 μL of reduced serum medium and mix well as reagent B, and let it stand for 5 minutes. Mix the above reagents A and B and pipette evenly, and let it stand for 20 minutes. After standing, add the mixed reagent dropwise to the 24-well plate cells to be transfected and return them to the 37°C incubator for culture. After 6 hours of transfection, the culture medium was replaced with DMEM medium containing 10% FBS. Cells were collected 48 hours after transfection to detect editing efficiency.
[0712] The collected cells were subjected to genomic extraction (TIANGEN, DP304-03), and primers were designed according to experimental requirements. The identification primer sequences used were hPCSK9-forward primer (SEQ ID NO: 38) and hPCSK9-reverse primer (SEQ ID NO: 39).
[0713] Using the genome as a template, PCR amplification was performed on the sequence near the target site. The system used for target site sequence amplification was as follows: 2×Taq Master Mix (Vazyme, P112-03) 25 μL; Primer-F (10 pmol / μL) 1 μL; Primer-R (10 pmol / μL) 1 μL; template 1 μL; ddH2O was added to make up to 50 μL.
[0714] The amplified PCR products were used for high-throughput deep sequencing (Genwizhi Biotechnology Co., Ltd.) or Sanger sequencing (Boshang Biotechnology (Shanghai) Co., Ltd.) to identify the editing efficiency.
[0715] Gene editing effect detection:
[0716] For the method of calculating gene editing efficiency, see Kluesner MG, Nedveck DA, Lahr WS, Garbe JR, Abrahante JE, Webber BR, Moriarity BS. EditR: A Method to Quantify Base Editing from Sanger Sequencing. CRISPR J. 2018 Jun; 1(3): 239-250. doi: 10.1089 / crispr.2018.0014. PMID: 31021262; PMCID: PMC6694769.
[0717] In the same way, the 005V1-nCas9 base editor plasmid and sgRNA plasmid were co-transfected into HEK293T cells, and their editing efficiency was calculated.
[0718] In this embodiment, the structure of the base editor comprising the adenosine deaminase and nCas9 herein is as follows:
[0719] NH2-[NLS]-[adenosine deaminase]-linker-[nCas9]-[NLS]-COOH, the amino acid sequence of this structure can be found in SEQ ID NO: 13. However, this is only used as an example and is not intended to limit the structure of the base editor.
[0720] Illustratively, the amino acid sequences of base editors TA6312-nCas9, TA9171-nCas9, TA6332-nCas9, and TA9999-nCas9 are shown in SEQ ID NOs: 23-26, respectively, and their nucleotide coding sequences are shown in SEQ ID NOs: 27-30.
[0721] This example statistically analyzes the editing efficiency of 005V1-nCas9 and various variant-nCas9 base editors for the PCSK9 target (A6 site, the efficiency of mutation from adenine A to guanine G). Among them, the editing activity of 005V1-nCas9 and various variant-nCas9 at this site is shown in Table 5. Compared with 005V1-nCas9, the activity of the base editor composed of multiple mutants is significantly improved.
[0722] Table 5
[0723] Example 3. Design of sgRNA targeting PCSK9 and selection of adenine base editor (ABE)
[0724] 1. hPCSK9-sgRNA was designed based on the human PCSK9 gene locus. The specific sequence and modification method are as follows:
[0725] CCCGCACCUUGGCGCAGCGGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO. 69), wherein the underlined sequence is the spacer sequence, and the other sequence parts are scaffold sequences.
[0726] The modified hPCSK9-sgRNA sequence is as follows:
[0727] mPCSK9-sgRNA was designed based on the mouse PCSK9 gene locus. The specific sequence and modification method are as follows:
[0728] CCCAUACCUUGGAGCAACGGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO. 34), wherein the underlined sequence is the spacer sequence, and the other sequence parts are scaffold sequences.
[0729] The modified mPCSK9-sgRNA sequence is as follows: mC*mC*mC*rArUrArCrCrUrUrGrGrArGrCrArArCrGrGrGrUrUrUrArGrArGrCrUrArGrArArArUrArGrCrArArGrUrUrArArArArGrGrCrUrArGrUrCrCrGrUrUrArArArArUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCrArArCrUrUrGrArArArArGrUrGrGrCrUrArGrUrCrCrGrUrUrArUrCrArArCrUrUrGrArArArArGrUrGrGrCrArCrGrArGrUrUrCrUrGrCrUrUrGrArArArArArGrUrGrGrCrUrUrGrCrUrGrCrUrUrGrCrUrUrCrArArArArGrUrGrGrCrArCrCrGrArGrUrCrGrUrGrCrUrGrCrU*mU*mU*mU.
[0730] In the above sequences, capital nucleotides (A, C, G, and U) indicate ribonucleotides, adenine, guanine, cytosine, and uracil, respectively; m indicates 2'oxymethyl; * indicates phosphorothioate; and r indicates ribonucleotide.
[0731] The hPCSK9-sgRNA was synthesized by Nanjing GenScript Biosynthesis using chemical synthesis.
[0732] 2. Selection of Adenine Base Editors (ABEs)
[0733] The base editor TA9999-nCas9 (amino acid sequence shown in SEQ ID NO: 26, and nucleotide coding sequence shown in SEQ ID NO: 30) obtained by fusing TA9999 with nCas9 is used for targeted hPCSK9 target gene editing.
[0734] 3. Synthesis of ionizable lipids Yoltech Lipid 1-5 (corresponding to compounds 10-14, respectively)
[0735] (1) Synthesis of Yoltech Lipid 1 (i.e., compound 10):
[0736] 7-Butyl-21-(10-butyl-3,9-dioxyidene-2,8-dioxahexadecan-1-yl)-19-[3-(diethylamino)propyl]-8-oxyidene-19-aza-9-oxadocosan-22-yl 5-[(2-butyl-1-oxyoctylene)oxy]pentanoate
[0737] Step 1: Synthesis of compound 1-2
[0738] To a 500 mL round-bottom flask, cyclohexyl ester (25.00 g, 249.70 mmol, 1.0 eq), distilled water (20 mL), ethanol (200 mL), and sodium hydroxide (10.99 g, 274.67 mmol, 1.1 eq) were added. After reacting at 70°C for 3 hours, the solvent was removed by concentration under reduced pressure. 200 mL of acetone, tetrabutylammonium iodide (4.61 g, 12.48 mmol, 0.05 eq), and benzyl bromide (51.25 g, 299.64 mmol, 1.2 eq) were then slowly added to the flask. The reaction was then allowed to react at 70°C overnight. The reaction was quenched by the addition of 500 mL of water and extracted twice with 500 mL of ethyl acetate. The organic phases were combined, washed with saturated brine, dried over anhydrous sodium sulfate, concentrated under reduced pressure, and purified by column chromatography to yield 5-hydroxyvalerate benzyl ester (37.00 g, 71.2% yield).
[0739] Step 2: Synthesis of Compounds 1-4
[0740] A 500 ml round-bottom flask was charged with benzyl 5-hydroxyvalerate (37.00 g, 177.67 mmol, 1.0 eq), 2-butyloctanoic acid (35.59 g, 177.67 mmol, 1.0 eq), 250 ml of dichloromethane, and 4-dimethylaminopyridine (21.70 g, 177.67 mmol, 1.0 eq). Finally, 1-(3-dimethylaminopropyl)-3-ethylcarbodiimide hydrochloride (51.09 g, 266.50 mmol, 1.5 eq) was added. The mixture was reacted at room temperature for 4 hours, diluted with 500 ml of water, and extracted twice with 500 ml of dichloromethane. The organic phases were combined, washed with saturated brine, dried over anhydrous sodium sulfate, concentrated under reduced pressure, and purified by column chromatography to obtain 5-(benzyloxy)-5-oxypentyl 2-butyloctanoate (64.00 g, 92.2% yield).
[0741] Step 3: Synthesis of Compounds 1-5
[0742] A 250 mL round-bottom flask was charged with 5-(benzyloxy)-5-oxyidenepentyl 2-butyloctanoate (64.00 g, 163.87 mmol, 1.0 eq), methanol (75 mL), and tetrahydrofuran (75 mL). Finally, Pd / C (3.49 g, 32.78 mmol, 0.2 eq, 10% purity) was added. The mixture was reacted at room temperature under an atmospheric pressure of hydrogen for 16 hours. The mixture was filtered and concentrated to yield 5-[(2-butyl-1-oxyoctylene)oxy]pentanoic acid (45.00 g, 91.4% yield).
[0743] Step 4: Synthesis of Compounds 1-7
[0744] At room temperature, 5-[(2-butyl-1-oxyoctyl)oxy]pentanoic acid (10.00 g, 33.29 mmol, 1.0 eq), 2-hydroxymethylpropane-1,3-diol (3.53 g, 33.29 mmol, 1.0 eq), 4-dimethylaminopyridine (0.81 g, 6.66 mmol, 0.2 eq), N-(3-dimethylaminopropyl)-N'-ethylcarbodiimide hydrochloride (9.57 g, 49.94 mmol, 1.5 eq) and N,N-diisopropylethylamine (8.60 g, 66.58 mmol, 2.0 eq) were added to a round-bottom flask containing 100 ml of dichloromethane and stirred at room temperature for 4 hours. The reaction solution was quenched by adding 200 ml of water, extracted twice with 200 ml of dichloromethane, and the organic phases were combined, washed with brine, dried over anhydrous sodium sulfate, filtered, concentrated, and purified by column chromatography to give 2-butyloctanoic acid-18-butyl-8-(hydroxymethyl)-5,11,17-trioxy-6,10,16-trioxatetracosane-1-yl ester (7.80 g, yield 69.9%).
[0745] Step 5: Synthesis of Compound 1-8
[0746] At room temperature, the compound 2-butyloctanoate-18-butyl-8-(hydroxymethyl)-5,11,17-trioxydeca-6,10,16-trioxa-tetracosane-1-yl ester (3.90 g, 5.81 mmol, 1.0 eq) and triethylamine (1.76 g, 17.43 mmol, 3.0 eq) were added to 30 ml of dichloromethane, and methylsulfonic anhydride (2.02 g, 11.62 mmol, 2.0 eq) was slowly added at zero degrees Celsius. The temperature was slowly restored to room temperature and the reaction was allowed to react for 4 hours. The reaction solution was quenched by adding 30 ml of water, extracted twice with 50 ml of dichloromethane respectively, the organic phases were combined, washed with brine, dried over anhydrous sodium sulfate, filtered, concentrated, and purified by column chromatography to give methanesulfonic acid-12-butyl-2-(10-butyl-3,9-dioxy-2,8-dioxahexadecan-1-yl)-5,11-dioxy-4,10-dioxahexadecan-1-yl ester (3.85 g, yield 88.4%).
[0747] Step 6: Synthesis of Compound 1-10
[0748] At room temperature, compound 1-8 (600.0 mg, 0.80 mmol, 1.0 eq), 3-amino-1-propanol (300.0 mg, 3.99 mmol, 5.0 eq), potassium carbonate (280.0 mg, 2.00 mmol, 2.5 eq), and potassium iodide (130.0 mg, 0.80 mmol, 1.0 eq) were added to 10 ml of acetonitrile, protected by nitrogen, heated to 90 degrees Celsius, and reacted for 16 hours. The reaction mixture was concentrated, diluted with water, and extracted three times with ethyl acetate. The organic phases were combined, washed with saturated brine, dried over anhydrous sodium sulfate, concentrated, and purified by column chromatography to obtain the compound 5-[(2-butyl-1-oxyoctyl)oxy]pentanoic acid-12-butyl-2-{[(3-hydroxypropyl)amino]methyl}-5,11-dioxyidene-4,10-dioxaoctadec-1-yl ester (210.0 mg, 36.11%). MS: m / z [M+H] + =728.6.
[0749] Step 7: Synthesis of compound 10
[0750] At room temperature, the compound 2-butyloctanoic acid-8-(10-butyl-3,9-dioxyidene-2,8-dioxahexadecan-1-yl)-14-ethyl-5-oxyidene-10,14-diaza-6-oxahexadecan-1-yl ester (500.0 mg, 0.64 mmol, 1.0 eq), 2-butyloctanoic acid-9-bromononyl ester (390.0 mg, 0.96 mmol, 1.5 eq), potassium carbonate (270.0 mg, 1.92 mmol, 3.0 eq), potassium iodide (110.0 mg, 0.64 mmol, 1.0 eq) were added to 20 ml of acetonitrile, protected by nitrogen, heated to 90 degrees Celsius, and reacted overnight. The reaction mixture was concentrated, diluted with water, and extracted three times with dichloromethane. The organic phases were combined, washed with saturated brine, dried over anhydrous sodium sulfate, concentrated, and purified by column chromatography to obtain 7-butyl-21-(10-butyl-3,9-dioxyidene-2,8-dioxahexadecan-1-yl)-19-[3-(diethylamino)propyl]-8-oxyidene-19-aza-9-oxadocosan-22-yl 5-[(2-butyl-1-oxyoctyl)oxy]pentanoate (132.8 mg, 18.8% yield). MS: m / z [M+H] + =1107.9. 1 HNMR(300MHz, CDCl3)δ4.15-4.01(m,10H),3.44-3.20(m,4H),2.71-2.50(m,6H), 2.39-2.23(m,12H),2.02-1.40(m,28H),1.38-1.22(m,48H),0.92-0.75(m,18H).
[0751] (2) Synthesis of Yoltech Lipid2 (i.e., compound 11)
[0752] ((2-((7-((2-butyloctanoyl)oxy)-N-(3-(diethylamino)propyl)heptylamino)methyl)propane-1,3-diyl)bis(oxy))bis(5-oxopentane-5,1-diyl)bis(2-butyloctanoate)
[0753] Compound 11 was prepared from 7-chloro-7-oxoheptyl 2-butyl octanoate according to the preparation steps of compound 6, and 7-chloro-7-oxoheptyl 2-butyl octanoate was prepared according to the preparation steps of 6-3.
[0754] Yield: 31.4%.
[0755] MS:m / z[M+H] + =1093.9. 1 H NMR (400MHz, CDCl3) δ4.16-4.02(m,10H),3.56-3.14(m,2H),2.53(q,J=7.1Hz,3H),2.42-2.31(m,10H),2.03(d,J=48.7Hz,1H),1.70(d,J=6. 5Hz,9H),1.66-1.57(m,9H),1.46(dd,J=13.0,5.3Hz,6H),1.40-1.36(m,4H),1.33-1.23(m,40H),1.03(t,J=7.2Hz,6H),0.96-0.83(m,20H).
[0756] Among them, the preparation method of compound 6 ((2-((5-((2-butyloctanoate)oxy)-N-(3-(diethylamino)propyl)pentanamide)methyl)propane-1,3-diol)bis(oxy))bis(5-oxopentanoic acid-5,1-diyl)bis(2-butyloctanoate)) is as follows:
[0757] Intermediate 6-3: 5-chloro-5-oxopentyl-2-butyloctanoate
[0758] 1-5 (2.00 g, 6.66 mmol, 1.0 equiv) and DMF (49.0 mg, 0.67 mmol, 0.1 equiv) were dissolved in DCM (20 mL) and oxalyl chloride (2.54 g, 19.98 mmol, 3.0 equiv) was slowly added. Under a nitrogen atmosphere, the reaction solution was stirred at 0° C. for 3 hours. It was then concentrated under reduced pressure to give the desired product as a colorless oil (2.00 g, 94.2% yield), which was used directly in the next step without further purification.
[0759] Intermediate 6-2: ((2-(((3-(diethylamino)propyl)amino)methyl)propane-1,3-diol)bis(oxy))bis(5-oxopentanoic acid-5,1-diyl)bis(2-butyloctanoate)
[0760] 1-9 (1.00 g, 1.34 mmol, 1.0 equiv), N,N-diethylpropane-1,3-diamine (870.0 mg, 6.70 mmol, 5.0 equiv), KCO (560.0 mg, 4.02 mmol, 3.0 equiv), and KI (220.0 mg, 1.34 mmol, 1.0 equiv) were dissolved in CHCN (20 mL) and stirred at 90°C for 16 hours. After the reaction, the reaction solution was diluted with EtOAc (20 mL x 2). The organic layers were combined, dried, and then dried over anhydrous NaSO. The resulting mixture was then concentrated under reduced pressure. The residue was purified by silica gel column chromatography (DCM / MeOH = 50 / 1 to 10 / 1) to yield the desired product as a colorless oil (700.0 mg, 66.9% yield). Mass spectrum (MS): m / z [M+H]+ = 783.6.
[0761] Compound 6: ((2-((5-((2-butyloctanoate)oxy)-N-(3-(diethylamino)propyl)pentanamide)methyl)propane-1,3-diol)bis(oxy))bis(5-oxopentanoic acid-5,1-diyl)bis(2-butyloctanoate)
[0762] 6-2 (500.0 mg, 0.64 mmol, 1.0 equiv) and TEA (190.0 mg, 1.92 mmol, 3.0 equiv) were dissolved in DCM (2 mL), and a solution of 5-chloro-5-oxopentyl-2-butyloctanoate (310.0 mg, 0.96 mmol, 1.5 equiv) in DCM (1 mL) was added. The reaction solution was stirred at 25° C. for 3 hours. After the reaction, it was diluted with H 2 O (10 mL), extracted with DCM (10 mL x 2), and the organic layers were combined and dried and then concentrated under reduced pressure. The residue was purified by silica gel column chromatography (DCM / MeOH=40 / 1 to 15 / 1) to give the desired product as a colorless oil (410.0 mg, 60.3% yield). Mass spectrum (MS): m / z[M+H]+=1065.9. 1 H NMR (400 MHz, CDCl3) δ 4.28-3.95 (m, 11H), 3.46-3.26 (m, 5H), 2.87-2.64 (m, 3H), 2.57 (dd, J = 14.2, 7.0 Hz, 3H), 2.50-2.25 (m, 12H), 1.80-1.52 (m, 22H), 1.44 (dd, J = 24.0, 18.7 Hz, 8H), 1.33-1.20 (m, 34H), 1.04 (t, J = 7.1 Hz, 4H), 0.88 (dt, J = 7.3, 3.6 Hz, 16H). (3) Synthesis of Yoltech Lipid3 (i.e., compound 12)
[0763] ((2-((9-((2-butyloctanoyl)oxy)-N-(3-(diethylamino)propyl)nonylamino)methyl)propane-1,3-diyl)bis(oxy))bis(5-oxopentane-5,1-diyl)bis(2-butyloctanoate)
[0764] Compound 12 was prepared according to the preparation steps of Compound 6 from 9-chloro-9-oxononyl 2-butyloctanoate, which was prepared according to Step 6-3.
[0765] Yield 41.2%.
[0766] MS:m / z[M+H] + =1121.9. 1H NMR (400MHz, CDCl3) δ4.15-4.04(m,10H),3.61-3.18(m,2H),2.53(q,J=7.1Hz,4H),2.43-2.31(m,10H),1.90(s,1H),1.70(d,J =4.2Hz,10H),1.61(dd,J=14.0,7.5Hz,9H),1.50-1.42(m,6H),1.35-1.24(m,47H),1.03(t,J=7.1Hz,6H),0.92-0.88(m,19H).
[0767] (4) Synthesis of Yoltech Lipid4 (i.e., compound 13)
[0768] ((2-((5-((2-butyloctanoyl)oxy)pentyl)(3-(diethylamino)propyl)amino)methyl)propane-1,3-diyl)bis(oxy))bis(5-oxopentane-5,1-diyl)bis(2-butyloctanoate)
[0769] According to the preparation procedure of compound 10, compound 13 was prepared from 5-bromopentyl 2-butyloctanoate.
[0770] Yield 44.4%.
[0771] MS:m / z[M+H] + =1051.9. 1 H NMR (400MHz, CDCl3) δ4.22-3.99(m,10H),2.81-2.50(m,7H),2.47-2.26(m,13H),2.23-2.07(m,2H),1.78-1.52 (m,18H),1.46(dd,J=14.6,6.9Hz,8H),1.38-1.19(m,37H),1.11(t,J=6.8Hz,6H),0.89(dd,J=9.4,4.4Hz,17H).
[0772] (5) Synthesis of Yoltech Lipid5 (i.e., compound 14)
[0773] ((2-((7-((2-butyloctanoyl)oxy)heptyl)(3-(diethylamino)propyl)amino)methyl)propane-1,3-diyl)bis(oxy))bis(5-oxopentane-5,1-diyl)bis(2-butyloctanoate)
[0774] According to the preparation procedure of compound 10, compound 14 was prepared from 7-bromoheptyl 2-butyloctanoate.
[0775] Yield: 23.6%.
[0776] MS:m / z[M+H] + =1079.9. 1 H NMR(300MHz, CDCl3)δ4.17-3.96(m,10H),3.44-3.15(m,4H),2.77-2.54(m,6H),2.40-2.30(m,8H), 2.28-2.19(m,4H),2.10-1.73(m,20H),1.70-1.40(m,8H),1.40-1.18(m,44H),0.92-0.75(m,18H).
[0777] Example 4. Cell-level PCSK9 gene base editing experiment
[0778] The DNA transcription template of TA9999-nCas9 mRNA (SEQ ID NO.31) was synthesized by Nanjing GenScript Biotechnology and used T7 High Yield RNA Synthesis Kit (NEB, E2040S) was used for in vitro transcription reaction to obtain TA9999-nCas9 mRNA.
[0779] In this example, the adenine base editor ABE8e (amino acid sequence shown in SEQ ID NO: 32, nucleotide coding sequence shown in SEQ ID NO: 33) was used as a comparison. The ABE8e plasmid was purchased from Addgene (Addgene, Plasmid #138489) and expressed and purified in the laboratory to obtain ABE8e mRNA. This example used the modified hPCSK9-sgRNA described in Example 3.
[0780] HepG2 cells (purchased from ATCC) were seeded in DMEM medium (Gibco, 11965092) supplemented with 10% FBS (v / v) and 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) and cultured in a 37°C cell culture incubator with 5% CO2. For transfection, cells were seeded in 96-well cell culture plates the day before and cultured. The cells were observed the next day. When the cells grew to a cell density of approximately 80%, TA9999-nCas9 mRNA and hPCSK9-sgRNA were delivered into the cells using LNP technology. The amounts of TA9999-nCas9 mRNA and hPCSK9-sgRNA used per well are shown in Table 6.
[0781] Table 6
[0782] Preparation of LNP@mRNA: A four-component LNP lipid, including Yoltech Lipid1 (i.e., compound 10), DSPC, cholesterol, and PEG-DMG, was used. Yoltech Lipid, DSPC, cholesterol, and PEG-DMG were dissolved in anhydrous ethanol at a molar ratio of 50:10:38.5:1.5. TA9999-nCas9 mRNA and hPCSK9-sgRNA (mass ratio of 1:1) prepared according to the dosage in Table 6 were dissolved in 100mM enzyme-free citrate buffer at pH 4 (RNA concentration of 0.2mg / mL). The ethanol solution of the lipid carrier was mixed with the mRNA buffer at a ratio of 1:3 (volume / volume) (wherein the mass ratio of total lipid to mRNA was 40:1), and nucleic acid lipid nanoparticles were obtained by microfluidic nanodrug manufacturing system (NanoAssemblr Ignite, Canada) at a flow rate of 12ml / min. The obtained nucleic acid lipid nanoparticles were immediately diluted into 1×DPBS buffer at 40 times the volume. Cells were collected 48 hours after LNP lipid transfection to detect editing efficiency.
[0783] The collected cells were subjected to genomic extraction (TIANGEN, DP304-03), and primers were designed according to experimental requirements. The identification primer sequences used were hPCSK9-F (SEQ ID NO: 38) and hPCSK9-R (SEQ ID NO: 39).
[0784] Using the genome as a template, PCR amplification of sequences near the target site was performed. The system used for target site sequence amplification was as follows: 25 μL of 2× Taq Master Mix (Vazyme, P112-03); 1 μL of Primer-F (10 pmol / μL); 1 μL of Primer-R (10 pmol / μL); 1 μL of template; and ddH2O was added to 50 μL. The amplified PCR products were used for high-throughput deep sequencing (Genwizhi Biotechnology Co., Ltd.) or Sanger sequencing (Boshang Biotechnology (Shanghai) Co., Ltd.) to assess editing efficiency (Figure 1). Analysis showed that the editing activity was much higher when the ABE mRNA and sgRNA were dosed at 0.5 ng per well than when both were dosed at 0.25 ng. At both doses, the editing activity of TA9999-nCas9 at the target site was significantly higher than that of ABE8e (blank control, ABE mRNA and msgRNA were replaced with the same dose of PBS).
[0785] Similarly, the editing activity of TA6312-nCas9, TA5160-nCas9, TA6166-nCas9, TA7904-nCas9, TA7694-nCas9, TA6332-nCas9, TA3154-nCas9, TA1574-nCas9, TA6547-nCas9, TA1274-nCas9, TA8952-nCas9, TA3119-nCas9, TA7730-nCas9, and TA4324-nCas9 targeting the PCSK9 gene in HepG2 cells was determined in the same manner. The mRNA sequences of the above base editors are shown in Table 7.
[0786] The DNA transcription templates for TA6312-nCas9, TA5160-nCas9, TA6166-nCas9, TA7904-nCas9, TA7694-nCas9, TA6332-nCas9, TA3154-nCas9, TA1574-nCas9, TA6547-nCas9, TA1274-nCas9, TA8952-nCas9, TA3119-nCas9, TA7730-nCas9, and TA4324-nCas9 mRNA were synthesized by Nanjing GenScript Biotechnology Co., Ltd., as shown in Table 7.
[0787] Table 7
[0788] LNP delivery was performed as described above, and the editing efficiency was measured under different doses (0.5 ng, 1 ng, ABE mRNA to sgRNA mass ratio was 1:1). The results are shown in Table 8.
[0789] Table 8
[0790] Next, the effect of delivery dose on editing efficiency was tested (TA9999-nCas9 was used as the test object). Different doses (see Table 9) were set for LNP delivery experiments, and the editing efficiency and protein expression were measured.
[0791] Table 9
[0792] A human PCSK9 ELISA kit (Abcam, Cat No. ab209884) was used to detect PCSK9 protein content in total cell protein. The results are shown in Table 9 and Figure 2. Analysis showed that the editing efficiency reached 60% at an addition dose of 2 ng / well, 75% at an addition dose of 4 ng / well, and nearly 90% at an addition dose of 8 ng / well. At an addition dose of 16 ng / well, a saturated editing efficiency of 95% was achieved, with a PCSK9 protein content of only 37 ng / ml.
[0793] Example 4. PCSK9 gene base editing experiment in animals
[0794] A lipid nanoparticle formulation (LNP formulation) loaded with TA9999-nCas9 and mPCSK9-sgRNA was obtained according to the method of Example 3. For comparison, a lipid nanoparticle formulation (LNP formulation) containing ABE8e mRNA and mPCSK9-sgRNA was also prepared according to the method of Example 3. The sgRNA in this example was modified according to the modification method of Example 3.
[0795] Nine C57BL / 6 mice aged 6-7 weeks and weighing about 20 g (purchased from Jicui Yaokang) were used as experimental subjects and randomly divided into an experimental group (n=3 in each of the TA9999-nCas9 and ABE8E groups) and a control group (n=3). The LNP preparation was administered to C57BL / 6 mice (purchased from Jicui Yaokang) by intravenous injection (IV) of 0.05mpk and 2mpk of total RNA. One week after administration, the mice were euthanized and liver tissue was collected. Genomic DNA was extracted, and deep sequencing analysis was performed on the mPCSK9 gene site to determine the base editing activity (blank control, ABE mRNA and mPCSK9-sgRNA were replaced with equal doses of PBS). The assay primers used are as follows: upstream primers mPCSK9-F (SEQ ID NO: 84) and mPCSK9-R (SEQ ID NO: 85).
[0796] After high-throughput deep sequencing (Genwizhi Biotechnology Co., Ltd.), the editing efficiency of each group was found as shown in Table 10.
[0797] Table 10
[0798] Analysis showed that TA9999-nCas9 and mPCSK9-sgRNA showed efficient editing activity on the PCSK9 site, with editing activities reaching 33.6% (0.05 mpk) and 45% (2 mpk) (Figure 3). Under the same dosage (0.05 mpk), the editing efficiency of TA9999-nCas9 was much higher than that of ABE8e.
[0799] Mice (0.05 mpk dose group) were also blood drawn for LDL-C and PCSK9 protein levels. Dosing occurred in the morning of the day of administration, with the LDL-C and PCSK9 protein levels in the mice's blood on the day of administration defined as 100%. All mice were fasted for 4-5 hours at 1-6, 12, and 15 months after the first dose. Blood was then collected from the eye sockets, and plasma was separated for LDL-C levels. PCSK9 protein levels were also measured using a mouse PCSK9 ELISA kit (Abcam, Cat No. ab215538). Analysis revealed that PCSK9 protein levels had decreased by 80% relative to baseline levels one month after dosing. At 15 months after dosing, PCSK9 protein levels remained just above 20%. One month after dosing, LDL-C levels had significantly decreased to approximately 40% relative to baseline, and after 12 months after dosing, LDL-C levels remained below 40% (Figure 4).
[0800] In addition, the same method was used to measure the editing efficiency of other base editors in mice, and the results were as follows: TA6312-nCas9 (8%), TA5160-nCas9 (5%), TA6166-nCas9 (19%), TA7904-nCas9 (12%), TA6547-nCas9 (29%), and TA1274-nCas9 (33%).
[0801] Yoltech Lipid2-5 and cationic lipid ALC-0315 (purchased from Avituo (Shanghai) Pharmaceutical Technology Co., Ltd.) were configured in the above proportions and manner with DSPC, cholesterol, PEG-DMG, and TA9999-nCas9 mRNA and mPCSK9-sgRNA (modified using the modification method of Example 1) as a comparison. C57BL / 6 mice were randomly divided into a control group (n=3), a YolTech-lipid group, and an ALC-0315 group. The mice were injected intravenously into the tail of the mice according to the above method (injection dose of 0.1 mpk), and the editing activity caused by the LNP preparations in each group was measured. The test results are shown in Figure 5. Analysis shows that the editing activity caused by the LNP preparation of YolTech-lipid1 can reach more than 60%, while the editing activity caused by the LNP preparation of ALC-0315 is only about 20%.
[0802] Example 5. PCSK9 gene editing in a non-human primate model
[0803] Six cynomolgus macaques weighing 3-4 kg (purchased from Lingkang Sinoco Biotechnology Co., Ltd.) were used. The animals were quarantined and acclimated for 14 days before use. The room temperature was maintained at 18-26°C, the relative humidity was 40-70%, and the light cycle was 12 hours per day with alternating light and dark. The animals had free access to water during the experiment.
[0804] According to the method of Example 3, a lipid nanoparticle formulation (LNP formulation) was prepared with YolTech-lipid1, DSPC, cholesterol, PEG-DMG, TA9999-nCas9 mRNA, and hPCSK9-sgRNA (SEQ ID NO. 69). The hPCSK9-sgRNA was chemically modified as in Example 3 and administered to cynomolgus monkeys by intravenous injection at a dose of 3 mpk. Two weeks later, a biopsy was performed to assess base editing, and the editing activity was detected by NGS sequencing, which showed an editing activity of 76.22% (Figure 6).
[0805] Taking the day of dosing as Day 0, the PCSK9 protein level in the plasma of cynomolgus macaques was measured at approximately 62 ng / ml three days before dosing. Blood was collected three days, 67 days, three months, six months, and nine months after the first dose. After plasma separation, PCSK9 protein levels were measured using a human PCSK9 ELISA kit (Abcam, Cat No. ab209884). The results showed that PCSK9 protein levels significantly decreased from three days to nine months after dosing. At nine months, the plasma PCSK9 protein level had decreased by more than 70% relative to the baseline concentration and remained at 20 ng / ml (Figure 7).
[0806] Example 6 Construction of new base editors and cell experiments
[0807] In order to detect the activity of the base editor constructed by the deaminase variant disclosed herein, TA7904 was also fused to the Cas12 protein to form a new base editor. Specifically, the deaminase TA7904 (amino acid sequence as shown in SEQ ID NO.55, the nucleotide sequence after human codon optimization is shown in SEQ ID NO.56), and the Cas12 protein is the dCas protein dC05440 (amino acid sequence as shown in SEQ ID NO.57, the human optimized nucleotide sequence is shown in SEQ ID NO.58).
[0808] The deaminase domain of the base editor fusion protein can be fused to the N-terminus or C-terminus of the dCas domain, and its structure includes NH2-[adenosine deaminase domain]-[dCas]-[NLS] and NH2-[dCas]-[adenosine deaminase domain]-[NLS], and "-" refers to an optional linker. In this embodiment, TA7904 was fused to the N-terminus and C-terminus of dC05440, respectively, to obtain TA7904-dC05440 (amino acid sequence as shown in SEQ ID NO.59, and the nucleotide sequence after human codon optimization as shown in SEQ ID NO.60) and dC05440-TA7904 (amino acid sequence as shown in SEQ ID NO.61, and the nucleotide sequence after human codon optimization as shown in SEQ ID NO.62).
[0809] TA7904-dC05440 mRNA (DNA coding sequence shown in SEQ ID NO. 60) and dC05440-TA7904 mRNA (DNA coding sequence shown in SEQ ID NO. 62) were synthesized by Nanjing GenScript Biotechnology.
[0810] crRNAs were designed for the human Hao1 gene and KLF4 gene targets, and the crRNAs were chemically modified, as shown in Table 11.
[0811] Table 11
[0812] Among them, the underlined sequence is the DR (Direct Repeat) sequence of crRNA, m indicates 2'oxymethyl, and * indicates thiophosphate.
[0813] HepG2 cells (purchased from ATCC) were seeded in DMEM medium (Gibco, 11965092) supplemented with 10% FBS (v / v) and 1% Penicillin Streptomycin (v / v) (Gibco, 15140122) and cultured in a 37°C cell culture incubator with 5% CO2. Cells for transfection were seeded in 96-well cell culture plates the day before and cultured. The cells were observed the next day and mRNA transfection was performed when the cells reached a cell density of approximately 80%.
[0814] Using Lipofectamine TM MessengerMAX TM Transfection reagent (Thermo Fisher Scientific, LMRNA015) was used and transfection was performed according to the product instructions (doses of 100 ng, 200 ng, and 400 ng were set, with a base editor mRNA and crRNA weight ratio of 1:1). 48 h after transfection, cells were collected and the genome was extracted. The target fragments were PCR amplified using the primers shown in Table 12, and high-throughput sequencing was performed (Genwizhi Biotechnology Co., Ltd.).
[0815] Table 12
[0816] After sequencing, the editing activity was calculated (Figure 8, AB). Analysis showed that at the hHao1 gene target site, both base editors mediated significant editing efficiency at sites +5, +7, +13, and +16 under multiple dosage conditions. At sites +13 and +16, the editing activity was no less than 26%. At doses of 200 ng and 400 ng, both TA7904-dC05440 and dC05440-TA7904 mediated higher editing efficiency, and dC05440-TA7904 mediated higher editing efficiency. 0-TA7904 had higher editing efficiency than TA7904-dC05440 at doses of 200ng and 400ng; at the hKLF4 gene target site, both base editors had significant editing efficiency at multiple sites under multiple dose conditions, while dC05440-TA7904 mediated higher editing activity at the +9 and +11 sites under different dose conditions, especially at a dose of 200ng, mediating no less than 25% editing activity.
[0817] Example 7. Used for the treatment of other base editing-related diseases.
[0818] The disease is obtained from the NCBI ClinVar database available on the NCBI ClinVar website, for example, can be selected from the base-edited disease targets shown in Table A1.
[0819] The sequences involved in the present disclosure are as follows:
[0820] All documents mentioned in this disclosure are incorporated herein by reference, just as if each document were incorporated herein by reference individually. It should also be understood that after reading the above teachings of this disclosure, those skilled in the art may make various changes or modifications to this disclosure, and that such equivalents also fall within the scope of the claims appended hereto.
Claims
1. A deaminase variant, characterized in that The variant is a non-natural protein, and the variant is mutated at the core amino acid position shown below corresponding to SEQ ID NO. 1, or is mutated at the corresponding position of another adenosine deaminase: (a)A46; (b)I47; (c) T48; (d) L49; (e) V104; (f)Q148; (g) P150; (h)E152; (i) V153; (j) F154; and (k)N155. Preferably, the deaminase variant further comprises mutations in the core amino acid positions shown below: (i) Q66; (ii) I67; (iii) V68; (iv) Q69; (v) C139; (vi) S140; (vii) M142; (viii) L163; (ix) N164; (x) Q165; and (xi)P166. Preferably, the deaminase variant further comprises mutations in the core amino acid positions shown below: (i) D167; (ii) R168; (iii) A169; and (iv) D170; and mutations in one or more core amino acid positions selected from the group consisting of: (a)A156; (b) E157; (c) R158; (d) S2; (e) E3; (f) L4; (g)N5; (h)A15; (i) L16; (j)Q18; (k) K19; (l)A20; (m)R21; (n)Q66; (o)I67; (p)V68; (q)Q69; (r)C139; (s)S140; (t)M142; (u)E141; (v)A111; (w)A112; (x)G113. Preferably, the deaminase variant further comprises mutations in the core amino acid positions shown below: (a) C144; (b) Q145; and (c)Q149. Preferably, the mutation is a combination of the following mutations in the amino acid sequence as shown in SEQ ID NO: 1: (1)A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; (4)S2+E3+L4+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; (7)A15+L16+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; (8)S2+E3+L4+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; (9)S2+E3+L4+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; (11)S2+E3+L4+N5+Q18+K19+A20+R21+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; (12)S2+E3+L4+N5+Q18+A20+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170; (16)S2+E3+N5+A46+I47+T48+L49+V104+Q148+P150+E152+V153+F154+N155+A156+E157+R158+D167+R168+A169+D170.
2. A fusion protein, characterized in that Comprising the variant of claim 1; and a nucleic acid programmable nucleotide binding domain. Preferably, the nucleic acid programmable nucleotide binding domain is a Cas protein or an Ago protein. Preferably, the Cas protein includes at least one of a type II CRISPR-Cas polypeptide, a type I CRISPR-Cas polypeptide, a type III CRISPR-Cas polypeptide, a type IV CRISPR-Cas polypeptide, a type V CRISPR-Cas polypeptide, a type VI CRISPR-Cas polypeptide, a type VII CRISPR-Cas polypeptide, an IscB polypeptide, a TnpB polypeptide, and an IsrB polypeptide. Preferably, the AGO protein is selected from pAgo, eAgo, Ago1, Ago2, Ago3 and Ago4.
3. A base editing system, characterized in that It includes: (i) the variant of claim 1, and a nucleic acid programmable nucleotide binding domain; or (ii) the fusion protein of claim 2; and guide RNA.
4. An isolated polynucleotide, characterized in that The polynucleotide encodes the variant according to claim 1 or the fusion protein according to claim 2 or the base editing system according to claim 3. Preferably, the polynucleotide encoding the fusion protein comprises a nucleotide sequence as shown in any one of SEQ ID NOs: 27-30.
5. A carrier, characterized in that Comprising the polynucleotide according to claim 4.
6. A delivery composition, characterized in that Comprising a delivery vector, and one or more selected from the following: the variant according to claim 1, the fusion protein according to claim 2, the system according to claim 3, the polynucleotide according to claim 4, and the vector according to claim 5. Preferably, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, nanoparticles, cell-penetrating peptides, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses). Preferably, the composition further comprises one or more lipid moieties selected from the group consisting of ionizable lipids, neutral lipids, PEG lipids, steroids or derivatives thereof, or combinations thereof. Preferably, the ionizable lipid comprises a compound represented by Formula I (described in PCT application number PCT / CN2023 / 106421): in, R1 is selected from -OH and R a (R b )N-, where R a and R b are independently hydrogen, C1-C 10 Alkyl or C1-C 10 alkyl halide; R2 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl; R3 is C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl or R c -(CH2)n-, wherein n is a positive integer of 1-20, preferably a positive integer of 1-14, more preferably a positive integer of 1-10; R c is selected from the following structures: in Represents the connection key; L1, L2, L3, L4, L5 are none or are independently selected from the following groups: X is none or -CH- or N; R4 is none or C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl; R5 is none or C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl; R6 is C1-C 30 Alkyl, C2-C 30 Alkenyl, C2-C 30 Alkynyl; R7 is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl; R8 and R9 are none or independently C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl or -R h -C1-C 14 Alkyl, -R h -C2-C 14 Alkenyl, -R h -C2-C 14 Alkynyl, where R h O or S; R 10 、R 11 Each independently is C1-C 14 Alkyl, C2-C 14 Alkenyl, C2-C 14 Alkynyl, or -R h -C1-C 14 Alkyl, -R h -C2-C 14 Alkenyl, -R h -C2-C 14 Alkynyl, where R h O or S.
7. A host cell, characterized in that Comprising the variant of claim 1, the fusion protein of claim 2, the system of claim 3, the polynucleotide of claim 4, the vector of claim 5 or the delivery composition of claim 6.
8. A kit, characterized in that Includes one or more components selected from the following: The variant of claim 1, the fusion protein of claim 2, the system of claim 3, the polynucleotide of claim 4, the vector of claim 5, the delivery composition of claim 6, and the host cell of claim 7.
9. A composition, characterized in that Including the variant of claim 1, the fusion protein of claim 2, the base editing system of claim 3, the polynucleotide of claim 4, the vector of claim 5, the delivery composition of claim 6 or the host cell of claim 7.
10. An enzyme preparation, characterized in that The enzyme preparation includes the variant of claim 1, the fusion protein of claim 2, the base editing system of claim 3, the polynucleotide of claim 4, the vector of claim 5, and the delivery composition of claim 6.
11. A medicine box, characterized in that: include: A first container, and the system according to claim 3, the polynucleotide according to claim 4, the vector according to claim 5, or the delivery composition according to claim 6, or a drug containing the system according to claim 3, the polynucleotide according to claim 4, the vector according to claim 5, or the delivery composition according to claim 6, located in the first container.
12. A medicine box, characterized in that: include: (a1) a first container, and the variant according to claim 1, the fusion protein according to claim 2, or a gene encoding the same or an expression vector thereof, or a drug containing the variant according to claim 1, the fusion protein according to claim 2, or a gene encoding the same or an expression vector thereof, located in the first container; (b1) an optional second container, and a guide RNA or its expression vector, or a drug containing the guide RNA or its expression vector, located in the second container.
13. A lipid nanoparticle composition, characterized in that The invention comprises an ionizable lipid, and a nucleotide sequence encoding the variant according to claim 1, a nucleotide sequence encoding the fusion protein according to claim 2, a nucleotide sequence encoding the system according to claim 3, the polynucleotide according to claim 4 or the vector according to claim 5.
14. A pharmaceutical composition, characterized in that The invention comprises an ionizable lipid, a variant according to claim 1, a fusion protein according to claim 2, a system according to claim 3, a polynucleotide according to claim 4 or a vector according to claim 5, and a pharmaceutically acceptable excipient, carrier or diluent.
15. A pharmaceutical preparation, characterized in that Comprising an ionizable lipid, and the variant of claim 1, the fusion protein of claim 2, the system of claim 3, the polynucleotide of claim 4 or the vector of claim 5, and a pharmaceutically acceptable excipient, carrier or diluent; or the pharmaceutical formulation comprises the lipid nanoparticle composition of claim 13, and a pharmaceutically acceptable excipient, carrier or diluent.
16. A method for targeting and editing a target gene, characterized in that include: The variant of claim 1, the fusion protein of claim 2, the base editing system of claim 3, the polynucleotide of claim 4, the vector of claim 5, the delivery composition of claim 6, or the host cell of claim 7, or the composition of claim 9, or the enzyme preparation of claim 10, or the lipid nanoparticle composition of claim 13, or the pharmaceutical composition of claim 14, or the pharmaceutical preparation of claim 15 is contacted with the target gene, or delivered to a cell containing the target gene, and the target sequence is present in the target gene.
17. A method for inducing a change in cell state, characterized in that: The method comprises contacting the variant of claim 1, the fusion protein of claim 2, the base editing system of claim 3, the polynucleotide of claim 4, the vector of claim 5, the delivery composition of claim 6, or the host cell of claim 7, or the composition of claim 9, or the enzyme preparation of claim 10, or the lipid nanoparticle composition of claim 13, or the pharmaceutical composition of claim 14, or the pharmaceutical preparation of claim 15 with a target gene in a cell.
18. A cell or progeny thereof obtained by the method of any one of claims 16-17, wherein the cell comprises a modification not present in its wild-type form.
19. A cell product of the cell of claim 18 or its progeny.
20. An in vitro, ex vivo or in vivo cell or cell line or progeny thereof, characterized in that The cell or cell line or their progeny comprises: the variant of claim 1, the fusion protein of claim 2, the base editing system of claim 3, the polynucleotide of claim 4, the vector of claim 5, the delivery composition of claim 6 or the composition of claim 9 or the lipid nanoparticle composition of claim 13.
21. Use of the variant of claim 1, the fusion protein of claim 2, the base editing system of claim 3, the polynucleotide of claim 4, the vector of claim 5, the delivery composition of claim 6, or the host cell of claim 7, or the composition of claim 9, or the enzyme preparation of claim 10, or the lipid nanoparticle composition of claim 13, or the pharmaceutical composition of claim 14, or the pharmaceutical preparation of claim 15 for preparing a medicament or preparation for treating a disease associated with or caused by a point mutation; Optionally, the disorder or disease is associated with one or more C>A point mutations or C>T point mutations; preferably, the disorder or disease includes the diseases shown in Table A1; Optionally, the disease or condition comprises one or more of hypercholesterolemia, transthyretin amyloidosis, beta-hemoglobinopathy.
22. A method of treating a condition or disease in a subject in need thereof, the method comprising administering to the subject the fusion protein of claim 2, the base editing system of claim 3, the polynucleotide of claim 4, the vector of claim 5, the delivery composition of claim 6, or the host cell of claim 7, or the composition of claim 9, or the enzyme preparation of claim 10, or the lipid nanoparticle composition of claim 13, or the pharmaceutical composition of claim 14, or the pharmaceutical preparation of claim 15; Optionally, the disorder or disease is associated with one or more C>A point mutations or C>T point mutations; preferably, the disorder or disease includes the diseases shown in Table A1; Optionally, the disease or condition comprises one or more of hypercholesterolemia, transthyretin amyloidosis, beta-hemoglobinopathy.
Citation Information
Patent Citations
Adenosine deaminase, base editor fusion protein, base editor system and application
CN114634923A
Novel nucleobase editor and use method thereof
CN114667149A
Adenine deaminase and use thereof in base editing
CN117187220A
Methods of editing a disease-associated gene using adenosine deaminase base editors, including for the treatment of genetic disease
WO2020168051A1
Adenosine deaminase, base editor, and use thereof
WO2023193536A1