Novel engineered adenine deaminase and adenine base editor
By performing amino acid substitution at specific amino acid sites on adenine deaminase, the scope of its editing window is broadened and catalytic efficiency is improved, and the problem of narrow editing window of existing adenine base editors is solved, achieving more flexible and efficient gene editing.
Patent Information
- Application Number
- CN202411735973.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-02
AI Technical Summary
The editing window of the existing adenine base editor is relatively fixed and narrow, which limits its application in gene editing, especially in scenarios where the editing window needs to be wide.
The scope of their editing activity window is broadened and the catalytic efficiency of the deaminase is improved by performing amino acid substitutions at specific amino acid sites on the adenine deaminase.
The editing activity window of adenine deaminase is broadened, and its catalytic efficiency is improved, enhancing the flexibility and efficiency of base editing.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present disclosure relates to engineered adenine deaminases, base editors based on these adenine deaminases, and complexes containing these base editors. The present disclosure also relates to polynucleotides encoding engineered adenine deaminases, codon-optimized polynucleotides, vectors containing these nucleotides, and cells containing these vectors. The present disclosure also relates to pharmaceutical compositions containing the base editors, complexes, vectors or cells, and methods and applications of using the base editors, complexes, vectors, cells, compositions or kits, and modifying target nucleic acids. Background Art
[0002] Traditionally, adenine base editors (ABE) are constructed by fusing adenine deaminase to the N-terminus or C-terminus of the Cas9 protein through a linker, but the main disadvantage is that its editing window is relatively fixed and narrow. For example, the SpCas9-mediated ABE editing window is positions 4-7. Obviously, its current narrow editing window hinders its practicality, such as limiting the research that requires the use of base editors with a wider editing window range, such as screening functional genes and locating key amino acid sites in protein domains. Therefore, there is an urgent need to develop base editors with a wider editing window range. Summary of the invention
[0003] One aspect of the present invention provides an adenine deaminase having an amino acid substitution at one, more or all of the following positions relative to the sequence shown in SEQ ID NO:1: Q20, M60, S97, Q101, S107, M118 and K129, and whose amino acid sequence has a sequence identity of about 70% to about 99.5% with the sequence shown in SEQ ID NO:1.
[0004] In some embodiments, the adenine deaminase: (1) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at the Q20 position; (2) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at the M60 position; (3) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at the S97 position; (4) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at the Q101 position; (5) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at the S107 position; (6) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at the M118 position; (7) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at the K129 position; (8) according to the sequence numbering shown in SEQ ID According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has amino acid substitutions at positions S97 and Q101; (9) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has amino acid substitutions at positions S97 and S107; (10) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has amino acid substitutions at positions S97, Q101 and M118; or (11) according to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has amino acid substitutions at positions S97, S107 and M118.
[0005] In some specific embodiments, the amino acid substitution is: (a) Q20 is substituted with Q20K; (b) M60 is substituted with M60I; (c) S97 is substituted with S97G; (d) Q101 is substituted with Q101R; (e) S107 is substituted with S107R; (f) M118 is substituted with M118I; or (g) K129 is substituted with K129R.
[0006] In some embodiments, the adenine deaminase comprises an amino acid sequence selected from any one of SEQ ID NOs: 2 to 12, or an amino acid sequence having at least about 80% sequence identity thereto.
[0007] " sequence identity " refers to the degree of sequence identity of a sequence ... Therefore, "percentage of sequence identity" is calculated as follows: by comparing two optimally aligned sequences within a comparison window, determining the number of positions where identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) or identical nucleic acid bases (e.g., A, T, C, G, I) occur in the two sequences to produce the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window (i.e., window size), and multiplying the result by 100 to obtain the percentage of sequence identity. In the present invention, when the compared sequences are two non-continuous sequences, the calculation of sequence identity is obtained based on the comparison result of the sequence. For example, the term "about 70% to about 99.5%" in the context of "having about 70% to about 99.5% sequence identity to the amino acid sequence set forth in SEQ ID NO. 1" refers in the present invention to any value from 70% to 99.5%, such as 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5% or 100%.
[0008] In some embodiments, the adenine deaminase comprises or is an amino acid sequence having at least about 80% sequence identity compared to the amino acid sequence shown in any one of SEQ ID NOs. 2 to 12. For example, the adenine deaminase comprises or is an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity compared to the amino acid sequence shown in any one of SEQ ID NOs. 2 to 12. The amino acid sequence of the adenine deaminase is shown in SEQ ID NO.1; the engineered adenine deaminase has an amino acid substitution at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33 or 34 positions relative to the amino acid sequence of SEQ ID NO.1, and the engineered adenine deaminase is mutated to have one or more of the following characteristics: broadening the editing activity window range and improving the catalytic efficiency of the deaminase.
[0009] In some embodiments, the mutation of the engineered adenine deaminase results in a broadening of the editing activity window of the adenine deaminase, for example, compared with SEQ ID NO.1, the editing activity window is broadened by 1bp, 2bp, 3bp, 4bp, 5bp, 6bp, 7bp, 8bp, 9bp, 10bp, 11bp, 12bp, 13bp, 14bp or 15bp.
[0010] In some embodiments, mutations in the engineered adenine deaminase result in an enhanced deaminase catalytic efficiency of the adenine deaminase, for example, an enhanced deaminase catalytic efficiency of at least 10%, for example, 10% to 500%, 10% to 100%, 10% to 200%, 10% to 300%, 10% to 50%, 10% to 30%, 10% to 20%, 50% to 100%, 50% to 200%, 50% to 300%, 100% to 200%, or 200% to 300%, compared to SEQ ID NO.1.
[0011] Specifically, the name of the mutant polypeptide in which the adenine deaminase is substituted with a specific amino acid at a specific position is named according to the mutation naming rules of the Human Genome Variation Society (HGVS), for example, "the adenine deaminase has an amino acid substitution at position Q20" "Q20K" means that K replaces the original Q at the 20th position of SEQ ID NO.1.
[0012] Another aspect of the present invention provides an adenine base editor comprising the adenine deaminase of the present invention; the adenine base editor comprises a DNA binding domain and an adenine deaminase; preferably, the adenine base editor comprises, from N-terminus to C-terminus: (i) a DNA binding domain and adenine deaminase; (ii) adenine deaminase and a DNA binding domain; or (iii) a DNA binding domain-N-terminal protein, adenine deaminase and a DNA binding domain-C-terminal protein.
[0013] In some embodiments, in the adenine base editor, the DNA binding domain can be fused to the adenine deaminase via one or more linker polypeptides (or connecting peptides, linkers); the DNA binding domain-N-terminal protein can be fused to the adenine deaminase via one or more linker polypeptides (or connecting peptides, linkers); the adenine deaminase can be fused to the DNA binding domain-C-terminal protein via one or more linker polypeptides (or connecting peptides, linkers).
[0014] Specifically, the linker polypeptide (or connecting peptide, linker) can have any of a variety of amino acid sequences. Proteins can be connected by spacer peptides, which are usually flexible, but other chemical bonds are not excluded. Suitable linkers include polypeptides with a length between 4 and 40 amino acids or a length between 4 and 25 amino acids. These linkers can be produced by using synthetic oligonucleotides encoding linkers to couple proteins, or can be encoded by a nucleic acid sequence encoding a fusion protein. Peptide linkers with a certain degree of flexibility can be used. The connecting peptide can actually have any amino acid sequence, and it should be remembered that preferred linkers will have sequences that produce generally flexible peptides. The use of small amino acids (such as glycine and alanine) is used to produce flexible peptides. For those skilled in the art, it is conventional to produce such sequences. A variety of different linkers are commercially available and are considered to be suitable for use.
[0015] Specifically, examples of linker polypeptides include glycine polymers ((G)n, wherein n is selected from 1, 2, 3, 4, 5 or 6), glycine-serine polymers ((GGGGS)n, (GGGS)n, (GGS)n, (GS)n or (G)n, wherein n is selected from 1, 2, 3, 4, 5 or 6), glycine-alanine polymers, alanine-serine polymers and α-helical linkers ((EAAAK)n, wherein n is selected from 1, 2, 3, 4, 5 or 6). Exemplary linkers can comprise an amino acid sequence including, but not limited to, EAAAK (SEQ ID NO: 18), SGGS (SEQ ID NO: 19), GGSG (SEQ ID NO: 20), GGSGG (SEQ ID NO: 21), GSGSG (SEQ ID NO: 22), GSGGG (SEQ ID NO: 23), GGGSG (SEQ ID NO: 24), GSSSG (SEQ ID NO: 25), SGGSSGGS (SEQ ID NO: 26), SGGSGGSGGS (SEQ ID NO: 27), GGGGSGGGGS (SEQ ID NO: 28), SGGSGGGGSGGGGS (SEQ ID NO: 29), SGSETPGTSESATPES (SEQ ID NO: 30), SGGSSGSETPGTSESATPESSGGS (SEQ ID NO: 31), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 32), and the like. The linker polypeptide can also be various XTEN linkers, etc. The length of the XTEN linker is about 16-80 amino acids, and the XTEN linker can be XTEN16 linker, XTEN18 linker, XTEN32 linker, XTEN80 linker (SEQ ID NO: 33). More specifically, the linker polypeptide includes but is not limited to the amino acid sequences shown in SEQ ID NO.18 to 33. Those skilled in the art will recognize that the design of the peptide conjugated to any desired element can include all or part of a flexible linker, so that the linker can include a flexible linker and one or more parts that confer a less flexible structure.
[0016] In some embodiments, the DNA binding domain comprises a CRISPR / Cas domain with single nickase activity, a zinc finger nuclease ZFN, a transcription activator-like effector nuclease TALEN, a large range of nucleases or any combination thereof; preferably, the DNA binding domain comprises a CRISPR / Cas domain with single nickase activity; optionally, it is nSpCas9 or nSaCas9; specifically, the nSpCas9 comprises or is an amino acid sequence shown in SEQ ID NO: 13 or an amino acid sequence having at least about 80% sequence identity with each of them; preferably, the DNA binding domain-N-terminal protein is nCas9-N-terminal protein; the DNA binding domain-C-terminal protein is nCas9-C-terminal protein; specifically, the nCas9-N-terminal protein comprises or is an amino acid sequence shown in SEQ ID NO: 14 or 16 or an amino acid sequence having at least about 80% sequence identity with each of them; the nCas9-C-terminal protein comprises or is an amino acid sequence shown in SEQ ID NO: 15 or 17 or an amino acid sequence having at least about 80% sequence identity with each of them.
[0017] In some specific embodiments, the nCas9-N-terminal protein, the adenine deaminase and the nCas9-C-terminal protein are each sequentially fused through one or more linker polypeptides (or connecting peptides, linkers) to form an adenine base editor.
[0018] Specifically, the nCas9-N-terminal protein and the nCas9-C-terminal protein are selected from those SaCas9-N-terminal proteins and SaCas9-C-terminal proteins disclosed in CN118703471A, and their disclosed contents are fully incorporated herein by reference, such as fusing the adenine deaminase of the present invention with those SaCas9-N-terminal proteins and SaCas9-C-terminal proteins disclosed in CN118703471A to form an adenine base editor, i.e., NH 2 -[SaCas9-N-terminal protein]-[adenine deaminase]-[SaCas9-C-terminal protein]-COOH.
[0019] In other embodiments, the adenine base editor provided by the present invention may further comprise a cytidine deaminase domain. Specifically, the adenine base editor further comprising a cytidine deaminase domain is selected from those disclosed in the literature (Zhong Jingli, Lin Jianxiang, Zhou Jiankui, Qiao Yunbo. Research progress of base editing system [J]. Journal of Biotechnology, 2024, 40(5): 1271-1292.).
[0020] In other embodiments, the adenine base editor provided by the present invention may further include a second adenine deaminase domain, which is the same as or different from the adenine deaminase provided by the present invention. Specifically, the adenine base editor further comprising a second adenine deaminase domain is selected from those disclosed in the literature (Zhong Jingli, Lin Jianxiang, Zhou Jiankui, Qiao Yunbo. Research progress of base editing systems [J]. Journal of Biotechnology, 2024, 40 (5): 1271-1292.).
[0021] In other embodiments, the adenine base editor further comprises one or more epitope tags, nuclear localization signals, reporter gene sequences, domains capable of binding to DNA molecules or intracellular molecules, enzymes capable of detecting signals, subcellular localization and protein transduction domains; preferably, a nuclear localization signal is contained at the C-terminus or / and the N-terminus of the adenine base editor; preferably, two nuclear localization signals are contained at the C-terminus or / and the N-terminus of the adenine base editor; preferably, three nuclear localization signals are contained at the C-terminus or / and the N-terminus of the adenine base editor.
[0022] In some embodiments, the epitope tag is an existing conventional tag, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select a suitable epitope tag according to the desired purpose (e.g., purification, detection or tracing).
[0023] In some embodiments, the reporter gene sequence is well known to those skilled in the art, and examples thereof include but are not limited to GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0024] In some embodiments, the domain capable of binding to a DNA molecule or an intracellular molecule is, for example, maltose binding protein (MBP), the DNA binding domain (DBD) of LexA, the DBD of GAL4, and the like.
[0025] In some embodiments, the enzyme that can detect the signal is a detectable enzyme well known to those skilled in the art, such as luciferase, fluorescent protein, etc. The adenine base editor provided by the present invention also includes radioactive isotopes, members of specific binding pairs, fluorophores or quantum dots (Quantum Dots, QDs: semiconductor nanoparticles that can receive excitation light to produce fluorescence) for detection, etc.
[0026] In some embodiments, the subcellular localization sequence (e.g., a nuclear localization signal (NLS) for targeting the cell nucleus, a sequence for keeping the adenine base editor outside the cell nucleus (e.g., a nuclear export sequence (NES)), a sequence for retaining the adenine base editor in the cytoplasm, a mitochondrial localization signal for targeting mitochondria, a chloroplast localization signal for targeting chloroplasts, an ER retention signal, etc.).
[0027] In some embodiments, the adenine base editor provided by the present invention is (fused with) a nuclear localization signal (NLS) (e.g., in some embodiments, 1 or more, 2 or more, 3 or more, 4 or more, or 5 or more NLS). In some embodiments, one or more NLSs are located at the N-terminus of the fusion protein, and one or more NLSs are located at the C-terminus of the adenine base editor. Preferably, the nuclear localization signal comprises or is an amino acid sequence shown in any one of SEQ ID NOs: 34 to 39 or an amino acid sequence having at least about 95% sequence identity with each of them.
[0028] In some embodiments, the adenine base editor provided by the present invention comprises (is fused with) 1 to 10 NLSs (e.g., 1-9, 1-8, 1-7, 1-6, 1-5, 2-10, 2-9, 2-8, 2-7, 2-6 or 2-5 NLSs). In some embodiments, the adenine base editor provided by the present invention is (is fused with) 2 to 5 NLSs (e.g., 2-4 or 2-3 NLSs), specifically, the amino acid sequence of the 2 NLSs is SEQ ID NO: 40.
[0029] In some embodiments, non-limiting examples of the NLS include, but are not limited to, NLS of SV40 virus large T antigen, nucleoplasmin NLS, nucleoplasmin bipartite NLS, c-myc NLS, hRNPA1 M9 NLS, IBB domain of importin-α, myoma T protein, human p53, mouse c-abl IV, influenza virus NS1, hepatitis virus delta antigen, mouse Mx1 protein, human poly (ADP-ribose) polymerase, steroid hormone receptor (human) glucocorticoid and other commonly used NLSs. Exemplary NLSs may comprise an amino acid sequence including, but not limited to, PKKKRKV (SEQ ID NO:34), PAAKRVKLD (SEQ ID NO:35), KRPAATKKAGQAKKKK (SEQ ID NO:36), RQRRNELKRSP (SEQ ID NO:37), KRTADGSEFESPKKKRKV (SEQ ID NO:38), KRTADGSEFEPKKKRKV (SEQ ID NO:39).
[0030] In some embodiments, a "protein transduction domain" or PTD (also known as a CPP-cell penetrating peptide) refers to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates traversal of a lipid bilayer, a micelle, a cell membrane, an organelle membrane, or a vesicle membrane. A PTD attached to another molecule (which may range from a small polar molecule to a large macromolecule and / or nanoparticle) facilitates the molecule to traverse the membrane, for example, from the extracellular space into the intracellular space or from the cytosol into the organelle. In some embodiments, the PTD is covalently linked to the amino terminus of an adenine base editor of the present invention to generate a fusion protein. In some embodiments, the PTD is covalently linked to the carboxyl terminus of an adenine base editor of the present invention to generate a fusion protein. In some embodiments, the PTD is inserted into the adenine base editor at a suitable insertion site (ie, not at the N-terminus or C-terminus of the fusion protein). In some embodiments, an adenine base editor comprises (conjugated to, fused to) one or more PTDs (eg, two or more, three or more, four or more PTDs). In some embodiments, a PTD comprises a nuclear localization signal (NLS) (eg, in some embodiments, 2 or more, 3 or more, 4 or more, or 5 or more NLS).
[0031] In some embodiments, the structure of the adenine base editor provided by the present invention is selected from: NH 2 -[Adenine base editor]-[NLS]-COOH; NH 2 -[Adenine base editor]-[NLS]-[NLS]-COOH; NH 2 -[Adenine base editor]-[NLS]-[NLS]-[NLS]-COOH; NH 2 -[NLS]-[Adenine base editor]-COOH; NH 2 -[NLS]-[NLS]-[Adenine Base Editor]-COOH; NH 2 -[NLS]-[NLS]-[NLS]-[Adenine Base Editor]-COOH; NH 2 -[NLS]-[Adenine base editor]-[NLS]-COOH; NH 2 -[NLS]-[NLS]-[Adenine base editor]-[NLS]-[NLS]-COOH; NH 2 -[NLS]-[NLS]-[NLS]-[Adenine base editor]-[NLS]-[NLS]-[NLS]-COOH; NH 2-[NLS]-[NLS]-[Adenine base editor]-[NLS]-COOH; NH 2 -[NLS]-[NLS]-[NLS]-[Adenine Base Editor]-[NLS]-COOH; NH 2 -[NLS]-[Adenine Base Editor]-[NLS]-[NLS]-COOH; NH 2 -[NLS]-[Adenine Base Editor]-[NLS]-[NLS]-[NLS]-COOH; NH 2 -[NLS]-[NLS]-[Adenine Base Editor]-[NLS]-[NLS]-[NLS]-COOH; or NH 2 -[NLS]-[NLS]-[NLS]-[Adenine Base Editor]-[NLS]-[NLS]-COOH; wherein]-[ represents an optionally present connecting peptide as defined below (the same below). Preferably, the adenine base editor is connected to one or more NLS via one or more linker polypeptides.
[0032] In some embodiments, in the adenine base editor provided by the present invention, the adenine base editor comprises an amino acid sequence selected from any one of SEQ ID NOs: 56 to 71 or an amino acid sequence having at least about 80% sequence identity with each of them.
[0033] Another aspect of the present invention provides a complex comprising the adenine base editor and a guide RNA, wherein the guide RNA complexes with the adenine base editor to guide the DNA binding domain of the adenine base editor to bind to the target nucleic acid; preferably, the guide RNA comprises a guide segment that hybridizes with the target nucleic acid and a repeat segment that binds to the adenine base editor; preferably, the repeat segment of the guide RNA comprises a nucleotide sequence shown in any one of SEQ ID NO.41 to 46 or a nucleotide sequence having 1 to 10 nucleotide substitutions, deletions or insertions compared with the nucleotide sequence shown in any one of SEQ ID NO.41 to 46; preferably, the repeat segment of the guide RNA comprises or is a nucleotide sequence shown in any one of SEQ ID NO.41 to 46.
[0034] Specifically, in the complex provided by the present invention, the gRNA is any one of the above "guide RNA (gRNA or sgRNA)", and these terms are used interchangeably herein.
[0035] Specifically, the adenine base editor provided by the present invention binds to the target gene target nucleic acid at a target sequence defined by the complementary region between the guide RNA targeting the target nucleic acid and the target gene target nucleic acid. Site-specific binding of the double-stranded target nucleic acid occurs at a position determined by the following two: (i) base pairing complementarity between the guide RNA and the target nucleic acid; and (ii) a protospacer adjacent motif (PAM) in the target nucleic acid.
[0036] In some embodiments, the complementarity percentage between the guide segment of the guide RNA and the target site of the target gene nucleic acid is 60% or more (e.g., 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the complementarity percentage between the guide segment of the guide RNA and the target site of the target gene nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the complementarity percentage between the guide segment of the guide RNA and the target site of the target gene nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the percent complementarity between the guide segment of the guide RNA and the target site of the target nucleic acid of the gene of interest is 100%.
[0037] In some embodiments, the guide segment of the guide RNA has a length of 14 to 50 nucleotides (e.g., 14 nucleotides (nt) to 20nt, 20nt to 25nt, 25nt to 30nt, 30nt to 35nt, 35nt to 40nt, 40nt to 45nt, or 45nt to 50nt). In some preferred embodiments, the guide segment of the guide RNA has a length in the range of 14-30 nucleotides (nt) (e.g., 14-25, 14-22, 14-20, 17-25, 17-22, 17-20, 19-30, 19-25, 19-22, 19-20, 20-30, 20-25, or 20-22 nt). In some embodiments, the guide segment of the guide RNA has a length in the range of 14-25 nucleotides (nt) (e.g., 14-22, 14-20, 17-22, 17-20, 19-25, 19-22, 19-20, 20-25, or 20-22 nt). In some embodiments, the guide segment of the guide RNA has a length of 14 or more nt (e.g., 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, or 22 or more nt; 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some preferred embodiments, the guide segment of the guide RNA has a length of 17 nt. In some preferred embodiments, the guide segment of the guide RNA has a length of 18 nt. In some preferred embodiments, the guide segment of the guide RNA has a length of 19 nt. In some preferred embodiments, the guide segment of the guide RNA has a length of 20 nt. In some preferred embodiments, the guide segment of the guide RNA has a length of 21 nt. In some preferred embodiments, the guide segment of the guide RNA has a length of 22 nt. In some preferred embodiments, the guide segment of the guide RNA has a length of 23 nt.
[0038] In some embodiments, in the complex provided by the present invention, the repeating segment (protein binding segment) of the guide RNA is a single nucleotide sequence, and the sequence length of the repeating segment can be 15 to 100 nt, for example, 20-100nt, 30-100nt, 40-100nt, 50-100nt, 60-100nt, 70-100nt, 80-100nt, 50-90nt, 60-90nt, 70-90nt. t, for example, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nt.
[0039] Another aspect of the present invention provides a fusion polypeptide of an adenine base editor, which comprises one or more heterologous polypeptides and an amino acid sequence having at least 80% sequence identity with the amino acid sequence shown in any one of SEQ ID NO.56 to 71; wherein the heterologous polypeptide is selected from an epitope tag, a nuclear localization signal, a reporter gene sequence, a domain capable of binding to a DNA molecule or an intracellular molecule, an enzyme that can detect a signal, a subcellular localization and a protein transduction domain.
[0040] In some specific embodiments, based on the sequence table, the fusion polypeptide comprises, from the N-terminus to the C-terminus: 1) NH 2 -DNA binding domain-linker-adenine deaminase-linker-NLS-COOH; 2) NH 2 -DNA binding domain-linker-adenine deaminase-linker-NLS-linker-NLS-COOH; 3) NH 2 -DNA binding domain-linker-adenine deaminase-linker-NLS-linker-NLS-linker-NLS-COOH; 4) NH 2 -NLS-linker-DNA binding domain-linker-adenine deaminase-COOH; 5) NH 2 -NLS-linker-NLS-linker-DNA binding domain-linker-adenine deaminase-COOH; 6) NH 2 -NLS-linker-NLS-linker-NLS-linker-DNA binding domain-linker-adenine deaminase-COOH; 7) NH 2 -NLS-linker-DNA binding domain-linker-adenine deaminase-linker-NLS-COOH; 8) NH 2-NLS-linker-NLS-linker-DNA binding domain-linker-adenine deaminase-linker-NLS-linker-NLS-COOH; 9) NH 2 -NLS-linker-NLS-linker-NLS-linker-DNA binding domain-linker-adenine deaminase-linker-NLS-linker-NLS-linker-NLS-COOH; 10) NH 2 -NLS-linker-NLS-linker-DNA binding domain-linker-adenine deaminase-linker-NLS-COOH; 11) NH 2 -NLS-linker-NLS-linker-NLS-linker-DNA binding domain-linker-adenine deaminase-linker-NLS-COOH; 12) NH 2 -NLS-Linker-DNA binding domain-Linker-Adenine deaminase-Linker-NLS-Linker-NLS-COOH; 13) NH 2 -NLS-Linker-DNA binding domain-Linker-Adenine deaminase-Linker-NLS-Linker-NLS-Linker-NLS-COOH; 14) NH 2 -NLS-linker-NLS-linker-DNA binding domain-linker-adenine deaminase-linker-NLS-linker-NLS-linker-NLS-COOH; 15) NH 2 -NLS-linker-NLS-linker-NLS-linker-DNA binding domain-linker-adenine deaminase-linker-NLS-linker-NLS-COOH; 16) NH 2 -Adenine deaminase-Linker-DNA binding domain-Linker-NLS-COOH; 17) NH 2 -Adenine deaminase-Linker-DNA binding domain-Linker-NLS-Linker-NLS-COOH; 18) NH 2 -Adenine deaminase-Linker-DNA binding domain-Linker-NLS-Linker-NLS-Linker-NLS-COOH; 19) NH 2 -NLS-linker-adenine deaminase-linker-DNA binding domain-COOH; 20) NH 2 -NLS-linker-NLS-linker-adenine deaminase-linker-DNA binding domain-COOH; 21) NH 2 -NLS-linker-NLS-linker-NLS-linker-adenine deaminase-linker-DNA binding domain-COOH; 22) NH 2 -NLS-linker-adenine deaminase-linker-DNA binding domain-linker-NLS-COOH; 23) NH 2-NLS-linker-NLS-linker-adenine deaminase-linker-DNA binding domain-linker-NLS-linker-NLS-COOH; 24) NH 2 -NLS-linker-NLS-linker-NLS-linker-adenine deaminase-linker-DNA binding domain-linker-NLS-linker-NLS-linker-NLS-COOH; 25) NH 2 -NLS-linker-NLS-linker-adenine deaminase-linker-DNA binding domain-linker-NLS-COOH; 26) NH 2 -NLS-linker-NLS-linker-NLS-linker-adenine deaminase-linker-DNA binding domain-linker-NLS-COOH; 27) NH 2 -NLS-Linker-Adenine Deaminase-Linker-DNA Binding Domain-Linker-NLS-Linker-NLS-COOH; 28) NH 2 -NLS-Linker-Adenine Deaminase-Linker-DNA Binding Domain-Linker-NLS-Linker-NLS-Linker-NLS-COOH; 29) NH 2 -NLS-linker-NLS-linker-adenine deaminase-linker-DNA binding domain-linker-NLS-linker-NLS-linker-NLS-COOH; 30) NH 2 -NLS-linker-NLS-linker-NLS-linker adenine deaminase-linker-DNA binding domain-linker-NLS-linker-NLS-COOH; 31) NH 2 -nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-COOH; 32) NH 2 -nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-linker-NLS-COOH; 33) NH 2 -nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-linker-NLS-linker-NLS-COOH; 34) NH 2 -NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-COOH; 35) NH 2 -NLS-linker-NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-COOH; 36) NH 2-NLS-linker-NLS-linker-NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-COOH; 37) NH 2 -NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-COOH; 38) NH 2 -NLS-linker-NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-linker-NLS-COOH; 39) NH 2 -NLS-linker-NLS-linker-NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-linker-NLS-linker-NLS-COOH; 40) NH 2 -NLS-linker-NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-COOH; 41) NH 2 -NLS-linker-NLS-linker-NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-COOH; 42) NH 2 -NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-linker-NLS-COOH; 43) NH 2 -NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-linker-NLS-linker-NLS-COOH; 44) NH 2 -NLS-linker-NLS-linker-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-linker-NLS-linker-NLS-COOH; 45) NH 2 -NLS-linker-NLS-linker-NLS-nCas9-N-terminal protein-linker-adenine deaminase-linker-nCas9-C-terminal protein-linker-NLS-linker-NLS-COOH.
[0041] Another aspect of the present invention provides a nucleic acid comprising a polynucleotide encoding the adenine base editor, the complex or the fusion polypeptide; in a preferred embodiment, the polynucleotide is codon optimized for expression in prokaryotic or eukaryotic cells.
[0042] In some embodiments, the present invention provides a nucleic acid comprising a guide RNA or a nucleotide sequence encoding the guide RNA, wherein the guide RNA comprises a repeated segment, comprising a nucleotide sequence shown in any one of SEQ ID NOs.41 to 46, or a nucleotide sequence having 1 to 10 nucleotide substitutions, deletions and / or insertions compared to the nucleotide sequence shown in any one of SEQ ID NOs.41 to 46; in a preferred embodiment, the repeated segment of the guide RNA is a nucleotide sequence shown in any one of SEQ ID NOs.41 to 46.
[0043] Another aspect of the present invention provides a vector comprising any one of the nucleic acids provided by the present invention. In a preferred embodiment, the vector is a plasmid or a viral vector. In a preferred embodiment, the viral vector is an adeno-associated viral vector, an adenoviral vector, a retroviral vector, a lentiviral vector or a herpes simplex virus vector.
[0044] Another aspect of the present invention provides a vector system, comprising a first vector and a second vector different from the first vector, wherein the first vector comprises a polynucleotide encoding any one of the adenine base editors provided by the present invention, the fusion polypeptide, or an amino acid sequence comprising any one of SEQ ID NO.56 to 71; the second vector comprises a guide RNA or a nucleotide sequence encoding the guide RNA. In a preferred embodiment, the first vector and the second vector are independently plasmids or viral vectors. In a preferred embodiment, the viral vector is an adeno-associated viral vector, an adenoviral vector, a retroviral vector, a lentiviral vector, or a herpes simplex viral vector.
[0045] Another aspect of the present invention provides a delivery system, comprising a polynucleotide encoding any adenine base editor provided by the present invention, the complex or the fusion polypeptide, any nucleic acid provided by the present invention, any vector provided by the present invention, or any vector system provided by the present invention. In a preferred embodiment, the delivery system comprises a liposome, a nanoparticle or an exosome.
[0046] Another aspect of the present invention provides a cell comprising the adenine base editor, the complex, the fusion polypeptide, the nucleic acid, the vector, the vector system, or the delivery system; preferably, the cell is a eukaryotic cell or a prokaryotic cell; more preferably, the cell is a human cell.
[0047] Another aspect of the present invention provides a composition or a kit comprising a polynucleotide encoding any adenine base editor provided by the present invention, the complex or the fusion polypeptide, any nucleic acid provided by the present invention, any vector provided by the present invention, any vector system provided by the present invention, or any delivery system provided by the present invention; and a pharmaceutically acceptable carrier.
[0048] The adenine base editor or its fusion polypeptide provided by the present invention can be used in a variety of methods (e.g., in combination with the guide RNA of the adenine base editor). For example, the adenine base editor or its fusion polypeptide of the present invention can be used for (i) modifying (e.g., amino acid substitution, etc.) a target nucleic acid (DNA or RNA; single-stranded or double-stranded); (ii) regulating the transcription of the target nucleic acid; (iii) base pair conversion of the target nucleic acid, etc. Preferably, the modification includes increasing or decreasing the expression of the target sequence in the target nucleic acid; preferably, the modification includes deaminating the target adenine or target cytosine in the target nucleic acid to achieve base pair conversion. In some embodiments, the method for modifying a target nucleic acid of the present invention comprises contacting the target nucleic acid with the following substances: a) an adenine base editor or its fusion polypeptide of the present invention; and b) one or more (e.g., two) guide RNAs of adenine base editors. In some embodiments, the contacting step is performed in in vitro cells. In some embodiments, the contacting step is performed in in vivo cells. In some embodiments, the contacting step is performed in ex vivo cells. For example, the present invention provides (but is not limited to) methods for editing a target nucleic acid; methods for regulating transcription of a target from a nucleic acid; methods for modifying a target nucleic acid, etc.
[0049] The present invention provides a method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with the adenine base editor of the present invention, the complex, the nucleic acid, the vector, the vector system or the delivery system, wherein the contact causes the target nucleic acid to be modified; preferably, the modification comprises increasing or decreasing the expression of a target sequence in the target nucleic acid; preferably, the modification comprises deaminating a target adenine in the target nucleic acid to achieve a base pair conversion.
[0050] The present invention provides an adenine base editor and its complex, fusion polypeptide, nucleic acid, vector, vector system, delivery system, cell, composition and kit for use in the preparation of a drug for modifying nucleic acid; preferably, the modification comprises increasing or decreasing the expression of a target sequence in the target nucleic acid; preferably, the modification comprises deaminating a target adenine or target cytosine in the target nucleic acid to achieve base pair conversion.
[0051] The present invention provides an adenine base editor having an amino acid sequence as shown in any one of SEQ ID NOs. 56 to 71, for use in preparing a drug for base pair conversion of a target nucleic acid; preferably, the adenine base editor regulates the expression of the target nucleic acid in the cell by converting the base pairs containing the target nucleic acid.
[0052] Specifically, the adenine base editor of the amino acid sequence shown in SEQ ID NO.56 to 71 does not contain NLS.
[0053] In some embodiments, the adenine base editors represented by the amino acid sequences shown in SEQ ID NOs. 55 to 71 are, in order, “MaTadA-linker 1-nSpCas9” (control group), “nSpCas9-N1247-linker 1-MaTadA-linker 1-nSpCas9-C1248” (experimental group 1), “nSpCas9-N1047-linker 1-MaTadA-linker 1-nSpCas9-C1064” (experimental group 2), “nSpCas9-N1247-linker 2-MaTadA-linker 2-nSpCas9-C1248” (experimental group 3), and “nSpCas9-N1247-linker 1 -MaTadAQ20K-linker 1-nSpCas9-C1248" (experimental group 4), "nSpCas9-N1247-linker 1-MaTadAM60I-linker 1-nSpCas9-C1248" (experimental group 5), "nSpCas9-N1247-linker 1-MaTadAS97G-linker 1-nSpCas9-C1248" (experimental group 6), "nSpCas9-N1247-linker 1-MaTadAQ101R-linker 1-nSpCas9-C1248" (experimental group 7), and "nSpCas9-N1247-linker 1-MaTadA S107R-linker 1-nSpCas9-C1248” (experimental group 8), “nSpCas9-N1247-linker 1-MaTadAM118I-linker 1-nSpCas9-C1248” (experimental group 9), “nSpCas9-N1247-linker 1-MaTadAK129R-linker 1-nSpCas9-C1248” (experimental group 10), “nSpCas9-N1247-linker 1-MaTadA S97G-Q101R-linker 1-nSpCas9-C1248” (experimental group 11), “nSpCas9-N1247-linker 1-MaTadAS97G-S107R-linker 1-nSpCas9-C1248” (experimental group 12), and “nSpCas9-N1247-linker 1-MaTadA S97G-Q101R-M118I-linker 1-nSpCas9-C1248” (experimental group 13), “nSpCas9-N1247-linker 1-MaTadA S97G-S107R-M118I-linker 1-nSpCas9-C1248” (experimental group 14), “MaTadA S97G-Q101R-M118I-linker 1-nSpCas9” (experimental group 15), and “MaTadAS97G-S107R-M118I-linker 1-nSpCas9” (experimental group 16).In the present invention, these adenine base editors are referred to as "base editors" or "ABEs", and these terms are used interchangeably herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A schematic diagram showing the structure of the adenine base editor provided by the present invention.
[0055] Figure 2 Schematic diagram of the gRNA vector structure used in an embodiment of the present invention.
[0056] Figure 3 A graph showing the sequencing results of the adenine base editor embedded in different sites provided by an embodiment of the present invention at the HBGp and HEK2 targets.
[0057] Figure 4 A bar graph showing the sequencing results of the adenine base editors embedded in different sites provided by an embodiment of the present invention at the HBGp and HEK2 targets.
[0058] Figure 5 A graph showing the sequencing results of the adenine base editors constructed with different linkers provided in an embodiment of the present invention at the HBGp and HEK2 targets.
[0059] Figure 6 A graph showing the sequencing results of the adenine base editor constructed with different MaTadA mutants provided in an embodiment of the present invention at the HEK2 target site.
[0060] Figure 7 A graph showing the sequencing results of the adenine base editor constructed with different MaTadA mutants provided in an embodiment of the present invention at the site 1 target site.
[0061] Figure 8 A graph showing the sequencing results of the adenine base editor constructed with different MaTadA mutants provided in an embodiment of the present invention at the site 4 target. DETAILED DESCRIPTION
[0062] Sequence Listing Example
[0063] Example 1. Adenine base editors formed by MaTadA embedded in different positions of nSpCas9 protein to screen for suitable embedded positions
[0064] The structural schematic diagrams of various adenine base editors provided by the present invention are as follows Figure 1 As shown in Table 1, the adenine base editor vector formed by the adenine deaminase MaTadA located at different positions of the nSpCas9 protein was constructed ( Figure 1 A to Figure 1 C) Specifically, Figure 1 A is the control group (labeled as MaTadA-linker 1-nSpCas9 in the sequence list), and the control protein is the C-terminus of MaTadA fused to the N-terminus of nSpCas9 through linker 1; Figure 1 B is experimental group 1 (labeled as nSpCas9-N1247-linker 1-MaTadA-linker 1-nSpCas9-C1248 in the sequence list), in which the two ends of the protein in experimental group 1 are fused to position 1247 of the nSpCas9 protein via linker 1; Figure 1 C is experimental group 2 (labeled as nSpCas9-N1047-linker 1-MaTadA-linker 1-nSpCas9-C1064 in the sequence list). The protein in experimental group 2 is MaTadA with two ends connected to linker 1. The MaTadA polypeptide containing two linkers 1 replaces the short peptide at positions 1047-1063 of the nSpCas9 protein. Figure 1 A to Figure 1 The amino acid sequences of C are SEQ ID NO.55 to SEQ ID NO.57, and an NLS (SEQ ID NO.34) is fused to both ends of each fusion protein through a linker. These eukaryotic codon-optimized adenine base editor sequences are constructed into mammalian cell expression vectors. These adenine base editors are expressed by the CAG promoter, and the eGFP gene (for cell sorting) is connected to the downstream of the adenine base editor through the self-cleaving polypeptide 2A (P2A); at the same time, a sgRNA expression cassette (HBGp expression cassette or HEK2 expression cassette) is constructed on each expression vector. The structure of the sgRNA expression cassette is as follows: Figure 2 As shown, the HBGp expression cassette ( Figure 2 A) includes U6 promoter, HBGp Target gRNA (SEQ ID NO.47) and SpCas9-gRNA scaffold (SEQ ID NO.41) in sequence, and the sequence of HBGp SpgRNA is as shown in SEQ ID NO.48; HEK2 expression cassette ( Figure 2B) includes U6 promoter, HEK2 TargetgRNA (SEQ ID NO.49) and SpCas9-gRNA scaffold (SEQ ID NO.41) in sequence, and the sequence of HEK2 SpgRNA is as SEQID NO.50. That is, each expression vector contains an adenine base editor sequence (control group protein, experimental group 1 protein and experimental group 2 protein) and an sgRNA expression cassette (HBGp expression cassette or HEK2 expression cassette).
[0065] Table 1
[0066] Then, the adenine base editor vectors constructed above were electroporated into human HEK293T cells respectively, and the electroporation program was DS150 (Lonza 4D electroporator). The cells were cultured at 37°C and 5% carbon dioxide concentration. The cell genome was extracted by lysis method 72 hours after transfection, and the target sequence was amplified by PCR using corresponding primers. The editing ability of these adenine base editors on the corresponding targets was determined by Sanger sequencing. The sequencing results are shown in the figure. Figure 3 (The red arrow in the figure indicates the target sequence and sequencing direction), the editing efficiency results are as follows Figure 4 The adenine base editor provided by the present invention can convert the base A of the target site into the base G (or the base T into the base C). Figure 3 and Figure 4 The results showed that embedding adenine deaminase MaTadA into the PI domain or Ruvc-III domain of nSpCas9 (proteins in experimental group 1 and protein in experimental group 2) can effectively broaden the editing window of the base editor and improve the efficiency of base conversion within the window. The editing window is the 3rd, 4th, 5th, 6th, 7th, 8th, 9th, 10th, 11th, 12th, 13th and 14th positions from the target. From the HBGp target sequence results ( Figure 3 A to Figure 3 C) and HEK2 target sequence results ( Figure 3 D to Figure 3 F) shows that the editing window of ABE mediated by the traditional ligation method SpCas9 is 3-7 ( Figure 3 A and Figure 3 D), while the editing window of MaTadA embedded in the ABE of nSpCas9 was widened to the 14th position, and the editing efficiency of base A was increased to 82% at most. The embedded site of the subsequent embodiments selected the 1247th position of nSpCas9.
[0067] Example 2. MaTadA is inserted into nSpCas9 to form adenine base editors through different linkers to screen for linkers of appropriate length
[0068] Referring to the method of Example 1, the adenine base editor vector of Table 2 was constructed ( Figure 1 D). Figure 1 D is experimental group 3 (sequence list marked as nSpCas9-N1247-linker 2-MaTadA-linker 2-nSpCas9-C1248), experimental group 3 protein is MaTadA fused to position 1247 of nSpCas9 protein through linker 2 at both ends, the amino acid sequence of experimental group 3 protein is SEQID NO.58, and the two ends of the protein are fused to an NLS (SEQ ID NO.34) through a linker. These eukaryotic codon-optimized adenine base editor sequences are constructed into mammalian cell expression vectors, and these adenine base editors are expressed by the CAG promoter, and the eGFP gene (for cell sorting) is connected to the downstream of the adenine base editor through a self-cleaving polypeptide 2A (P2A); at the same time, the expression vector contains an adenine base editor sequence (experimental group 3 protein) and an sgRNA expression cassette (HBGp expression cassette or HEK2 expression cassette).
[0069] Table 2
[0070] The above-constructed protein vector encoding experimental group 1 and protein vector encoding experimental group 3 were transiently transfected into human HEK293T cells using PEI 40K transfection reagent, and a blank control (i.e., no vector was transfected) was set up. The cells were then cultured at 37°C and 5% carbon dioxide concentration. 48 hours after transfection, flow cytometry sorting was performed, and eGFP-positive cells were collected by fluorescence activated cell sorting (FACS). The cells were cultured for 72 hours after sorting, and then the cell genome was extracted by lysis method. The target sequence was amplified by PCR using corresponding primers, and the editing ability of adenine base editors with different linkers on the corresponding targets was determined by Sanger sequencing. The results are shown in the figure. Figure 5 From the results of HBGp target sequence and HEK2 target sequence, the adenine base editor (experimental group 1 protein and experimental group 3 protein) in which MaTadA is embedded into nSpCas9 through a long linker (SEQ ID NO.32) or a short linker (SEQ ID NO.26) can effectively broaden the editing window of the base editor and improve the base conversion efficiency within the window. The linker 1 (SEQ ID NO.32) is selected in the subsequent examples.
[0071] Example 3. Testing the editing efficiency of adenine base editors constructed with different MaTadA mutants
[0072] Referring to the method of Example 1, based on the experimental group 1 protein of SEQ ID NO.56, different MaTadA mutants were replaced with MaTadA of experimental group 1 to construct the adenine base editor vector of Table 3 ( Figure 1 E to Figure 1 O). Figure 1 E is experimental group 4 (labeled as nSpCas9-N1247-linker 1-MaTadA Q20K-linker 1-nSpCas9-C1248 in the sequence list), Figure 1 F is experimental group 5 (labeled as nSpCas9-N1247-linker 1-MaTadAM60I-linker 1-nSpCas9-C1248 in the sequence list), Figure 1 G is experimental group 6 (labeled as nSpCas9-N1247-linker 1-MaTadA S97G-linker 1-nSpCas9-C1248 in the sequence list), Figure 1 H is experimental group 7 (labeled as nSpCas9-N1247-linker 1-MaTadA Q101R-linker 1-nSpCas9-C1248 in the sequence list), Figure 1 I is experimental group 8 (the sequence list is marked as nSpCas9-N1247-connector 1-MaTadA S107R-connector 1-nSpCas9-C1248), Figure 1 J is experimental group 9 (labeled as nSpCas9-N1247-linker 1-MaTadAM118I-linker 1-nSpCas9-C1248 in the sequence list), Figure 1 K is experimental group 10 (labeled as nSpCas9-N1247-linker 1-MaTadAK129R-linker 1-nSpCas9-C1248 in the sequence list), Figure 1 L is experimental group 11 (labeled as nSpCas9-N1247-linker 1-MaTadAS97G-Q101R-linker 1-nSpCas9-C1248 in the sequence list), Figure 1 M is experimental group 12 (labeled as nSpCas9-N1247-linker 1-MaTadAS97G-S107R-linker 1-nSpCas9-C1248 in the sequence list), Figure 1 N is experimental group 13 (labeled as nSpCas9-N1247-linker 1-MaTadA S97G-Q101R-M118I-linker 1-nSpCas9-C1248 in the sequence list), Figure 1O is experimental group 14 (sequence list marked as nSpCas9-N1247-linker 1-MaTadA S97G-S107R-M118I-linker 1-nSpCas9-C1248), the amino acid sequences of experimental group 4 protein to experimental group 14 protein are SEQ ID NO.59 to SEQ ID NO.69, and an NLS (SEQ ID NO.34) is fused to both ends of the protein through a linker. These eukaryotic codon-optimized adenine base editor sequences were constructed into mammalian cell expression vectors, and these adenine base editors were expressed by the CAG promoter, and the eGFP gene (for cell sorting) was connected to the downstream of the adenine base editor through the self-cleaving polypeptide 2A (P2A); at the same time, each expression vector contains an adenine base editor sequence (experimental group 4 protein to experimental group 14 protein) and a HEK2 expression cassette.
[0073] Table 3
[0074] The above-constructed vectors encoding protein 4 of experimental group 1 to protein 14 of experimental group 1 were transiently transfected into human HEK293T cells using PEI 40K transfection reagent, and a blank control (i.e., no vector was transfected) was set up at the same time. The cells were then cultured at 37°C and 5% carbon dioxide concentration. 48 hours after transfection, flow cytometry sorting was performed, and eGFP-positive cells were collected by fluorescence activated cell sorting (FACS). The cells were cultured for 72 hours after sorting, and then the cell genome was extracted by lysis method. The target sequence was amplified by PCR using corresponding primers, and the editing efficiency of adenine base editors with different linkers on the corresponding targets was determined by Sanger sequencing. The results are as follows: Figure 6 And Table 4. From the results of the HEK2 target sequence, different MaTadA mutants embedded in the adenine base editors inside nSpCas9 (experimental group 4 protein and experimental group 14 protein) can effectively broaden the editing window of the base editor and improve the efficiency of base conversion within the window. Specifically, the single-site mutation MaTadA (experimental group 4 protein to experimental group 10 protein) has a high editing efficiency for the 5th, 7th and 12th bases, and the double-site mutation MaTadA (experimental group 11 protein and experimental group 12 protein) has a high editing efficiency for the 5th, 7th, 8th and 12th bases, and the triple-site mutation MaTadA (experimental group 13 protein and experimental group 14 protein) has a high editing efficiency for the 5th, 7th, 8th, 9th, 12th and 14th bases. The subsequent examples are tested using experimental group 13 protein and experimental group 14 protein.
[0075] Table 4
[0076] Example 4. Testing the editing efficiency of adenine base editors constructed with different MaTadA mutants
[0077] Referring to the method of Example 1, based on the control protein of SEQ ID NO.55, the three-point mutated MaTadA (MaTadA S97G-S107R-M118I and MaTadA S97G-S107R-M118I) replaced the MaTadA of the control group to construct the adenine base editor vector of Table 5 ( Figure 1 P to Figure 1 Q). Figure 1 P is experimental group 15 (labeled as MaTadAS97G-Q101R-M118I-linker 1-nSpCas9 in the sequence list), Figure 1 Q is experimental group 16 (labeled as MaTadA S97G-S107R-M118I-linker 1-nSpCas9 in the sequence list), the amino acid sequences of the experimental group 15 protein and the experimental group 16 protein are SEQ ID NO.70 to SEQ ID NO.71, and an NLS (SEQ ID NO.34) is fused to both ends of the protein through a linker. These eukaryotic codon-optimized adenine base editor sequences were constructed into mammalian cell expression vectors. These adenine base editors were expressed by the CAG promoter, and the eGFP gene (for cell sorting) was connected to the downstream of the adenine base editor through the self-cleaving polypeptide 2A (P2A); at the same time, a site1 expression cassette ( Figure 2 C), the structure of the site4 expression cassette is as follows Figure 2 As shown, the site1 expression cassette ( Figure 2 C) includes U6 promoter, site1 Target gRNA (SEQ ID NO.51) and SpCas9-gRNA scaffold (SEQ ID NO.41) in sequence, and the sequence of site1 SpgRNA is as SEQ ID NO.52. That is, each expression vector contains an adenine base editor sequence (experimental group 15 protein or experimental group 16 protein) and a site1 expression cassette.
[0078] At the same time, the sgRNA expression cassette of the control protein vector of Example 1 constructed in the above embodiment was replaced with the site1 expression cassette, the sgRNA expression cassette of the protein vector of experimental group 1 was replaced with the site1 expression cassette, the sgRNA expression cassette of the protein vector of experimental group 13 was replaced with the site1 expression cassette, and the sgRNA expression cassette of the protein vector of experimental group 14 was replaced with the site1 expression cassette.
[0079] Table 5
[0080] The above-constructed protein vectors encoding control group protein vector, experimental group 15 protein vector, experimental group 16 protein vector, experimental group 1 protein vector, experimental group 13 protein vector and experimental group 14 protein vector were transiently transfected into human HEK293T cells using PEI 40K transfection reagent, and then the cells were cultured at 37°C and 5% carbon dioxide concentration. After transfection for 48 hours, flow cytometry sorting was performed, and eGFP-positive cells were collected by fluorescence activated cell sorting (FACS). After cell sorting, the cells were cultured for 72 hours, and then the cell genome was extracted by lysis method. The target sequence was amplified by PCR using corresponding primers, and the editing efficiency of adenine base editors with different linkers on the corresponding targets was determined by Sanger sequencing. The results are shown in Figure 7 And Table 6. From the site1 target sequence results, the control group (unmutated MaTadA) fused to the N-terminus of nSpCas9 has a high editing efficiency for the 4th, 5th, 6th and 9th bases. The three-point mutated MaTadA (MaTadA S97G-S107R-M118I and MaTadA S97G-S107R-M118I) of the present invention is fused to the N-terminus of nSpCas9. The adenine base editor has a high editing efficiency for the 4th, 5th, 6th, 9th and 11th bases, and the adenine base editor with these three point mutations MaTadA embedded in nSpCas9 has a high editing efficiency for the 4th, 5th, 6th, 9th, 11th and 12th bases.
[0081] Table 6 ABE-Editing Efficiency %-Target: site1 A2 A4 A5 A6 A9 A11 A12 Control group 3 91 18 85 57 7 0 Experimental Group 15 2 91 24 88 72 14 3 Experimental Group 16 3 91 24 86 78 7 4 Experimental Group 1 1 31 15 90 89 59 5 Experimental Group 13 1 37 17 95 95 89 9 Experimental Group 14 1 41 20 94 93 77 15
[0082] Example 5. Testing the editing efficiency of adenine base editors constructed with different MaTadA mutants
[0083] Referring to the method of Example 4, the site1 expression cassette encoding the control group protein vector, the experimental group 15 protein vector, the experimental group 16 protein vector, the experimental group 1 protein vector, the experimental group 13 protein vector and the experimental group 14 protein vector constructed in Example 4 was replaced with the site4 expression cassette. The structure of the site4 expression cassette is as follows: Figure 2 As shown in D, the site4 expression cassette ( Figure 2 D) includes U6 promoter, site4 Target gRNA (SEQ ID NO.53) and SpCas9-gRNA scaffold (SEQ ID NO.41) in sequence, and the sequence of site4 SpgRNA is as SEQ ID NO.54.
[0084] The above-constructed protein vectors encoding the control group protein vector, experimental group 15 protein vector, experimental group 16 protein vector, experimental group 1 protein vector, experimental group 13 protein vector and experimental group 14 protein vector were transiently transfected into human HEK293T cells using PEI 40K transfection reagent, and a blank control (i.e., no vector was transfected) was set up at the same time. Then, the cells were cultured at 37°C and 5% carbon dioxide concentration. After 48 hours of transfection, flow cytometry sorting was performed, and eGFP-positive cells were collected by fluorescence activated cell sorting (FACS). After cell sorting, the cells were cultured for 72 hours, and then the cell genome was extracted by lysis method. The target sequence was amplified by PCR using the corresponding primers, and the editing efficiency of adenine base editors with different linkers on the corresponding targets was determined by Sanger sequencing. The results are as follows Figure 8 And Table 7. From the site4 target sequence results, the control group (unmutated MaTadA) fused to the N-terminus of nSpCas9 has a high editing efficiency for the 3rd, 5th, 8th and 14th bases. The three-point mutated MaTadA of the present invention is fused to the N-terminus of nSpCas9. The adenine base editor has a high editing efficiency for the 3rd, 5th, 8th and 14th bases. The three-point mutated MaTadA embedded in the nSpCas9 has a high editing efficiency for the 5th, 8th, 9th, 12th and 14th bases.
[0085] It can be seen from the above embodiments that the adenine base editor formed by the fusion of the MaTadA mutant of the present invention at the end of the nSpCas9 protein or the embedding of the MaTadA protein at a specific site of the nSpCas9 protein can effectively broaden the editing window of the base editor and improve the base conversion efficiency within the window.
[0086] Table 7 ABE-Editing Efficiency%-Target Site4 A2 A3 A5 A8 A9 A12 A14 A16 Blank control 3 6 1 0 3 3 3 1 Control group 3 49 89 79 6 3 14 4 Experimental Group 15 4 65 87 81 8 6 18 0 Experimental Group 16 4 2 2 3 3 3 4 1 Experimental Group 1 4 3 46 87 49 23 83 1 Experimental Group 13 4 5 49 86 58 32 83 1 Experimental Group 14 5 4 40 87 78 25 83 2
[0087] It can be seen from the above embodiments that the adenine base editor formed by the fusion of the MaTadA mutant of the present invention at the end of the nSpCas9 protein or the embedding of the MaTadA protein at a specific site of the nSpCas9 protein can effectively broaden the editing window of the base editor and improve the base conversion efficiency within the window.
Claims
1. An adenine deaminase having an amino acid substitution at one, more or all of the following positions relative to the sequence shown in SEQ ID NO: 1: Q20, M60, S97, Q101, S107, M118 and K129, and having an amino acid sequence with a sequence identity of about 70% to about 99.5% with the sequence shown in SEQ ID NO:
1.
2. Adenine deaminase according to claim 1, wherein: (1) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at position Q20; (2) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at position M60; (3) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at position S97; (4) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at position Q101; (5) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at position S107; (6) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at position M118; (7) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has an amino acid substitution at position K129; (8) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has amino acid substitutions at positions S97 and Q101; (9) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has amino acid substitutions at positions S97 and S107; (10) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has amino acid substitutions at positions S97, Q101 and M118; or (11) According to the sequence numbering shown in SEQ ID NO.1, the adenine deaminase has amino acid substitutions at positions S97, S107 and M118.
3. Adenine deaminase according to claim 1 or 2, wherein: (a) Q20 is substituted with Q20K; (b) substitution of M60 with M60I; (c) S97 is substituted with S97G; (d) Q101 is substituted with Q101R; (e) S107 is substituted with S107R; (f) M118 is substituted by M118I; or (g) K129 is substituted with K129R.
4. The adenine deaminase according to any one of claims 1 to 3, wherein the adenine deaminase comprises or is an amino acid sequence as shown in any one of SEQ ID NOs: 2 to 12 or an amino acid sequence having at least about 80% sequence identity with each of them.
5. An adenine base editor comprising the adenine deaminase according to any one of claims 1 to 4; The adenine base editor comprises a DNA binding domain and an adenine deaminase; preferably, the adenine base editor comprises, from the N-terminus to the C-terminus: (i) DNA binding domain and adenine deaminase; (ii) an adenine deaminase and DNA binding domain; or (iii) DNA binding domain-N-terminal protein, adenine deaminase and DNA binding domain-C-terminal protein.
6. The adenine base editor according to claim 5, wherein the DNA binding domain comprises a CRISPR / Cas domain with single nickase activity, a zinc finger nuclease ZFN, a transcription activator-like effector nuclease TALEN, a large range of nucleases or any combination thereof; preferably, the DNA binding domain comprises a CRISPR / Cas domain with single nickase activity; optionally, it is nSpCas9 or nSaCas9; specifically, the nSpCas9 comprises or is an amino acid sequence shown in SEQ ID NO: 13 or an amino acid sequence having at least about 80% sequence identity with each of them; Preferably, the DNA binding domain-N-terminal protein is nCas9-N-terminal protein; the DNA binding domain-C-terminal protein is nCas9-C-terminal protein; specifically, the nCas9-N-terminal protein comprises or is an amino acid sequence as shown in SEQ ID NO: 14 or 16, or an amino acid sequence having at least about 80% sequence identity with each of them; the nCas9-C-terminal protein comprises or is an amino acid sequence as shown in SEQ ID NO: 15 or 17, or an amino acid sequence having at least about 80% sequence identity with each of them; More preferably, the adenine base editor comprises an amino acid sequence selected from any one of SEQ ID NOs: 56 to 71 or an amino acid sequence having at least about 80% sequence identity with each of them.
7. The adenine base editor of claim 5, further comprising a cytidine deaminase domain.
8. The adenine base editor of claim 5, further comprising a second adenine deaminase domain, wherein the second adenine deaminase domain is the same as or different from the adenine deaminase of any one of claims 1 to 4.
9. The adenine base editor according to any one of claims 5 to 8, further comprising one or more epitope tags, nuclear localization signals, reporter gene sequences, domains capable of binding to DNA molecules or intracellular molecules, enzymes capable of detecting signals, subcellular localization and protein transduction domains; preferably, a nuclear localization signal is contained at the C-terminus or / and N-terminus of the adenine base editor; preferably, two nuclear localization signals are contained at the C-terminus or / and N-terminus of the adenine base editor; preferably, three nuclear localization signals are contained at the C-terminus or / and N-terminus of the adenine base editor; preferably, the nuclear localization signal comprises or is an amino acid sequence shown in any one of SEQ ID NOs: 34 to 39 or an amino acid sequence having at least about 95% sequence identity with each of them; Preferably, the structure of the adenine base editor is selected from: NH2-[Adenine base editor]-[NLS]-COOH; NH2-[Adenine base editor]-[NLS]-[NLS]-COOH; NH2-[Adenine base editor]-[NLS]-[NLS]-[NLS]-COOH; NH2-[NLS]-[Adenine base editor]-COOH; NH2-[NLS]-[NLS]-[Adenine Base Editor]-COOH; NH2-[NLS]-[NLS]-[NLS]-[Adenine Base Editor]-COOH; NH2-[NLS]-[Adenine base editor]-[NLS]-COOH; NH2-[NLS]-[NLS]-[Adenine base editor]-[NLS]-[NLS]-COOH; NH2-[NLS]-[NLS]-[NLS]-[Adenine base editor]-[NLS]-[NLS]-[NLS]-COOH; NH2-[NLS]-[NLS]-[Adenine base editor]-[NLS]-COOH; NH2-[NLS]-[NLS]-[NLS]-[Adenine base editor]-[NLS]-COOH; NH2-[NLS]-[Adenine base editor]-[NLS]-[NLS]-COOH; NH2-[NLS]-[Adenine base editor]-[NLS]-[NLS]-[NLS]-COOH; NH2-[NLS]-[NLS]-[Adenine Base Editor]-[NLS]-[NLS]-[NLS]-COOH; or NH2-[NLS]-[NLS]-[NLS]-[Adenine base editor]-[NLS]-[NLS]-COOH.
10. A complex comprising the adenine base editor of any one of claims 5 to 9 and a guide RNA, wherein the guide RNA complexes with the adenine base editor to guide the DNA binding domain of the adenine base editor to bind to a target nucleic acid.
11. A nucleic acid comprising a polynucleotide encoding an adenine base editor as described in any one of claims 5 to 9 or a complex as described in claim 10; preferably, the polynucleotide is codon optimized for expression in prokaryotic or eukaryotic cells; preferably, the nucleic acid is DNA or mRNA.
12. A vector comprising the nucleic acid of claim 11; preferably, the vector is a plasmid or a viral vector; preferably, the viral vector is an adeno-associated viral vector, an adenoviral vector, a retroviral vector, a lentiviral vector or a herpes simplex viral vector.
13. A vector system, comprising a first vector and a second vector different from the first vector, wherein the first vector comprises a polynucleotide encoding an adenine base editor as described in any one of claims 5 to 9 or an adenine base editor comprising an amino acid sequence shown in any one of SEQ ID NOs. 56 to 71; the second vector comprises a guide RNA or a nucleotide sequence encoding the guide RNA; preferably, the first vector and the second vector are independently plasmids or viral vectors; preferably, the viral vector is an adeno-associated viral vector, an adenoviral vector, a retroviral vector, a lentiviral vector or a herpes simplex viral vector.
14. A delivery system comprising the adenine base editor of any one of claims 5 to 9, the complex of claim 10, the nucleic acid of claim 11, the vector of claim 12, or the vector system of claim 13; preferably, the delivery system comprises liposomes, nanoparticles or exosomes.
15. A cell comprising the adenine base editor of any one of claims 5 to 9, the complex of claim 10, the nucleic acid of claim 11, the vector of claim 12, the vector system of claim 13, or the delivery system of claim 14; preferably, the cell is a eukaryotic cell or a prokaryotic cell; more preferably, the cell is a human cell.
16. A composition or kit comprising the adenine base editor of any one of claims 5 to 9, the complex of claim 10, the nucleic acid of claim 11, the vector of claim 12, the vector system of claim 13, the delivery system of claim 14 or the cell of claim 15; and a pharmaceutically acceptable carrier.
17. A method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with an adenine base editor according to any one of claims 5 to 9, the complex according to claim 10, the nucleic acid according to claim 11, the vector according to claim 12, the vector system according to claim 13, or the delivery system according to claim 14, wherein the contact results in the target nucleic acid being modified; preferably, the modification comprises increasing or decreasing the expression of a target sequence in the target nucleic acid; preferably, the modification comprises deaminating a target adenine in the target nucleic acid to achieve a base pair conversion.