Base editing tool and application thereof

By developing adenosine deaminase with specific amino acid sequence variations and fusing it with Cas protein to form a base editing tool, the problem of the lack of single-base editing tools in CRISPR/Cas technology has been solved, and the accuracy and efficiency of gene editing have been improved.

CN121825946APending Publication Date: 2026-04-10SHANDONG SHUNFENG BIOTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing CRISPR/Cas technologies lack effective single-base editing tools for gene editing, especially since there are few types of adenosine deaminases, making it difficult to achieve precise conversion of A>G.

Method used

A series of adenosine deaminases (GZ-71, GZ-68, GZ-65, GZ-64, GZ-46, GZ-74, GZ-67, GZ-13, and GZ-76) have been developed. These adenosine deaminases have specific amino acid sequence variations that can fuse with Cas proteins to form base editing tools, enabling single-base editing.

Benefits of technology

This expands the application scope of single-base editing technology, enables precise modification of specific gene sites, and improves the accuracy and efficiency of gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121825946A_ABST
    Figure CN121825946A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of nucleic acid editing, and particularly relates to the technical field of regularly clustered interval short palindromic repeat (CRISPR). Specifically, the invention provides adenosine deaminase which is fused with DNA binding protein, can be used for single base editing of target nucleic acid, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. CN202311769067.7, filed on December 21, 2023, and Chinese Patent Application No. CN202410713487.1, filed on June 4, 2024. This application incorporates the entire text of the above-mentioned Chinese patent applications.

[0002] This application is a divisional application of Chinese Patent Application No. 202411887884.7, filed on December 20, 2024, with the title of “A Base Editing Tool and Its Application”. TECHNICAL FIELD

[0003] The present application relates to the field of gene editing, particularly the field of Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) technology. Specifically, the present application relates to a base editing tool, particularly a base editing tool based on adenosine deaminase. BACKGROUND

[0004] CRISPR / Cas technology is a widely used gene editing technology that specifically binds to target sequences on the genome through RNA guidance and cuts DNA to produce double-strand breaks, using biological non-homologous end joining or homologous recombination for site-directed gene editing.

[0005] The development of single-base gene editing tools enables precise modification of specific gene sites without causing DNA double-strand breaks. Adenine base editor (ABE) based on adenosine deaminase achieves A>G (from A to G) conversion. Currently, there are few types of adenosine deaminase, therefore, the present application screens multiple adenosine deaminases and constructs a single-base editing tool, expanding the application range of single-base editing technology. SUMMARY

[0006] In one aspect, the present application provides an adenosine deaminase, which is referred to as GZ-71, GZ-68, GZ-65, GZ-64, GZ-46, GZ-74, GZ-67, GZ-13 and GZ-76 in the present application, and the amino acid sequences of the above adenosine deaminases are shown in SEQ ID Nos. 3-11, respectively.

[0007] In one embodiment, the amino acid sequence of the adenosine deaminase has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to any one of SEQ ID Nos. 3-11, and substantially retains the biological function of the sequence from which it is derived. Preferably, the adenosine deaminase is derived from the same species as GZ-71, GZ-68, GZ-65, GZ-64, GZ-46, GZ-74, GZ-67, GZ-13, or GZ-76. More preferably, the adenosine deaminase is derived from Enterobacteriaceae, Escherichia, Citrobacter, Rahnella aceris, Shigella flexneri, or Salmonella enterica subsp.

[0008] In one embodiment, the amino acid sequence of the adenosine deaminase has one or more substitutions, deletions, or additions of amino acids compared to any one of SEQ ID Nos. 3-11; and substantially retains the biological function of the sequence from which it is derived; the one or more substitutions, deletions, or additions of amino acids comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 substitutions, deletions, or additions of amino acids. Preferably, the adenosine deaminase is derived from the same species as GZ-71, GZ-68, GZ-65, GZ-64, GZ-46, GZ-74, GZ-67, GZ-13, or GZ-76. More preferably, the adenosine deaminase is derived from Enterobacteriaceae, Escherichia, Citrobacter, Rahnella aceris, Shigella flexneri, or Salmonella enterica subsp.

[0009] In one embodiment, the adenosine deaminase comprises the amino acid sequence set forth in any one of SEQ ID Nos. 3-11.

[0010] In one embodiment, the amino acid sequence of the adenosine deaminase is set forth in any one of SEQ ID Nos. 3-11.

[0011] In one embodiment, the amino acid sequence of the adenosine deaminase comprises a mutation at any one or any number of the following amino acid positions corresponding to the amino acid sequence set forth in SEQ ID No. 7: position 55, position 110, position 113, position 123, position 126, position 127, position 153, position 156, position 159, position 160, position 161, and position 167, as compared to SEQ ID No. 7.

[0012] In one embodiment, the amino acid sequence of the adenosine deaminase comprises a mutation at positions 110, 153, 161, and 167 of the amino acid sequence set forth in SEQ ID No. 7, as compared to SEQ ID No. 7.

[0013] In one embodiment, the amino acid sequence of the adenosine deaminase comprises a mutation at positions 110, 153, 161, 167, and 113 of the amino acid sequence set forth in SEQ ID No. 7, as compared to SEQ ID No. 7.

[0014] In one embodiment, the amino acid sequence of the adenosine deaminase comprises a mutation at positions 110, 153, 161, 167, and 126 of the amino acid sequence set forth in SEQ ID No. 7, as compared to SEQ ID No. 7.

[0015] In one embodiment, the amino acid sequence of the adenosine deaminase comprises a mutation at positions 110, 153, 161, 167, and 127 of the amino acid sequence set forth in SEQ ID No. 7, as compared to SEQ ID No. 7.

[0016] In one embodiment, the amino acid sequence of the adenosine deaminase comprises a mutation at positions 110, 153, 161, 167, and 159 of the amino acid sequence set forth in SEQ ID No. 7, as compared to SEQ ID No. 7.

[0017] In one embodiment, the amino acid sequence of the adenosine deaminase has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to SEQ ID No. 7; and the amino acid sequence of the adenosine deaminase has a mutation at any one or any number of the following amino acid positions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 amino acids) of the amino acid sequence set forth in SEQ ID No. 7: position 55, position 110, position 113, position 123, position 126, position 127, position 153, position 156, position 159, position 160, position 161, position 167.

[0018] In one embodiment, the amino acid sequence of the adenosine deaminase has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to SEQ ID No. 7; and the amino acid sequence of the adenosine deaminase has a mutation at any one or any number of the following amino acid positions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 amino acids) of the amino acid sequence set forth in SEQ ID No. 7: position 55, position 110, position 113, position 123, position 126, position 127, position 153, position 156, position 159, position 160, position 161, position 167.

[0019] In one embodiment, the amino acid at position 55 is mutated to an amino acid other than L, e.g., G, A, V, I, P, F, Y, W, S, T, C, M, N, Q, D, E, K, R, H; preferably, T.

[0020] In one embodiment, the amino acid at position 110 is mutated to an amino acid other than V, e.g., G, A, L, I, P, F, Y, W, S, T, C, M, N, Q, D, E, K, R, H; preferably, W, K, R, or H; more preferably, K.

[0021] In one embodiment, the amino acid at position 113 is mutated to an amino acid other than S, e.g., G, A, L, I, P, F, Y, W, V, T, C, M, N, Q, D, E, K, R, H; preferably, N or R; more preferably, N.

[0022] In one embodiment, the amino acid at position 123 is mutated to an amino acid other than N, e.g., G, A, L, I, P, F, Y, W, V, T, C, M, S, Q, D, E, K, R, H; preferably, D.

[0023] In one embodiment, amino acid at position 126 is mutated to an amino acid other than N, e.g., G, A, L, I, P, F, Y, W, V, T, C, M, S, Q, D, E, K, R, H; preferably Q.

[0024] In one embodiment, amino acid at position 127 is mutated to an amino acid other than Y, e.g., G, A, L, I, P, F, N, W, V, T, C, M, S, Q, D, E, K, R, H; preferably H, L, or M; more preferably H.

[0025] In one embodiment, amino acid at position 153 is mutated to an amino acid other than Y, e.g., L, A, V, I, P, F, N, W, S, T, C, M, G, Q, D, E, K, R, H; preferably F.

[0026] In one embodiment, amino acid at position 156 is mutated to an amino acid other than P, e.g., L, A, V, I, Y, F, N, W, S, T, C, M, G, Q, D, E, K, R, H; preferably N.

[0027] In one embodiment, amino acid at position 159 is mutated to an amino acid other than V, e.g., L, A, P, I, Y, F, N, W, S, T, C, M, G, Q, D, E, K, R, H; preferably A, H, K, or R; more preferably H or R.

[0028] In one embodiment, amino acid at position 160 is mutated to an amino acid other than F, e.g., L, A, V, I, Y, P, N, W, S, T, C, M, G, Q, D, E, K, R, H; preferably Q.

[0029] In one embodiment, amino acid at position 161 is mutated to an amino acid other than N, e.g., L, A, V, I, P, F, Y, W, S, T, C, M, G, Q, D, E, K, R, H; preferably R.

[0030] In one embodiment, amino acid at position 167 is mutated to an amino acid other than E, e.g., G, A, L, I, P, F, Y, W, S, T, C, M, N, Q, D, V, K, R, H; preferably R.

[0031] In one embodiment, the adenosine deaminase is a mutated adenosine deaminase having mutations at any one or any number of amino acid positions corresponding to amino acid positions 55, 110, 113, 123, 126, 127, 153, 156, 159, 160, 161, 167 of the amino acid sequence set forth in SEQ ID No. 7, as compared to the amino acid sequence of a parent adenosine deaminase (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 amino acids).

[0032] In one embodiment, the mutated adenosine deaminase has mutations at amino acid positions 110, 153, 161, and 167 of the amino acid sequence set forth in SEQ ID No. 7, as compared to the amino acid sequence of a parent adenosine deaminase.

[0033] In one embodiment, the mutated adenosine deaminase has mutations at amino acid positions 110, 153, 161, 167, and 113 of the amino acid sequence set forth in SEQ ID No. 7, as compared to the amino acid sequence of a parent adenosine deaminase.

[0034] In one embodiment, the mutated adenosine deaminase has mutations at amino acid positions 110, 153, 161, 167, and 126 of the amino acid sequence set forth in SEQ ID No. 7, as compared to the amino acid sequence of a parent adenosine deaminase.

[0035] In one embodiment, the mutated adenosine deaminase has mutations at amino acid positions 110, 153, 161, 167, and 127 of the amino acid sequence set forth in SEQ ID No. 7, as compared to the amino acid sequence of a parent adenosine deaminase.

[0036] In one embodiment, the mutated adenosine deaminase has mutations at amino acid positions 110, 153, 161, 167, and 159 of the amino acid sequence set forth in SEQ ID No. 7, as compared to the amino acid sequence of a parent adenosine deaminase.

[0037] In one embodiment, the mutated adenosine deaminase has mutations at amino acid positions 110, 153, 161, 167, 55, 113, 123, 126, 127, 156, 159, and 160 of the amino acid sequence set forth in SEQ ID No. 7, as compared to the amino acid sequence of a parent adenosine deaminase.

[0038] In one embodiment, the adenosine deaminase has a mutation at an amino acid position corresponding to amino acid position 106 and / or 157 of the amino acid sequence set forth in SEQ ID No. 2.

[0039] In one embodiment, the mutated adenosine deaminase has a mutation at an amino acid position corresponding to amino acid position 106 and / or 157 of the amino acid sequence set forth in SEQ ID No. 2, as compared to the amino acid sequence of the parent adenosine deaminase.

[0040] In one embodiment, the amino acid at position 106 is mutated to an amino acid other than V, e.g., G, A, L, I, P, F, Y, W, S, T, C, M, N, Q, D, E, K, R, H; preferably, W, K, or H; more preferably, H.

[0041] In one embodiment, the amino acid at position 157 is mutated to an amino acid other than N, e.g., L, A, V, I, P, F, Y, W, S, T, C, M, G, Q, D, E, K, R, H; preferably, R.

[0042] In one embodiment, the parent adenosine deaminase is a naturally-occurring wild-type adenosine deaminase; in other embodiments, the parent adenosine deaminase is an engineered adenosine deaminase.

[0043] In one embodiment, the parent adenosine deaminase is TadA, GZ-71, GZ-68, GZ-65, GZ-64, GZ-46, GZ-74, GZ-67, GZ-13, or GZ-76.

[0044] In one embodiment, the amino acid sequence of the parent adenosine deaminase has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the sequence set forth in any one of SEQ ID Nos. 2-11.

[0045] In one embodiment, the amino acid sequence of the parent adenosine deaminase is set forth in any one of SEQ ID Nos. 2-11.

[0046] In one embodiment, the amino acid sequence of the parent adenosine deaminase is set forth in any one of SEQ ID No. 2 or SEQ ID No. 7.

[0047] In one embodiment, the adenosine deaminase is selected from any one of the following groups I-III:

[0048] I. an adenosine deaminase resulting from a mutation of the amino acid sequence set forth in SEQ ID No. 7 at any one or any combination of the following amino acid positions: position 55, position 110, position 113, position 123, position 126, position 127, position 153, position 156, position 159, position 160, position 161, position 167; or an adenosine deaminase resulting from a mutation of the amino acid sequence set forth in SEQ ID No. 2 at the following amino acid positions: position 106 and / or position 157;

[0049] II. an adenosine deaminase having the mutation sites described in I; and having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the adenosine deaminase described in I;

[0050] III. an adenosine deaminase having the mutation sites described in I; and having a sequence with one or more substitutions, deletions, or additions of amino acids as compared to the adenosine deaminase described in I; the one or more amino acids include 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions, deletions, or additions.

[0051] In the present disclosure, adenosine deaminase, also known as adenine deaminase, catalyzes the hydrolytic deamination of adenine or adenosine. The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as a bacterium.

[0052] In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase, or a variant that has been mutated but still has adenosine deaminase activity.

[0053] In some embodiments, the adenosine deaminase is from a prokaryote. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Enterobacteriaceae, Escherichia, Citrobacter rodentium or Citrobacter, Rahnella aceris, Shigella flexneri, or Salmonella enterica subsp.

[0054] In preferred embodiments, the adenosine deaminase of the present application is TadA, which comprises any of the TadA or variants thereof in WO2021158921A2, such as TadA7.10, TadA-8e, TadA-8e V106W.

[0055] In preferred embodiments, the adenosine deaminase of the present application is TadA8e, with the amino acid sequence as shown in SEQ ID No. 2.

[0056] In preferred embodiments, the adenosine deaminase of the present application is GZ-71, with the amino acid sequence as shown in SEQ ID No. 3, which is derived from Enterobacteriaceae.

[0057] In preferred embodiments, the adenosine deaminase of the present application is GZ-68, with the amino acid sequence as shown in SEQ ID No. 4, which is derived from Escherichia.

[0058] In preferred embodiments, the adenosine deaminase of the present application is GZ-65, with the amino acid sequence as shown in SEQ ID No. 5, which is derived from Citrobacter rodentium.

[0059] In preferred embodiments, the adenosine deaminase of the present application is GZ-64, with the amino acid sequence as shown in SEQ ID No. 6, which is derived from Citrobacter.

[0060] In preferred embodiments, the adenosine deaminase of the present application is GZ-46, with the amino acid sequence as shown in SEQ ID No. 7, which is derived from Rahnella aceris.

[0061] In a preferred embodiment, the adenosine deaminase GZ-74 according to the present application, whose amino acid sequence is shown as SEQ ID No. 8, is derived from Shigella flexneri.

[0062] In a preferred embodiment, the adenosine deaminase GZ-67 according to the present application, whose amino acid sequence is shown as SEQ ID No. 9, is derived from Salmonella enterica subsp.

[0063] In one aspect, the present application provides a fusion protein comprising a DNA binding protein and the above-mentioned adenosine deaminase.

[0064] In one embodiment, the DNA binding protein is a nucleic acid programmable DNA binding protein (napDNAbp) domain, which refers to any protein that can locate and bind to a specific target DNA nucleotide sequence (e.g. a genomic locus). In one embodiment, the DNA binding protein is a Cas protein selected from any one of Cas9, Cas9n, dCas9, CasX, CasY, C2cl, C2c2, C2c3, GeoCas9, CjCas9, Casl2a, Casl2b, Casl2g, Casl2h, Casl2i, Cas12j, Casl3b, Casl3c, Casl3d, Casl4, Csn2, xCas9, Cas9-NG, LbCasl2a, enAsCasl2a, Cas9-KKH, circularly permutated Cas9, Argonaute (Ago) domain, SmacCas9, Spy-macCas9, SpGas9-NRRH, SpaCas9-NRTH, SpaCas9-NRCH, Cas9-NG-CP1041, Cas9-NG-VRQR, dCasl2i, nCasl2i, and variants thereof.

[0065] In one embodiment, the Cas protein is Casl2i.

[0066] In one embodiment, the Cas protein can be naturally occurring or non-naturally occurring and engineered.

[0067] In one embodiment, the Cas protein is a nuclease-inactivated Cas protein (dCas).

[0068] In an embodiment, the Cas protein is a dCas12i3 protein, and the amino acid sequence of the wild-type Cas12i is shown as SEQ ID No. 1; the dCas12i protein has a mutation of E to A at position 844 or a mutation of D to A at position 619 compared to SEQ ID No. 1.

[0069] In an embodiment, the Cas protein is a Cas12i mutant protein, and the amino acid sequence of the wild-type Cas12i is shown as SEQ ID No. 1; the Cas12i mutant protein has a mutation of S to R at position 7, a mutation of D to R at position 233, a mutation of D to R at position 267, a mutation of N to R at position 369, and a mutation of S to R at position 433 compared to SEQ ID No. 1.

[0070] In an embodiment, the Cas protein is a dCas12i mutant protein, and the amino acid sequence of the wild-type Cas12i is shown as SEQ ID No. 1; the dCas12i mutant protein has a mutation of E to A at position 844, a mutation of S to R at position 7, a mutation of D to R at position 233, a mutation of D to R at position 267, a mutation of N to R at position 369, and a mutation of S to R at position 433 compared to SEQ ID No. 1.

[0071] In an embodiment, the DNA binding protein in the fusion protein is fused to the N-terminus of the adenosine deaminase, and in other embodiments, the DNA binding protein in the fusion protein is fused to the C-terminus of the adenosine deaminase.

[0072] In an embodiment, the DNA binding protein and the adenosine deaminase in the fusion protein are connected by a linker.

[0073] In an embodiment, the linker is an XTEN linker.

[0074] In an embodiment, the fusion protein of the present application further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In other embodiments, both the N-terminus and the C-terminus of the fusion protein are connected with an NLS.

[0075] In some embodiments, the NLS is fused to the N-terminus of the Cas protein. In some embodiments, the NLS is fused to the C-terminus of the Cas protein. In some embodiments, the NLS is fused to the N-terminus of the deaminase. In some embodiments, the NLS is fused to the C-terminus of the deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker.

[0076] It is clear to those skilled in the art that the structure of a protein can be changed without adversely affecting its activity and functionality, e.g., one or more conservative amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples of conservative amino acid substitutions and embodiments are clear to those skilled in the art. Specifically, an amino acid residue can be replaced with another amino acid residue belonging to the same group, i.e., a nonpolar amino acid residue is replaced with another nonpolar amino acid residue, a polar uncharged amino acid residue is replaced with another polar uncharged amino acid residue, a basic amino acid residue is replaced with another basic amino acid residue, and an acidic amino acid residue is replaced with another acidic amino acid residue. Such substituted amino acid residues can or can not be encoded by the genetic code. Conservative substitutions in which one amino acid is replaced with another amino acid from the same group are within the scope of the present application, provided that the substitution does not result in inactivation of the biological activity of the protein. Thus, the proteins of the present application can contain one or more conservative substitutions in the amino acid sequence, which are preferably made according to Table 1. In addition, the present application also encompasses proteins which further contain one or more other non-conservative substitutions, provided that the non-conservative substitutions do not significantly affect the desired function and biological activity of the proteins of the present application.

[0077] Conservative amino acid substitutions can be made at one or more predicted nonessential amino acid residues. A "nonessential" amino acid residue is a residue that can be altered without a significant loss of activity, i.e., a residue that is not required for biological activity. A "conservative amino acid substitution" is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain. Amino acid substitutions can be made in non-conserved regions of the Cas mutant proteins or fusion proteins described above. In general, such substitutions will not be made in conserved amino acid residues, or in active site residues, where such residues are required for protein activity. However, it will be appreciated by one skilled in the art that functional variants can have fewer than the conservative or non-conservative alterations in conserved regions.

[0078] Table 1

[0079]

[0080] It is well known in the art that one or more amino acid residues can be altered (substituted, deleted, truncated, or inserted) from the N and / or C terminus of a protein while still retaining its functional activity. Thus, proteins having one or more amino acid residues altered from the N and / or C terminus of a Cas mutant protein while retaining its desired functional activity are also within the scope of the present application. These alterations can include alterations introduced by modern molecular methods, such as PCR, including PCR amplification of protein-encoding sequences by means of inclusion of amino acid-encoding sequences in the oligonucleotides used in the PCR amplification to alter or extend the protein-encoding sequence.

[0081] It is recognized that proteins can be altered in a variety of ways including amino acid substitutions, deletions, truncations and insertions, and that methods for such manipulations are generally known in the art. For example, amino acid sequence variants of the above-described proteins can be prepared by mutations of the DNA. They can also be accomplished by other forms of mutagenesis and / or by directed evolution, e.g., using known mutagenic, recombination and / or shuffling methods in combination with relevant screening methods to make single or multiple amino acid substitutions, deletions and / or insertions.

[0082] As will be appreciated by those of ordinary skill in the art, these minor amino acid changes in the Cas proteins of the application can occur (e.g., naturally-occurring mutations) or be produced (e.g., using r-DNA technology) without loss of protein function or activity. If the mutations occur in the catalytic domain, active site, or other functional domain of the protein, the properties of the polypeptide can change, but the polypeptide can retain its activity. If the mutations are not near the catalytic domain, active site, or other functional domain, less impact can be expected.

[0083] As will be appreciated by those of ordinary skill in the art, essential amino acids of the Cas mutant proteins of the application can be identified according to methods known in the art, such as by site-directed mutagenesis or protein evolution or bioinformatic analysis. The catalytic domain, active site, or other functional domain of the protein can also be determined by physical analysis of the structure, such as by nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling in combination with mutations of putative key site amino acids.

[0084] In the present application, the amino acid residues can be represented by single letter or three letter, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamine (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine (Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), arginine (Arg, R).

[0085] The term "AxxB" means the amino acid A at xx position is mutated to amino acid B, for example, E844A means the E at position 844 is mutated to A. When multiple amino acid positions are mutated simultaneously, the mutations can be represented in the form of S7R-D233R-D267R-N369R-S433R, for example, S7R-D233R-D267R-N369R-S433R means the S at position 7 is mutated to R, the D at position 233 is mutated to R, the D at position 267 is mutated to R, the N at position 369 is mutated to R and the S at position 433 is mutated to R.

[0086] The specific amino acid positions (numbering) within the protein of the present application are determined by aligning the amino acid sequence of the protein of interest with the target sequence (e.g. SEQ ID No. 1) using standard sequence alignment tools, such as aligning the two sequences using the Smith-Waterman algorithm or using the CLUSTALW2 algorithm, wherein the sequences are considered aligned when the alignment score is highest. The alignment score can be calculated according to the method described in Wilbur, W. J. and Lipman, D. J. (1983) Rapid similarity searches of nucleic acid and protein data banks. Proc. Natl. Acad. Sci. USA, 80:726-730. In the ClustalW2 (1.82) algorithm, the use of the default parameters is preferred: Protein gap open penalty = 10.0; Protein gap extension penalty = 0.2; Protein matrix = Gonnet; Protein / DNA end gap = -1; Protein / DNA GAP DIST = 4. Preferably, the AlignX program (part of the vector NTI suite) is used to determine the positions of the specific amino acids within the protein of the present application by aligning the amino acid sequence of the protein with SEQ ID No. 1 using the default parameters for multiple alignment (gap open penalty: 10, gap extension penalty 0.05). One skilled in the art can use commonly used software in the art, such as Clustal Omega, to compare and align the amino acid sequence of any parent Cas protein with SEQ ID NO. 1 for sequence identity, and thus obtain the amino acid positions in the parent Cas protein that correspond to the amino acid positions defined based on SEQ ID NO. 1 in the present application.

[0087] The fusion protein of the present application is not limited by the way it is produced, for example, it can be produced by genetic engineering methods (recombinant technology), or it can be produced by chemical synthesis methods.

[0088] The present application also provides a base editing tool comprising the above-mentioned fusion protein, for example, a single base editing tool.

[0089] Nucleic acid

[0090] In another aspect, the present application provides an isolated polynucleotide comprising:

[0091] (a) a polynucleotide sequence encoding the adenosine deaminase or the fusion protein of the present application;

[0092] (b) a polynucleotide having a sequence as set forth in any one of SEQ ID Nos. 12-20.

[0093] (c) a sequence having one or more substitutions, deletions, or additions of one or more bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 substitutions, deletions, or additions of bases) compared to the sequence set forth in any one of SEQ ID Nos. 12-20;

[0094] (d) a nucleotide sequence having a homology of >80% (preferably >90%, more preferably >95%, most preferably >98%) to the sequence set forth in any one of SEQ ID Nos. 12-20, and encoding a polypeptide set forth in any one of SEQ ID Nos. 3-11; or,

[0095] (e) a polynucleotide complementary to the polynucleotide of any one of (a)-(d).

[0096] In one embodiment, the nucleotide sequence is codon-optimized for expression in a prokaryotic cell. In one embodiment, the nucleotide sequence is codon-optimized for expression in a eukaryotic cell.

[0097] In one embodiment, the cell is an animal cell, e.g., a mammalian cell.

[0098] In one embodiment, the cell is a human cell.

[0099] In one embodiment, the cell is a plant cell, e.g., a cell of a cultivated plant (such as cassava, corn, sorghum, wheat, or rice), an alga, a tree, or a vegetable.

[0100] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.

[0101] Guide RNA (gRNA)

[0102] In another aspect, the present application provides a gRNA, which comprises a protein-binding sequence and a targeting sequence targeting a target nucleic acid.

[0103] The protein-binding sequence of the gRNA can interact with the DNA-binding protein (Cas protein) in the fusion protein of the present application, so that the DNA-binding protein (Cas protein) and the gRNA form a complex.

[0104] It is routine technical knowledge in the art to design a gRNA that interacts with a Cas protein of the present application. For example, it does not require inventive labor to design a gRNA that interacts with a Cas9 protein, a Cas12 family of Cas proteins (including but not limited to, Cas12a, Cas12b, Cas12i, etc.).

[0105] The targeting sequence of the targeting nucleic acid of the present application comprises a nucleotide sequence that is complementary to a sequence in the target nucleic acid. In other words, the targeting sequence of the targeting nucleic acid of the present application or the targeting segment of the targeting nucleic acid interacts with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Thus, the targeting sequence of the targeting nucleic acid or the targeting segment of the targeting nucleic acid can be altered, or can be modified to hybridize to any desired sequence within the target nucleic acid. The nucleic acid is selected from DNA or RNA.

[0106] The percentage of complementarity between the targeting sequence of the targeting nucleic acid or the targeting segment of the targeting nucleic acid and the target sequence of the target nucleic acid can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).

[0107] The "protein binding sequence" of the gRNA of the present application can interact with a CRISPR protein (or, alternatively, a Cas protein). The gRNA of the present application, through the targeting sequence of the targeting nucleic acid, directs the DNA binding protein it interacts with to a specific nucleotide sequence within the target nucleic acid.

[0108] The gRNA of the present application is capable of forming a complex with the DNA binding protein (Cas protein).

[0109] Vector

[0110] The present application also provides a vector comprising the adenosine deaminase, the fusion protein, the isolated nucleic acid molecule or the polynucleotide as described above; preferably, it further comprises a regulatory element operably linked thereto.

[0111] In one embodiment, the regulatory element is selected from one or more of the group consisting of enhancer, transposon, promoter, terminator, leader sequence, polyadenylation sequence, marker gene.

[0112] In one embodiment, the vector comprises a cloning vector, an expression vector, a shuttle vector, an integration vector.

[0113] In some embodiments, the vector included in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a bacteriophage, and the like, which are well known to those skilled in the art.

[0114] Base editing system

[0115] The present application provides an engineered non-naturally occurring base editing system, which comprises the fusion protein described above or a nucleic acid sequence encoding the fusion protein, and a nucleic acid encoding one or more guide RNAs, wherein the DNA binding protein in the fusion protein is a Cas protein, and the gRNA is capable of binding to the Cas protein.

[0116] In one embodiment, the nucleic acid sequence encoding the fusion protein and the nucleic acid encoding one or more guide RNAs are artificially synthesized.

[0117] In one embodiment, the nucleic acid sequence encoding the fusion protein and the nucleic acid encoding one or more guide RNAs are not naturally co-existing.

[0118] The one or more guide RNAs target one or more target sequences in a cell. The one or more target sequences are hybridized to a genomic locus of a DNA molecule encoding one or more gene products, and direct the fusion protein to the genomic locus of the DNA molecule encoding the one or more gene products, and the fusion protein modifies or edits the target sequence after reaching the target sequence, thereby the expression of the one or more gene products is altered or modified.

[0119] The cell of the present application comprises one or more of an animal, a plant, or a microorganism.

[0120] In some embodiments, the fusion protein is codon-optimized for expression in a cell.

[0121] The present application also provides an engineered non-naturally occurring vector system, which can comprise one or more vectors, the one or more vectors comprising:

[0122] a) a first regulatory element operably linked to a gRNA,

[0123] b) a second regulatory element operably linked to the fusion protein;

[0124] wherein components (a) and (b) are located on the same or different vectors of the system.

[0125] The first and second regulatory elements comprise a promoter (e.g., a constitutive promoter or an inducible promoter), an enhancer (e.g., 35S promoter or 35S enhanced promoter), an internal ribosome entry site (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences).

[0126] In some embodiments, the vectors in the system are viral vectors (e.g., retroviral vectors, lentiviral vectors, adenoviral vectors, adeno-associated vectors, and herpes simplex vectors), and can also be plasmids, viruses, cosmids, bacteriophages, and the like, which are well known to those skilled in the art.

[0127] In some embodiments, the system provided herein is in a delivery system. In some embodiments, the delivery system is a nanoparticle, a liposome, an exosome, a microvesicle, and a gene gun.

[0128] In one embodiment, the target sequence is a DNA or RNA sequence from a prokaryotic cell or a eukaryotic cell. In one embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0129] In one embodiment, the target sequence is present within a cell. In one embodiment, the target sequence is present within a nucleus or within a cytoplasm (e.g., organelle). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.

[0130] Protein-nucleic acid complex / composition

[0131] In another aspect, the present application provides a complex or composition comprising:

[0132] (i) a protein component selected from the group consisting of: the fusion protein described above, wherein the DNA-binding protein in the fusion protein is a Cas protein;

[0133] (ii) a nucleic acid component comprising (a) a guide sequence capable of hybridizing to a target sequence; and (b) a protein-binding sequence capable of binding to the DNA-binding protein in the fusion protein of the present application.

[0134] The protein component and the nucleic acid component are capable of binding to each other to form a complex.

[0135] In one embodiment, the nucleic acid component is a guide RNA in a CRISPR-Cas system or a base editing system.

[0136] In one embodiment, the complex or composition is non-naturally occurring or modified. In one embodiment, at least one component in the complex or composition is non-naturally occurring or modified. In one embodiment, the first component is non-naturally occurring or modified; and / or, the second component is non-naturally occurring or modified.

[0137] Delivery and delivery composition

[0138] The fusion proteins, gRNAs, nucleic acid molecules, vectors, systems, complexes and compositions of the application can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipofection, nucleofection, microinjection, sonoporation, biolistics, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection, magnetofection, lipofection, perforation transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial virions, and the like.

[0139] Thus, in another aspect, the present application provides a delivery composition comprising a delivery vehicle, and one or any number of the following selected from the group consisting of: a fusion protein, gRNA, nucleic acid molecule, vector, system, complex and composition of the application.

[0140] In one embodiment, the delivery vehicle is a particle.

[0141] In one embodiment, the delivery vehicle is selected from the group consisting of a lipid particle, a sugar particle, a metal particle, a protein particle, a liposome, an exosome, a microvesicle, a biolistic particle or a viral vector (e.g. a replication-defective retrovirus, a lentivirus, an adenovirus or an adeno-associated virus).

[0142] Host cell

[0143] The present application also relates to an in vitro, ex vivo or in vivo cell or cell line or a progeny thereof, said cell or cell line or a progeny thereof comprising a fusion protein, a nucleic acid molecule, a protein-nucleic acid complex, a vector, a delivery composition of the application.

[0144] In certain embodiments, the cell is a prokaryotic cell.

[0145] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, e.g. a cell of a non-human primate, a bovine, an ovine, a porcine, a canine, a monkey, a rabbit, a rodent (e.g. a rat or a mouse). In certain embodiments, the cell is a non-mammalian eukaryotic cell, e.g. a cell of an avian bird (e.g. a chicken), a fish or a crustacean (e.g. a clam, a shrimp). In certain embodiments, the cell is a plant cell, e.g. a cell of a monocotyledonous or dicotyledonous plant or of a cultivated plant or a food crop such as cassava, maize, sorghum, soybean, wheat, oat or rice, e.g. a cell of an alga, a tree or a production plant, fruit or vegetable (e.g. a tree such as a citrus tree, a nut tree; a solanaceous plant, cotton, tobacco, tomato, grape, coffee, cocoa, etc.).

[0146] In certain embodiments, the cell is a stem cell or stem cell line.

[0147] In certain cases, the host cell of the present application comprises a modification of a gene or genome that is not present in its wild type.

[0148] Gene editing methods and applications

[0149] The fusion protein, nucleic acid, composition, base editing system, vector system, delivery composition, or host cell of the present application can be used for any one or any combination of the following purposes: gene editing; targeting and / or editing a target nucleic acid; specifically editing a double-stranded nucleic acid; base editing a double-stranded nucleic acid; base editing a single-stranded nucleic acid. In other embodiments, it can also be used for preparing a reagent or kit for any one or any combination of the above purposes.

[0150] The present application also provides a method of editing a nucleic acid, comprising the step of contacting a target region of a nucleic acid (e.g., a double-stranded DNA sequence) with a complex comprising the fusion protein and a gRNA; wherein the target region comprises a targeted base pair, and the targeted base pair in the target region is replaced by base substitution. In one embodiment, the deaminase in the fusion protein is an adenosine deaminase, and the targeted base pair is replaced from A:T to G:C.

[0151] A:T refers to the paired bases that make up a base pair are A and T; similarly, G:C refers to the paired bases that make up a base pair are G and C.

[0152] The present application also provides the use of the fusion protein, nucleic acid, composition, base editing system, vector system, delivery composition, or host cell of the present application in gene editing; or, in the preparation of a reagent or kit for gene editing.

[0153] In one embodiment, the gene editing is performed in and / or outside of a cell.

[0154] In one embodiment, the gene editing is single base editing of a target gene.

[0155] The present application also provides a method of editing a target nucleic acid, comprising contacting the target nucleic acid with the fusion protein, nucleic acid, composition, base editing system, vector system, or delivery composition of the present application. In one embodiment, the method is editing the target nucleic acid in or outside of a cell.

[0156] The gene editing or editing a target nucleic acid comprises the step of editing a single base of a target gene.

[0157] The editing can be performed in prokaryotic cells and / or eukaryotic cells.

[0158] In another aspect, the present application also provides a kit for gene editing, the kit comprising the above-mentioned adenosine deaminase, fusion protein, gRNA, nucleic acid, above-mentioned composition, above-mentioned base editing system, above-mentioned vector system, above-mentioned delivery composition or above-mentioned host cell.

[0159] In another aspect, the present application provides use of the above-mentioned adenosine deaminase, fusion protein, nucleic acid, above-mentioned composition, above-mentioned base editing system, above-mentioned vector system, above-mentioned delivery composition or above-mentioned host cell in the manufacture of a medicament or a kit for:

[0160] (i) gene or genome editing;

[0161] (ii) editing a target sequence in a target locus to modify an organism;

[0162] (iii) single base editing;

[0163] (iv) treatment of a disease.

[0164] Preferably, the above-mentioned gene or genome editing is gene or genome editing in or out of a cell.

[0165] Preferably, the above-mentioned treatment of a disease is treatment of a disorder caused by a defect in a target sequence in a target locus.

[0166] Methods of specifically modifying target nucleic acids

[0167] In another aspect, the present application also provides a method of specifically modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with the above-mentioned fusion protein, nucleic acid, above-mentioned composition, above-mentioned base editing system, above-mentioned vector system or above-mentioned delivery composition.

[0168] The specific modification can occur in vivo or in vitro.

[0169] The specific modification can occur in or out of a cell.

[0170] In some cases, the cell is selected from a prokaryotic cell or a eukaryotic cell, for example, an animal cell, a plant cell or a microbial cell.

[0171] Adenosine deaminases

[0172] As used herein, the term "adenosine deaminase" catalyzes the hydrolytic deamination of the nucleobase adenine. The adenosine deaminases provided herein (e.g., engineered adenosine deaminases or optimized adenosine deaminases) can be enzymes that convert adenosine (A) in DNA to inosine (I), and such adenosine deaminases can cause conversion of A:T to G:C base pairs. For example, in some embodiments, the amino acid sequence of the adenosine deaminase has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any one of SEQ ID Nos. 2-11.

[0173] DNA binding proteins

[0174] As used herein, the term "DNA-binding protein" or "DNA-binding protein domain" refers to any protein that is located at and binds to a particular target DNA nucleotide sequence (e.g., a locus of a genome). The term includes RNA-programmable proteins that bind (e.g., form a complex) with one or more nucleic acid molecules (i.e., that include, for example, a guide RNA in the case of a Cas system) that direct or otherwise program the protein to localize to a particular target nucleotide sequence (e.g., a DNA sequence) that is complementary to the one or more nucleic acid molecules (or portions or regions thereof) to which the protein binds. Exemplary RNA-programmable proteins are CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and can include Cas9 equivalents from any type of CRISPR system (e.g., Type II, Type V, Type VI), including Cpfl (Type V CRISPR-Cas system), C2cl (Type V CRISPR-Cas system), C2c2 (Type VI CRISPR-Cas system), C2c3 (Type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas 12a, Cas 12b, Casl2c, Casl2d, Casl2g, Casl2h, Casl2i, dCasl2i, and further Cas equivalents.

[0175] CRISPR system

[0176] As used herein, the term "Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meaning generally understood by those skilled in the art, which generally includes a transcription product or other element related to the expression of a CRISPR-associated ("Cas") gene, or a transcription product or other element capable of directing the activity of the Cas gene. The Cas protein in the present invention is Crispr associated protein.

[0177] CRISPR / Cas complex

[0178] As used herein, the term "CRISPR / Cas complex" refers to a complex formed by the binding of a guide RNA or a mature crRNA to a Cas protein, which comprises a direct repeat sequence that is hybridized to a guide sequence of a target sequence and bound to a Cas protein, the complex being capable of recognizing and cleaving a polynucleotide that can hybridize to the guide RNA or the mature crRNA.

[0179] Guide RNA (gRNA)

[0180] As used herein, the terms "guide RNA" (gRNA), "mature crRNA", "guide sequence" are used interchangeably and have the meaning generally understood by those skilled in the art. Generally, a guide RNA can comprise or essentially consist of or consist of a direct repeat sequence and a guide sequence.

[0181] In certain instances, a guide sequence is any polynucleotide sequence that has sufficient complementarity to a target sequence to hybridize to the target sequence and direct specific binding of a CRISPR / Cas complex to the target sequence. In one embodiment, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned, is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining optimal alignment is within the capabilities of those of ordinary skill in the art. For example, there are published and commercially available alignment algorithms and programs, such as, but not limited to, ClustalW, Smith-Waterman in matlab, Bowtie, Geneious, Biopython, and SeqMan.

[0182] Target sequence

[0183] A "target sequence" refers to a polynucleotide targeted by a guide sequence in a gRNA, e.g., a sequence having complementarity to the guide sequence, wherein hybridization between the target sequence and the guide sequence will facilitate formation of a CRISPR / Cas complex, including a Cas protein and a gRNA. Perfect complementarity is not required, so long as there is sufficient complementarity to cause hybridization and facilitate formation of a CRISPR / Cas complex.

[0184] A target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located within a cell or outside of a cell. In some cases, the target sequence is located in the nucleus or cytoplasm of a cell. In some cases, the target sequence can be located within an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast. A sequence or template that can be used for recombination into a target locus comprising the target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence." In one embodiment, the editing template is an exogenous nucleic acid. In one embodiment, the recombination is homologous recombination.

[0185] In the present invention, a "target sequence" or "target polynucleotide" or "target nucleic acid" can be any endogenous or exogenous polynucleotide to a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).

[0186] Base editing

[0187] The term "base editing" refers to a genome editing technology that involves the conversion of a particular nucleic acid base to another base at a targeted genomic locus. In some embodiments, this can be achieved without the need for a double-stranded DNA break (DSB) or single-stranded break (nicking). Other genome editing technologies, including CRISPR-based systems, to date, start with the introduction of a DSB at the locus of interest. Subsequently, cellular DNA repair enzymes patch the break, often resulting in the random insertion or deletion of bases (indels) at the site of the DSB. However, these genome editing technologies are not suitable when one desires to introduce or correct a point mutation at a target locus rather than the random destruction of an entire gene, as the rate of correction is low (typically 0.1% to 5%), with indels being the major genome editing product. To increase the efficiency of gene correction without introducing Rando indels, the present inventors used an adenosine deaminase in conjunction with a CRISPR system to directly convert one DNA base to another DNA base without forming a DSB.

[0188] Wild type

[0189] As used herein, the term "wild type" has the meaning generally understood by those skilled in the art and refers to the typical form of an organism, strain, gene or characteristic that distinguishes it from mutant or variant forms when it occurs in nature, which can be isolated from a source in nature and has not been intentionally modified by man.

[0190] Non-naturally occurring

[0191] As used herein, the terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of man's art, when these terms are used to describe a nucleic acid molecule or polypeptide, it indicates that the nucleic acid molecule or polypeptide is at least substantially isolated from at least another component with which it is associated in nature or as found in nature.

[0192] Identity

[0193] As used herein, the term "identity" is used in reference to the matching of sequences between two polypeptides or between two nucleic acids. When a position in each of two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percentage of identity" between two sequences is a function of the number of matching positions shared by the sequences divided by the number of positions compared x 100. For example, if 6 of 10 positions in two sequences are matched then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (3 of 6 positions are matched). Typically, the comparison is made over the length of the sequences being compared, after aligning the two sequences to produce maximum identity. Such alignment can be achieved by methods such as that of Needleman et al. (1970) J. Mol. Biol. 48:443-453, using, for example, the computer program Align (DNAstar, Inc.), conveniently. The percentage of identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)) as integrated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. In addition, the percentage of identity between two amino acid sequences can be determined using the algorithm of Needleman and Wunsch (J MoI Biol. 48:444-453 (1970)) as implemented in the GAP program, using either a Blossum 62 matrix or a PAM250 matrix, and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a gap length weight of 1, 2, 3, 4, 5, or 6, as incorporated in the GCG software package (available at www.gcg.com).

[0194] Vector

[0195] The term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules that comprise one or more free ends, no free ends (e.g., circular), nucleic acid molecules that comprise DNA, RNA, or both; and other varieties of nucleic acids known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection, and the resulting genetically modified host cell can be used to express the genetic material elements carried by the vector. A vector can be introduced into a host cell to thereby produce a transcript, protein, or peptide, including a protein, fusion protein, isolated nucleic acid molecule, etc. (e.g., a CRISPR transcript, such as a nucleic acid transcript, protein, or enzyme) as described herein. A vector can contain a variety of control elements, including but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, a vector can contain a replication origin.

[0196] One type of vector is a "plasmid," which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques.

[0197] Another type of vector is a viral vector, wherein virally-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by a virus for transfection into a host cell. Certain vectors (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors) are capable of autonomous replication in a host cell into which they are introduced.

[0198] Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors."

[0199] Host cell

[0200] As used herein, the term "host cell" refers to a cell that can be used to introduce a vector, including but not limited to, prokaryotic cells such as E. coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.

[0201] Those of skill in the art will appreciate that the design of the expression vector can depend on such factors as the choice of the host cell to be transformed, the level of expression of the desired gene, etc.

[0202] Regulatory element

[0203] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, CA (1990). In certain instances, regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells as well as those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired tissue of interest, such as muscle, neuronal, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). In certain instances, regulatory elements can also direct expression in a temporal-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which can or can not be tissue- or cell type-specific. In certain instances, the term "regulatory element" encompasses enhancer elements, such as the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., vol. 8(1), pp. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit beta-globin (Proc. Natl. Acad. Sci. USA., vol. 78(3), pp. 1527-31, 1981).

[0204] Promoter

[0205] As used herein, the term "promoter" has its art-understood meaning and refers to a non-coding nucleotide sequence located upstream of a gene that initiates expression of the downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell essentially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked with a polynucleotide encoding or specifying a gene product, results in production of the gene product in a cell essentially only when the cell is of the tissue type to which the promoter corresponds.

[0206] NLS

[0207] A "nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" a protein for import into the nucleus by nuclear transport, i.e., a protein with an NLS is transported to the nucleus. Typically, an NLS comprises positively charged Lys or Arg residues that are exposed on the surface of the protein. Exemplary nuclear localization sequences include, but are not limited to, NLS from SV40 large T antigen, EGL-13, c-Myc, and TUS protein. In some embodiments, the NLS comprises the MKRTADGSEFESPKKKRKV sequence. In some embodiments, the NLS comprises the KRPAATKKAGQAKKKK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNPAl, the sequence KIPIK and PY-NLS in the yeast transcriptional repressor Matα2.

[0208] Operably linked

[0209] As used herein, the term "operably linked" is intended to mean that a nucleotide sequence of interest is linked to the one or more regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or, when the vector is introduced into a host cell, in the host cell).

[0210] Complementarity

[0211] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence by virtue of traditional Watson-Crick or other non-traditional types of bonding. Percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 is 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Perfect complementarity" indicates that all consecutive residues of a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. "Substantially complementary" as used herein refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to hybridization under stringent conditions of two nucleic acids.

[0212] Stringent conditions

[0213] As used herein, "stringent conditions" for hybridization refer to conditions under which a nucleic acid having complementarity to a target sequence will hybridize primarily to that target sequence and not to non-target sequences. Stringent conditions are often sequence dependent, and are varied depending on many factors. In general, the longer the sequence, the higher the temperature at which the sequence will specifically hybridize to its target sequence.

[0214] Hybridization

[0215] The terms "hybridize" or "complementary" or "substantially complementary" refer to the non-covalent binding of nucleotides of a nucleic acid (e.g., RNA, DNA) to another nucleic acid by means of base pairing and / or G / U base pairing in a sequence-specific, anti-parallel manner ("annealing" or "hybridizing") to another nucleic acid.

[0216] Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases can occur. Suitable conditions for hybridization between two nucleic acids depend on the length and complementarity of the nucleic acids, which are variables known in the art. Typically, the length of the hybridizable nucleic acid is 8 nucleotides or more (e.g., 10 nucleotides or more, 12 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 22 nucleotides or more, 25 nucleotides or more, or 30 nucleotides or more).

[0217] It is understood that the sequence of a polynucleotide need not be 100% complementary to the sequence of its target nucleic acid to hybridize specifically thereto. A polynucleotide can comprise 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% complementarity to the sequence of the target region in the sequence of the target nucleic acid to which it hybridizes.

[0218] Hybridization of a target sequence to a gRNA represents that at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize, forming a complex; or represents that at least 12, 15, 16, 17, 18, 19, 20, 21, 22, or more bases of the nucleic acid sequences of the target sequence and the gRNA can base pair, hybridizing to form a complex.

[0219] Expression

[0220] As used herein, the term "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., transcribed into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. Transcripts and encoded polypeptides can be collectively referred to as "gene product." If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell.

[0221] Linker

[0222] As used herein, the term "linker" refers to a linear polypeptide formed by the linkage of a plurality of amino acid residues via peptide bonds. The linkers of the present application can be artificial synthetic amino acid sequences, or naturally occurring polypeptide sequences, such as polypeptides having a hinge region function. Such linker polypeptides are well known in the art (see, e.g., Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, R. J. et al. (1994) Structure 2:1121-1123).

[0223] Treatment

[0224] As used herein, the term "treatment" refers to therapy or healing of a disorder, delay in the onset of symptoms of a disorder, and / or delay in the progression of a disorder.

[0225] Subject

[0226] As used herein, the term "subject" includes, but is not limited to, various animals, plants, and microorganisms.

[0227] Animal

[0228] For example, a mammal, such as a bovine, equine, ovine, porcine, canine, feline, leporid, rodent (e.g., a mouse or rat), non-human primate (e.g., a macaque or cynomolgus monkey), or a human. In certain embodiments, the subject (e.g., a human) has a disorder (e.g., a disorder resulting from a disease-associated gene defect).

[0229] Plant

[0230] The term "plant" is to be understood as any differentiated multicellular organism capable of performing photosynthesis, including crop plants at any stage of maturation or development, in particular monocotyledonous or dicotyledonous plants, vegetable crops, including artichokes, Brussels sprouts, cress, leeks, asparagus, lettuce (e.g., iceberg lettuce, leaf lettuce, long-leaf lettuce), bok choy, yellow fleshed yam, melons (e.g., muskmelons, watermelons, crenshaw melons, cantaloupes, Roman melons), oilseed crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, collards, kale, Chinese cabbage, bok choy), artichokes, carrots, napa, okra, onions, celery, parsley, chickpeas, parsnips, chicory, peppers, potatoes, gourds (e.g., zucchini, cucumbers, crookneck squash, calabaza, squash), radishes, dry bulb onions, rutabagas, eggplants (also known as aubergines), burdock, endive, green onions, endive, garlic, spinach, green onions, squash, greens, sugar beets (sugar beet and fodder beet), sweet potatoes, Swiss chard, wasabi, tomatoes, turnips, and spices; fruits and / or vine crops, such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, aronia, almonds, chestnuts, hazelnuts, pecans, pistachios, walnuts, citrus, blueberries, boysenberries, cranberries, currants, goji berries, raspberries, strawberries, blackberries, grapes, avocados, bananas, kiwis, persimmons, pomegranates, pineapples, tropical fruits, pomes, melons, mangoes, papayas, and lychees; field crops, such as clover, alfalfa, milkweed, meadow grass, corn / maize (fodder corn, sweet corn, popcorn), hops, jojoba, peanuts, rice, safflower, small grain crops (barley, oats, rye, wheat, etc.), sorghum, tobacco, cotton, legumes (beans, lentils, peas, soybeans), oil plants (rape, mustard, olives, sunflowers, coconuts, castor oil plants, cocoa beans, groundnuts), Arabidopsis, fiber plants (cotton, flax, jute), lauraceae (cinnamon, camphor), or a plant such as coffee, sugar cane, tea, and natural rubber plants; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants, and trees such as forests (broad-leaved trees and evergreens, such as conifers), fruit trees, ornamental trees, and nut-bearing trees, and shrubs and other young plants.

[0231] Beneficial effects of the invention

[0232] The present application obtains a variety of adenosine deaminases by screening, fuses them with DNA binding proteins, and can be used for single base editing of target nucleic acids, having a wide application prospect.

[0233] Embodiments of the present application will be described in detail below with reference to the attached drawings and examples, but those skilled in the art will understand that the following drawings and examples are only used to illustrate the present application, and are not a limitation on the scope of the present application. According to the following detailed description of the drawings and preferred embodiments, various objects and advantages of the present application will become apparent to those skilled in the art.

[0234] The sequence information related to the present application is as follows:

[0235] BRIEF DESCRIPTION OF DRAWINGS

[0236] Figure 1 Structure diagram of ABE single base editing tool.

[0237] Figure 2 Comparison chart of editing efficiency of ABE single base editing tool composed of different adenosine deaminases and dCas12i3.

[0238] Figure 3 Structure diagram of ABE single base editing tool.

[0239] Figure 4 Verification results of base editing efficiency of ABE vector composed of different adenosine deaminases GZ-46 and dCas12i3.

[0240] Figure 5 Verification results of base editing efficiency of ABE vector composed of different adenosine deaminases with combination mutations and dCas12i3.

[0241] Figure 6 Verification results of base editing efficiency of ABE vector composed of different adenosine deaminases with combination mutations and dCas12i3.

[0242] Figure 7 Verification results of base editing efficiency of ABE vector composed of different adenosine deaminases and dCas12i3. DETAILED DESCRIPTION

[0243] The following examples merely illustrate the application and are not intended to limit the application in any way. The experiments and methods described in the examples were performed essentially as described in the art and as described in various references unless otherwise indicated. For example, the general techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, among others, used in the present application can be found in Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd ed. (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F. M. Ausubel et al. eds., (1987)); the series METHODS IN ENZYMOLOGY (Academic Press, Inc.): PCR 2: A PRACTICAL APPROACH (M. J. MacPherson, B. D. Hames, and G. R. Taylor eds. (1995), Harlow and Lane, ANTIBODIES, A LABORATORY MANUAL, (1988), and ANIMAL CELL CULTURE (R. I. Freshney, ed. (1987)).

[0244] In addition, where particular conditions are not specified in the examples, those conditions were performed under routine conditions or as suggested by the manufacturer. Where the manufacturer of reagents or instruments is not indicated, it is intended that the reagents or instruments used were of a routine variety available from a commercial vendor. Those skilled in the art will recognize that the examples describe the present application in terms of preferred embodiments and are not intended to limit the scope of the application as claimed. All publications and other references mentioned herein are incorporated by reference in their entirety.

[0245] Example 1. Editing activity of single base editing tools constructed from different deaminases

[0246] The applicant constructed ABE single-base editing tools by using dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein with various adenosine deaminases (GZ-71, GZ-68, GZ-65, GZ-64, GZ-46, GZ-74, GZ-67, GZ-13, GZ-76) with inactivated nuclease activity. Among them, dCas12i3 is a Cas12i3 (Cas12i3 is Cas12f.4 in CN111757889B, and the amino acid sequence of the wild-type Cas12i3 is shown in SEQ ID No. 1) with an E844A mutation, i.e., the amino acid sequence of dCas12i3 has an E mutation to A at position 844 compared with SEQ ID No. 1; the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is a Cas12i3 protein with E844A, S7R, D233R, D267R, N369R and S433R mutations, and the amino acid sequence of the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein has an E mutation to A at position 844, an S mutation to R at position 7, a D mutation to R at position 233, a D mutation to R at position 267, an N mutation to R at position 369, and an S mutation to R at position 433 compared with SEQ ID No. 1. In other embodiments, the Cas12i3 protein can also have a D mutation to A at position 619 of the sequence shown in SEQ ID No. 1 to obtain a Cas mutant protein with inactivated nuclease activity (i.e., dCas12i3); in other embodiments, the Cas12i3 protein can also have a D mutation to A at position 619 of the sequence shown in SEQ ID No. 1 to obtain a Cas mutant protein with inactivated nuclease activity (i.e., dCas12i3); in other embodiments, the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein can be replaced by other Cas mutant proteins with inactivated nuclease activity, such as dCas9, nCas9.

[0247] The amino acid sequence of adenosine deaminase TadA8e is shown in SEQ ID No. 2; the amino acid sequence of adenosine deaminase GZ-71 is shown in SEQ ID No. 3, and the nucleotide sequence is shown in SEQ ID No. 12; the amino acid sequence of adenosine deaminase GZ-68 is shown in SEQ ID No. 4, and the nucleotide sequence is shown in SEQ ID No. 13; the amino acid sequence of adenosine deaminase GZ-65 is shown in SEQ ID No. 5, and the nucleotide sequence is shown in SEQ ID No. 14; the amino acid sequence of adenosine deaminase GZ-64 is shown in SEQ ID No. 6, and the nucleotide sequence is shown in SEQ ID No. 15; the amino acid sequence of adenosine deaminase GZ-46 is shown in SEQ ID No. 7, and the nucleotide sequence is shown in SEQ ID No. 16; the amino acid sequence of adenosine deaminase GZ-74 is shown in SEQ ID No. 8, and the nucleotide sequence is shown in SEQ ID No. 17; the amino acid sequence of adenosine deaminase GZ-67 is shown in SEQ ID No. 9, and the nucleotide sequence is shown in SEQ ID No. 14. The amino acid sequence of adenosine deaminase GZ-13 is shown in SEQ ID No. 10, and the nucleotide sequence is shown in SEQ ID No. 19; the amino acid sequence of adenosine deaminase GZ-76 is shown in SEQ ID No. 11, and the nucleotide sequence is shown in SEQ ID No. 20.

[0248] A schematic diagram of the ABE single-base editing tool constructed using dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein and adenosine deaminase is shown below. Figure 1 As shown, adenosine deaminase ( Figure 1 TadA in dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein ( Figure 1 The N-terminus of dCas12i3 is linked via an XTEN linker, the amino acid sequence of which is SGGSSGGSSGSETPGTSESATPESSGGSSGGS; the other end of the adenosine deaminase and dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is also linked with an NLS, the amino acid sequence of the N-terminus of the adenosine deaminase is MKRTADGSEFESPKKKRKV, and the amino acid sequence of the C-terminus of the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is KRPAATKKAGQAKKKK. EGFP is a tag designed to screen for positive cells. Figure 1This is merely an example of the linkage between adenosine deaminase and the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein. To verify the editing activity of single-base editing tools constructed from different adenosine deaminases, the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein can be constructed with different adenosine deaminases using the same linkage method, i.e., [linking the linked protein]. Figure 1 The adenosine deaminase TadA in the above-mentioned components can be replaced with other adenosine deaminases. In other embodiments, those skilled in the art can adjust the position or connection order of the above-mentioned components.

[0249] To verify the activity of the ABE single-base editing tool, referring to experimental methods well-known to those skilled in the art, the target DNA sequence was set as the tga terminator region between the EGFP gene and the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein, with PAM of TTG. EGFP cannot emit light normally under these conditions; only when the ABE reporter system, through single-base editing, transforms the tga terminator into cga(R), can EGFP emit green fluorescence. The editing efficiency of the ABE base editing tool can be evaluated by detecting the proportion of green fluorescence produced. The constructed expression vector was transfected into the 293T cell line, and after 48 hours, the EGFP proportion was detected by flow cytometry. The single-base editing efficiency was calculated as the number of EGFP-positive cells divided by the total number of cells.

[0250] The editing activity of ABE single-base editing tools constructed from dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein and different adenosine deaminases is as follows: Figure 2 As shown: Using adenosine deaminase GZ-71 ( Figure 2 The ABE single-base editing tool constructed using GZ0109598-71 (from GZ0109598-71) achieved an editing efficiency of 51.90%; the adenosine deaminase GZ-68 (from GZ0109598-71) was used. Figure 2 The ABE single-base editing tool constructed using GZ0109598-68 (from GZ0109598-68) achieved an editing efficiency of 35.44%; the adenosine deaminase GZ-65 (from GZ0109598-68) was used. Figure 2 The ABE single-base editing tool constructed using GZ0109598-65 (from GZ0109598-65) achieved an editing efficiency of 27.45%; the adenosine deaminase GZ-64 (from GZ0109598-65) was used. Figure 2 The ABE single-base editing tool constructed using GZ0109598-64 (from GZ0109598-64) had an editing efficiency of 21.60%; the adenosine deaminase GZ-46 (from GZ0109598-64) was used. Figure 2 The ABE single-base editing tool constructed using GZ0109598-46 (from GZ0109598-46) had an editing efficiency of 16.62%; the adenosine deaminase GZ-74 (from GZ0109598-46) was used. Figure 2The editing efficiency of the ABE single base editing tool constructed by GZ0109598-74 in the above is 16.53%; the editing efficiency of the ABE single base editing tool constructed by GZ-67 (GZ0109598-67) in the above is 15.74%; the editing efficiency of the ABE single base editing tool constructed by GZ-13 (GZ0112253-13) in the above is 3.77%; the editing efficiency of the ABE single base editing tool constructed by GZ-76 (GZ0112253-76) in the above is 3.23%; Figure 2 The editing efficiency of the ABE single base editing tool constructed by GZ0109598-74 in the above is 16.53%; the editing efficiency of the ABE single base editing tool constructed by GZ-67 (GZ0109598-67) in the above is 15.74%; the editing efficiency of the ABE single base editing tool constructed by GZ-13 (GZ0112253-13) in the above is 3.77%; the editing efficiency of the ABE single base editing tool constructed by GZ-76 (GZ0112253-76) in the above is 3.23%; Figure 2 The editing efficiency of the ABE single base editing tool constructed by GZ0109598-74 in the above is 16.53%; the editing efficiency of the ABE single base editing tool constructed by GZ-67 (GZ0109598-67) in the above is 15.74%; the editing efficiency of the ABE single base editing tool constructed by GZ-13 (GZ0112253-13) in the above is 3.77%; the editing efficiency of the ABE single base editing tool constructed by GZ-76 (GZ0112253-76) in the above is 3.23%; Figure 2 The editing efficiency of the ABE single base editing tool constructed by GZ0109598-74 in the above is 16.53%; the editing efficiency of the ABE single base editing tool constructed by GZ-67 (GZ0109598-67) in the above is 15.74%; the editing efficiency of the ABE single base editing tool constructed by GZ-13 (GZ0112253-13) in the above is 3.77%; the editing efficiency of the ABE single base editing tool constructed by GZ-76 (GZ0112253-76) in the above is 3.23%; Figure 2 GFP of the above is a positive control group, specifically, the tga terminator region is not contained in the fluorescence reporter system, that is, the single base editing is not required, and the EGFP can normally emit light; Figure 2 NC of the above is a negative control group, specifically, the ABE base editing tool without any adenosine deaminase is used for the experiment, and since there is no adenosine deaminase, the base editing cannot be performed, and the fluorescence cannot be emitted.

[0251] It can be known from the above results that the single base editing tools constructed by the above various adenosine deaminases can all achieve single base editing.

[0252] Example 2, Screening of mutated adenosine deaminases GZ-46 and verification of editing activity

[0253] The structure of the adenosine deaminase GZ-46 (the amino acid sequence is shown in SEQ ID No. 7, and the nucleotide sequence is shown in SEQ ID No. 16) in Example 1 is predicted, the key amino acid sites that may affect the biological function thereof are predicted by bioinformatics, the related sites are selected to establish a mutant library, and the screening is performed by using a fluorescence reporter system. In this embodiment, the amino acid site mutation is performed on the 110th, 153rd, 161st and 167th positions from the N terminus of SEQ ID No. 7.

[0254] The variants of adenosine deaminase GZ-46 were generated by PCR-based site-directed mutagenesis. Specifically, the DNA sequence of the adenosine deaminase was designed to be divided into two parts centered on the mutation site, primers were designed for the mutation site, each two pairs of primers corresponded to one amino acid mutation, two pairs of primers were designed to amplify the two parts of the DNA sequence, and the desired mutation sequence was introduced into the primers, and finally the two fragments were loaded into the pcDNA3.3-eGFP vector by Gibson cloning. The combination of the mutants was achieved by splitting the DNA of the adenosine deaminase into multiple segments, using PCR and Gibson clone for construction. Fragment amplification kit: TransStart FastPfu DNA Polymerase (containing 2.5 mM dNTPs), and the specific experimental procedures are described in the instruction manual. Gel recovery kit: FastPure® Gel DNA Extraction Mini Kit, and the specific experimental procedures are described in the instruction manual. Kit for vector construction: pEASY-Basic Seamless Cloning and Assembly Kit (CU201-03), and the specific experimental procedures are described in the instruction manual.

[0255] Based on the above-mentioned amino acid mutation sites, the wild-type protein of adenosine deaminase GZ-46, and adenosine deaminases with single-site mutations at positions 110, 153, 161, and 167 (named according to the mutation type) were obtained: V110K, V110H, Y153F, N161R, and E167R. The mutant adenosine deaminase V110K has a K mutation at position 110 compared to SEQ ID No. 7. The mutant adenosine deaminase V110H has a H mutation at position 110 compared to SEQ ID No. 7. The mutant adenosine deaminase Y153F has a F mutation at position 153 compared to SEQ ID No. 7. The mutant adenosine deaminase N161R has a R mutation at position 161 compared to SEQ ID No. 7. The mutant adenosine deaminase E167R has a R mutation at position 167 compared to SEQ ID No. 7.

[0256] In addition, adenosine deaminases with combined mutations at positions 110, 153, 161, and 167 (named according to the mutation type) were obtained: V110K-Y153F-N161R-E167R. The mutant adenosine deaminase V110K-Y153F-N161R-E167R has a K mutation at position 110, a F mutation at position 153, a R mutation at position 161, and a R mutation at position 167 compared to SEQ ID No. 7.

[0257] The applicant constructed an ABE single-base editing tool using the nuclease-inactivated dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein from Example 1, along with the adenosine deaminase obtained above. A schematic diagram of the ABE single-base editing tool constructed using the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein and adenosine deaminase is shown below. Figure 3 As shown, adenosine deaminase in dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein ( Figure 3 The N-terminus of dCas12i3 is linked via an XTEN linker, the amino acid sequence of which is SGGSSGGSSGSETPGTSESATPESSGGSSGGS; the other end of the adenosine deaminase and dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is also linked with an NLS, the amino acid sequence of the N-terminus of the adenosine deaminase is MKRTADGSEFESPKKKRKV, and the amino acid sequence of the C-terminus of the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is KRPAATKKAGQAKKKK. EGFP is a tag designed to screen for positive cells. Figure 3 This is merely an example of the linkage between adenosine deaminase and the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein. To verify the editing activity of single-base editing tools constructed from different adenosine deaminases, the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein can be constructed with different adenosine deaminases using the same linkage method, i.e., [linking the linked protein]. Figure 3 The adenosine deaminase in the above-mentioned components can be replaced with other adenosine deaminases. In other embodiments, those skilled in the art can adjust the position or connection order of the above-mentioned components.

[0258] The editing activity of the ABE single-base editing tool was verified by referring to the method for verifying the activity of the ABE single-base editing tool in Example 1.

[0259] The editing activity of ABE single-base editing tools constructed from dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein and different adenosine deaminases is as follows: Figure 4 As shown. Figure 4 The NC was the negative control group. Specifically, the experiment was conducted using the ABE base editing tool, which does not contain any adenosine deaminase. Since there is no adenosine deaminase, base editing cannot be performed, and therefore no fluorescence can be emitted.

[0260] In Figure 4 the editing efficiency of the ABE single base editing tool constructed by using wild-type adenosine deaminase GZ-46 (46 RAHSY) in Figure 4 was 14.4%; the editing efficiency of the ABE single base editing tool constructed by using the mutated adenosine deaminase V110K was 44.1%; the editing efficiency of the ABE single base editing tool constructed by using the mutated adenosine deaminase V110H was 31.5%; the editing efficiency of the ABE single base editing tool constructed by using the mutated adenosine deaminase Y153F was 20.2%; the editing efficiency of the ABE single base editing tool constructed by using the mutated adenosine deaminase N161R was 19.8%; the editing efficiency of the ABE single base editing tool constructed by using the mutated adenosine deaminase E167R was 24.2%; and the editing efficiency of the ABE single base editing tool constructed by using the mutated adenosine deaminase V110K-Y153F-N161R-E167R (V110K Y153F N161R E167R) in Figure 4 was 61.4%.

[0261] From the above results, it can be seen that after the single-site mutation and combined-site mutation of the 110th, 153rd, 161st or 167th amino acid of adenosine deaminase GZ-46, the editing efficiency is significantly improved.

[0262] Example 3, Further obtaining of mutated adenosine deaminases with improved activity on the basis of mutated adenosine deaminase GZ-46 (V110K-Y153F-N161R-E167R) Figure 5

[0263] Based on the mutated adenosine deaminase GZ-46 (V110K-Y153F-N161R-E167R) obtained in Example 2, this example refers to it as KFRR (compared with SEQ ID No. 7, the 110th amino acid is mutated to K, the 153rd amino acid is mutated to F, the 161st amino acid is mutated to R, and the 167th amino acid is mutated to R).

[0264] In this example, the following sites are mutated based on adenosine deaminase KFRR: 55th, 113th, 123rd, 126th, 127th, 156th, 159th, 160th amino acids, and the mutation method is referred to Example 2: KFRR (S113N), KFRR (S113R), KFRR (N126Q), KFRR (Y127H), KFRR (Y127L), KFRR (Y127M), KFRR (V159A), KFRR (V159H), KFRR (V159K), KFRR (V159R), TKNDQHFNHQRR, TKNDQHFNRQRR.

[0265] KFRR (S113N) is mutated at amino acid 110 to K, amino acid 153 to F, amino acid 161 to R, amino acid 167 to R and amino acid 113 to N compared with SEQ ID No. 7;

[0266] KFRR (S113R) is mutated at amino acid 110 to K, amino acid 153 to F, amino acid 161 to R, amino acid 167 to R and amino acid 113 to R compared with SEQ ID No. 7;

[0267] KFRR (N126Q) is mutated at amino acid 110 to K, amino acid 153 to F, amino acid 161 to R, amino acid 167 to R and amino acid 126 to Q compared with SEQ ID No. 7;

[0268] KFRR (Y127H) is mutated at amino acid 110 to K, amino acid 153 to F, amino acid 161 to R, amino acid 167 to R and amino acid 127 to H compared with SEQ ID No. 7;

[0269] KFRR (Y127L) is mutated at amino acid 110 to K, amino acid 153 to F, amino acid 161 to R, amino acid 167 to R and amino acid 127 to L compared with SEQ ID No. 7;

[0270] KFRR (Y127M) is mutated at amino acid 110 to K, amino acid 153 to F, amino acid 161 to R, amino acid 167 to R and amino acid 127 to M compared with SEQ ID No. 7;

[0271] KFRR (V159A) is mutated at amino acid 110 to K, amino acid 153 to F, amino acid 161 to R, amino acid 167 to R and amino acid 159 to A compared with SEQ ID No. 7;

[0272] KFRR (V159H) is mutated at amino acid 110 to K, amino acid 153 to F, amino acid 161 to R, amino acid 167 to R and amino acid 159 to H compared with SEQ ID No. 7;

[0273] Mutated adenosine deaminase KFRR (V159K) has mutations of K at the 110th amino acid, F at the 153rd amino acid, R at the 161st amino acid, R at the 167th amino acid and K at the 159th amino acid compared with SEQ ID No. 7;

[0274] Mutated adenosine deaminase KFRR (V159R) has mutations of K at the 110th amino acid, F at the 153rd amino acid, R at the 161st amino acid, R at the 167th amino acid and R at the 159th amino acid compared with SEQ ID No. 7;

[0275] Mutated adenosine deaminase TKNDQHFNHQRR has mutations of T at the 55th amino acid, K at the 110th amino acid, N at the 113th amino acid, D at the 123rd amino acid, Q at the 126th amino acid, H at the 127th amino acid, F at the 153rd amino acid, N at the 156th amino acid, H at the 159th amino acid, Q at the 160th amino acid, R at the 161st amino acid and R at the 167th amino acid compared with SEQ ID No. 7;

[0276] Mutated adenosine deaminase TKNDQHFNHQRR has mutations of T at the 55th amino acid, K at the 110th amino acid, N at the 113th amino acid, D at the 123rd amino acid, Q at the 126th amino acid, H at the 127th amino acid, F at the 153rd amino acid, N at the 156th amino acid, H at the 159th amino acid, Q at the 160th amino acid, R at the 161st amino acid and R at the 167th amino acid compared with SEQ ID No. 7.

[0277] The dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein in Example 1 was used to construct an ABE single-base editing tool with the adenosine deaminase obtained above, and the ABE single-base editing tool was constructed in the same way as in Example 2. The editing activity of the single-base editing tool was verified by referring to the method for verifying the activity of the ABE single-base editing tool in Example 1.

[0278] The editing activities of the ABE single-base editing tools constructed by the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein and different adenosine deaminases are shown in Figure 6 and Figure 5 Figure 6 and Figure 5 ​NC is a negative control group, specifically, an experiment is performed using an ABE base editing tool without any adenosine deaminase, and since there is no adenosine deaminase, base editing cannot be performed, and thus fluorescence cannot be emitted.

[0279] In the Figure 5 , the editing efficiency of the ABE single base editing tool constructed by using the wild type adenosine deaminase GZ-46 (46 RAHSY in Figure 6 ) is 13.09%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR is 57.70%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (S113N) is 68.02%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (S113R) is 61.95%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (N126Q) is 62.62%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (Y127H) is 63.39%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (Y127L) is 60.63%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (Y127M) is 67.72%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (V159A) is 67.00%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (V159H) is 68.50%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (V159K) is 72.49%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase KFRR (V159R) is 73.57%.

[0280] In the Figure 6 , the editing efficiency of the ABE single base editing tool constructed by using the wild type adenosine deaminase GZ-46 (46 RAHSY in Figure 6 ) is 14.59%; the editing efficiency of the ABE single base editing tool constructed by using the adenosine deaminase TadA8e.1 (Tad8e.1 in Example 4, Screening of mutated adenosine deaminases and verification of editing activity ) is 54.22%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase TKNDQHFNHQRR is 60.69%; the editing efficiency of the ABE single base editing tool constructed by using the mutant adenosine deaminase TKNDQHFNRQRR is 56.11%. Among them, the amino acid sequence of TadA8e.1 is compared with SEQ ID No. 2, the 106th amino acid V is mutated to W.

[0281] From the above results, it can be seen that, on the basis of the combination site mutation of the 110th, 153rd, 161st and 167th amino acids of adenosine deaminase GZ-46 (V110K-Y153F-N161R-E167R), the editing efficiency is significantly improved after unit point mutation and combination site mutation of the 55th, 113th, 123rd, 126th, 127th, 156th, 159th or 160th amino acid of adenosine deaminase GZ-46, and even higher than the editing efficiency of adenosine deaminase TadA8e.1.

[0282] That is, after unit point mutation and combination site mutation of the 55th, 110th, 113th, 123rd, 126th, 127th, 153rd, 156th, 159th, 160th, 161st or 167th amino acid of adenosine deaminase GZ-46, the editing efficiency is significantly improved.

[0283] Figure 7

[0284] Adenosine deaminase TadA8e (the amino acid sequence is shown in SEQ ID No. 2) was mutated by the method of adenosine deaminase mutation in Example 2, and adenosine deaminase TadA8e of the 106th (homologous site with the 110th amino acid of adenosine deaminase GZ-46) and the 157th (homologous site with the 161st amino acid of adenosine deaminase GZ-46) amino acid unit point mutation of adenosine deaminase TadA8e (named by mutation type) was obtained: TadA8e (V106W), TadA8e (V106K), TadA8e (V106H), TadA8e (N157R); the mutated adenosine deaminase TadA8e (V106W) has the 106th amino acid mutated to W compared with SEQ ID No. 2; the mutated adenosine deaminase TadA8e (V106K) has the 106th amino acid mutated to K compared with SEQ ID No. 2; the mutated adenosine deaminase TadA8e (V106H) has the 106th amino acid mutated to H compared with SEQ ID No. 2; the mutated adenosine deaminase TadA8e (N157R) has the 157th amino acid mutated to R compared with SEQ ID No. 2.

[0285] The dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein with nuclease activity inactivated in Example 1 was constructed with the adenosine deaminase obtained above to form an ABE single-base editing tool. The schematic diagram and construction method of the ABE single-base editing tool are the same as in Example 2. Referring to the method for verifying the activity of the ABE single-base editing tool in Example 1, the editing activity of the single-base editing tool was verified.

[0286] The editing activity of the ABE single-base editing tool constructed by the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein and different adenosine deaminases TadA8e is shown in Figure 7 Figure 7 NC is a negative control group, specifically, an experiment is performed using an ABE base editing tool without any adenosine deaminase. Since there is no adenosine deaminase, base editing cannot be performed, and thus fluorescence cannot be emitted.

[0287] In Figure 7 , the editing efficiency of the ABE single-base editing tool constructed by the adenosine deaminase TadA8e (Tad8e in ​ ) is 48.86%; the editing efficiency of the ABE single-base editing tool constructed by the mutated adenosine deaminase TadA8e (V106W) is 50.16%; the editing efficiency of the ABE single-base editing tool constructed by the mutated adenosine deaminase TadA8e (V106K) is 56.32%; the editing efficiency of the ABE single-base editing tool constructed by the mutated adenosine deaminase TadA8e (V106H) is 63.08%; and the editing efficiency of the ABE single-base editing tool constructed by the mutated adenosine deaminase TadA8e (N157R) is 52.26%.

[0288] From the above results, it can be seen that the editing efficiency is significantly improved after the single-point mutation and combined-site mutation of the 106th or 157th amino acid of the adenosine deaminase TadA8e.

[0289] Since the 106th amino acid of the adenosine deaminase TadA8e is a homologous site of the 110th amino acid of the adenosine deaminase GZ-46, and the 157th amino acid of the adenosine deaminase TadA8e is a homologous site of the 161st amino acid of the adenosine deaminase GZ-46, it is indicated that the homologous site mutation of the adenosine deaminase TadA8e or the adenosine deaminase GZ-46 can improve the editing efficiency or activity of the deaminase.

[0290] ​While the specific embodiments of the application have been described in detail, those skilled in the art will appreciate that various modifications and alterations to the details can be made within the scope of the application as disclosed in the teachings of the present application. The entire disclosure of the application is set out in the accompanying claims and any equivalents thereof.

Claims

1. An adenosine deaminase, characterized in that, The adenosine deaminase is the following I- Any of the adenosine deaminases mentioned above: I. The amino acid sequence of the adenosine deaminase has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of the sequences in SEQ ID No. 3-11, and substantially retains the biological function of the sequence from which it originated; preferably, the adenosine deaminase is derived from Enterobacteriaceae, Escherichia coli, Citrobacter, Rahnella aceris, Shigella flexneri, or Salmonella enterica subsp. II. The amino acid sequence of the adenosine deaminase, compared with any sequence in SEQ ID No. 3-11, has one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 amino acids) substitutions, deletions or additions, and substantially retains the biological function of its derived sequence; preferably, the adenosine deaminase is derived from Enterobacteriaceae, Escherichia coli, Citrobacter, Rahnella aceris, Shigella flexneri or Salmonella enterica subsp. III. The adenosine deaminase comprises any of the amino acid sequences shown in SEQ ID No. 3-11; The adenosine deaminase is a mutant adenosine deaminase, which, compared with the parental adenosine deaminase, has mutations at the following amino acid sites corresponding to the amino acid sequence shown in SEQ ID No. 2: position 106 and / or position 157.

2. The adenosine deaminase according to claim 1, characterized in that, The adenosine deaminase is a mutant adenosine deaminase, and the amino acid sequence of the parental adenosine deaminase has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of the sequences shown in SEQ ID No. 2-11.

3. A fusion protein, characterized in that, The fusion protein comprises a DNA-binding protein and the adenosine deaminase according to any one of claims 1-2; Preferably, the DNA-binding protein is a Cas protein; preferably, the Cas protein is selected from any one of Cas9, Cas9n, dCas9, CasX, CasY, C2cl, C2c2, C2c3, GeoCas9, CjCas9, Casl2a, Casl2b, Casl2g, Casl2h, Casl2i, Cas12j, Casl3b, Casl3c, Casl3d, Casl4, Csn2, xCas9, Cas9-NG, LbCasl2a, enAsCasl2a, Cas9-KKH, cyclic substitution Cas9, Argonaute (Ago) domain, SmacCas9, Spy-macCas9, SpGas9-NRRH, SpaCas9-NRTH, SpaCas9-NRCH, Cas9-NG-CP1041, Cas9-NG-VRQR, dCas12i, nCas12i, and their variants. More preferably, the DNA-binding protein is a nuclease-inactivated Cas protein.

4. An isolated polynucleotide, characterized in that, The polynucleotide encodes the adenosine deaminase according to any one of claims 1-2, or the polynucleotide encodes the fusion protein according to claim 3.

5. A carrier, characterized in that, The carrier includes the polynucleotide of claim 4 and a regulatory element operatively linked thereto.

6. A base editing system, characterized in that, The system includes the fusion protein of claim 3 and at least one gRNA; the DNA-binding protein in the fusion protein is a Cas protein, and the gRNA is capable of binding the Cas protein.

7. A composition, characterized in that, The composition comprises: (i) A protein component selected from the fusion protein of claim 3, wherein the DNA-binding protein in the fusion protein is a Cas protein; (ii) A nucleic acid component, which is gRNA, said gRNA being able to bind said Cas protein.

8. An engineered host cell, characterized in that, The host cell comprises the adenosine deaminase of any one of claims 1-2, or the fusion protein of claim 3, or the polynucleotide of claim 4, or the vector of claim 5, or the base editing system of claim 6, or the composition of claim 7.

9. The use of the adenosine deaminase of any one of claims 1-2, or the fusion protein of claim 3, or the polynucleotide of claim 4, or the vector of claim 5, or the base editing system of claim 6, or the composition of claim 7, or the host cell of claim 8 in gene editing; or, in the preparation of reagents or kits for gene editing; Preferably, the gene editing is a single-base editing of the target gene.

10. A kit for gene editing, characterized in that, The kit comprises the adenosine deaminase of any one of claims 1-2, or the fusion protein of claim 3, or the polynucleotide of claim 4, or the vector of claim 5, or the base editing system of claim 6, or the composition of claim 7, or the host cell of claim 8.

11. A method for editing nucleic acids, the method comprising the step of contacting a target region of the nucleic acid with a fusion protein of claim 3 and a gRNA, wherein the gRNA comprises a sequence capable of binding a Cas protein in the fusion protein of claim 3 and a sequence capable of binding a target region of the nucleic acid; wherein, The target region contains targeted base pairs, and the fusion protein is capable of replacing the base pairs. Preferably, the targeted base pair is replaced by G:C instead of A:T.

Citation Information

Patent Citations

  • Novel CRISPR / Cas12f enzymes and systems

    CN111757889B

  • Adenine base editors and uses thereof

    WO2021158921A2