Mutational base editing tool and application thereof
By mutation of specific amino acid sites on adenosine deaminase and fusing with DNA binding proteins, an efficient single-base editing tool was constructed, solving the compatibility limitations of the adenine base editor and achieving precise modification of specific gene loci.
Patent Information
- Application Number
- CN202510586800.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-04
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-08
AI Technical Summary
The existing adenine base editor is limited by the compatibility of adenosine deaminase domain and DNA binding proteins, making it difficult to achieve efficient single-base editing.
By mutation of adenosine deaminase at specific amino acid sites, a fusion protein with DNA binding protein is constructed, expanding the application range of single-base editing tools and improving editing efficiency.
Accurate modification of specific gene loci without causing DNA double-strand breaks, expanding the application range of single-base editing and improving editing efficiency.
Smart Images

Figure CN120442605A_ABST
Abstract
Description
[0001] This application claims priority to Chinese patent application CN202410713487.1, filed on June 4, 2024. This application incorporates the entire text of the aforementioned Chinese patent application. Technical Field
[0002] The present invention relates to the field of gene editing, in particular to the field of clustered regularly interspaced short palindromic repeats (CRISPR). Specifically, the present invention relates to a base editing tool, in particular, base editing based on a mutant adenosine deaminase. Background Art
[0003] CRISPR / Cas technology is a widely used gene editing technology that uses RNA to guide specific binding to target sequences on the genome and cut DNA to produce double-strand breaks, and uses biological non-homologous end joining or homologous recombination to perform site-directed gene editing.
[0004] The development of single-base gene editing tools can achieve precise modification of specific gene sites without causing double-strand breaks in DNA. Adenine base editor (ABE) based on adenosine deaminase realizes the conversion of A>G (from A to G). At present, the application of adenine base editor is limited by the limited compatibility of adenosine deaminase domain and DNA binding protein. Therefore, the present invention is based on a mutant adenosine deaminase to construct a single-base editing tool, which expands the application range of single-base editing technology and improves the efficiency of single-base editing. Summary of the Invention
[0005] In one aspect, the present invention provides a mutant adenosine deaminase.
[0006] In one embodiment, the mutated adenosine deaminase has a mutation at any one or any several (e.g., any two, any three) amino acid positions corresponding to the amino acid sequence shown in SEQ ID No. 1, compared with the amino acid sequence of the parent adenosine deaminase: position 110, position 153, position 161, and position 167.
[0007] In one embodiment, the mutant adenosine deaminase has mutations at the following amino acid positions corresponding to the amino acid sequence shown in SEQ ID No. 1: position 110, position 153, position 161 and position 167 compared to the amino acid sequence of the parent adenosine deaminase.
[0008] In one embodiment, the amino acid at position 110 is mutated to a non-V amino acid, for example, G, A, L, I, P, F, Y, W, S, T, C, M, N, Q, D, E, K, R, H; preferably, K or H; more preferably, K.
[0009] In one embodiment, the amino acid at position 153 is mutated to an amino acid other than Y, for example, L, A, V, I, P, F, N, W, S, T, C, M, G, Q, D, E, K, R, H; preferably, F.
[0010] In one embodiment, the amino acid at position 161 is mutated to a non-N amino acid, for example, L, A, V, I, P, F, Y, W, S, T, C, M, G, Q, D, E, K, R, H; preferably, R.
[0011] In one embodiment, the amino acid at position 167 is mutated to an amino acid other than E, for example, G, A, L, I, P, F, Y, W, S, T, C, M, N, Q, D, V, K, R, H; preferably, R.
[0012] In one embodiment, the amino acid sequence of the parent adenosine deaminase has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to SEQ ID No. 1.
[0013] In one embodiment, the mutant adenosine deaminase is selected from any one of the following groups I-III:
[0014] I. An adenosine deaminase obtained by generating mutations in the amino acid sequence of SEQ ID No. 1 at any one or more (e.g., any two, any three, or any four) of the following amino acid positions: position 110, position 153, position 161, and position 167;
[0015] II. an adenosine deaminase mutant having the mutation site described in I compared to the mutant adenosine deaminase described in I; and an adenosine deaminase mutant having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to the mutant adenosine deaminase described in I;
[0016] III. Compared with the mutant adenosine deaminase described in I, it has the mutation site described in I; and, compared with the mutant adenosine deaminase described in I, it has a sequence with one or more amino acid substitutions, deletions or additions; the one or more amino acids include 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions.
[0017] In one embodiment, the amino acid sequence of the parent adenosine deaminase is shown as SEQ ID No. 1.
[0018] In the present invention, adenosine deaminase, also known as adenine deaminase, catalyzes the hydrolytic deamination of adenine or adenosine. The adenosine deaminase provided herein (e.g., engineered adenosine deaminase, evolved adenosine deaminase) can be derived from any organism, such as bacteria.
[0019] In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase, or a variant thereof that is mutated but still has adenosine deaminase activity.
[0020] In some embodiments, the adenosine deaminase is from a prokaryotic organism. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis.
[0021] In some embodiments, the adenosine deaminase is GZ0109598-46 deaminase.
[0022] In a preferred embodiment of the present invention, the parent adenosine deaminase is derived from Rahnella aceris.
[0023] In one aspect, the present invention provides a fusion protein comprising a DNA binding protein and the above-mentioned adenosine deaminase.
[0024] In one embodiment, the DNA binding protein has a nucleic acid programmable DNA binding protein (napDNAbp) domain, and the "DNA binding protein" or "DNA binding protein domain" refers to any protein that can locate and bind to a specific target DNA nucleotide sequence (e.g., a genomic locus). In one embodiment, the DNA binding protein is a Cas protein, and the Cas protein is selected from Cas9, Cas9n, dCas9, CasX, CasY, C2cl, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas12j, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, Cas9-NG, LbCas12a, enAsCas12a, Cas9-KKH, circular displacement Cas9, Argonaute (Ago) domain, SmacCas9, Spy-macCas9, SpGas9-NRRH, SpaCas9-NRTH, SpaCas9-NRCH, Cas9-NG-CP1041, Cas9-NG-VRQR, dCas12i, nCas12i and variants thereof.
[0025] In one embodiment, the Cas protein is Cas12i.
[0026] In one embodiment, the Cas protein may be naturally occurring or non-naturally occurring and engineered.
[0027] In one embodiment, the Cas protein is a nuclease-inactivated Cas protein (dCas).
[0028] In one embodiment, the Cas protein is a dCas12i3 protein, and the amino acid sequence of the wild-type Cas12i is shown in SEQ ID No. 3; compared with SEQ ID No. 3, the dCas12i protein has an E mutation at position 844 to an A mutation or a D mutation at position 619 to an A mutation.
[0029] In one embodiment, the Cas protein is a Cas12i mutant protein, and the amino acid sequence of the wild-type Cas12i is shown in SEQ ID No. 3; compared with SEQ ID No. 3, the Cas12i mutant protein is mutated from S at position 7 to R, D at position 233 to R, D at position 267 to R, N at position 369 to R, and S at position 433 to R.
[0030] In one embodiment, the Cas protein is a dCas12i mutant protein, and the amino acid sequence of the wild-type Cas12i is shown in SEQ ID No. 3; compared with SEQ ID No. 3, the dCas12i mutant protein has E at position 844 mutated to A, S at position 7 mutated to R, D at position 233 mutated to R, D at position 267 mutated to R, N at position 369 mutated to R, and S at position 433 mutated to R.
[0031] In one embodiment, the DNA binding protein in the fusion protein is fused to the N-terminus of adenosine deaminase. In other embodiments, the DNA binding protein in the fusion protein is fused to the C-terminus of deaminase.
[0032] In one embodiment, the DNA binding protein and adenosine deaminase in the fusion protein are connected via a linker.
[0033] In one embodiment, the linker is an XTEN linker.
[0034] In one embodiment, the fusion protein of the present invention further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In other embodiments, both the N-terminus and the C-terminus of the fusion protein are connected to the NLS.
[0035] In some embodiments, the NLS is fused to the N-terminus of the Cas protein. In some embodiments, the NLS is fused to the C-terminus of the Cas protein. In some embodiments, the NLS is fused to the N-terminus of the deaminase. In some embodiments, the NLS is fused to the C-terminus of the deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker.
[0036] Those skilled in the art will appreciate that protein structure can be altered without adversely affecting its activity and functionality. For example, one or more conservative amino acid substitutions can be introduced into a protein's amino acid sequence without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples and implementations of conservative amino acid substitutions will be apparent to those skilled in the art. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the substituted residue, i.e., a non-polar amino acid residue can be substituted for another non-polar amino acid residue, a polar uncharged amino acid residue can be substituted for another polar uncharged amino acid residue, a basic amino acid residue can be substituted for another basic amino acid residue, and an acidic amino acid residue can be substituted for another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, where one amino acid is replaced with another amino acid belonging to the same group, fall within the scope of the present invention, as long as the substitution does not inactivate the biological activity of the protein. Therefore, the proteins of the present invention may contain one or more conservative substitutions in their amino acid sequences, preferably generated by substitutions according to Table 1. Furthermore, the present invention also encompasses proteins containing one or more other non-conservative substitutions, as long as such non-conservative substitutions do not significantly affect the desired function and biological activity of the proteins of the present invention.
[0037] Conservative amino acid replacement can be carried out at the non-essential amino acid residue of one or more predictions.A "non-essential" amino acid residue is an amino acid residue that can change (deletion, substitution or replacement) and does not change biological activity, while an "essential" amino acid residue is required for biological activity.A "conservative amino acid replacement" is a replacement in which an amino acid residue is replaced by an amino acid residue with a similar side chain.Amino acid replacement can be carried out in the non-conserved region of above-mentioned mutans or fusion protein. Generally speaking, this type of replacement is not carried out to a conserved amino acid residue, or is not carried out to an amino acid residue positioned within a conserved motif, whereby this type of residue is required for protein activity.However, it will be appreciated by those skilled in the art that functional variants can have less conservative or non-conservative changes in a conserved region.
[0038] Table 1
[0039] Initial residue Representative replacement Preferred substitutions Ala(A) Val; Leu; Ile Val Arg(R) Lys; Gln; Asn Lys Asn(N) Gln; His; Lys; Arg Gln Asp(D) Glu Glu Cys(C) Ser Ser Gln(Q) Asn Asn Glu(E) Asp Asp Gly(G) Pro; Ala Ala His(H) Asn; Gln; Lys; Arg Arg Ile(I) Leu; Val; Met; Ala; Phe Leu Leu(L) Ile; Val; Met; Ala; Phe Ile Lys(K) Arg; Gln; Asn Arg Met(M) Leu; Phe; Ile Leu Phe(F) Leu; Val; Ile; Ala; Tyr Leu Pro(P) Ala Ala Ser(S) Thr Thr Thr(T) Ser Ser Trp(W) Tyr; Phe Tyr Tyr(Y) Trp; Phe; Thr; Ser Phe Val(V) Ile;Leu;Met;Phe;Ala Leu
[0040] It is well known in the art that one or more amino acid residues can be altered (replaced, deleted, truncated or inserted) from the N- and / or C-terminus of a protein while still retaining its functional activity. Thus, proteins wherein one or more amino acid residues are altered from the N- and / or C-terminus of such mutant proteins while retaining their desired functional activity are also within the scope of the present invention. These alterations may include those introduced by modern molecular methods such as PCR, which involves PCR amplification of a protein coding sequence by altering or extending the amino acid coding sequence by including the amino acid coding sequence in the oligonucleotides used in the PCR amplification.
[0041] It will be appreciated that proteins can be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such manipulations are generally known in the art. For example, amino acid sequence variants of the above-described proteins can be prepared by mutations in the DNA. Other forms of mutagenesis and / or directed evolution can also be accomplished, for example, using known mutagenesis, recombination, and / or shuffling methods, in combination with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.
[0042] Those skilled in the art will appreciate that these minor amino acid changes in the proteins of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not close to the catalytic domain, active site, or other functional domains, lesser effects can be expected.
[0043] Those skilled in the art can identify the essential amino acids of the mutant proteins of the present invention according to methods known in the art, such as site-directed mutagenesis or protein evolution or analysis by bioinformatics systems. The catalytic domain, active site or other functional domains of the protein can also be determined by physical analysis of the structure, such as by techniques such as nuclear magnetic resonance, crystallography, electron diffraction or photoaffinity labeling, combined with mutations in putative key amino acids.
[0044] In the present invention, amino acid residues can be represented by single letters or three letters, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamine (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine (Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), arginine (Arg, R).
[0045] The term "AxxB" indicates that the amino acid A at position xx is changed to amino acid B. For example, V110K indicates that the V at position 110 is mutated to a K. When multiple amino acid positions are mutated simultaneously, the expression can be expressed in the form of V110K-Y153F-N161R-E167R, V110K Y153FN161R E167R, etc. For example, V110K-Y153F-N161R-E167R indicates that the V at position 110 is mutated to a K, the Y at position 153 is mutated to an F, the N at position 161 is mutated to an R, and the E at position 167 is mutated to an R.
[0046] Specific amino acid positions (numbers) within the proteins of the present invention are determined by aligning the amino acid sequence of the target protein with the target sequence (e.g., SEQ ID No. 1) using standard sequence alignment tools, such as the Smith-Waterman algorithm or the CLUSTALW2 algorithm, wherein the sequences are considered aligned when the alignment score is the highest. The alignment score can be calculated according to the method described in Wilbur, WJ and Lipman, DJ (1983) Rapid similarity searches of nucleic acid and protein databanks. Proc. Natl. Acad. Sci. USA, 80:726-730. Preferably, the default parameters are used in the ClustalW2 (1.82) algorithm: protein gap open penalty = 10.0; protein gap extension penalty = 0.2; protein matrix = Gonnet; protein / DNA end gap = -1; protein / DNAGAPDIST = 4. The positions of specific amino acids within the protein of the present invention are preferably determined by comparing the amino acid sequence of the protein with SEQ ID No. 1 using the AlignX program (part of the vectorNTI suite) with default parameters suitable for multiple alignment (gap opening penalty: 10 log, gap extension penalty: 0.05). Those skilled in the art can use commonly used software in the art, such as Clustal Omega, to compare and align the amino acid sequence of any parent adenosine deaminase with SEQ ID No. 1 for sequence identity, thereby determining the amino acid sites in the parent adenosine deaminase that correspond to the amino acid sites defined in this application based on SEQ ID No. 1.
[0047] The fusion protein of the present invention is not limited by the method of production. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.
[0048] The present invention also provides a base editing tool comprising the above-mentioned fusion protein, for example, a single base editing tool.
[0049] Nucleic Acids
[0050] In another aspect, the present invention provides an isolated polynucleotide comprising:
[0051] (a) a polynucleotide sequence encoding a mutant adenosine deaminase or fusion protein of the present invention;
[0052] or,
[0053] (b) A polynucleotide complementary to the polynucleotide described in (a).
[0054] In one embodiment, the nucleotide sequence is codon optimized for expression in prokaryotes. In one embodiment, the nucleotide sequence is codon optimized for expression in eukaryotic cells.
[0055] In one embodiment, the cell is an animal cell, eg, a mammalian cell.
[0056] In one embodiment, the cell is a human cell.
[0057] In one embodiment, the cell is a plant cell, such as a cell from a cultivated plant (such as cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.
[0058] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.
[0059] Guide RNA (gRNA)
[0060] On the other hand, the present invention provides a gRNA, which includes a protein binding sequence and a targeting sequence that targets a target nucleic acid.
[0061] The protein binding sequence of the gRNA can interact with the DNA binding protein (Cas protein) in the fusion protein of the present invention, thereby forming a complex between the DNA binding protein (Cas protein) and the gRNA.
[0062] Designing gRNAs that interact with the Cas proteins of the present invention is conventional technical knowledge in the art. For example, for Cas9 protein, Cas proteins of the Cas12 family (including but not limited to Cas12a, Cas12b, Cas12i, etc.), designing gRNAs that interact with them does not require creative work.
[0063] The targeting sequence of the targeting nucleic acid of the present invention comprises a nucleotide sequence that is complementary to the sequence in the target nucleic acid. In other words, the targeting sequence of the targeting nucleic acid of the present invention or the targeting segment of the targeting nucleic acid interacts with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the targeting sequence of the targeting nucleic acid or the targeting segment of the targeting nucleic acid can be changed, or can be modified to hybridize any desired sequence in the target nucleic acid. The nucleic acid is selected from DNA or RNA.
[0064] The percent complementarity between the targeting sequence of a targeting nucleic acid or the targeting segment of a targeting nucleic acid and the target sequence of a target nucleic acid can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).
[0065] The "protein binding sequence" of the gRNA of the present invention can interact with the CRISPR protein (or, Cas protein). The gRNA of the present invention guides its interacting DNA binding protein to a specific nucleotide sequence within the target nucleic acid through the action of the targeting sequence of the target nucleic acid.
[0066] The gRNA of the present invention is capable of forming a complex with the DNA binding protein (Cas protein).
[0067] carrier
[0068] The present invention also provides a vector comprising the adenosine deaminase, fusion protein, isolated nucleic acid molecule or polynucleotide as described above; preferably, it further comprises a regulatory element operably linked thereto.
[0069] In one embodiment, the regulatory element is selected from one or more of the following groups: enhancer, transposon, promoter, terminator, leader sequence, polyadenylation sequence, marker gene.
[0070] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.
[0071] In some embodiments, the vector included in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a phage, etc., which are well known to those skilled in the art.
[0072] Base editing system
[0073] The present invention provides an engineered non-naturally occurring base editing system, which includes the above-mentioned fusion protein or a nucleic acid sequence encoding the fusion protein and a nucleic acid encoding one or more of the above-mentioned guide RNAs, wherein the DNA binding protein in the fusion protein is a Cas protein, and the gRNA is capable of binding to the Cas protein.
[0074] In one embodiment, the nucleic acid sequence encoding the fusion protein and the nucleic acid encoding one or more guide RNAs are artificially synthesized.
[0075] In one embodiment, the nucleic acid sequence encoding the fusion protein and the nucleic acid encoding one or more guide RNAs do not naturally occur together.
[0076] The one or more guide RNAs target one or more target sequences in the cell. The one or more target sequences hybridize to the genomic loci of DNA molecules encoding one or more gene products, and guide the fusion protein to the genomic loci of the DNA molecules encoding the one or more gene products. After the fusion protein reaches the target sequence position, it modifies or edits the target sequence, thereby changing or modifying the expression of the one or more gene products.
[0077] The cells of the present invention include one or more of animals, plants, or microorganisms.
[0078] In some embodiments, the fusion protein is codon-optimized for expression in a cell.
[0079] The present invention also provides an engineered non-naturally occurring vector system, which may include one or more vectors, wherein the one or more vectors include:
[0080] a) a first regulatory element, which is operably linked to the gRNA,
[0081] b) a second regulatory element, the second regulatory element being operably linked to the fusion protein;
[0082] Components (a) and (b) are located on the same or different carriers of the system.
[0083] The first and second regulatory elements include a promoter (e.g., a constitutive promoter or an inducible promoter), an enhancer (e.g., a 35S promoter or a 35S enhanced promoter), an internal ribosome entry site (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences).
[0084] In some embodiments, the vector in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a phage, etc., which are well known to those skilled in the art.
[0085] In some embodiments, the systems provided herein are in a delivery system. In some embodiments, the delivery system is a nanoparticle, a liposome, an exosome, a microbubble, and a gene gun.
[0086] In one embodiment, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell. In one embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0087] In one embodiment, the target sequence is present in a cell. In one embodiment, the target sequence is present in the nucleus or in the cytoplasm (e.g., an organelle). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.
[0088] Protein-nucleic acid complexes / compositions
[0089] In another aspect, the present invention provides a compound or composition comprising:
[0090] (i) a protein component selected from the group consisting of: the above-mentioned fusion protein, wherein the DNA binding protein in the fusion protein is a Cas protein;
[0091] (ii) a nucleic acid component comprising (a) a guide sequence capable of hybridizing to a target sequence; and (b) a protein binding sequence capable of binding to a DNA binding protein in a fusion protein of the present invention.
[0092] The protein component and the nucleic acid component can bind to each other to form a complex.
[0093] In one embodiment, the nucleic acid component is a guide RNA in a CRISPR-Cas system.
[0094] In one embodiment, the complex or composition is non-naturally occurring or modified. In one embodiment, at least one component of the complex or composition is non-naturally occurring or modified. In one embodiment, the first component is non-naturally occurring or modified; and / or the second component is non-naturally occurring or modified.
[0095] Delivery and delivery compositions
[0096] The fusion proteins, gRNAs, nucleic acid molecules, vectors, systems, complexes, and compositions of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipofection, nucleofection, microinjection, sonoporation, gene guns, calcium phosphate-mediated transfection, cationic transfection, lipofection, dendritic transfection, heat shock transfection, nucleofection, magnetofection, lipofection, puncture transfection, optical transfection, agent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial virions, and the like.
[0097] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: a fusion protein, gRNA, nucleic acid molecule, vector, system, complex and composition of the present invention.
[0098] In one embodiment, the delivery vehicle is a particle.
[0099] In one embodiment, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).
[0100] host cells
[0101] The present invention also relates to an in vitro, ex vivo or in vivo cell or cell line or their progeny, wherein the cell or cell line or their progeny comprises: the fusion protein, nucleic acid molecule, protein-nucleic acid complex, vector, and delivery composition of the present invention.
[0102] In certain embodiments, the cell is a prokaryotic cell.
[0103] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, such as a cell of a non-human primate, cattle, sheep, pig, dog, monkey, rabbit, rodent (such as rat or mouse). In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as a cell of poultry (such as chicken), fish or crustacean (such as clams, shrimp). In certain embodiments, the cell is a plant cell, such as a cell or cultivated plant or food crop such as cassava, corn, sorghum, soybean, wheat, oat or rice that a monocot or dicot has, such as an algae, tree or production plant, fruit or vegetable (for example, trees such as citrus trees, nut trees; Solanaceae, cotton, tobacco, tomato, grape, coffee, cocoa, etc.).
[0104] In certain embodiments, the cell is a stem cell or a stem cell line.
[0105] In certain cases, the host cells of the invention comprise genetic or genomic modifications that are not present in their wild-type form.
[0106] Gene Editing Methods and Applications
[0107] The fusion proteins, nucleic acids, compositions, CIRSPR / Cas systems, vector systems, delivery compositions, or host cells of the present invention can be used for any one or more of the following purposes: targeting and / or editing target nucleic acids; specifically editing double-stranded nucleic acids; base editing double-stranded nucleic acids; base editing single-stranded nucleic acids. In other embodiments, they can also be used to prepare reagents or kits for any one or more of the above purposes.
[0108] The present invention also provides a method for editing nucleic acids, comprising contacting a target region of a nucleic acid (e.g., a double-stranded DNA sequence) with a complex comprising the above-mentioned fusion protein and gRNA; wherein the target region comprises a targeted base pair, and base substitution is performed on the targeted base pair in the target region. In one embodiment, the deaminase in the fusion protein is adenosine deaminase, and the targeted base pair is substituted from A:T to G:C.
[0109] The above A:T means that the bases paired together are A and T; similarly, G:C means that the bases paired together are G and C.
[0110] The present invention also provides the use of the fusion protein, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition or the above-mentioned host cell in gene editing; or, the use in preparing a reagent or kit for gene editing.
[0111] In one embodiment, the gene editing is performed inside and / or outside the cell.
[0112] In one embodiment, the gene editing is single-base editing of the target gene.
[0113] The present invention also provides a method for editing a target nucleic acid, comprising contacting the target nucleic acid with the above-mentioned fusion protein, nucleic acid, composition, CIRSPR / Cas system, vector system or delivery composition. In one embodiment, the method is for editing the target nucleic acid inside or outside the cell.
[0114] The gene editing or editing of target nucleic acid includes the step of editing a single base of the target gene.
[0115] The editing can be performed in prokaryotic cells and / or eukaryotic cells.
[0116] On the other hand, the present invention also provides a kit for gene editing, which includes the above-mentioned adenosine deaminase, fusion protein, gRNA, nucleic acid, the above-mentioned composition, the above-mentioned CIRSPR / Cas system, the above-mentioned vector system, the above-mentioned delivery composition or the above-mentioned host cell.
[0117] In another aspect, the present invention provides the use of the adenosine deaminase, fusion protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition, or host cell in preparing a preparation or kit for:
[0118] (i) gene or genome editing;
[0119] (ii) editing a target sequence in a target locus to modify an organism;
[0120] (iii) single-base editing;
[0121] (iv) Treatment of disease.
[0122] Preferably, the above-mentioned gene or genome editing is performed inside or outside the cell.
[0123] Preferably, the treatment of the disease is the treatment of a condition caused by a defect in the target sequence in the target locus.
[0124] Method for specifically modifying target nucleic acid
[0125] On the other hand, the present invention also provides a method for specifically modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with the above-mentioned fusion protein, nucleic acid, composition, CIRSPR / Cas system, vector system or delivery composition.
[0126] The specific modification can occur in vivo or in vitro.
[0127] The specific modification can occur inside or outside the cell.
[0128] In some cases, the cell is selected from a prokaryotic cell or a eukaryotic cell, eg, an animal cell, a plant cell, or a microbial cell.
[0129] Adenosine deaminase
[0130] As used herein, the term "adenosine deaminase" catalyzes the hydrolytic deamination of the nucleobase adenine. Provided herein are enzymes that can convert adenosine (A) in DNA into inosine (I), such as engineered adenosine deaminase or optimized adenosine deaminase. Such adenosine deaminase can cause the conversion of A:T to G:C base pairs. In some embodiments, adenosine deaminase is a variant of a naturally occurring adenosine deaminase from an organism. In some embodiments, adenosine deaminase does not exist in nature. For example, in some embodiments, adenosine deaminase has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identity to a naturally occurring adenosine deaminase.
[0131] DNA-binding proteins
[0132] As used herein, the term "DNA binding protein" or "DNA binding protein domain" is any protein designated to locate and bind to a specific target DNA nucleotide sequence (e.g., a genomic locus). The term includes RNA-programmable proteins that bind (e.g., form a complex) with one or more nucleic acid molecules (i.e., including, for example, guide RNAs in the case of Cas systems) that guide or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a DNA sequence) that is complementary to the one or more nucleic acid molecules (or portions or regions thereof) to which the protein binds. Exemplary RNA programmable proteins are CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and can include Cas9 equivalents from any type of CRISPR system (e.g., type II, type V, type VI), including Cpf1 (type V CRISPR-Cas system), C2cl (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), C2c3 (type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Cas 12a, Cas 12b, Cas12c, Cas12d, Cas12g, Cas12h, Cas12i, dCas12i, and the like, and further Cas equivalents.
[0133] CRISPR system
[0134] As used herein, the terms "Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-CRISPR-Associated (Cas) (CRISPR-Cas) System" or "CRISPR System" are used interchangeably and have the meaning commonly understood by those skilled in the art, which generally include transcripts or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of directing the activity of the Cas genes. The Cas protein in the present invention is Crisprassociated protein.
[0135] CRISPR / Cas complex
[0136] As used herein, the term "CRISPR / Cas complex" refers to a complex formed by the binding of guide RNA or mature crRNA to the Cas protein, which comprises a direct repeat sequence that hybridizes to the guide sequence of the target sequence and binds to the Cas protein, and the complex is capable of recognizing and cleaving a polynucleotide that can hybridize to the guide RNA or mature crRNA.
[0137] Guide RNA (gRNA)
[0138] As used herein, the terms "guide RNA (gRNA)", "mature crRNA", and "guide sequence" are used interchangeably and have meanings commonly understood by those skilled in the art. The guide RNA comprises a protein binding sequence and a targeting sequence that targets a target nucleic acid.
[0139] The guide RNA capable of binding to Cas12i may comprise a direct repeat sequence and a guide sequence, or may consist essentially of or consist of a direct repeat sequence and a guide sequence.
[0140] In some cases, a targeting sequence or guide sequence is any polynucleotide sequence that has sufficient complementarity to a target sequence to hybridize with the target sequence and guide the specific binding of the CRISPR / Cas complex to the target sequence. In one embodiment, when optimally aligned, the degree of complementarity between a targeting sequence or guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Determining optimal alignment is within the capabilities of those of ordinary skill in the art. For example, there are publicly available and commercially available alignment algorithms and programs, such as, but not limited to, ClustalW, Smith-Waterman algorithm in matlab, Bowtie, Geneious, Biopython, and SeqMan.
[0141] Target sequence
[0142] "Target sequence" refers to a polynucleotide targeted by a targeting sequence or guide sequence in a gRNA, such as a sequence having complementarity with the targeting sequence or guide sequence, wherein hybridization between the target sequence and the targeting sequence or guide sequence promotes the formation of a CRISPR / Cas complex (including Cas protein and gRNA). Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR / Cas complex.
[0143] The target sequence can comprise any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the nucleus or cytoplasm of the cell. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast. A sequence or template that can be used to recombine into a target locus comprising the target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence". In one embodiment, the editing template is an exogenous nucleic acid. In one embodiment, the recombination is homologous recombination.
[0144] In the present invention, a "target sequence" or "target polynucleotide" or "target nucleic acid" can be any polynucleotide that is endogenous or exogenous to a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide that is present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).
[0145] Base editing
[0146] The term "base editing" refers to a genome editing technique that involves converting a specific nucleic acid base into another base at a targeted genomic locus. In certain embodiments, this can be achieved without the need for double-stranded DNA breaks (DSBs) or single-strand breaks (nicking). So far, other genome editing techniques including systems based on CRISPR have been used to introduce DSBs at the seat of interest. Subsequently, cellular DNA repair enzymes repair the break, typically resulting in random insertions or deletions (indels) of bases at the DSB site. However, when it is desired to introduce or correct point mutations rather than random destruction of the entire gene at the target locus, these genome editing techniques are not suitable because the correction rate is low (typically 0.1% to 5%), and wherein the main genome editing product is indel. In order to increase the efficiency of gene correction without introducing Rando indel, the inventors use adenosine deaminase in combination with the CRISPR system to convert one DNA base directly into another DNA base without forming DSBs.
[0147] wild type
[0148] As used herein, the term "wild type" has the meaning generally understood by those skilled in the art to refer to the typical form of an organism, strain, gene, or characteristic as it exists in nature, as distinguished from mutant or variant forms, which can be isolated from a source in nature and has not been intentionally modified by man.
[0149] Non-naturally occurring
[0150] As used herein, the terms "non-naturally occurring" or "engineered" are used interchangeably and indicate the involvement of human effort. When these terms are used to describe a nucleic acid molecule or polypeptide, they indicate that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is associated in nature or as found in nature.
[0151] Identity
[0152] As used herein, the term "identity" refers to the match between two polypeptides or between two nucleic acids. When a position in both sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of the two DNA molecules is occupied by adenine, or a position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared x 100. For example, if 6 out of 10 positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT share 50% identity (3 out of 6 positions match). Typically, two sequences are compared when they are aligned for maximum identity. Such an alignment can be achieved, for example, by using the method of Needleman et al. (1970) J. Mol. Biol. 48:443-453, which can be conveniently performed using a computer program such as the Align program (DNAstar, Inc.). The percent identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)), which has been incorporated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. In addition, the percent identity between two amino acid sequences can be determined using the Needleman and Wunsch (J Mol. Biol. 48:444-453 (1970)) algorithm, which has been incorporated into the GAP program in the GCG software package (available at www.gcg.com), using either a Blossum 62 matrix or a PAM250 matrix and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6.
[0153] carrier
[0154] The term "vector" refers to a nucleic acid molecule that is capable of transporting another nucleic acid molecule to which it is attached. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends, or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other various polynucleotides known in the art. A vector can be introduced into a host cell by transformation, transduction, or transfection so that the genetic material elements it carries are expressed in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc. as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector can contain a variety of elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. In addition, the vector may also contain a replication initiation site.
[0155] One type of vector is a "plasmid," which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, eg, by standard molecular cloning techniques.
[0156] Another type of vector is a viral vector, in which a virally derived DNA or RNA sequence is present in a vector for packaging a virus (e.g., a retrovirus, a replication-defective retrovirus, adenovirus, a replication-defective adenovirus, and adeno-associated virus). The viral vector also comprises a polynucleotide carried by a virus for transfection into a host cell. Some vectors (e.g., bacterial vectors and episomal mammalian vectors with a bacterial origin of replication) can replicate autonomously in the host cell into which they are introduced.
[0157] Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell upon introduction into the host cell and are thereby replicated along with the host genome. Furthermore, some vectors are capable of directing the expression of genes to which they are operably linked. Such vectors are referred to herein as "expression vectors."
[0158] host cells
[0159] As used herein, the term "host cell" refers to cells that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.
[0160] Those skilled in the art will appreciate that the design of the expression vector may depend on factors such as the choice of the host cell to be transformed, the level of expression desired, and the like.
[0161] Regulatory elements
[0162] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, CA (1990). In some cases, regulatory elements include those that direct the constitutive expression of a nucleotide sequence in many types of host cells and those that direct the expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or special cell types (e.g., lymphocytes). In some cases, regulatory elements can also direct expression in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell-type-specific. In some cases, the term "regulatory element" encompasses enhancer elements such as WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I ((Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).
[0163] promoter
[0164] As used herein, the term "promoter" has a meaning well known to those skilled in the art and refers to a non-coding nucleotide sequence located upstream of a gene that can initiate expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to the promoter.
[0165] NLS
[0166] A "nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" a protein for import into the cell nucleus via nuclear transport, i.e., proteins with an NLS are transported to the cell nucleus. Typically, an NLS comprises a positively charged Lys or Arg residue exposed on the surface of the protein. Exemplary nuclear localization sequences include, but are not limited to, NLSs from: SV40 large T antigen, EGL-13, c-Myc, and TUS protein. In some embodiments, the NLS comprises a PKKKRKV sequence. In some embodiments, in some embodiments, the NLS comprises a MKRTADGSEFESPKKKRKV sequence. In some embodiments, the NLS comprises a KRPAATKKAGQAKKKK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the sequence KIPIK in the yeast transcription repressor Matα2, and PY-NLS.
[0167] operably connected
[0168] As used herein, the term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to the one or more regulatory elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).
[0169] Complementarity
[0170] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence by means of traditional Watson-Crick or other non-traditional types. Percent complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues of a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, "substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.
[0171] Stringent conditions
[0172] As used herein, "stringent conditions" for hybridization refer to conditions under which a nucleic acid having complementarity with a target sequence predominantly hybridizes to the target sequence and does not substantially hybridize to non-target sequences. Stringent conditions are generally sequence-dependent and vary depending on many factors. Generally speaking, the longer the sequence, the higher the temperature at which the sequence specifically hybridizes to its target sequence.
[0173] hybridization
[0174] The terms "hybridize" or "complementary" or "substantially complementary" refer to a nucleic acid (e.g., RNA, DNA) comprising a nucleotide sequence that enables it to non-covalently bind, i.e., form base pairs and / or G / U base pairs, "anneal" or "hybridize" with another nucleic acid in a sequence-specific, antiparallel manner (i.e., a nucleic acid specifically binds to a complementary nucleic acid).
[0175] Hybridization requires that the two nucleic acids contain complementary sequences, although there may be mismatches between the bases. Suitable conditions for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, which are variables well known in the art. Typically, the length of a hybridizable nucleic acid is 8 nucleotides or more (e.g., 10 nucleotides or more, 12 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 22 nucleotides or more, 25 nucleotides or more, or 30 nucleotides or more).
[0176] It is understood that the sequence of a polynucleotide need not be 100% complementary to the sequence of its target nucleic acid to hybridize specifically. A polynucleotide may comprise 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% sequence complementarity to the target region in the target nucleic acid sequence with which it hybridizes.
[0177] The hybridization of the target sequence and the gRNA represents that at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the nucleic acid sequences of the target sequence and the gRNA can hybridize to form a complex; or represents that at least 12, 15, 16, 17, 18, 19, 20, 21, 22 or more bases of the nucleic acid sequences of the target sequence and the gRNA can complement each other and hybridize to form a complex.
[0178] Express
[0179] As used herein, the term "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0180] connector
[0181] As used herein, the term "linker" refers to a linear polypeptide formed by connecting multiple amino acid residues via peptide bonds. The linker of the present invention can be an artificially synthesized amino acid sequence, or a naturally occurring polypeptide sequence, such as a polypeptide having a hinge region function. Such linker polypeptides are well known in the art (see, for example, Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90: 6444-6448; Poljak, RJ et al. (1994) Structure 2: 1121-1123).
[0182] treat
[0183] As used herein, the term "treat" refers to treating or curing a disorder, delaying the onset of symptoms of a disorder, and / or delaying the progression of a disorder.
[0184] Subjects
[0185] As used herein, the term "subject" includes, but is not limited to, various animals, plants, and microorganisms.
[0186] animal
[0187] For example, mammals, such as bovines, equines, ovines, porcines, canines, felines, lagomorphs, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In certain embodiments, the subject (e.g., a human) has a disorder (e.g., a disorder caused by a disease-associated gene defect).
[0188] plant
[0189] The term "plant" is to be understood as meaning any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development, in particular monocotyledonous or dicotyledonous plants, vegetable crops including artichokes, Brussels sprouts, rocket, leeks, asparagus, lettuce (e.g., head lettuce, leaf lettuce, romaine lettuce), bok choy, yellow taro, melons (e.g., cantaloupe, watermelon, Crenshaw melon, honeydew melon, cantaloupe), oilseed crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, kale, kale, Chinese cabbage, bok choy), cardoon, carrot, napa, okra, onion, celery, parsley, chickpeas, parsnips, endive, peppers, potatoes, cucurbits (e.g., zucchini, cucumber, courgette, squash, pumpkin), radish, cabbage, Onions, rutabagas, eggplant (also known as eggplant), salsify, lettuce, shallots, endive, garlic, spinach, green onions, squash, greens, beets (sugar beets and fodder beets), sweet potatoes, Swiss chard, horseradish, tomatoes, turnips, and spices; fruits and / or vines such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, quince, almonds, chestnuts, hazelnuts, pecans, pistachios, walnuts, citrus, blueberries, boysenberries, erry), cranberries, currants, loganberries, raspberries, strawberries, blackberries, grapes, avocados, bananas, kiwis, persimmons, pomegranates, pineapples, tropical fruits, pome fruits, melons, mangoes, papayas, and lychees; field crops such as clover, alfalfa, evening primrose, meadowsweet, corn / maize (feed corn, sweet corn, popcorn), hops, jojoba, peanuts, rice, safflower, small grain cereals (barley, oats, rye, wheat, etc.), sorghum, tobacco, kapok, legumes (beans, lentils, peas beans, soybeans), oil plants (rapeseed, mustard, olive, sunflower, coconut, castor oil plant, cocoa bean, peanut), Arabidopsis, fiber plants (cotton, flax, jute), Lauraceae (cinnamon, camphor), or a plant such as coffee, sugar cane, tea, and natural rubber plant; and / or bedding plants, such as flowering plants, cacti, succulents and / or ornamental plants, as well as trees such as forests (broadleaf trees and evergreen trees, such as conifers), fruit trees, ornamental trees, and nut-bearing trees, as well as shrubs and other seedlings.
[0190] Advantageous Effects of the Invention
[0191] The present invention improves adenosine deaminase and fuses it with a DNA-binding protein, which can be used for single-base editing of target nucleic acids, thereby improving the efficiency of single-base editing and having broad application prospects.
[0192] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples, but it will be understood by those skilled in the art that the following drawings and examples are intended only to illustrate the present invention and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art based on the following detailed description of the accompanying drawings and preferred embodiments.
[0193] The sequence information involved in this application is as follows:
[0194]
[0195] BRIEF DESCRIPTION OF THE DRAWINGS
[0196] Figure 1 .Schematic diagram of the ABE fluorescence reporter system.
[0197] Figure 2 .Verification results of base editing efficiency of ABE vectors composed of different adenosine deaminases and dCas12i3. DETAILED DESCRIPTION
[0198] The following examples are only used to describe the present invention, but are not intended to limit the present invention. Unless otherwise specified, the experiment and method described in the embodiment are carried out substantially according to conventional methods well known in the art and described in various references. For example, the conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA used in the present invention can be found in Sambrook, Fritsch and Maniatis, " molecular cloning: laboratory manual " (MOLECULAR CLONING:A LABORATORY MANUAL), 2nd edition (1989); " current protocols in molecular biology " (FM Ausubel et al. edit, (1987)); " methods in enzymes " (METHODS IN ENZYMOLOGY) series (Academic Publishing Company): " PCR 2: practical methods " (PCR 2:A PRACTICAL APPROACH) (MJ MacPherson, BD Hames and GR Taylor, eds. (1995)), ANTIBODIES: A LABORATORY MANUAL (Harlow and Lane, eds. (1988)), and ANIMAL CELL CULTURE (RI Freshney, ed. (1987)).
[0199] In addition, if specific conditions are not specified in the examples, the experiments were performed under conventional conditions or the conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments is not specified, they are all conventional products that can be obtained commercially. It is understood that the examples describe the present invention by way of example and are not intended to limit the scope of the present invention. All publications and other references mentioned herein are incorporated herein by reference in their entirety.
[0200] Example 1. Screening of mutant adenosine deaminase
[0201] The structure of adenosine deaminase GZ0109598-46 (amino acid sequence shown in SEQ ID No. 1, nucleotide sequence shown in SEQ ID No. 2) from Rahnella aceris was predicted. Using bioinformatics, the applicants predicted key amino acid sites that may affect its biological function. They then selected relevant sites to construct a mutant library, which was then screened using a fluorescent reporter system. In this embodiment, amino acid mutations were made at positions 110, 153, 161, and 167 of SEQ ID No. 1 from the N-terminus.
[0202] A variant of adenosine deaminase GZ0109598-46 was generated by PCR-based site-directed mutagenesis. The specific method is to divide the DNA sequence of adenosine deaminase into two parts with the mutation site as the center, design primers for the mutation site, and each two pairs of primers correspond to an amino acid mutation. Two pairs of primers are designed to amplify these two parts of the DNA sequence respectively. At the same time, the sequence to be mutated is introduced into the primers, and finally the two fragments are loaded into the pcDNA3.3-eGFP vector by Gibson cloning. The combination of mutants is constructed by splitting the DNA of adenosine deaminase into multiple segments and using PCR and Gibson clone. Fragment amplification kit: TransStart FastPfu DNA Polymerase (containing 2.5mM dNTPs), please refer to the instructions for the specific experimental process. Gel recovery kit: Gel DNA Extraction Mini Kit. Detailed experimental procedures are provided in the instructions. Vector construction kit: pEASY-Basic Seamless Cloning and Assembly Kit (CU201-03). Detailed experimental procedures are provided in the instructions.
[0203] Based on the above amino acid mutation sites, the wild-type protein of adenosine deaminase GZ0109598-46 and adenosine deaminase with single-point mutations at amino acids 110, 153, 161, and 167 (named after the mutation type) were obtained: V110K, V110H, Y153F, N161R, and E167R; compared with SEQ ID No. 1, the 110th amino acid of the mutant adenosine deaminase V110K mutated to K; compared with SEQ ID No. 1, the 110th amino acid of the mutant adenosine deaminase V110H mutated to H; compared with SEQ ID No. 1, the 153rd amino acid of the mutant adenosine deaminase Y153F mutated to F; compared with SEQ ID No. 1, the 161st amino acid of the mutant adenosine deaminase N161R mutated to R; compared with SEQ ID No. 1, the 161st amino acid of the mutant adenosine deaminase E167R mutated to Compared with ID No.1, the 167th amino acid mutated to R.
[0204] In addition, an adenosine deaminase with combined mutations at amino acids 110, 153, 161, and 167 was obtained (named after the mutation type): V110K-Y153F-N161R-E167R; compared with SEQ ID No. 1, the mutant adenosine deaminase V110K-Y153F-N161R-E167R had the amino acid at position 110 mutated to K, the amino acid at position 153 mutated to F, the amino acid at position 161 mutated to R, and the amino acid at position 167 mutated to R.
[0205] Example 2. Verification of base editing activity of mutant adenosine deaminase
[0206] The applicant used the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein with inactivated nuclease activity and the adenosine deaminase obtained in Example 1 to construct the ABE single-base editing tool. Among them, dCas12i3 is Cas12i3 with E844A mutation (Cas12i3 is Cas12f.4 in CN111757889B, and the amino acid sequence of wild-type Cas12i3 is shown in SEQ ID No. 3), and its nuclease activity is inactivated, that is, the amino acid sequence of dCas12i3 is mutated to A at position 844 compared with SEQ ID No. 3; dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is Cas12i3 protein with E844A, S7R, D233R, D267R, N369R and S433R mutations, and the amino acid sequence of dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is the same as SEQ ID No. Compared with No.3, E at position 844 mutated to A, S at position 7 mutated to R, D at position 233 mutated to R, D at position 267 mutated to R, N at position 369 mutated to R, and S at position 433 mutated to R. In other embodiments, the D at position 619 of the Cas12i3 protein corresponding to the sequence shown in SEQ ID No.3 can also be mutated to A to obtain a Cas mutant protein (i.e., dCas12i3) in which nuclease activity is inactivated; in other embodiments, the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein can also be replaced with other Cas mutant proteins in which nuclease activity is inactivated, such as dCas9 and nCas9.
[0207] The schematic diagram of the ABE single-base editing tool constructed using dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein and adenosine deaminase is shown in the figure. Figure 1 As shown, adenosine deaminase is expressed in dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein ( Figure 1The N-terminus of the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is connected through an XTEN linker, and the amino acid sequence of the XTEN linker is SGGSSGGSSGSETPGTSESATPESSGGSSGGS; the other end of the adenosine deaminase and the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein is also connected to an NLS, and the amino acid sequence of the NLS at the N-terminus of the adenosine deaminase is MKRTADGSEFESPKKKRKV, and the amino acid sequence of the NLS at the C-terminus of the dCas12i3 (S7R-D233R-D2 67R-N369R-S433R) protein is KRPAATKKAGQAKKKK, and EGFP is a tag designed to screen positive cells. Figure 1 It is just an exemplary connection method of adenosine deaminase and dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein. In order to verify the editing activity of single-base editing tools constructed with different adenosine deaminases, dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein can be constructed with different adenosine deaminases in the same connection method, that is, Figure 1 The adenosine deaminase in the expression vector can be replaced with other adenosine deaminase. In other embodiments, those skilled in the art can adjust the position or connection order of the above elements.
[0208] In order to verify the activity of the above-mentioned ABE single-base editing tool, referring to the experimental methods familiar to those skilled in the art, the target DNA sequence was set to the tga terminator region between the EGFP gene and the dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein, whose PAM is TTG, and EGFP cannot emit light normally. Only when the ABE reporter system converts the tga terminator to cga (R) through single-base editing can EGFP emit green fluorescence. The editing efficiency of the ABE base editing tool can be evaluated by detecting the ratio of green fluorescence. The constructed expression vector is transfected into the 293T cell line, and the EGFP ratio is detected by flow cytometry after 48 hours; the single-base editing efficiency is calculated as the number of EGFP-positive cells / total number of cells.
[0209] The editing activity of ABE single-base editing tools constructed by dCas12i3 (S7R-D233R-D267R-N369R-S433R) protein and different adenosine deaminases is shown in Figure 2. Figure 2 shown. Figure 2 NC is the negative control group, which specifically uses the ABE base editing tool that does not contain any adenosine deaminase for experiments. Due to the absence of adenosine deaminase, base editing cannot be performed and fluorescence cannot be emitted.
[0210] exist Figure 2 In the experiment, wild-type adenosine deaminase GZ0109598-46 ( Figure 2 The editing efficiency of the ABE single-base editing tool constructed using the 46RAHSY in the PCR reaction was 14.4%; the editing efficiency of the ABE single-base editing tool constructed using the mutant adenosine deaminase V110K was 44.1%; the editing efficiency of the ABE single-base editing tool constructed using the mutant adenosine deaminase V110H was 31.5%; the editing efficiency of the ABE single-base editing tool constructed using the mutant adenosine deaminase Y153F was 20.2%; the editing efficiency of the ABE single-base editing tool constructed using the mutant adenosine deaminase N161R was 19.8%; the editing efficiency of the ABE single-base editing tool constructed using the mutant adenosine deaminase E167R was 24.2%; the editing efficiency of the ABE single-base editing tool constructed using the mutant adenosine deaminase V110K-Y153F-N161R-E167R ( Figure 2 The editing efficiency of the ABE single-base editing tool constructed by using V110K Y153F N161R E167R in the genome is 61.4%.
[0211] From the above results, it can be seen that the editing efficiency is significantly improved after single-site mutation and combined site mutation of amino acid 110, 153, 161 or 167 of adenosine deaminase GZ0109598-46.
[0212] Although the specific embodiments of the present invention have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details based on all the teachings published, and these changes are all within the scope of protection of the present invention. The entire invention is given by the appended claims and any equivalents thereof.
Claims
1. A mutant adenosine deaminase, wherein the mutant adenosine deaminase has a mutation at any one or more of the following amino acid positions corresponding to the amino acid sequence of SEQ ID No. 1: position 110, position 153, position 161, and position 167, compared to the amino acid sequence of the parent adenosine deaminase; Preferably, the amino acid at position 110 mutates to K or H; or, the amino acid at position 153 mutates to F; or, the amino acid at position 161 mutates to R; or, the amino acid at position 167 mutates to R; More preferably, the amino acid sequence of the parent adenosine deaminase has at least 80% sequence identity with SEQ ID No.
1.
2. A fusion protein, characterized in that The fusion protein comprises a DNA binding protein and the adenosine deaminase according to claim 1; Preferably, the DNA binding protein is a Cas protein; preferably, the Cas protein is selected from Cas9, Cas9n, dCas9, CasX, CasY, C2cl, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12g, Cas12h, Cas12i, Cas12j, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, Cas9-NG, LbCas12a, enAsCas12a, Cas9-KKH, circular displacement Cas9, Argonaute (Ago) domain, SmacCas9, Spy-macCas9, SpGas9-NRRH, SpaCas9-NRTH, SpaCas9-NRCH, Cas9-NG-CP1041, Cas9-NG-VRQR, dCas12i, nCas12i and variants thereof; More preferably, the Cas protein is a nuclease-inactivated Cas protein.
3. An isolated polynucleotide, characterized in that The polynucleotide encodes the adenosine deaminase according to claim 1, or the polynucleotide encodes the fusion protein according to claim 2.
4. A carrier, characterized in that The vector comprises the polynucleotide according to claim 3 and a regulatory element operably linked thereto.
5. A base editing system, characterized in that The system comprises the fusion protein of claim 2 and at least one gRNA; the DNA binding protein in the fusion protein is a Cas protein, and the gRNA is capable of binding to the Cas protein.
6. A composition, characterized in that The composition comprises: (i) a protein component selected from: the fusion protein of claim 2, wherein the DNA binding protein in the fusion protein is a Cas protein; (ii) a nucleic acid component, which is a gRNA, and the gRNA is capable of binding to the Cas protein.
7. An engineered host cell, characterized in that The host cell comprises the adenosine deaminase of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the base editing system of claim 5, or the composition of claim 6.
8. Use of the adenosine deaminase according to claim 1, or the fusion protein according to claim 2, or the polynucleotide according to claim 3, or the vector according to claim 4, or the base editing system according to claim 5, or the composition according to claim 6, or the host cell according to claim 7 in gene editing; or use in the preparation of a reagent or kit for gene editing; Preferably, the gene editing is single-base editing of the target gene.
9. A kit for gene editing, characterized in that: The kit comprises the adenosine deaminase of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the base editing system of claim 5, or the composition of claim 6, or the host cell of claim 7.
10. A method for editing a nucleic acid, comprising the step of contacting a target region of the nucleic acid with the fusion protein of claim 2 and a gRNA, wherein the gRNA comprises a segment capable of binding to the Cas protein in the fusion protein of claim 2 and a segment capable of binding to the target region of the nucleic acid; wherein, The target region comprises a targeted base pair, and the fusion protein is capable of performing base substitution on the targeted base pair; Preferably, the targeted base pair is replaced by A:T to G:C.
Citation Information
Patent Citations
Novel CRISPR / Cas12f enzymes and systems
CN111757889B