Targeted mutagenesis system based on adenine and cytosine dual base editors
By building a dual-base editor, combining Cas9 nuclease, adenine deaminase and cytosine deaminase, the existing base editing technology is inefficient and Indels problems, and efficient mutations on multiple bases are achieved, suitable for directed evolutionary proteins and genetic diseases treatment.
Patent Information
- Application Number
- CN202110925521.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-12
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-08-12
AI Technical Summary
The existing base editing technology has a low mutation rate at the targeted sites, and the diversified sites are limited to C. A simple replacement of Cas9 will introduce a large number of Indels, resulting in limited base editing efficiency.
A two-base editor based on adenine and cytosine was constructed, combining Cas9 nuclease, adenine deaminase and cytosine deaminase, and a multiple sgRNA guidance to achieve A>G and C>T conversion on the DNA strands, avoid double-strand breaks, and introduce mutations using intracellular repair mechanisms.
It has achieved efficient introduction of mutations on four bases: A, T, C and G, which has improved mutagenesis activity and has better application prospects in directed evolution of proteins and the treatment of genetic diseases.
Smart Images

Figure CN115704015B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a targeted mutagenesis system based on adenine and cytosine double base editors and methods for using the mutagenesis system to direct the evolution of proteins or treat genetic diseases. Background Art
[0002] CRISPR–Cas9 has been used for targeted mutation, modification, and regulation in a variety of organisms (Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science, 339, 819-823 (2013); Liu, P., Chen, M., Liu, Y., Qi, LS & Ding, S. CRISPR-Based Chromatin Remodeling of the Endogenous Oct4 or Sox2 Locus Enables Reprogramming to Pluripotency. Cell Stem Cell, 22, 252-261. e254 (2018); Ran, FA et al. Doublenicking by RNA-guided CRISPR Cas9 for enhanced genome editing specificity. Cell, 154, 1380-1389 (2013); Jinek, M. et al. A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial Immunity. Science, 337, 816-821 (2012). Of particular importance is the ability of CRISPR–Cas9 and deaminases to function as base editors. Two types of base editors, CBE and ABE, are currently available, which can respectively change C to T or A to G within a certain range of the sgRNA targeting region (Komor, AC, Kim, YB, Packer, MS, Zuris, JA & Liu, DR. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature, 533, 420-424 (2016); Gaudelli, N. Met al. Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature 551, 464-471 (2017)).The first-generation cytosine base editors are composed of a fusion of rat-derived cytosine deaminase rAPOBEC1 and the inactive dCas9. Under the guidance of sgRNA, dCas9 binds to DNA and exposes single-stranded DNA. Cytosine deaminase then deaminates cytosine within a certain range on the single-stranded DNA to uracil. Finally, the uracil is converted to thymine by the cell's DNA repair or replication system, thus achieving the cytosine-to-thymine conversion. Because uracil DNA glycosylase (UDG) in cells removes uracil from DNA, the efficiency of the first-generation cytosine base editors is low. By fusing uracil DNA glycosylase inhibitors (UGIs) to inhibit uracil DNA glycosylase in cells, a more efficient second-generation cytosine base editor is formed. After further replacing dCas9 with nCas9 with single-stranded DNA cleavage activity, a more efficient third-generation cytosine base editor was developed by utilizing the intracellular repair characteristics (Komor, AC, Kim, YB, Packer, MS, Zuris, JA & Liu, DR Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature, 533, 420-424 (2016)). Subsequently, different laboratories obtained cytosine base editors with different activities by replacing different cytosine deaminases (Ma, Y. et al. Targeted AID-mediated mutagenesis (TAM) enables efficient genomic diversification in mammalian cells. Nat Methods 13, 1029-1035 (2016); Nishida, K. et al. Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. Science 353, aaf8729-aaf8729 (2016)). David R. Liu's laboratory evolved a single-stranded DNA adenine deaminase TadA-TadA* based on the RNA adenine deaminase (Escherichia coli TadA, ecTadA) from Escherichia coli.Based on a principle similar to that of the cytosine base editor, adenine deaminase was used to replace the cytosine deaminase in the cytosine base editor, and an adenine base editor was obtained that can convert adenine to guanine within a certain range of the sgRNA targeting sequence (Gaudelli, NMet al. Programmable base editing of A*T to G*C ingenomic DNA without DNA cleavage. Nature 551, 464-471 (2017)).
[0003] The Cas9 (dCas9) without cutting activity is combined with cytosine deaminase to diversify the targeted gene sites under the guidance of multiple sgRNAs. However, the diversified sites are limited to C. In addition, since the component used is dCas9 instead of nCas9 in conventional base editors, the efficiency of the complete base editor is not brought into play, resulting in a low mutation rate. However, if dCas9 is simply replaced with nCas9, it is found that a large number of Indels (insertions and deletions) are introduced instead of base changes (Hess, GT et al. Directed evolution using dCas9-targeted somatic hypermutation in mammalian cells. Nat Methods, 13, 1036-1042 (2016); Ma, Y. et al. Targeted AID-mediated mutagenesis (TAM) enables efficient genomic diversification in mammalian cells. Nat Methods, 13, 1029-1035 (2016)). By recombining the components of the two types of base editors, a dual base editor was produced that can simultaneously achieve A>G and C>T in the targeted site, but its A>G function was impaired (Li, C. et al. Targeted, random mutagenesis of plant genes with dual cytosine andadenine base editors. Nat Biotechnol, (2020)).
[0004] The present invention aims to solve the above-mentioned problems existing in the prior art. Summary of the Invention
[0005] The present invention provides a targeted mutagenesis system based on adenine and cytosine double base editors that has high mutagenic activity and can trigger transitions between any bases within the targeted range.
[0006] One aspect of the present invention relates to a dual-base editor, which comprises: Cas9 nuclease or its encoding nucleic acid sequence, adenine deaminase or its encoding nucleic acid sequence, and cytosine deaminase or its encoding nucleic acid sequence.
[0007] According to the aforementioned dual-base editor of the present invention, it further comprises a nuclear localization signal (NLS) sequence or its encoding nucleic acid sequence, and / or comprises or does not comprise a UGI component or its encoding nucleic acid sequence.
[0008] According to the aforementioned dual-base editor of the present invention, it further comprises a guide polynucleotide or its encoding nucleic acid sequence.
[0009] According to the aforementioned dual-base editor of the present invention, two or more of the Cas9 nuclease, adenine deaminase, cytosine deaminase, UGI component (if present), nuclear localization signal (NLS) and the coding nucleic acid sequence of the guide polynucleotide are connected by a linker sequence or its coding nucleic acid sequence.
[0010] According to the aforementioned dual-base editor of the present invention, one or more of the components of the dual-base editor, namely, Cas9 nuclease, adenine deaminase, cytosine deaminase, UGI component (if present), nuclear localization signal (NLS) and guide polynucleotide encoding nucleic acid sequence are respectively located in one or more vectors.
[0011] According to the aforementioned dual-base editor of the present invention, the Cas9 nuclease, adenine deaminase, cytosine deaminase, UGI component (if present), and nuclear localization signal (NLS) in the dual-base editor are located in one vector, and the coding nucleic acid sequence of the guide polynucleotide is located in another vector.
[0012] According to the aforementioned dual-base editor of the present invention, the guide polynucleotide is one or more, for example, at least 2, at least 5, at least 10, at least 20, at least 30, at least 50, and the multiple guide polynucleotides can be arranged in series, preferably separated by repeat sequences.
[0013] According to the aforementioned dual-base editor of the present invention, the multiple guide polynucleotides are placed under the control of a single promoter, for example, an RNA polymerase II promoter.
[0014] According to the aforementioned dual-base editor of the present invention, the guide polynucleotide targets both strands or one strand of the double-stranded target DNA.
[0015] According to the aforementioned dual-base editor of the present invention, the dual-base editor comprises dCas9 or its encoding nucleic acid sequence, and wherein the guide polynucleotide or its encoding nucleic acid sequence targets both strands of the double-stranded target DNA.
[0016] According to the aforementioned dual-base editor of the present invention, the Cas9 nuclease is an inactive Cas9 nuclease, i.e., dCas9, or a Cas9 nickase, i.e., nCas9.
[0017] According to the aforementioned double-base editor of the present invention, the encoding nucleic acid sequence of the adenine deaminase is the TadA-TadA* sequence at positions 1263-2354 in SEQ ID NO: 1, or a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity thereto, the encoding nucleic acid sequence of the Cas9 nuclease is the nCas9 sequence at positions 2451-6551 in SEQ ID NO: 1, or a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity thereto, and the encoding nucleic acid sequence of the cytosine deaminase is SEQ ID NO: NO:1, or a PmCDA1 sequence at positions 6711-7337, or a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity thereto; the encoding nucleic acid sequence of the UGI component is a sequence at positions 7368-7895, or a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity thereto.
[0018] According to the aforementioned dual-base editor of the present invention, its nucleic acid sequence is SEQ ID NO: 1, or it is the amino acid sequence encoded by SEQ ID NO: 1.
[0019] According to the aforementioned dual-base editor of the present invention, the dual-base editor is obtained by replacing the sequence corresponding to PmCDA1 in SEQ ID NO: 1 with AncAPOBEC1 of SEQ ID NO: 10 or a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity thereto.
[0020] According to the aforementioned dual-base editor of the present invention, the dual-base editor is obtained by exchanging the sequences corresponding to AncAPOBEC1 and TadA-TadA* in the aforementioned dual-base editor.
[0021] According to the aforementioned dual-base editor of the present invention, the dual-base editor is obtained by assembling the sequences corresponding to AncAPOBEC1 and TadA-TadA* in the aforementioned dual-base editor into single-base editors, respectively, and combining them into the same plasmid in the form of different open reading frames.
[0022] According to the aforementioned dual-base editor of the present invention, the dual-base editor is obtained by replacing the sequence of nCas9 in the aforementioned dual-base editor of the present invention with dCas9 of SEQ ID NO: 9 or a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity thereto.
[0023] Another aspect of the present invention relates to a vector comprising a nucleic acid sequence encoding the guide polynucleotide contained in the aforementioned double-base editor according to the present invention and / or a nucleic acid sequence encoding other components of the double-base editor other than the nucleic acid sequence encoding the guide polynucleotide contained in the aforementioned double-base editor according to the present invention.
[0024] Yet another aspect of the present invention relates to a tool cell, into which the aforementioned dual-base editor according to the present invention or the aforementioned vector according to the present invention is transfected.
[0025] According to the aforementioned tool cell of the present invention, it is a HEK293T cell or a mESC cell. Optionally, the mESC cell is knocked out of the AP enzyme Apex1.
[0026] Another aspect of the present invention relates to a targeted mutagenesis system for targeted mutagenesis of proteins, which comprises the aforementioned dual-base editor according to the present invention or the aforementioned vector according to the present invention, wherein the guide polynucleotide contained in the targeted mutagenesis system targets the target region of the protein coding sequence to be mutagenized.
[0027] Another aspect of the present invention relates to a method for targeted mutagenesis of a protein, which comprises the aforementioned dual-base editor according to the present invention or the aforementioned vector according to the present invention, or the aforementioned tool cell according to the present invention, or the aforementioned targeted mutagenesis system according to the present invention, wherein the guide polynucleotide contained in the dual-base editor, vector, tool cell or targeted mutagenesis system targets the target region of the protein coding sequence to be mutagenized.
[0028] According to the aforementioned method for targeted mutagenesis of a protein of the present invention, it is used for directed evolution of a protein.
[0029] Another aspect of the present invention relates to a kit for mutagenesis or directed evolution of proteins, comprising: (1) the aforementioned dual-base editor according to the present invention or its encoding nucleic acid sequence or the aforementioned vector according to the present invention, or the aforementioned tool cell according to the present invention, or the aforementioned targeted mutagenesis system according to the present invention.
[0030] The targeted mutagenesis system based on the adenine and cytosine dual-base editors of the present invention has stronger mutagenic activity than existing mutagenesis systems and can introduce mutations at four bases: A, T, C, and G, and can convert them to any other base at varying ratios. Therefore, it has great application prospects in directed protein evolution, such as mutagenesis to generate high-affinity antibodies. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1. Establishment of a dual-base editor-based targeted mutagenesis system
[0032] a. Recombining the components of the two types of base editors to construct different forms of dual-base editors to select the one with the best activity;
[0033] b. The principle of TMBEs for diversified targeting regions is to place all sgRNAs on the same DNA strand. Assuming that nCas9 nicks multiple sites on a DNA strand, the complementary single-stranded DNA will be exposed. The A and C within the single-stranded DNA can serve as substrates for adenine deaminase and cytosine deaminase, respectively, and be deaminated to I and U. Mutations can then be introduced through DNA repair or replication mechanisms without causing double-strand breaks due to cleavage of double-stranded DNA.
[0034] c. Frequency distribution of base substitutions triggered by different base editors within the targeted range.
[0035] The substitution frequency at each site is equal to the number of mutations at that site divided by the coverage number. Small black arrows represent sgRNAs. Each dot represents a single nucleotide. Data are mean ± SD of three independent replicates.
[0036] Figure 2. Placing multiple sgRNAs on the same DNA strand and using nCas9 to nick the non-editing strand can efficiently achieve targeted mutagenesis while avoiding a large number of double-strand breaks.
[0037] ab Mutation frequency (a) and indel rate (b) caused by TMBEs-1d and TMBEs-1 targeting one DNA strand of EGFP under the guidance of 11 sgRNAs (multi-sgRNA expression vector 2, One) or targeting both DNA strands of EGFP under the guidance of 21 sgRNAs (multi-sgRNA expression vector 4, Two).
[0038] Data are the mean ± SD of three independent experiments. **** indicates P < 0.0001 by two-tailed t test.
[0039] Figure 3. Establishing an sgRNA expression system and testing its function
[0040] a. The sgRNA with the Csy4 recognition site is serially connected to the 3'UTR of Csy4. After transcription and translation, Csy4 cuts the serial sgRNA into a single sgRNA.
[0041] b. Both the adenine base editor (ABEmax) and the cytosine base editor (AncBE4max) are compatible with multi-sgRNA expression systems, editing A to G and C to T within the targeted range, respectively.
[0042] Figure 4 Frequency distribution of base substitutions triggered by different base editors within the targeted range.
[0043] The substitution frequency at each site is equal to the number of mutations at that site divided by the coverage number. Small black arrows represent sgRNAs. Each dot represents a single nucleotide. Data are mean ± SD of three independent replicates.
[0044] Figure 5. Dual-base editors achieve A>G and C>G changes at multiple different sites in HEK293T cells
[0045] a. Heatmap of TMBEs-1 editing efficiency at 15 different sites. The letters on the left of each panel represent the corresponding sgRNA sequence, and the last three letters represent the PAM sequence. The first column of each panel represents on-target editing efficiency, the second column represents off-target editing efficiency, and the third column represents the untreated group. Data are the average of three independent replicates.
[0046] be.TMBEs-1 editing efficiency for A and C at 15 targeted sites (the horizontal axis represents the base position, with PAM sequences counting positions 21–23) and the distribution of mutations. Data are the average of three independent replicates.
[0047] Figure 6. Types of mutations on extensions A and C
[0048] Possible results after aC is deaminated to U: U is read as T; U is cut by uracil glycosylase in the cell to form an abasic site (AP) and then randomly read as A, T, C or G; U is removed by uracil glycosylase in the cell to form an abasic site, and AP is then removed by AP enzyme in the cell to introduce Indels.
[0049] b, c. Under the guidance of four sgRNAs targeting the same DNA chain, the mutation distribution and Indel rate caused by TMBEs-1 and TMBEs-1B on A and C.
[0050] d. Knockout of the endogenous major AP enzyme Apex1 can reduce Indels triggered by TMBEs-1B.
[0051] The frequency and distribution of mutations induced by TMBEs-1B and TMBEs-1, using 11 sgRNAs targeting the same DNA strand of EGFP, were analyzed. Data are mean ± SD from three independent replicates. ** indicates p < 0.01 by a two-tailed t-test, and **** indicates p < 0.0001.
[0052] Figure 7. TMBEs-1B and TMBEs-1 induced endogenous gene mutations in cells
[0053] ac. Mutation frequency (a), mutation distribution (b), and indel rate (c) induced by TMBEs-1B and TMBEs-1 targeting the same Mecp2 DNA strand under the guidance of 11 sgRNAs. Data are mean ± SD of three independent replicates. ** indicates p < 0.01 by two-tailed t test.
[0054] Figure 8. Relationship between the mutagenic properties of TMBEs-1B and TMBEs-1 and the mutagenic time
[0055] a, b. Relationship between the mutation rate induced by TMBEs-1B and TMBEs-1 and time;
[0056] c, d. Distribution of TMBEs-1B and TMBEs-1-induced mutations during the 3-15 day detection period.
[0057] e. Relationship between the mutation combination rate (number of mutation combinations / number of reads) generated by TMBEs-1B and TMBEs-1 and the mutagenesis time; f. Relationship between the indel rate generated by TMBEs-1B and TMBEs-1 and the mutagenesis time;
[0058] g. Comparison of the number of mutation combinations between TMBEs-1B and TMBEs-1 at uniform mutation rate.
[0059] Data are mean ± SD from two independent replicates. * indicates p < 0.05 by two-tailed t test. Some samples with poor sequencing quality were excluded when calculating the TMBEs-1B mutation combination rate in panel e.
[0060] Figure 9. Further expansion of mutations to four bases: A, T, C, and G
[0061] a. The nCas9 in TMBEs-1B and TMBEs-1 was replaced by the dCas9 without cleavage activity to obtain TMBEs-1Bd and TMBEs-1d.
[0062] b. Mutation rates triggered by TMBEs-1Bd and TMBEs-1d when sgRNA targets both DNA strands.
[0063] c, e. The relationship between the mutation rate caused by TMBEs-1Bd and TMBEs-1d and time;
[0064] d, f. Distribution of mutations induced by TMBEs-1Bd and TMBEs-1d during the detection period of 3-15 days.
[0065] g. Indels caused by TMBEs-1Bd and TMBEs-1 mutagenesis for 7 days.
[0066] The data in panels b and g are the mean ± SD of three independent replicates. The other data are the mean ± SD of two independent replicates.
[0067] Figure 10 . Mutagenesis of DNA topoisomerase 1 in HEK293T cells to obtain a Topotecan-resistant mutant
[0068] Figure 11. Directed evolution of EGFP to obtain EGFP with enhanced fluorescence intensity (SA)
[0069] a. HEK293T cells were transfected with equal and excess amounts of plasmids containing EGFP and EGFP(SA). 24 hours later, the green fluorescence intensity was analyzed by flow cytometry. The peak fluorescence intensities of EGFP and EGFP(SA) were 5K and 14K, respectively, indicating that the green fluorescence of EGFP(SA) was stronger than that of EGFP.
[0070] b. EGFP and EGFP(SA) were induced to express in BL21 E. coli, and the green fluorescence (excitation light 488 nm, emission light 510 nm) of equal amounts of bacterial suspension was observed.
[0071] c. Excitation spectra of EGFP and EGFP(SA)
[0072] d. Emission spectra of EGFP and EGFP(SA)
[0073] e. EGFP and EGFP(SA) were induced to express in BL21 Escherichia coli, and the visual image of equal amounts of bacterial liquid under blue light with a peak value of 470nm. Detailed Description of the Invention
[0075] definition
[0076] Unless otherwise specified, the terms used herein have the same definitions as those generally understood by one of ordinary skill in the art to which the invention belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the invention.
[0077] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound comprising a base and an acidic portion, such as a nucleoside, a nucleotide, or a polymer of nucleotides. As used herein, the terms "oligonucleotide" and "polynucleotide" are used interchangeably to refer to a polymer of nucleotides (e.g., at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA and single-stranded DNA and / or double-stranded DNA. In some embodiments, RNA is RNA associated with the Cas9 system. For example, RNA can be CRISPR RNA (crRNA), trans-micromolecule RNA (tracrRNA), single guide RNA (sgRNA), or guide RNA (gRNA).
[0078] The term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. A protein may be located at the amino terminus (N-terminus) or at the carboxyl terminus (C-terminus) of the fusion protein, thereby generating an amino-terminal fusion protein or a carboxyl-terminal fusion protein, respectively. The protein may comprise different domains, for example, a nucleic acid binding domain (e.g., a gRNA binding domain of Cas9 that guides protein binding to a target site) and a nucleic acid cleavage domain, or a catalytic domain of a nucleic acid editing protein. In some embodiments, the protein comprises a protein-containing portion, for example, an amino acid sequence constituting a nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid splitting agent. In some embodiments, the protein is complexed or associated with nucleic acids such as RNA or DNA. Any protein provided herein can be manufactured by any method known in the industry. For example, the protein provided herein can be manufactured by recombinant protein expression and purification, which is particularly suitable for fusion proteins comprising peptide linker sequences. Methods for recombinant protein expression and purification are well known, including those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the disclosure of which is incorporated herein by reference in its entirety.
[0079] The term "recombinant" refers to a protein or nucleic acid that does not occur in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0080] The terms "coding sequence" or "protein coding sequence" are used interchangeably herein to refer to a polynucleotide segment that encodes a protein.
[0081] The term "fragment" refers to a portion of a polypeptide or nucleic acid molecule. This portion preferably contains at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of a reference nucleic acid molecule or polypeptide. A fragment can contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0082] As used herein, the term "mutation" refers to a sequence, such as a residue within a nucleic acid sequence or an amino acid sequence, which is replaced by another residue, or a deletion or insertion of one or more residues within the sequence. Typically, mutations herein are described by characterizing the original residue, followed by the residue position within the sequence, and by characterizing the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known, such as those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). In some embodiments, the base editor herein can effectively generate "planned mutations" such as point mutations in nucleic acids (e.g., nucleic acids within the genome of an individual), without generating a significant number of unplanned mutations, such as unplanned point mutations. In some embodiments, the planned mutation is a mutation generated by a specific base editor (e.g., cytosine base editor or adenine base editor) bound to a guide polynucleotide (e.g., gRNA) (which is specifically designed to generate a planned mutation). Typically, mutations in a sequence are numbered relative to a reference (or wild-type) sequence (ie, a sequence without the mutation). Those skilled in the art readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence.
[0083] The term "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid with another amino acid having the same properties. A functional approach to defining the same properties between individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins from homologous organisms (Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer Verlag, New York (1979)). Based on this analysis, groups of amino acids can be defined, within which substitutions of amino acids have the most similar effects on the overall protein structure (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include amino acid substitutions, e.g., lysine for arginine and vice versa to maintain a positive charge; glutamic acid for aspartic acid and vice versa to maintain a negative charge; serine for threonine to maintain a free OH group; and glutamine for asparagine to maintain a free NH2 group.
[0084] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a base editor. In one embodiment, the target site is deaminated by a deaminase or a fusion protein comprising a deaminase (e.g., cytosine deaminase or adenine deaminase).
[0085] The terms "base" or "nitrogenous base" are used interchangeably herein to refer to nitrogen-containing biological compounds that generate nucleosides, which in turn are building blocks of nucleotides. Examples of nucleosides include adenine, guanine, uracil, cytosine, 5-methyluracil (m5U), deoxyadenine, deoxyguanine, thymine, deoxyuracil, and deoxycytosine. Examples of nucleosides with modified bases include inosine (I), xanthosine (X), 7-methylguanine (m7G), dihydrouracil (D), and 5-methylcytosine (m5C).
[0086] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase, which catalyzes the hydrolytic deamination of cytosine or deoxycytosine to uracil or deoxyuracil, respectively. In one embodiment, the cytosine deaminase converts 5-methylcytosine into thymine. The lamprey cytosine deaminase 1, i.e., PmCDA1, derived from the lamprey (Petromyzon marinus), AID (activation-induced cytosine deaminase, AICDA) derived from mammals, and APOBEC are examples of cytosine deaminases. In some embodiments, the deaminase is an adenine deaminase, which catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase or deaminase domain is non-naturally present in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase.
[0087] The terms "Cas9 protein" or "Cas9 nuclease" or "Cas9" are used interchangeably herein to refer to an RNA-guided nuclease that targets a DNA site through RNA:DNA hybridization, so in principle, the nuclease can be targeted to any sequence determined by the guide RNA. When bound to the target, it cuts the complementary strand of the target DNA. The final result of Cas9-mediated DNA breakage is a double-strand break (DSB) within the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). The DSB is then repaired through one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway, or (2) the less efficient but highly fidelity homologous recombination repair (HDR) pathway. It is known that the DNA cleavage domain of Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cuts the complementary strand of the gRNA, while the RuvC1 subdomain cuts the non-complementary strand. Mutations within these subdomains can inhibit the activity of Cas9. Therefore, in some embodiments, Cas9 or Cas9 domains can have active, inactivated, or partially inactivated DNA cleavage domains, and / or gRNA binding domains. For example, nCas9 (Cas9 nickase) is a Cas9 variant that can cause single-strand breaks and BER repair (a repair that does not cause mutations), which can only break one of the two chains in a double-stranded nucleic acid molecule (e.g., DNA) without causing DNA double-strand breaks and NHEJ repair; inactive Cas9 proteins are interchangeably referred to as "dCas9" proteins, which do not have the activity of breaking DNA chains. Methods for generating Cas9 proteins (or fragments thereof) that do not have or partially have DNA cleavage activity are known.
[0088] In some embodiments, the Cas9 or Cas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences as set forth herein. In some embodiments, the Cas9 or Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations compared to any one of the amino acid sequences set forth herein. In some embodiments, the Cas9 or Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical contiguous amino acid residues compared to any of the amino acid sequences set forth herein.
[0089] The term "guide polynucleotide" refers to a polynucleotide that is specific for a target sequence and can form a complex with a nuclease (e.g., Cas9). In one embodiment, the guide polynucleotide is a guide RNA. As used herein, the term "guide RNA (gRNA)" and its grammatical synonyms may refer to an RNA that is specific for a target DNA and can form a complex with a Cas protein. The RNA / Cas complex can help guide the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA cuts the linear or circular target dsDNA complementary to the spacer sequence in an endonucleolytic manner. The target strand that is not complementary to the crRNA is first cut in an endonucleolytic manner and then modified from 3' to 5' in an exonucleolytic manner. In some embodiments, the guide polynucleotide is at least one single guide RNA ("sgRNA" or "gRNA"). In some embodiments, the guide polynucleotide is at least one tracrRNA. Typically, a gRNA present as a single RNA comprises two domains: (1) a domain that shares homology with the target nucleic acid (guiding the Cas9 complex to bind to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA, which comprises a stem-loop structure. In some embodiments, domain (2) is identical or homologous to tracrRNA. In some embodiments, a gRNA comprises two or more of domain (1) and domain (2), which may be referred to as an "extended gRNA". For example, an extended gRNA will bind two or more Cas9 proteins and a target nucleic acid in two or more regions. The gRNA comprises a nucleotide sequence complementary to the target site that mediates binding of the nuclease / RNA complex to the target site, providing sequence specificity to the nuclease:RNA complex. In nature, DNA binding and cleavage typically require a protein and two RNAs. Cas9 recognizes a short motif (PAM or pre-spacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self.
[0090] The guide RNA or guide polynucleotide can be an expression product. For example, the DNA encoding the guide RNA can be a vector comprising a sequence encoding the guide RNA. The guide RNA or guide polynucleotide can be transferred into the cell by transfecting the cell with an isolated guide RNA or a plasmid DNA comprising a sequence encoding the guide RNA and a promoter. The guide RNA or guide polynucleotide can also be transferred into the cell in other ways, such as by viral-mediated gene delivery. The guide polynucleotide can be chemically synthesized, enzymatically synthesized, or a combination thereof.
[0091] The term "base editor" refers to a reagent comprising a polypeptide that can modify a base (e.g., A, T, C, G, or U) inside a nucleic acid sequence (e.g., DNA or RNA). In several embodiments, the base editor is a cytosine base editor (CBE). In several embodiments, the base editor is an adenine base editor (ABE). In several embodiments, the base editor comprises a Cas9 protein fused to a deaminase domain (e.g., adenine deaminase or cytosine deaminase). In several embodiments, the base editor comprises a catalytically dead Cas9 (dCas9) fused to a deaminase domain. In several embodiments, the base editor comprises a Cas9 nickase (nCas9) fused to a deaminase domain. In several embodiments, the base editor comprises a base excision repair (BER) inhibitor fused to a deaminase domain. In several embodiments, the base excision repair inhibitor is a uracil DNA glycosylation inhibitor (UGI).
[0092] The term "base editor system" refers to a system for editing the bases of a target nucleotide sequence. In some embodiments, the base editor system comprises (1) a base editor (BE) comprising a Cas9 nuclease or its encoding nucleic acid sequence and a deaminase domain for deaminating a base; and (2) a guide polynucleotide (e.g., a guide RNA) together with the Cas9 nuclease or its encoding nucleic acid sequence. In some embodiments, the base editor system may include more than one base editing component. For example, the base editor system may include more than one deaminase. In some embodiments, the nuclease base editor system may include one or more cytosine deaminases and / or one or more adenine deaminases.
[0093] The term "base editing activity" refers to the action of chemically changing the bases within a polynucleotide. In one embodiment, the base editing activity is cytosine deaminase activity, for example, converting the target CG to TA. In another embodiment, the base editing activity is adenine deaminase activity, for example, converting the target AT to GC.
[0094] In some embodiments, the base editor system may include multiple guide polynucleotides, such as gRNA. For example, gRNA can target one or more target loci contained in the base editor system (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 g RNAs, at least 50 gRNAs). Multiple gRNA sequences can be arranged in series and preferably separated by repeat sequences. The DNA sequence encoding the guide RNA or the guide polynucleotide can also be part of the vector. The vector may include additional expression control sequences (e.g., enhancer sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., GFP or antibiotic resistance genes, such as puromycin), replication origins, etc. The DNA sequence encoding the guide RNA can be linear or circular. In some embodiments, the nuclease Cas9 or Cas9 domain is used together with one or more gRNAs.
[0095] In some embodiments, one or more components of the base editor system can be encoded by a DNA sequence. These DNA sequences can be introduced into an expression system, such as a cell, together or separately. For example, the coding sequences for each component can be located in different vectors or in the same vector.
[0096] The term "precursor spacer adjacent motif (PAM)" or PAM motif refers to the 26 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. In some embodiments, the PAM can be a 5' PAM (i.e., located upstream of the 5' end of the precursor spacer). In other embodiments, the PAM can be a 3' PAM (i.e., located downstream of the 5' end of the precursor spacer). The PAM sequence is required for target binding, but the exact sequence depends on the type of Cas protein. The base editors provided herein may comprise a CRISPR protein-derived domain that can bind to a nucleotide sequence containing a canonical or non-canonical spacer precursor adjacent motif (PAM) sequence. The PAM site is a nucleotide sequence adjacent to a target polynucleotide sequence. Several aspects of the present invention provide base editors comprising all or part of a CRISPR protein with different PAM specificities. For example, a typical Cas9 protein, such as Cas9 from Streptococcus pyogenes (spCas9), requires a canonical NGG PAM sequence to bind to a specific nucleic acid region, where the N in NGG is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. The PAM can be CRISPR protein specific and can be different between different base editors comprising different CRISPR protein-derived domains. The PAM can be 5' or 3' to the target sequence. The PAM can be upstream or downstream of the target sequence. The PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides in length. Often the PAM is between 2 and 6 nucleotides in length.
[0097] The term "exonuclease" refers to a protein or polypeptide that can digest a nucleic acid (e.g., RNA or DNA) from its free end. The term "endonuclease" refers to a protein or polypeptide that can catalyze (e.g., break) the internal cleavage of a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease can cut a single strand of a double-stranded nucleic acid. In some embodiments, an endonuclease can cut both strands of a double-stranded nucleic acid molecule.
[0098] The terms "nuclear localization sequence," "nuclear localization signal," or "NLS" refer to an amino acid sequence that facilitates the import of a protein into the nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al., International PCT Application No. PCT / EP2000 / 011690, filed November 23, 2000, and published as WO / 2001 / 038547 on May 31, 2001. Fusion proteins comprising a nuclear localization sequence (NLS) can utilize vectors encoding a CRISPR enzyme comprising one or more nuclear localization sequences (NLSs). For example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 NLSs can be used. The CRISPR enzyme can comprise an NLS at or near the amino terminus and about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 NLSs at or near the carboxyl terminus, or any combination thereof (e.g., one or more NLSs at the amino terminus and one or more NLSs at the carboxyl terminus). When more than one NLS is present, each can be independently selected such that a single NLS can be present in more than one copy and / or in combination with one or more other NLSs to be present in more than one copy. The CRISPR enzyme used in the method can comprise about 6 NLSs. An NLS is considered to be near the N-terminus or C-terminus when the amino acid closest to the NLS is within about 50 amino acids along the polypeptide chain from the N-terminus or C-terminus, e.g., within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, or 50 amino acids.
[0099] The term "base excision repair inhibitor" refers to a protein that can inhibit the activity of a nucleic acid repair enzyme (e.g., a base excision repair enzyme). Non-limiting examples of base excision repair (BER) inhibitors include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7 Endol, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the base excision repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI refers to a protein that can inhibit the uracil DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or a fragment thereof. In some embodiments, the base excision repair inhibitor is an inosine base excision repair inhibitor.
[0100] In some embodiments, the base editor system further includes a base excision repair inhibitor component. The components of the base editor system can be associated with each other by covalent bonds, non-covalent interactions, or any combination of their association and interaction. In some embodiments, the base excision repair inhibitor can be targeted to the target nucleotide sequence. In some embodiments, the nuclease is fused or linked to the base excision repair inhibitor. In some embodiments, the nuclease is fused or linked to the deaminase domain and the base excision repair inhibitor. In some embodiments, the base excision repair inhibitor can be targeted to the target nucleotide sequence by guiding the polynucleotide. For example, in some embodiments, the base excision repair inhibitor may include an additional heterologous portion or domain (e.g., a polynucleotide binding domain, such as an RNA or DNA binding protein), which can interact with a portion or a segment (e.g., a polynucleotide motif) of the guide polynucleotide, be associated with, or can form a complex with it. In some embodiments, the additional heterologous portion or domain (e.g., a polynucleotide binding domain, such as an RNA or DNA binding protein) of the guide polynucleotide can be fused or linked to the base excision repair inhibitor. In some embodiments, the additional heterologous portion can be bound to a polypeptide linking sequence. In some embodiments, the additional heterologous moiety can be bound to the linker sequence.
[0101] As used herein, the term "linker" may refer to a covalent linker sequence (e.g., a covalent bond), a non-covalent linker sequence, a chemical group, or a molecule that links two molecules or parts, such as two components of a protein complex or a ribonucleic acid complex, or two domains of a fusion protein. The linker sequence can connect different components of a base editor system or different parts of a component. For example, in some embodiments, the linker sequence can connect a CRISPR polypeptide to a deaminase. In some embodiments, the linker sequence can connect Cas9 to a deaminase. In some embodiments, the linker sequence can connect dCas9 to a deaminase. In some embodiments, the linker sequence can connect nCas9 to a deaminase. The linker sequence may be located between two groups, molecules, or other parts, or flanked by two groups, molecules, or other parts, and connected to each other by covalent bonds or non-covalent interactions. In some embodiments, the linker sequence may be a polynucleotide. In some embodiments, the linker sequence may be a DNA linker sequence. In some embodiments, the polypeptide linker sequence may be part of a base editor system component. For example, a base editing component may include a deaminase domain and an RNA recognition motif.
[0102] In some embodiments, the linking sequence can be a peptide or protein. In some embodiments, the linking sequence can be about 5100 amino acids in length, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90 or 90-100 amino acids in length. In some embodiments, the linking sequence can be about 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450 or 450-500 amino acids in length. Longer or shorter linking sequences are also contemplated. In some embodiments, the linker sequence comprises multiple proline residues and is 5-21, 5-14, 5-9, or 5-7 amino acids long, for example, PAPAP, PAPAPA, PAPAPAP, PAPAPAPA, P(AP)4, P(AP)7, or P(AP)10. Such proline-rich linkers are also known as "rigid" linkers. In some embodiments, the domains of the base editor are fused via a linker sequence.
[0103] In some embodiments, a reporter system can be used to detect base editing activity and test candidate guide polynucleotides. In some embodiments, the reporter system may include an analytical test based on a reporter gene, in which the base editing activity results in the expression of a reporter gene. Non-limiting examples of reporter genes include genes encoding green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secreted alkaline phosphatase (SEAP) genes, or any other genes. The reporter system can be used to test many different gRNAs to determine which residues the deaminase will target relative to the target DNA sequence. In some embodiments, the guide polynucleotide may include at least one detectable label. The detectable label can be a fluorescent group (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo tags, or suitable fluorescent dyes), a detection label (e.g., biotin, digoxin, etc.), a quantum particle, or a gold particle.
[0104] According to the nucleic acid encoding the base editor disclosed herein, it can be administered to an individual or delivered into a cell in vitro by methods known in the industry or as described herein. In one embodiment, the base editor is selectively delivered to cells of the liver, lungs, or any other organ and their progenitor cells. In a specific embodiment, the edited cells can be used to analyze the functional effects of test gene editing on the function of the encoded protein. In one embodiment, the base editor can be delivered by, for example, a vector (e.g., a viral or non-viral vector), a non-vector-based method (e.g., using naked DNA, DNA complexes, lipid nanoparticles), or a combination thereof. The nucleic acid encoding the base editor can be delivered directly to the cells of the liver, lungs, or any other organ as naked DNA or naked RNA, for example, by transfection or electrophoresis; or it can be connected to a molecule that promotes uptake by the target cells (e.g., N-acetylgalactosamine). Nucleic acid vectors, such as those described herein, can also be used.
[0105] The base editors of the present invention can be used to correct point mutations in disease-related genes and alleles, and thus applied to therapeutics and basic research. In this case, site-specific mutation residues that cause inactivation of protein mutations, or mutations that inhibit protein function, can be used to eliminate or inhibit protein function. The present invention provides a method for treating an individual suffering from a disease associated with or caused by a point mutation, which can be corrected by the base editor system provided herein. For example, in some embodiments, a method is provided, which comprises administering an effective amount of a base editor (e.g., an adenine deaminase base editor or a cytosine deaminase base editor) to an individual suffering from such a disease, such as a disease caused by a gene mutation, which introduces an inactivating mutation into a disease-related gene. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a tumor disease. In some embodiments, the disease is a metabolic disease.
[0106] The base editor of the present invention can also be used to evolve proteins by introducing mutations into proteins, thereby changing the function of the protein or improving the original function of the protein.
[0107] The term "effective amount" refers to the amount of an agent or active compound (e.g., a base editor as described herein) required to improve the symptoms of a disease compared to an untreated patient. The effective amount of the active compound used in the present invention for therapeutic treatment of a disease varies depending on the mode of administration, individual age, weight, and general health. Ultimately, the clinician or veterinarian will determine the appropriate usage and dosage. This amount is referred to as an "effective" amount. In one embodiment, the effective amount is an amount of the base editor of the present invention sufficient to introduce a genetic change into a cell (e.g., a cell in vitro or in vivo).
[0108] The term "patient" or "individual" or "subject" refers to a mammalian individual or person diagnosed with, at risk of developing, or suspected of having or developing a disease or condition. In some embodiments, the term "patient" refers to a mammalian individual with a higher than average chance of developing a disease or condition. Examples of patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, alpacas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. DETAILED DESCRIPTION
[0109] The present invention is further described in detail below with reference to the examples. These examples are intended to illustrate the present invention only and are not to be construed as limiting the scope of the present invention. Modifications or improvements made without departing from the present invention are intended to fall within the scope of the present invention. Unless otherwise specified, the experimental methods used in the following examples are conventional methods; the reagents and materials used are all commercially available.
[0110] Example
[0111] Example 1 Construction of a targeted mutagenesis system based on a dual base editor
[0112] The inventors used adenine base editors and cytosine base editors (Komor, AC, Kim, YB, Packer, MS, Zuris, JA & Liu, DR Programmable editing of a target base ingenomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016); Gaudelli, NMet al. Programmable base editing of A*T to G*C in genomic DNAwithout DNA cleavage. Nature 551, 464-471 (2017); Nishida, K. et al. Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. Science 353, aaf8729-aaf8729 (2016); Koblan, LWet al. Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat Biotechnol 36,843-846(2018)) based on which various forms of double base editors were constructed, named TMBEs ( T argeted M utagenesis system based on B ase E ditors) as the basis for further improvement to obtain dual-base editors with high mutagenic activity. These TMBEs were constructed using NEBuilder HiFi DNA Assembly Master Mix (NEB#E2621X), and the specific construction process was carried out according to the product instructions.
[0113] The sequence of the constructed base editor TMBEs-1 is shown in SEQ ID NO: 1, wherein positions 1-1179 in SEQ ID NO: 1 are EF-1-alpha promoter, positions 1206-1262 are bpNLS, positions 1263-2354 are TadA-TadA*, positions 2355-2450 are linker, positions 2451-6551 are nCas9, positions 6552-6614 are bpNLS, positions 6615-6710 are linker, positions 6711-7337 are PmCDA1, positions 7368-7895 are 2*UGI, and positions 7896-7958 are bpNLS. PmCDA1 in TMBEs-1 was replaced with AncAPOBEC1 (SEQ ID NO: 10) to obtain TMBEs-3; AncAPOBEC1 and TadA-TadA* in TMBEs-3 were exchanged to obtain TMBEs-2; AncBE4max and ABEmax (Koblan, LW et al. Improving cytidine and adenine base editors by expression optimization and ancestralreconstruction. Nat Biotechnol 36, 843-846 (2018)) were combined into the same plasmid in the form of different open reading frames to obtain TMBEs-4 ( Figure 1a ).
[0114] In the dual-base editor-based targeted mutagenesis system constructed in the present invention (TMBEs-1d is shown in Example 7), multiple sgRNAs are set on the same DNA strand, and nCas9 is used to cut the non-editing strand. The present invention proves that such a setting can achieve efficient targeted mutagenesis and avoid a large number of double-strand breaks ( Figure 1b , Figure 2a ,b).
[0115] Example 2 Construction of multi-sgRNA expression vector
[0116] The Csy4 RNase in Pseudomonas aeruginosa can process CRISPR-derived RNA in bacteria (Haurwitz, R.E., Jinek, M., Wiedenheft, B., Zhou, K. & Doudna, J.A. Sequence-and structure-specific RNA processing by a CRISPR endonuclease. Science 329, 1355-1358 (2010)). Therefore, the inventors used Csy4 to construct a system for expressing multiple sgRNAs from a single RNA polymerase II promoter ( Figure 3a ). The NEB Golden Gate Assembly Kit (BsaI-HFv2) (NEB#E1601) was used and the product instructions were followed to construct a multi-sgRNA expression vector 1 capable of expressing four sgRNAs (wherein the Golden Gate site in SEQ ID NO: 2 was replaced with SEQ ID NO: 3, and the sequences at positions 1-20, 117-136, 233-252, and 349-368 in SEQ ID NO: 3 are Protospacer sequences, i.e., targeting sequences, and the remaining sequences are gRNA scaffold sequences with a Csy4 site). The sequence of the multi-sgRNA expression vector skeleton used was SEQ ID NO: 2 (positions 832-849 are Golden Gate sites).
[0117] AncBE4max was co-transfected with the base editor in ABEmax and the multi-sgRNA expression vector 1 into HEK293T cells. The specific transfection method is as follows: one day in advance, an appropriate amount of cells were seeded so that the cell density at the time of transfection was about 70%, and fresh culture medium was replaced 1 hour before transfection. Transfection was performed using 3000 Reagent (Thermo Fisher Scientific, L3000) according to the product instructions. Fresh culture medium was replaced 12 hours after transfection before subsequent culture and experiments.
[0118] Two days after co-transfection into HEK293T cells, cells expressing green fluorescence were sorted using flow cytometry (indicating successful plasmid transfection). Sorted cells were cultured until day 7, after which Sanger sequencing of the targeted region was performed. Culture conditions: 5% CO2, 37°C, basal medium (DMEM (Invitrogen) plus 10% FBS (Invitrogen)).
[0119] Experiments have shown that with the help of Cys4 elements, a single RNA polymerase II promoter can express multiple sgRNAs and be used for base editing ( Figure 3b , all multi-sgRNAs in this experiment were expressed using this system).
[0120] Example 3 Mutagenic Activity of Dual-Base Editors
[0121] The multi-sgRNA expression vector 1 and the PB transposase vector (System Biosciences, PB210PA-1) were co-transfected into mESC (R1 mouse embryonic stem cells, ATCC) to construct an mESC cell line stably expressing four sgRNAs. The various dual-base editors constructed in Example 1 were transfected into this cell line according to the method of Example 2. After 2 days, cells containing green fluorescence were sorted by flow cytometry (indicating successful plasmid transfection). The sorted cells were cultured for 7 days, and then high-throughput amplicon sequencing of the targeted region was performed.
[0122] The cell culture conditions were: 5% carbon dioxide, 37°C, mouse embryonic stem cell culture medium (Knock-Out-DMEM (Invitrogen) supplemented with 15% FBS (Invitrogen), 1% GlutaMAX™, 1% NEAA, 0.1 mM 2-mercaptoethanol (Sigma Aldrich), 10 ng / ml leukemia inhibitory factor (LIF, Millipore), 3 mM CHIR99021 (Selleck), and 1 mM PD0325901 (Selleck)).
[0123] The specific process of high-throughput amplicon sequencing is as follows:
[0124] After obtaining mutagenized genomic DNA, a high-throughput sequencing library was constructed using a two-step PCR method. All PCRs were performed in a 50 μl system using 500-1000 ng of DNA as template.
[0125] For the first round of PCR, specific primers were designed based on the sequence of the targeted region. In addition to binding to the target region, these primers also contain a partial adapter sequence at their 5' end (forward: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3'; reverse: 5'-GTCTCGTGGGCTCGGAG ATGTGTATAAGAGACAG-3'). After approximately 26 cycles of PCR amplification, purification and recovery were performed using Ampure Beads (Beckman Coulter) according to the product instructions, with the exception of using 0.8 times the volume of Ampure Beads relative to the sample volume to facilitate the removal of unreacted primers.
[0126] The second round of PCR incorporated the Illumina sequencing adapters and indexes into the amplicon. The first-round PCR product was amplified at a low cycle rate using the following primers (forward: 5'-AATGATACGGCGACCACCGAGATCTACACNNNNNNNNTCGTCGGCAGCGTC-3'; reverse: 5'-CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTCTCGTGGGCTCGG-3', where NNNNNNNN represents the Illumina index). Six cycles of amplification were performed using 10 μl of the first-round PCR product as template. The product was then purified using Ampure Beads (Beckman Coulter) as in the first round of PCR and subjected to PE250 high-throughput sequencing.
[0127] The sequencing data was first processed using fastp software to remove reads with quality scores below Q30. Double-ended 250bp reads were then fused into a single sequence based on the overlapping regions. Nonspecific reads were removed using samtools, and preliminary data preparation and quality control were performed using fastqc software. The reference sequence was indexed using bwa software, and the quality-checked sequences were aligned. Aligned data were generated using samtools, and a Python script was used to statistically analyze the mutation rate and diversity within the targeted region.
[0128] The experimental results showed that different forms of TMBEs all have strong mutagenic activity, but the types of mutagenesis vary greatly. Directly fusing the adenine deaminase TadA-TadA* and the cytosine deaminase AncAPOBEC1 with nCas9 (TMBEs-2, TMBEs-3) can better exert the function of cytosine base editors, but the function of adenine base editors is severely impaired ( Figure 4 ,mutation rates<10 -1Using different open reading frames to express two base editors (TMBEs-4) or using PmCDA1 cytosine deaminase instead of AncAPOBEC1 cytosine deaminase (TMBEs-1) can more completely express the characteristics of dual base editors ( Figure 1c , Figure 4 TMBEs-1 and TMBEs-4 can simultaneously and efficiently induce mutations in A and C (mutation rates>10 -1 ).
[0129] TMBEs-1 has a simpler structure, so the inventors chose TMBEs-1 for subsequent experiments.
[0130] Example 4 Editing efficiency of dual-base editors in human endogenous genomes
[0131] To further verify the function of the dual-base editor obtained previously, the inventors transfected TMBEs-1 with different sgRNAs (see Table 1) into HEK293T cells to target 15 different endogenous sites.
[0132] The process was as follows: The dual-base editor with the corresponding sgRNA was transfected into HEK293T cells as described in Example 2. After 2 days, cells expressing green fluorescence (indicating successful plasmid transfection) were sorted using flow cytometry. The sorted cells were cultured for 7 days, and then high-throughput amplicon sequencing of the targeted region was performed as described in Example 2. Culture conditions: 5% CO2, 37°C, basal medium.
[0133] TMBEs-1 induced A>G (maximum, 70.6%) and C>T (maximum, 81.6%) editing in all targeted gene sites, and no detectable changes occurred in T and G within the targeted range. TMBEs-1 with an off-target sgRNA of SEQ ID NO: 8 (TMBEs-1-U6-sgRNA, a vector expressing TMBEs-1 and a single sgRNA (off-target sgRNA) formed by inserting SEQ ID NO: 8 into the TMBEs-1 vector) also did not cause detectable mutations ( Figure 5a ), indicating that the obtained dual-base editor can target and specifically edit A and C. Among the 15 target sites tested, the mutations of A and C induced by TMBEs-1 mainly occurred at positions 4-8 and 1-7, respectively, and a small number of C changes also occurred at positions 11-14 ( Figure 5b,c), which is not much different from the activity window of ABEmax and Target-AID (Nishida, K. et al. Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. Science 353, aaf8729-aaf8729 (2016)) (PAM is counted as 21-23). Among the mutations caused by TMBEs-1, 99.63% of A mutated to G and 94.14% of C mutated to T ( Figure 5d ,e).
[0134] Table 1: Protospacer sequences
[0135] DDB2-1 aatattcaagcagcaggcac GRIN2B-1 ggcattgctgtcatcctcgt DDB2-2 ctcgcgcaggaggctgcagc GRIN2B-2 tgacagcaatgccaatgctg FANCF-1 tggaggcaagagggcggctt GRIN2B-3 ttccgacgaggtggccatca FANCF-2 cgctccagagccgtgcgaat GRIN2B-4 tgaccggaagatccaggggg FES ccagctgctgccttgcctcc DYRK1A-1 gccaaacataagtgaccaac EMX1-1 caaacggcagaagctggagg DYRK1A-2 tcagcaacctctaactaacc EMX1-2 tgagtccgagcagaagaaga DYRK1A-3 ggtcactgtactgatgtgaa EMX1-3 agggctcccatcacatcaac
[0136] Example 5: Improving the dual-base editor to induce diverse mutations
[0137] As verified in Examples 3 and 4, although the TMBEs-1 has good programmability, specificity and efficient editing function, it mainly induces mutation types of C>T and A>G. In order to further increase the diversity of mutations, the inventors tried to increase the diversity of mutations by manipulating the DNA repair pathway in the cell. According to existing literature, the possible results after C is deaminated to U are: U is read as T; U is cut by uracil glycosylase in the cell to form an abasic site (AP) and then randomly read as A, T, C or G; U is removed by uracil glycosylase in the cell to form an abasic site, and AP is then removed by AP enzyme in the cell to introduce Indels ( Figure 6a ). Removing the UGI component within TMBEs-1 to generate TMBEs-1B can promote the removal of U to generate more APs. Therefore, the inventors removed the UGI component in TMBEs-1 to test whether the resulting dual-base editor (called TMBEs-1B) can increase mutation diversity.
[0138] TMBEs-1B and TMBEs-1 were transfected into mESCs stably expressing four sgRNAs (same as in Example 3, using multi-sgRNA expression vector 1). After two days, cells expressing green fluorescence were sorted using flow cytometry (indicating successful plasmid transfection). The sorted cells were cultured until day 7, after which high-throughput amplicon sequencing of the targeted region was performed. Culture conditions: 5% CO2, 37°C, in mouse embryonic stem cell culture medium.
[0139] High-throughput amplicon sequencing revealed that although TMBEs-1B introduced a large number of Indels compared to TMBEs-1 ( Figure 6c ), but did increase the C>G and C>A mutation types ( Figure 6b ).
[0140] Next, in order to reduce the occurrence of Indels, the inventors stably transferred TMBEs-1B and TMBEs-1 into mESC cells and mESC-Apex1 (an mESC cell line in which the AP enzyme Apex1 was knocked out by methods well known in the art), respectively, to obtain two tool cells for mutagenesis. The multi-sgRNA expression vector 2 targeting EGFP and expressing 11 sgRNAs (obtained by replacing the Golden Gate site in SEQ ID NO: 2 with SEQ ID NO: 4, where the sequences at positions 1-20, 117-136, 233-252, 349-368, 465-484, 581-600, 697-716, 813-832, 929-948, 1045-1064, and 1161-1180 in SEQ ID NO: 4 are protospacer sequences, i.e., targeting sequences, and the remaining sequences are gRNA scaffold sequences with a Csy4 site) and a PB transposase vector (System Biosciences, PB210PA-1) were co-transfected into the above-mentioned cells. Puromycin was added at a final concentration of 1 μg / ml for selection for 2 days to ensure that the plasmid entered the cells, and then the cells were replaced with a 0.5 μg / ml plasmid containing 5 μg / ml daptomycin. Puromycin-containing culture medium was continued to culture until day 7, and high-throughput amplicon sequencing of the targeted region was performed.
[0141] Culture conditions: 5% carbon dioxide, 37°C, mouse embryonic stem cell culture medium.
[0142] High-throughput amplicon sequencing results showed that although indels were not completely eliminated, knockout of Apex1 reduced indels triggered by TMBEs-1B. Figure 6d ). Both TMBEs-1B and TMBEs-1 can efficiently introduce mutations at the target site ( Figure 6e ), and TMBEs-1B has more C>G and C>A mutation types than TMBEs-1 ( Figure 6f ), in addition, TMBEs-1B also increased the A>C and A>T changes at the A site ( Figure 6f ).
[0143] To further confirm the mutagenic properties of TMBEs-1B in mESC-Apex1, the inventors used a multi-sgRNA expression vector 5 targeting the same chain of Mecp2 (obtained by replacing the Golden Gate site in SEQ ID NO: 2 with SEQ ID NO: 7, and the sequences at positions 1-20, 117-136, 233-253, 350-369, 466-485, 582-601, 698-717, 814-833, 930-949, 1046-1065, and 1162-1181 in SEQ ID NO: 7 are Protospacer sequences, i.e., targeting sequences, and the remaining sequences are gRNA scaffold sequences with Csy4 sites) to express 11 sgRNAs to target and mutagenesis the endogenous Mecp2 gene in cells. The experimental procedure was as follows: TMBEs-1B and TMBEs-1 were stably introduced into mESC-Apex1 and mESC cells, respectively, to generate two tool cells for mutagenesis. These cells were then co-transfected with a multi-sgRNA expression vector 5 targeting Mecp2 and a PB transposase vector. Puromycin was added for selection at a final concentration of 1 μg / ml for two days to ensure plasmid entry. The cells were then cultured in a medium supplemented with 0.5 μg / ml of Puromycin and cultured for 7 days. High-throughput amplicon sequencing of the targeted region was then performed. Culture conditions: 5% CO2, 37°C, in mouse embryonic stem cell culture medium.
[0144] Consistent with the previous results, although a small number of Indels (such as Figure 7c ), but TMBEs-1B can induce more mutation types on A and C than TMBEs-1 (such as Figure 7a ,b).
[0145] Therefore, the inventors successfully expanded the diversity of mutations on A and C by manipulating the intracellular DNA repair mechanism (removing UGI and knocking out Apex1).
[0146] Example 6 Effect of mutagenesis time on the mutagenic properties of dual-base editors
[0147] In order to determine the optimal mutagenesis duration, the inventors detected the characteristics of TMBEs-1B and TMBEs-1 mutagenesis within 3 to 15 days. TMBEs-1B and TMBEs-1 were stably transferred into mESC-Apex1 and mESC cells, respectively, to obtain two tool cells for mutagenesis. The multi-sgRNA expression vector 2 targeting EGFP and the PB transposase vector (System Biosciences, PB210PA-1) were co-transfected into the above cells, and Puromycin with a final concentration of 1ug / ml was added for screening for 2 days to ensure that the plasmid entered the cells. The culture medium was then changed to 0.5ug / ml Puromycin and cultured until the 15th day. Sampling was performed every two days starting from the 3rd day, and high-throughput amplicon sequencing was performed on the targeted region. Culture conditions: 5% carbon dioxide, 37°C, mouse embryonic stem cell culture medium.
[0148] The mutation rate, mutation combination rate, and indel rate induced by TMBEs-1B and TMBEs-1 increased with the increase of mutagenesis time and reached the maximum value on the 7th day of mutagenesis (TMBEs-1B: mutation rate 3.02x10 -3 , mutation combination rate 4.23%, Indel rate 27.95%; TMBEs-1: mutation rate 2.99x10 -2 , mutation combination rate 13.48%, Indel rate 9.01%). If the mutagenesis time is further extended, the mutation rate, mutation combination rate and Indel rate will decrease instead ( Figure 8a ,b,e,f). This test further confirmed the above results, namely that TMBEs-1B increased the number of C and A mutation types relative to TMBEs-1 at all time points ( Figure 8c ,d), and TMBEs-1B also induced a small number of T and G mutations in all test points ( Figure 8a ,c). Because the mutation rate induced by TMBEs-1B is smaller than that of TMBEs-1, the corresponding mutation combination rate produced by TMBEs-1B is also smaller than that produced by TMBEs-1 ( Figure 8e ), but because TMBEs-1B can induce more types of mutations ( Figure 8c ,d), TMBEs-1B can produce more mutation combinations under the same mutation rate ( Figure 8g ).
[0149] Example 7 Further expansion of mutation types
[0150] Since all sgRNAs are set on the same DNA chain, although TMBEs-1B can introduce a small number of mutations on T and G, the mutations introduced by TMBEs-1B and TMBEs-1 are mainly on A and C. In order to further expand the mutation range and enable mutations in all bases within the target range, the inventors tried to replace the nCas9 in TMBEs-1B and TMBEs-1 with dCas9 (SEQ ID NO: 9) without cutting activity, and obtained TMBEs-1Bd and TMBEs-1d ( Figure 9a The inventors designed a multi-sgRNA expression vector 4 that can express 21 sgRNAs (obtained by replacing the Golden Gate site in SEQ ID NO: 2 with SEQ ID NO: 6, SEQ ID The sequences at positions 1-20, 117-136, 233-252, 349-368, 465-484, 581-600, 697-716, 813-832, 929-948, 1045-1064, 1161-1180, 1317-1336, 1433-1452, 1549-1568, 1665-1684, 1781-1800, 1897-1916, 2013-2032, 2129-2148, 2245-2263, and 2360-2379 in NO:6 are protospacer sequences, i.e., targeting sequences, and the remaining sequences are gRNAs with Csy4 sites. scaffold sequence) to target the two DNA chains of EGFP, and both TMBEs-1Bd and TMBEs-1d can introduce mutations at the four bases of A, T, C and G ( Figure 9b ).
[0151] To further understand the relationship between the mutagenic properties of TMBEs-1Bd and TMBEs-1d and the duration of mutagenesis, the inventors examined the mutagenic characteristics of TMBEs-1Bd and TMBEs-1d over a period of 3 to 15 days. The specific process involved stably introducing TMBEs-1Bd and TMBEs-1d into mESC-Apex1 and mESC cells, respectively, to obtain two tool cells for mutagenesis. The EGFP-targeting multi-sgRNA expression vector 4 and the PB transposase vector were co-transfected into these cells. Puromycin was added at a final concentration of 1 μg / ml for two days to ensure plasmid entry into the cells. The cells were then switched to a medium containing 0.5 μg / ml Puromycin and cultured until day 15. Starting on day 3, samples were collected every two days, and high-throughput amplicon sequencing of the targeted region was performed. Culture conditions: 5% carbon dioxide, 37°C, in mouse embryonic stem cell culture medium.
[0152] The results showed that the mutation rate of TMBEs-1Bd first increased and then decreased with the extension of mutagenesis time, reaching a maximum of 1.89x10 -3 ( Figure 9c ), while the TMBEs-1d mutation rate reached a maximum of 4.66x10 on day 11. -3 , the mutation rate will no longer decrease if the mutagenesis time is extended ( Figure 9e The inventors speculate that the reason is that TMBEs-1d induces less DNA breakage ( Figure 9g ), so the survival of cells undergoing mutagenesis will not be significantly reduced compared to cells not undergoing mutagenesis. Both TMBEs-1Bd and TMBEs-1d can introduce mutations at the four bases of A, T, C, and G ( Figure 9c ,e), and can be converted into any other base in different proportions ( Figure 9d ,f).
[0153] In summary, the inventors obtained four mutagenesis systems, namely TMBEs-1B, TMBEs-1, TMBEs-1Bd and TMBEs-1d, with different mutation spectra.
[0154] Example 8 Obtaining Resistant Mutants Using TMBEs
[0155] The foregoing demonstrates that the present invention has generated mutagenesis systems with diverse mutational profiles. The inventors subsequently tested the ability of these mutagenesis systems to generate Topotecan-resistant DNA topoisomerase 1 (Top1) mutants, using the DNA topoisomerase 1 (Top1) gene as an example. DNA topoisomerase 1 in eukaryotes is the specific target of clinically used anticancer drugs such as Topotecan and its analog, Camptothecin. Topotecan binds to Top1, stabilizing the Top1-DNA complex, leading to double-strand breaks and ultimately cell death.
[0156] The inventors constructed a multi-sgRNA expression vector 3 expressing 11 sgRNAs (obtained by replacing the Golden Gate site in SEQ ID NO: 2 with SEQ ID NO: 5, where the sequences at positions 1-20, 117-136, 233-252, 349-368, 465-484, 581-600, 697-716, 813-832, 929-948, 1045-1064, and 1161-1180 in SEQ ID NO: 5 are Protospacer sequences, i.e., targeting sequences, and the remaining sequences are gRNAs with Csy4 sites. The scaffold sequence (the 11 sgRNAs target the same strand of TOP1) was used to target the area surrounding the reported Topotecan-resistant mutation. This multi-sgRNA expression vector was co-transfected with TMBEs-1 and TMBEs-1B vectors into HEK293T cells. Puromycin was added at a final concentration of 1 μg / ml for 3 days to ensure plasmid entry. Cells were switched to normal culture medium and cultured until no cells survived in the negative control group and resistant mutants were obtained in the mutagenesis group. Single clones of resistant mutants were isolated and expanded. Drug resistance was further verified by adding 50 nM Topotecan HCl or 50 nM Camptothecin to the resistant mutants. The DNA topoisomerase 1 gene within the mutants was amplified by PCR, and the presence of the DNA topoisomerase 1 mutation was confirmed by Sanger sequencing. Culture conditions: 5% CO2, 37°C, basal medium.
[0157] Hundreds of resistant mutant clones were obtained, and 6 clones were selected for sequencing verification of DNA topoisomerase 1. These 6 clones contained 4 mutation combinations (m1, m2, m3 and m4) ( Figure 10 After re-culturing these resistant mutants and adding Camptothecin, it was found that m1, m2, and m3 were resistant to Camptothecin, while m4 was sensitive to Camptothecin ( Figure 10 ).
[0158] The results of this study confirm that the TMBEs of the present invention can indeed induce the generation of corresponding drug-resistant mutants. Preemptive research on these drug-resistant mutants can provide timely and feasible treatment options when resistant mutants emerge clinically. Furthermore, the results of this study also demonstrate that even drugs with the same target can have different effects on the same mutant, suggesting that preemptive classification of resistant mutants can facilitate precision medicine.
[0159] Example 9 Directed Evolution of Proteins Using TMBEs
[0160] In directed protein evolution, the function of the target protein is often coupled to fluorescence or adsorption. However, screening methods that rely on fluorescence or adsorption are often inefficient and cumbersome, making many mutagenesis methods difficult to implement. To test whether TMBEs are also applicable to these unselectable phenotypes, the inventors attempted to use TMBEs to generate several new fluorescent proteins, using EGFP as an example.
[0161] The inventors transformed the previously constructed EGFP-targeting multi-sgRNA expression vectors 2 and 4 into mESCs and mESC-Apex1 cells constitutively expressing TMBEs (TMBEs-1, TMBEs-1d, and TMBEs-1B, TMBEs-1Bd), respectively. Puromycin was added at a final concentration of 1 μg / ml for three days to ensure plasmid entry into the cells, and then cultured in fresh, antibiotic-free medium. Seven days after mutagenesis, all cells expressing fluorescence that differed from the EGFP spectrum or exhibited enhanced fluorescence intensity were sorted using flow cytometry. These cells were then expanded and sorted a second time. The cells obtained from the second sort were individually expanded, the corresponding genes were amplified, and subsequent validation was performed. Culture conditions: 5% CO2, 37°C, mouse embryonic stem cell culture medium.
[0162] After two rounds of continuous sorting, the inventors obtained a blue fluorescent protein derived from EGFP mutation (not shown), and fluorescence-enhanced EGFP (SA) (an EGFP mutant in which serine at position 72 was converted to alanine) ( Figure 11a ,b,e). The excitation light and emission light of EGFP(SA) are red-shifted by about 6nm and 2nm respectively relative to EGFP ( Figure 11c ,d),Adjusting the excitation and emission light can further enhance the fluorescence intensity of EGFP(SA).,EGFP(SA) will be more suitable for labeling low-level expressed proteins.
[0163] In summary, the targeted mutagenesis system based on the adenine and cytosine double base editor proposed in the present invention has strong mutagenic activity and has good application prospects in the directed evolution of proteins, such as the generation of high-affinity antibodies by mutagenizing the variable regions of antibodies.
Claims
1. A dual-base editor, which is (1) based on a dual-base editor having a nucleic acid sequence of SEQ ID NO: 1, or an amino acid sequence encoded by SEQ ID NO: 1, the nCas9 component at positions 2451-6551 is replaced with a dCas9 component having a sequence of SEQ ID NO: 9; or (2) Based on a dual-base editor having a nucleic acid sequence of SEQ ID NO: 1 or an amino acid sequence encoded by SEQ ID NO: 1, the UGI component at positions 7368-7895 is removed and the nCas9 component at positions 2451-6551 is replaced with a dual-base editor having a dCas9 component having a sequence of SEQ ID NO:
9.
2. A vector comprising a nucleic acid sequence encoding the dual-base editor of claim 1.
3. A tool cell, which is transfected with the dual-base editor according to claim 1 or the vector according to claim 2, wherein the tool cell is not an embryonic stem cell. The tool cell according to claim 3 , which is a HEK293T cell.
5. A targeted mutagenesis system for targeted mutagenesis of proteins, the targeted mutagenesis system comprising the dual-base editor of claim 1 or the vector of claim 2, wherein the guide polynucleotide contained in the targeted mutagenesis system targets a target region of the protein coding sequence to be mutagenized.
6. The targeted mutagenesis system according to claim 5, wherein the guide polynucleotide is constructed in a multi-sgRNA expression vector with a sequence of SEQ ID NO:
2.
7. A kit for mutagenesis or directed evolution of proteins, comprising: (1) the dual-base editor of claim 1 or its encoding nucleic acid sequence or the vector of claim 2, or the tool cell of claim 3 or 4, or the targeted mutagenesis system of claim 5 or 6.
Citation Information
Patent Citations
Polypeptides comprising multimers of nuclear localization signals or of protein transduction domains and their use for transferring molecules into cells
WO2001038547A2
Composition for modifying nucleotide sequence, method and application
CN109517841A
Novel base conversion editing system and application thereof
CN110835634A
Multi-effector nucleobase editors and methods of using same to modify a nucleic acid target sequence
CN112805379A
Base editing system and use method therefor
WO2021032155A1