Single base editors and deaminases used therewith and applications
By constructing a cytosine base editor that integrates the protein cytidine deaminase L8, Cas protein, and uracil glycosylation inhibitor, the problem of low editing efficiency of DddA deaminase was solved, and efficient CG to TA base editing was achieved in maize cells.
Patent Information
- Application Number
- CN202410776475.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-06-17
AI Technical Summary
Existing DddA deaminase-based cytosine base editing technologies have low editing efficiency in maize organelles or nuclei, making it difficult to achieve efficient CG to TA base editing.
A fusion protein containing cytidine deaminase L8, Cas protein, and uracil glycosylation inhibitor was constructed. By linking with nuclear localization signals, a cytosine base editor was formed, improving editing efficiency and scope.
It significantly improves the efficiency of cytosine base editing technology, expands the editing activity window, and enables precise base mutations in maize and other plant cells.
Smart Images

Figure BDA0004895819110000081 
Figure BDA0004895819110000091 
Figure BDA0004895819110000101
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of genetic engineering, and particularly relates to a single-base editor and a deaminase used by the same and application thereof. BACKGROUND
[0002] Genome editing technology is a kind of genetic engineering technology for modifying a specific position of a genome of an organism through a sequence-specific nuclease. The process is to generate a double-strand break (DSB) on the genome, thereby triggering an endogenous repair mechanism of the cell, such as a non-homologous end joining repair or a homologous recombination repair, and generating sequence insertion, deletion and replacement, etc. At present, commonly used genome editing technology platforms include ZFN, TALEN and CRISPR / Cas system. Among them, the CRISPR / Cas genome editing technology platform is the most efficient.
[0003] Single-nucleotide polymorphism (SNP) is the genetic basis of the agronomic traits of crops, and many important changes in crop traits are usually caused by a single base variation. At present, new target gene site-directed modification base editing systems based on the CRISPR system, such as cytosine base editing system and adenine base editing system, are developed and widely used in various organisms such as animals and plants. The cytidine deaminases that have been successfully applied to plant base editing systems include APOBEC1, APOBEC3, DddA, etc., and the adenosine deaminases include TadA7.10, TadA8e, etc.
[0004] Since the DddA cytidine deaminase is toxic, the cytosine base editing system based on the DddA deaminase mainly uses two different Cas9 or two TALEN, combined with a split DddA, to form a fusion protein. Under the guidance of sgRNA, only when the target gene region is combined, the two halves will restore catalytic activity, deaminating the C base in the editing window to dU, and through DNA replication and repair, the base editing of C-G to T-A in the organelle and the nucleus of the target gene is finally realized. However, the cytosine base editing technology based on the DddA deaminase still has the disadvantage of low editing efficiency in the organelle or nucleus of corn at present. SUMMARY
[0005] The technical problem to be solved by the present application is to mine natural sequence-unpreference cytosine deaminases, construct a base editor, and improve the base editing efficiency and range.
[0006] To solve the above technical problems, the present application provides a cytosine base editor, which is a fusion protein.
[0007] The fusion protein is a protein containing cytidine deaminase (cytosine deaminase), Cas protein, and uracil glycosylase inhibitor, the cytidine deaminase is protein L8, the protein L8 is any one of the following,
[0008] A1) is a protein with an amino acid sequence of SEQ ID NO. 2 at positions 28-164;
[0009] A2) is a protein with 80% or more identity to the protein of A1) and related to deaminase obtained by substitution, deletion, and / or addition of amino acid residues of the amino acid sequence of A1);
[0010] A3) is a fusion protein obtained by connecting a tag to the N-terminus and / or C-terminus of A1) or A2).
[0011] Further, the connection of A3) can be linked by a peptide bond.
[0012] Further, the fusion protein is a protein connected by the cytidine deaminase, the Cas protein, the uracil glycosylase inhibitor, and the nuclear localization signal.
[0013] Further, in the fusion protein, the Cas protein can be nCas9.
[0014] Further, in the fusion protein, the nCas9 can be a protein with an amino acid sequence of SEQ ID NO. 2 at positions 182-1548.
[0015] Further, in the fusion protein, the uracil glycosylase inhibitor can be a protein with an amino acid sequence of SEQ ID NO. 2 at positions 1559-1641 and / or SEQ ID NO. 2 at positions 1652-1734.
[0016] Further, in the fusion protein, the amino acid sequence of the nuclear localization signal can be SEQ ID NO. 2 at positions 3-9 and / or SEQ ID NO. 2 at positions 1766-1781.
[0017] Further, the fusion protein can be a protein with an amino acid sequence of SEQ ID NO. 2.
[0018] The present application also provides a biological material related to the above-mentioned fusion protein, and the biological material can be at least one of the following D1) to D7):
[0019] D1) is a nucleic acid molecule encoding the fusion protein;
[0020] D2) is an expression cassette containing the nucleic acid molecule of D1);
[0021] D3) A recombinant vector containing the nucleic acid molecule described in D1) or a recombinant vector containing the expression cassette described in D2);
[0022] D4) Recombinant microorganisms containing the nucleic acid molecules described in D1), recombinant microorganisms containing the expression cassette described in D2), or recombinant microorganisms containing the recombinant vector described in D3);
[0023] D5) A transgenic plant cell line containing the nucleic acid molecule described in D1), a transgenic plant cell line containing the expression cassette described in D2), or a transgenic plant cell line containing the recombinant vector described in D3);
[0024] D6) Transgenic plant tissue containing the nucleic acid molecules described in D1), transgenic plant tissue containing the expression cassette described in D2), or transgenic plant tissue containing the recombinant vector described in D3);
[0025] D7) Transgenic plant organs containing the nucleic acid molecules described in D1), transgenic plant organs containing the expression cassette described in D2), or transgenic plant organs containing the recombinant vector described in D3).
[0026] Furthermore, in the biological material, the DNA molecule in D1) may be at least one of the following:
[0027] D11) Its coding strand has a nucleotide sequence that is the cDNA molecule or DNA molecule shown in positions 3589-8899 of SEQ ID NO.1;
[0028] D12) Its coding strand has a nucleotide sequence that is the cDNA molecule or DNA molecule shown in positions 3514-9040 of SEQ ID NO.1;
[0029] The nucleotide sequence defined by D12 has 80% or more identity with D11) and / or D12) and is a cDNA molecule or DNA molecule encoding the fusion protein.
[0030] To address the aforementioned technical problems, this application also provides the aforementioned protein L8 and related biological materials.
[0031] Protein L8 is described in any of the following:
[0032] A1) The amino acid sequence is the protein consisting of positions 28-164 of SEQ ID NO.2;
[0033] The proteins obtained by substituting and / or deleting and / or adding amino acid residues of the amino acid sequences shown in A2) and A1) have more than 80% identity with the protein shown in A1) and are related to deaminases.
[0034] A3) a fusion protein obtained by connecting a tag at the N-terminus and / or C-terminus of A1) or A2).
[0035] Further, the connection of A3) can be linked by a peptide bond.
[0036] Further, the tag described herein includes, but is not limited to, a GST (glutathione S-transferase) tag protein, a His tag protein, a MBP (maltose binding protein) tag protein, a Flag tag protein, a SUMO tag protein, a HA tag protein, a Myc tag protein, an eGFP (enhanced green fluorescent protein), an eCFP (enhanced cyan fluorescent protein), an eYFP (enhanced yellow green fluorescent protein), an mCherry (monomeric red fluorescent protein), or an AviTag tag protein.
[0037] In some embodiments of the present application, the tag is a His tag.
[0038] The present application also provides a biological material related to the above-mentioned protein L8, which can be at least one of the following:
[0039] B1) a nucleic acid molecule encoding the protein;
[0040] B2) an expression cassette containing the nucleic acid molecule of B1);
[0041] B3) a recombinant vector containing the nucleic acid molecule of B1) or the expression cassette of B2);
[0042] B4) a recombinant microorganism containing the nucleic acid molecule of B1), the expression cassette of B2), or the recombinant vector of B3);
[0043] B5) a transgenic plant cell line containing the nucleic acid molecule of B1), the expression cassette of B2), or the recombinant vector of B3);
[0044] B6) a transgenic plant tissue containing the nucleic acid molecule of B1), the expression cassette of B2), or the recombinant vector of B3);
[0045] B7) a transgenic plant organ containing the nucleic acid molecule of B1), the expression cassette of B2), or the recombinant vector of B3).
[0046] Further, in the biological material, the nucleic acid molecule of B1) can be at least one of the following:
[0047] B11) a cDNA molecule or a DNA molecule whose coding sequence is the one shown in positions 2-412 of SEQ ID NO. 4;
[0048] B12) a cDNA molecule or a DNA molecule whose coding sequence is the one shown in positions 3589-4189 of SEQ ID NO. 1;
[0049] B13) a cDNA molecule or a DNA molecule having 80% or more identity with the nucleotide sequence defined in B11) and / or B11) and encoding said protein L8.
[0050] In the present application, the expression cassette refers to DNA that is capable of expressing the protein in a host cell (e.g., a plant cell). The DNA can include not only a promoter that initiates transcription of the protein gene, but also a terminator that terminates transcription of the protein gene. Further, the expression cassette can also include an enhancer sequence. The promoters that can be used in the present application include, but are not limited to, constitutive promoters, tissue-, organ-, and development-specific promoters, and inducible promoters. Examples of promoters include, but are not limited to, the Ubiquitin promoter of maize, the constitutive 35S promoter of cauliflower mosaic virus; the wound-inducible promoter from tomato, the leucine aminopeptidase ("LAP", Chao et al. (1999) Plant Physiology 120:979-992); the chemically inducible promoter from tobacco, pathogenesis-related protein 1 (PR1) (induced by salicylic acid and BTH (benzothiadiazole-7-thiohydroxy acid S-methyl ester)); the tomato protease inhibitor II promoter (PIN2) or the LAP promoter (both of which are inducible by jasmonate acid methyl ester); the heat shock promoter (U.S. Patent 5,187,267); the tetracycline-inducible promoter (U.S. Patent 5,057,422); seed-specific promoters, such as the millet seed-specific promoter pF128 (CN101063139B (Chinese Patent 2007 1 0099169.7)), seed storage protein-specific promoters (e.g., the promoters for phaseolin, napin, oleosin, and soybean beta conglycin (Beachy et al. (1985) EMBO J. 4:3047-3053)). They can be used alone or in combination with other plant promoters. All references cited herein are incorporated in their entirety. Suitable transcription terminators include, but are not limited to, the Agrobacterium nopaline synthase terminator (NOS terminator), the cauliflower mosaic virus CaMV 35S terminator, the tml terminator, the pea rbcS E9 terminator, and the nopaline synthase and octopine synthase terminators (see, e.g., Odell et al. (1985) Nature 313:810; Rosenberg et al. (1987) Gene, 56:125; Guerineau et al. (1991) Mol. Gen. Genet, 262:141; Proudfoot (1991) Cell, 64:671; Sanfacon et al. Genes Dev., 5:141; Mogen et al. (1990) Plant Cell, 2:1261; Munroe et al. (1990) Gene, 91:151; Ballad et al. (1989) Nucleic Acids Res. 17:7891; Joshi et al. (1987) Nucleic Acid Res., 15:9627).
[0051] In some embodiments of the application, the expression cassette utilizes the Zmubi promoter (nucleotide sequence is SEQ ID NO. 1 from position 1506-3498) to initiate the expression of the deaminase L8, DddA, linker, nCas9(D10A), and glycosylase inhibitor (UGI) fusion protein, and the expression is terminated by the RBCS E9T terminator (nucleotide sequence is SEQ ID NO. 1 from position 9044-9678). The expression cassette is connected by the Zmubi promoter, the nucleic acid molecule encoding the cytosine base editor, and the RBCS E9T terminator.
[0052] The term "identity" as used herein refers to sequence similarity between amino acid or nucleic acid sequences. Identity can be evaluated by eye or by computer software. Using computer software, identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate identity between related sequences. The above 80% identity can be at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more.
[0053] The above 80% identity can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity. The above 85% identity can be at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity. The above 90% identity can be at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity. The above 95% identity can be at least 95%, 96%, 97%, 98%, or 99% identity.
[0054] In the above, the recombinant microorganism can be specifically bacteria, yeast, algae, and fungi.
[0055] In some embodiments of the application, the bacterial solution has the same meaning as "culture", and the term "culture" refers to the general term for the liquid or solid product (all substances in the culture vessel) that grows after artificial inoculation and cultivation of a microbial population. That is, the product obtained by growing and / or amplifying microorganisms, which can be a biologically pure culture of microorganisms, or can contain a certain amount of culture medium, metabolites or other components produced during the culture process. It can also be a mixture containing a certain amount of culture medium, microbial metabolites, and removing microbial bodies.
[0056] The application also provides the use of the protein L8, the biomaterial related to the protein L8, the fusion protein and / or the biomaterial related to the fusion protein in single base editing and / or in the preparation of a product of single base editing.
[0057] Further, in the application, the single base editing can occur outside or inside a cell.
[0058] Further, in the application, the single base editing can occur inside a plant cell.
[0059] Further, in the application, the single base editing can occur inside an organelle of a plant cell.
[0060] The application also provides a method for mutating C:G to T:A on a plant genome.
[0061] The application provides a method for mutating C:G to T:A on a base pair of a plant genome, which comprises the following steps: introducing a DNA molecule expressing the cytosine base editor and sgRNA into a recipient plant to obtain a target plant with C:G mutated to T:A; the target sequence of the sgRNA is 5'-N19-20PAM-3', wherein N19-20 is 19-20 N.
[0062] In the above method, the Cas protein in the cytosine base editor can be nCas9, and the PAM (protospacer adjacent motif) is NGG; the N is A, G, C or T.
[0063] In the process of introducing the DNA molecule expressing the cytosine base editor and sgRNA into the recipient plant, the PEG-mediated transformation method can be used, or one of the gene gun method or the Agrobacterium infection method can be used to introduce the gene editing tool kit into the maize protoplast or callus, which is easily understood by those skilled in the art. It is well known to those skilled in the art that the maize genomic DNA is composed of two strands, therefore, the target nucleotide sequence can be on any one of the complementary strands. For example, when the target nucleotide sequence is located in the sense strand of a certain functional gene, if the C at a specific site of the functional gene is mutated to T and / or the G is mutated to A, and if one of the mutations can obtain the expected amino acid in the corresponding functional protein, this system can also be used to achieve, that is, the C in the triple codon can be replaced by T and / or the G by directly replacing the base on the sense strand, thereby obtaining a maize gene functional "correction" mutant; or when the target nucleotide sequence is located in the antisense strand of a certain functional gene, if the C at a specific site of the functional gene is mutated to T, and if one of the mutations can obtain the expected amino acid in the corresponding functional protein, this system can be used to achieve, that is, the G in the antisense strand is mutated to A, and then the corresponding complementary C in the sense strand is replaced by T to change the amino acid encoded by the triple codon in the sense strand, thereby obtaining a maize gene functional "correction" mutant.
[0064] In the above, the plant can be a dicotyledonous plant or a monocotyledonous plant. The monocotyledonous plant can be maize. The single base editing can be replacing cytosine C with uracil.
[0065] The technical effects obtained by the present application are as follows:
[0066] The present application uses macrogenomic big data to conduct a comprehensive and systematic bioinformatics mining analysis on new cytidine deaminases, selects a protein L8 similar in structure to known DddA proteins but with low sequence similarity (identity <63.04%) for function verification, and develops a brand new CRISPR / Cas gene editing tool kit. The editing efficiency of the cytosine base editing technology mediated by the deaminase L8 fused with nCas9 (D10A) is significantly higher than that of the cytosine base editing technology mediated by DddA, and the editing activity window can be expanded, thereby improving the efficiency of the cytosine base editing technology and expanding its application range. The present application provides a set of important gene function research and correction tools for researchers in the fields of plant research and crop genetic improvement. The single base editor of the present application can improve the efficiency of cytosine base editing and accurately mediate base mutations at target sites, and can be widely used in maize and even other plant cells. BRIEF DESCRIPTION OF DRAWINGS
[0067] Figure 1 For the deaminase in vitro deamination activity analysis system.
[0068] Figure 2 For the deamination activity of deaminase L8 in vitro.
[0069] Figure 3 For the DNA substrate structure in the system of verifying the target sequence preference in vitro.
[0070] Figure 4 For the verification results of the target sequence preference of deaminase L8 in vitro.
[0071] Figure 5 For the position and structure of the target fragment in the L8 vector.
[0072] Figure 6 For the fluorescence detection of L8 and DddA in maize protoplasts.
[0073] Figure 7 For the single base editing vector of deaminase L8 and DddA system.
[0074] Figure 8 For the editing window and editing efficiency of deaminase L8 and DddA on the Bx9 gene.
[0075] Figure 9 For the sequence preference of editing of deaminase L8 and DddA on the Bx9 gene. DETAILED DESCRIPTION
[0076] The present application will be further described in conjunction with the specific embodiments. The examples given are only to illustrate the present application, and are not intended to limit the scope of the present application. The examples provided below can serve as a guide for further improvement by those skilled in the art, and do not in any way constitute a limitation on the present application.
[0077] In the following examples, the experimental methods are conventional methods, and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions, unless otherwise specified. The materials, reagents, etc. used in the following examples can be obtained commercially, unless otherwise specified.
[0078] In the quantitative test in the following examples, three replicates were set, and the results were averaged, unless otherwise specified.
[0079] In the following examples, the data were processed using GraphPad Prism statistical software, and the experimental results were expressed as mean ± standard deviation. Two-way ANOVA was used, * represents P<0.05, ** represents P<0.01, *** represents P<0.001, and **** represents P<0.0001.
[0080] Example 1, New Cytidine Deaminase Mining Process
[0081] 1.1, Access and download all sequenced metagenome assembly nucleic acid sequence information from JGI and NCBI biological databases; annotate protein sequences using Prodigal software to obtain all proteins; use Pfam database hidden Markov model to annotate the structure domain of the predicted coding genes.
[0082] 1.2, Because DddA belongs to the SCP1.201 family, its full length is >1400 aa, and the protein back half - DddAtox, which is annotated as SCP1.201-deam, is responsible for deamination. Therefore, based on this term as a keyword for screening; all candidate proteins screened by Mafft software alignment are aligned, and according to prior knowledge, it is known that the active site of DddAtox is glutamic acid E at position 1347, so proteins with the same site are screened; Because DddA is involved in the Type VI secretion system (T6SS) of the intercellular protein delivery system, it has downstream immune protein DddIA, which interacts with each other to inhibit the toxicity of DddA protein and prevent cell inactivation. In order to facilitate the purification of DddA protein in subsequent experiments, therefore, in the screened DddA homologous proteins, whether there is a corresponding immune protein DddIA within the range of 10 kb upstream and downstream is found; The length of the proteins screened is not the same, and only the DddAtox part is needed, so the sequences are aligned again to obtain the truncated sequences of the DddA homologous proteins that play a role, a total of 128.
[0083] 1.3, Use IQ-TREE software to construct an evolutionary tree of all proteins mined and other deaminase family protein sequences reported in the literature (Evolution of the deaminase fold and multiple origins of eukaryotic editing and mutagenic nucleic acid deaminases from bacterial toxin systems) together, and visualize it through ITOL software. From the figure, it can be concluded that the 128 new deaminases mined through the established bioinformatics process are divided into the SCP1.201 evolutionary branch, so these proteins indeed belong to the SCP1.201 family.
[0084] 1.4、To further screen new deaminases with activity, AlphaFold 2 software was used to predict the structure of all the proteins mined, and TM-align structure comparison was performed between each of them and known DddA proteins. Finally, 27 proteins were selected as new deaminases with similar structures to known proteins for subsequent experimental verification analysis. The 27 proteins are numbered as deaminases L8, L64, L65, L67, L70, L71, L72, L74, L78, L79, L80, L81, L83, L84, L85, L86, L88, L89, L90, L91, L115, L135, L136, L138, S45, S57, S76.
[0085] 1.5、The deaminase sequences mined were optimized for maize codons, and the company was commissioned to synthesize the coding genes of the deaminases and downstream immune proteins. Taking deaminase L8 as an example, the deaminase L8 and immune protein coding genes were connected to the pET-duet-1 vector to construct a co-expression vector of deaminase and downstream immune protein, which was transformed into E. coli TSR2566 to induce expression and purify the protein. The coding genes of deaminase L8 and immune protein are the DNA molecules of SEQ ID NO. 4.
[0086] The co-expression vector of deaminase and downstream immune protein is named pETduet1-L8 vector. The structure of pETduet1-L8 vector is: the fragment between BamHI and XhoI sites (small fragment between BamHI and XhoI sites) of pET-duet-1 vector is replaced with the DNA molecule with nucleotide sequence of SEQ ID NO. 4, and the other nucleotide sequences of pET-duet-1 vector remain unchanged. Among them, the 2-412th of SEQ ID NO. 4 is the coding gene of deaminase L8, the 478th-496th is the T7 promoter, and the 564th-944th is the coding gene of immune protein.
[0087] The pETduet1-L8 vector can express L8 fusion protein with vector sequence and His tag and L8 immune protein (amino acid sequence see Table 2), the amino acid sequence of the L8 fusion protein is the amino acid sequence with His tag encoded by the start codon on the vector to the sequence, the 11th-14th is a linker, and the 15th-151th is the amino acid sequence of deaminase L8.
[0088] The steps of inducing expression and protein purification are as follows: the constructed pETduet1-L8 vector is transformed into E. coli competent TSR2566 to obtain recombinant bacteria TSR2566 / pETduet1-L8. The bacteria are shaken in LB liquid medium to OD600 = 0.6-0.7, IPTG is added to induce protein expression, and the bacterial solution is collected. Then, the nickel column is used to purify the bacterial solution to obtain the protein complex of L8 fusion protein (see Table 2 for the amino acid sequence) and L8 immunoprotein (see Table 2 for the amino acid sequence). Then, different concentrations of 8M, 6M, 4M, 2M, and 1M urea solution are used to denature and renature the protein complex, and finally only the eluate containing deaminase L8 is purified by molecular sieve to obtain the L8 fusion protein.
[0089] The preparation of the other 26 deaminases is similar to that of deaminase L8, except that the coding sequences of the deaminases are different.
[0090] Deaminase DddA is used as a control, and the preparation of deaminase DddA is similar to that of deaminase L8, except that the DNA molecule of SEQ ID NO. 4 in the pETduet1-L8 vector is replaced by the DNA molecule shown in SEQ ID NO. 5. SEQ ID NO. 5 is the coding gene of deaminase DddA, the 2-415th position is the T7 promoter, and the 570-1001th position is the coding gene of the immunoprotein of deaminase DddA. The deaminase DddA fusion protein (see Table 2 for the amino acid sequence) obtained by inducing expression is complexed with the DddA immunoprotein (see Table 2 for the amino acid sequence) to form a protein complex.
[0091] Example 2, in vitro activity verification
[0092] 2.1, in vitro deamination activity
[0093] The 5'-FAM fluorophore-labeled DNA substrate is prepared into a deamination reaction system with 20nM concentration of deaminase to analyze the deamination activity. If the deaminase has deamination activity, the DNA substrate will be cleaved from the deaminated cytosine, and the reaction product of the deamination reaction system has a cleavage band near 18nt. Figure 1 )。
[0094] The deamination reaction system is prepared by mixing 50μL L8 protein (20nM) and 10μL deamination buffer solution containing DNA substrate to obtain a deamination reaction mixture. The composition of the deamination buffer solution containing DNA substrate is as follows: 20m M MES, 200mM NaCl, 1mM DTT, 80g / L Ficoll 70, 1μM DNA substrate.
[0095] The deamination reaction conditions are as follows: the deamination reaction mixture is deaminated at 37°C for 1 h; then 3 μL of UDG (Uracil-DNA Glycosylase) and 7 μL of 10xUDG buffer are added to the system, and the enzyme cutting reaction is carried out at 37°C for 30 min; 7 μL of 1M NaOH is added to the system, and the reaction is terminated by being placed at 95°C for 3 min, and the obtained deamination reaction sample is used for subsequent electrophoresis verification.
[0096] Electrophoresis verification: before electrophoresis, the marker system is as follows: 1 μL of 36 nt substrate, 1 μL of 18 nt substrate, 5 μL of 2xRNA loading buffer, and 13 μL of ddH2O; the deamination reaction sample 77 μL is mixed with 10 μL of 2xRNA loading buffer and denatured at 100°C for 5 min. During gel electrophoresis, the voltage is first set to 100V for 30 min, and then the voltage is adjusted to 150V for 44 min.
[0097] Result analysis: the DNA fragments labeled with 5'-FAM fluorophore can be displayed by scanning with a gel scanner (model: iBright FL1500 imaging system). The DNA fragments without 5'-FAM fluorophore labeling cannot be displayed by scanning with a gel scanner. That is, if the deaminase has in vitro deamination activity, the DNA substrate with fluorescent labeling is deaminated and denatured, and the product with fluorescent labeling of 18 nt DNA single strand is obtained. The reaction product of the deamination reaction system is scanned by a gel scanner, and a band is displayed near 18 nt; if the deaminase has no in vitro deamination activity (or the activity is very weak), the reaction product of the deamination reaction system is scanned by a gel scanner, and no band is displayed near 18 nt.
[0098] The deamination activities of the above-mentioned 28 candidate proteins are verified, and a total of 18 deaminases (L8, L70, L71, L72, L78, L79, L80, L91, L115, L138, S57, L83, L84, L85, L86, L89, L90, S45) with deamination activity are successfully verified. After repeated verification, the results of two verifications are combined, and finally five deaminases L8, L70, L85, L91, L90, S45 with higher deamination activity are selected for subsequent verification. The deamination activity verification results of deaminase L8 are shown in Figure 2 .
[0099] 2.2, verification of target sequence preference
[0100] The deamination reaction system is prepared according to 2.1, and the deamination reaction is carried out, and the difference from 2.2 is only that the DNA substrate is different, and N represents ATGC four different bases Figure 3), to determine the preference of the candidate deaminase for the target sequence C in vitro. The results show that deaminase L8 can effectively deaminate cytosine (C) in the NC (N is any base) background Figure 4 ) in which deaminase L8 showed better activity. Therefore, subsequent experiments were continued with deaminase L8.
[0101] Example 3, functional verification of deaminase L8 in eukaryotes
[0102] 3.1, expression verification of deaminase L8 in eukaryotic cells
[0103] Since the mined deaminase L8 is a double-stranded cytidine deaminase, it does not need to produce a break in the DNA strand during deamination, so there is a problem that cannot be ignored, that is, L8 is cytotoxic, which can cause vector mutation or directly cause cell death during vector construction. In order to solve this problem, a cat1 intron sequence is inserted between the L8 gene and the DddA gene, which can ensure that the L8 and DddA genes are not expressed during vector construction, so that the vector construction can be successfully carried out.
[0104] In order to verify that the intron sequence inserted in L8 and DddA can be correctly spliced in corn LH244 protoplast, a vector expressing L8 and eGFP protein with intron under the transcription of 35s promoter is constructed Figure 5 ), named L8 vector. The L8 vector is used for corn protoplast transformation, and the positive corn protoplast obtained by transformation is represented by L8. The steps of protoplast transformation refer to the reference "Targeted large fragment deletion in plants using paired crRNAs with type I CRISPR system".
[0105] A vector expressing DddA and eGFP protein with intron under the transcription of 35s promoter is constructed, named DddA vector. The pUC-eGFP vector and the DddA vector are used for corn protoplast transformation, as a control. The positive corn protoplast obtained by transformation is represented by pUC-eGFP and DddA, respectively.
[0106] Wherein, the L8 carrier structure is that the fragment between the BamHI and XhoI enzyme recognition sites of pUC-eGFP is replaced by a DNA molecule with the nucleotide sequence of SEQ ID NO. 6, and the other nucleotide sequences of the skeleton carrier remain unchanged. The 4th-21st of SEQ ID NO. 6 is the nucleotide sequence of 6xHis tag, the 22nd-33rd is the nucleotide sequence of linker, the 34th-334th is the coding gene of N-terminal of deaminase L8, the 335th-524th is the nucleotide sequence of cat1 intron, and the 525th-634th is the coding gene of C-terminal of deaminase L8. The nucleotide sequence of the pUC-eGFP carrier is shown in Table 2, wherein the 1137th-1142nd of the nucleotide sequence of the pUC-eGFP carrier is the BamHI enzyme recognition site, the 1161st-1166th is the XhoI enzyme recognition site, the 744th-1089th is the 35s promoter, and the 1182nd-1901st is eGFP.
[0107] The structure of the DddA carrier is only different from the L8 carrier in that the DNA molecule with the nucleotide sequence of SEQ ID NO. 6 (the coding gene of N-terminal of deaminase L8, cat1 intron and the coding gene of C-terminal of deaminase L8) in the L8 carrier is replaced by a DNA molecule with the nucleotide sequence of SEQ ID NO. 3. The 1st-217th of SEQ ID NO. 3 is the nucleotide sequence of the coding gene of N-terminal of deaminase DddA, the 218th-407th is the nucleotide sequence of cat1 intron, and the 408th-604th is the nucleotide sequence of the coding gene of C-terminal of deaminase DddA.
[0108] If the intron sequence can be correctly spliced, the GFP protein can be expressed, and can emit light under a fluorescence microscope. On the contrary, the GFP will be shifted, and no green fluorescence can be observed under a fluorescence microscope. The results show that the corn protoplasts L8, DddA and pUC-eGFP can emit green fluorescence under a microscope Figure 6 ), indicating that the introns in the L8 gene and the DddA gene can be spliced in the corn protoplasts.
[0109] 3.2, Construction of single base editing vector
[0110] Subsequently, a single base editing vector of deaminase L8 and DddA system was designed, and deaminase DddA was used as a control to verify whether deaminase L8 had deamination activity in plants. The expression of sgRNA was driven by the Osu3 promoter, and the expression of the fusion protein of deaminase L8 or DddA, linker, nCas9(D10A) and glycosylase inhibitor protein (UGI) was driven by the Zmubi promoter. They were integrated into the pUC19 vector (NEB product, product number N3041S) by homologous recombination, and the obtained recombinant vectors were named pUC19-L8 vector and pUC19-DddA vector, respectively.
[0111] The nucleotide sequence of the pUC19-L8 vector is SEQ ID NO. 1. The pUC19-L8 can express a fusion protein with the amino acid sequence of SEQ ID NO. 2.
[0112] The nucleotide sequence of the sgRNA gene expression cassette is SEQ ID NO. 1 57-1493. Among them: the nucleotide sequence of the OsU3 promoter is SEQ ID NO. 1 57-436; the BsaI enzyme digestion recognition site is SEQ ID NO. 1 438-444 and 1120-1127; the sgRNA scaffold is SEQ ID NO. 1 1127-1202; the nucleotide sequence of the OsU3 terminator is SEQ ID NO. 1 1203-1493.
[0113] SEQ ID NO. 1 nucleotide sequence of expression cassette of deaminase L8, nCas9(D10A) and glycosylase inhibitor protein (UGI) fusion protein from 3508 to 9043. Among them: SEQ ID NO. 1 from 1506 to 3498 is the nucleotide sequence of Zmubi promoter; SEQ ID NO. 1 from 3514 to 3534 is the nucleotide sequence of SV40 NLS, from 3535 to 3558 is the nucleotide sequence of linker1, from 3559 to 3576 is the nucleotide sequence of 6xHis tag, from 3577 to 3588 is the nucleotide sequence of linker2, from 3589 to 3889 is the nucleotide sequence of deaminase L8 N-terminal coding gene, from 3890 to 4079 is the nucleotide sequence of cat1 intron, from 4080 to 4189 is the nucleotide sequence of deaminase L8 C-terminal coding gene, from 4190 to 4240 is the nucleotide sequence of A(EAAAK)3A linker, from 4241 to 8341 is the nucleotide sequence of nCas9 coding gene, from 8342 to 8371 is the nucleotide sequence of 10aa linker, from 8372 to 8620 is the nucleotide sequence of UGI coding gene, from 8621 to 8650 is the nucleotide sequence of 10aa linker, from 8651 to 8899 is the nucleotide sequence of UGI coding gene, from 8900 to 8911 is the nucleotide sequence of 4aa linker, from 8912 to 8992 is the nucleotide sequence of 3xHA(81), from 8993 to 9040 is the nucleotide sequence of NLS. From 9044 to 9678 is the nucleotide sequence of RBCS E9T terminator.
[0114] SEQ ID NO. 2 amino acid sequence of SV40 NLS from 3 to 9, linker1 amino acid sequence from 10 to 17, 6xHis tag amino acid sequence from 18 to 23, linker2 amino acid sequence from 24 to 27, deaminase L8 amino acid sequence from 28 to 164, A(EAAAK)3A linker amino acid sequence from 165 to 181, nCas9 amino acid sequence from 182 to 1548, 10aa linker amino acid sequence from 1549 to 1558, UGI amino acid sequence from 1559 to 1641, 10aa linker amino acid sequence from 1642 to 1651, UGI amino acid sequence from 1652 to 1734, 4aa linker amino acid sequence from 1735 to 1738, 3xHA(81) amino acid sequence from 1739 to 1765, NLS amino acid sequence from 1766 to 1781.
[0115] The structure of the pUC19-DddA vector is the same as that of the pUC19-L8 vector, and the only difference is that the DNA molecule with the nucleotide sequence of SEQ ID NO. 1 3589-4189 in the pUC19-L8 vector is replaced by a DNA molecule with the nucleotide sequence of SEQ ID NO. 3. The pUC19-DddA vector expresses a base editor containing a DddA protein (a fusion protein containing a DddA protein), and the amino acid sequence of the base editor containing the DddA protein is shown in Table 2.
[0116] 3.3, In order to detect the single base editing activity of L8, a target of NGG (SEQ ID NO. 7: CTGGCCCGCACCGTCACAG) was designed on the maize Bx9 (Zm00001eb033030) gene. The annealing primer was designed to synthesize a DNA fragment with sticky ends, which was ready for use. The pUC19-L8 vector was digested with BsaI enzyme to obtain the digestion product, and the obtained DNA fragment and digestion product were ligated with T4 enzyme to obtain the recombinant vector pUC19-L8-Bx9 vector targeting the Bx9 gene. The structure of the pUC19-L8-Bx9 vector is that the fragment between the BsaI enzyme recognition sites of the pUC19-L8 vector (the small fragment between the BsaI enzyme recognition sites) is replaced by a DNA molecule with the nucleotide sequence of SEQ ID NO. 7, and the other nucleotide sequences of the pUC19-L8 vector remain unchanged. After the pUC19-L8-Bx9 vector was subjected to transfection level plasmid large extraction, it was used to transform maize protoplasts, and after 48 h of culture, the DNA of the transformed maize protoplasts was extracted. The target was designed upstream and downstream of the target, and a fragment of about 250 bp was amplified to construct a DNA library for second-generation sequencing.
[0117] The nucleotide sequence of the annealing primer is as follows:
[0118] Bx9T-F: ATATATGGTCTCTGGCG CTGGCCCGCACCGTCACAG GTT;
[0119] Bx9T-R: ATTATTGGTCTCTAAAC CTGTGACGGTGCGGGCCAG C.
[0120] A cytosine base editor constructed with deaminase DddA was used as a control. The preparation method of the pUC19-DddA-Bx9 vector refers to that of the pUC19-L8-Bx9 vector, and the only difference is that the pUC19-L8 vector is replaced by the pUC19-DddA vector.
[0121] The C to T and G to A ratios within the target range were counted by analyzing the sequencing results, and the N of PAM NGG was defined as 0, the upstream was recorded as negative, and the downstream was recorded as positive. The results showed that the editing window range of deaminase L8 and DddA on Bx9 gene was about 140 bp; the editing efficiency of deaminase L8 was significantly higher than that of deaminase DddA at most sites, especially at G-74 site, L8 had the highest editing efficiency of 19.25%, and the average value was 18.2% Figure 8 ).
[0122] In summary, it can be proved that the mined L8 has activity, and shows higher editing activity than DddA at multiple sites of Bx9 target gene.
[0123] Table 1 L8 and DddA editing efficiency statistical results
[0124]
[0125] Note: CK is protoplast cell without any treatment.
[0126] 3.4, In order to detect the sequence preference of L8 and DddA editing on Bx9 gene, 3bp upstream and downstream of the effective editing site were taken, the sequences were aligned, and the base distribution was counted using weblogo software. The results showed that DddA had TC sequence preference on Bx9 gene as described by the previous person. L8 can deaminate HC (H is A, C or T) ( Figure 9 ).
[0127] Table 2 part of the sequence in the application
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134] The application has been described in detail. For those skilled in the art, the application can be implemented in a wide range of equivalent parameters, concentrations and conditions without departing from the spirit and scope of the application, and without unnecessary experiments. Although the application gives a special example, it should be understood that the application can be further improved. In general, according to the principle of the application, the application is intended to include any modification, use or improvement of the application, including changes made by conventional techniques known in the art, which deviates from the range disclosed in the application.
Claims
1. A fusion protein, characterized by: The fusion protein is a protein with the amino acid sequence SEQ ID NO.
2.
2. A nucleic acid molecule encoding the fusion protein of claim 1.
3. An expression cassette containing the nucleic acid molecule of claim 2.
4. A recombinant vector containing the nucleic acid molecule of claim 2 or a recombinant vector containing the expression cassette of claim 3.
5. A recombinant microorganism containing the nucleic acid molecule of claim 2, a recombinant microorganism containing the expression cassette of claim 3, or a recombinant microorganism containing the recombinant vector of claim 4.
6. The nucleic acid molecule according to claim 2, characterized in that: The nucleic acid molecule is at least one of the following: D11) The nucleotide sequence of the coding strand is the cDNA molecule or DNA molecule shown in positions 3508-9043 of SEQ ID NO.1; The nucleotide sequences defined by D12 and D11 have 80% or more identity and encode a cDNA molecule or DNA molecule of the fusion protein of claim 1.
7. The use of the fusion protein of claim 1 in single-base editing, wherein the single-base editing occurs within plant cells.
8. The use of the fusion protein of claim 1 in the preparation of products with single-base editing.
9. The application of the nucleic acid molecule of claim 2 in single-base editing, wherein the single-base editing occurs within plant cells.
10. The use of the nucleic acid molecule of claim 2 in the preparation of products with single-base editing.
11. The application of the expression cassette of claim 3 in single-base editing, wherein the single-base editing occurs within plant cells.
12. The use of the expression cassette of claim 3 in the preparation of products with single-base editing.
13. The application of the recombinant vector of claim 4 in single-base editing, wherein the single-base editing occurs within plant cells.
14. The use of the recombinant vector of claim 4 in the preparation of products with single-base editing.
15. The application of the recombinant microorganism of claim 5 in single-base editing, wherein the single-base editing occurs within plant cells.
16. The use of the recombinant microorganism of claim 5 in the preparation of single-base edited products.
17. A protein, characterized by: The protein is at least one of the following: A1) The amino acid sequence is the protein consisting of positions 28-164 of SEQ ID NO.2; A2) The fusion protein obtained by attaching a tag to the N-terminus and / or C-terminus of A1).
18. A nucleic acid molecule encoding the protein of claim 17.
19. An expression cassette containing the nucleic acid molecule of claim 18.
20. A recombinant vector containing the nucleic acid molecule of claim 18 or a recombinant vector containing the expression cassette of claim 19.
21. A recombinant microorganism containing the nucleic acid molecule of claim 18, a recombinant microorganism containing the expression cassette of claim 19, or a recombinant microorganism containing the recombinant vector of claim 20.
22. The nucleic acid molecule according to claim 18, characterized in that: The nucleic acid molecule is at least one of the following: B11) The nucleotide sequence of the coding strand is the cDNA molecule or DNA molecule shown at positions 3589-4189 of SEQ ID NO. 1; The nucleotide sequence defined in B12) has 80% or more identity with that defined in B11) and encodes a cDNA molecule or DNA molecule of the protein of claim 17.
Citation Information
Patent Citations
Seed specificity highly effective promoter and its application
CN101063139A
Seed specific highly effective promoter and its application
CN101063139B
Recombinant DNA: transformed microorganisms, plant cells and plants: a process for introducing an inducible property in plants, and a process for producing a polypeptide or protein by means of plants or plant cells
US5057422A
Plant proteins, promoters, coding sequences and use
US5187267A
Cytidine deaminase and base editor
CN118147120A