Modified aminoacyl-trna synthetase, nucleic acid construct and genetically engineered strain
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2026-08-13
Smart Images

Figure CN2024142051_13082026_PF_FP_ABST
Abstract
Description
A modified aminoacyl-tRNA synthetase, nucleic acid constructs, and genetically engineered strains Technical Field
[0001] This invention relates to the field of biotechnology, and more preferably to an aminoacyl-tRNA synthetase and its nucleic acid constructs, genetically engineered strains, and its application in cell-free synthesis of non-natural amino acid proteins. Background Technology
[0002] Whether in protein structure and function research or antibody-drug conjugate (ADC) production, there is an increasing need to introduce new functional groups into protein polypeptide chains. This can be achieved using orthogonal protein translation systems, which encode non-natural amino acids with specific chemical functional groups into the peptide chain according to a pre-defined genetic codon. Although the introduction of non-natural amino acids can be achieved intracellularly, competition with classical protein translation systems and the fact that the special chemically active side chain groups of non-natural amino acids often prevent them from freely crossing the cell membrane make cell-free in vitro protein translation systems irreplaceable for the large-scale production of proteins containing non-natural amino acids.
[0003] Cell-free in vitro protein translation systems, lacking the constraint of cell membranes, theoretically allow for reaction advancement by increasing the concentration of key components. Non-natural amino acids, catalyzed by specific aminoacyl-tRNA synthetases, form aminoacyl-tRNA, which is then transported to ribosomes. Through codon pairing, the non-natural amino acids are assembled into the peptide chain under ribosome catalysis. However, the binding affinity of non-natural aminoacyl-tRNA to transport proteins and its ribosome compatibility are weaker than those of classical aminoacyl-tRNA, hindering the reaction. Therefore, increasing substrate concentration is necessary to promote the reaction. The orthogonal aminoacyl-tRNA synthetases used for non-natural amino acid introduction are typically derived by mutating the substrate binding site of natural aminoacyl-tRNA synthetases. This alteration of substrate specificity often comes at the cost of decreased catalytic activity. To compensate for this activity reduction, high concentrations of enzymes, tRNA, or non-natural amino acids are required to steer the reaction towards peptide chain synthesis.
[0004] The commonly used orthogonal translation system pylrs-tRNA CUA pyl PylRS (MmPylRS, MbPylRS) originated from the archaea *Methanosarcina mazei* and *Methanosarcina barkeri*. Due to their unique structure, pylrs recognize tRNA independently of anticodon loops and exhibit substrate recognition plasticity. Through directed evolution, pylrs–tRNA… CUA pyl It can already introduce more than 200 kinds of non-natural amino acids.
[0005] However, both MmPylRS and MbPylRS contain an N-terminal domain with low solubility, resulting in poor solubility. Even after codon optimization, the expression levels of MmPylRS and MbPylRS in E. coli are difficult to increase, and the protein concentration after concentration still cannot exceed 8 mg / ml. This fails to fully utilize the advantages of cell-free in vitro protein translation systems, as the introduction efficiency of non-natural amino acids is very low, making them unsuitable for practical production applications.
[0006] Methanomethylophilus alvus PylRS (PylRS) is an aminoacyl-tRNA synthetase. This synthetase belongs to the nucleotide class and its full name is pyrrolysyl-tRNA synthetase, commonly abbreviated as PylRS. The main function of PylRS is to assist prokaryotes in synthesizing a specific amino acid—pyrrolysine (Pyl), which is absent in all known eukaryotes. [1] .
[0007] Methanomethylophilus alvus is a methanogenic archaea originally isolated from the stomachs of bovine ruminants. PylRS was discovered through sequence analysis of the M. alvus genome and can be used to modify foreign proteins and express novel biological activities, thus having significant application value in the field of synthetic biology.
[0008] Recently, Chin et al. discovered pylrs (MaPylRS) from Methanomethylophilus alvus, which have similar catalytic and tRNA-binding domains to MmPylRS and MbPylRS, but lack the N-terminal domain.
[0009] In current scientific research, MaPylRS is mainly used for the modification of exogenous proteins in prokaryotes such as Escherichia coli. [2] There are also reports of successful applications in eukaryotes. [3] According to the latest research findings, MaPylRS in *Saccharomyces cerevisiae* [4] When expressed in [a specific culture / organism], MaPylRS can achieve efficient and site-specific insertion of non-natural amino acids into the target protein. However, in both *E. coli* and yeast, MaPylRS exists in cells as an exogenous plasmid, and there is currently no eukaryotic cell system that can stably and efficiently express MaPylRS by directly recombining it into their genome. [5] .
[0010] In vitro protein synthesis systems refer to protein synthesis performed under non-in vivo conditions, such as cell lysis buffers. Only DNA or RNA templates, RNA polymerase, and necessary amino acids, ATP, and cofactors need to be added to the reaction system to synthesize the target protein. Because it eliminates the need for the complex metabolic processes of the entire cell, this system enables rapid and efficient large-scale protein production. Due to its high flexibility and simplified procedures, it has become one of the most widely used tools in the biomedical field and scientific research. [6] Currently, commonly used commercial in vitro protein expression systems include the E. coli system (E. coli extract, ECE). [7] Rabbit reticulocyte lysate (RRL) [8] Wheat germ extract (WGE) [9] Insects (Insect cell extract, ICE)
[0010] and human source system
[0011] .
[0011] Due to the absence of the N-terminal domain with low solubility, mapylrs are more easily expressed in E. coli, and the final expression product can be concentrated to 40 mg / ml without precipitation. It is suitable for cell-free non-natural amino acid delivery systems.
[0012] CRISPR / Cas (Clustered Regularly Interspaced Short Palindromic Repeats / CRISPR associated) is an immune system widely found in bacteria and archaea. As part of an organism's defense mechanism, it can recognize and destroy foreign DNA or RNA that invades host cells, and has been developed into a highly efficient and accurate gene-editing tool.
[0012] This technology can precisely cut DNA, allowing researchers to add or remove specific parts of genes to study their function, and it can also be used to treat certain genetic diseases.
[0013] When using the CRISPR / Cas9 system for gene editing, a specific gRNA sequence needs to be paired with the Cas9 nuclease protein and delivered to the target cell. Once the gRNA binds to Cas9, it is delivered to the target genome. During the formation of the DNA double-strand nick, the gRNA regulates the Cas9 enzyme to discover the PAM sequence (protospacer adjacent motif) on the genome and recognize the sequence 20 bp upstream of it. Then, the Cas9 enzyme creates a double-strand nick three bases upstream of the PAM. Simultaneously, if donor DNA is provided, the broken DNA strands can be joined together through homologous recombination (HDR), thereby achieving the purpose of gene modification.
[0014] There are many examples of using the CRISPR / Cas9 system to modify the genome of Saccharomyces cerevisiae, including point mutations, gene knockouts, and gene insertions. [15,16] For example, the pCAS plasmid, which is widely used in existing technologies, simultaneously contains the Cas9 gene sequence and gRNA elements.
[0017] Kluyveromyces is a type of yeast fungus, belonging to the Saccharomyces cerevisiae family. It has wide industrial applications, particularly in the food and beverage industries. Kluyveromyces is widely used in the fermentation processes of dairy products, wine, beer, and other food and beverage products. It can break down sugars and produce various enzymes, promoting fermentation and imparting good taste and flavor to food. Compared to other Saccharomyces cerevisiaes, Kluyveromyces has some unique characteristics, such as high temperature adaptability and acid tolerance. This makes it more suitable for certain industrial applications under specific conditions, such as high-temperature fermentation and acidic environments. Furthermore, Kluyveromyces offers many advantages as a host system for expressing pharmaceutical proteins. First, through appropriate gene expression vector and promoter sequence design, Kluyveromyces can achieve high-level expression of exogenous genes. Second, Kluyveromyces possesses a complete protein folding and modification system, capable of correctly folding complex pharmaceutical proteins and performing necessary glycosylation modifications. Third, Kluyveromyces is easy to cultivate and can be expressed solubleally intracellularly, facilitating the extraction and purification of target proteins and resulting in lower engineering production costs. Overall, Kluyveromyces is a host system with great potential for efficient expression of medicinal proteins.
[0013] References
[0014] 1.Yanagisawa T,Seki E,Tanabe H,Fujii Y,Sakamoto K,Yokoyama S.Crystal Structure of Pyrrolysyl-tRNA Synthetase from a Methanogenic Archaeon ISO4-G1 and Its Structure-Based Engineering for Highly-Productive Cell-Free Genetic Code Expansion with Non-Canonical Amino Acids.Int J Mol Sci.2023Mar 26;24(7):6256.doi:10.3390 / ijms24076256.PMID:37047230;PMCID:PMC10094482.
[0015] 2.Kobayashi T,Yanagisawa T,Sakamoto K,Yokoyama S.Recognition of non-alpha-amino substrates by pyrrolysyl-tRNA synthetase.J Mol Biol.2009Feb 6;385(5):1352-60.doi:10.1016 / j.jmb.2008.11.059.Epub 2008Dec 11.PMID:19100747.
[0016] 3.Beránek V,Willis JCW,Chin JW.An Evolved Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase / tRNA Pair Is Highly Active and Orthogonal in Mammalian Cells.Biochemistry.2019 Feb 5;58(5):387-390.doi:10.1021 / acs.biochem.8b00808.Epub 2018 Sep 27.PMID:30260626;PMCID:PMC6365905.
[0017] 4.Stieglitz JT,Lahiri P,Stout MI,Van Deventer JA.Exploration of Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase Activity in Yeast.ACS Synth Biol.2022 May 20;11(5):1824-1834.doi:10.1021 / acssynbio.2c00001.Epub 2022 Apr 13.PMID:35417129;PMCID:PMC10112046.
[0018] 5.Avila-Crump S,Hemshorn ML,Jones CM,Mbengi L,Meyer K,Griffis JA,Jana S,Petrina GE,Pagar VV,Karplus PA,Petersson EJ,Perona JJ,Mehl RA,Cooley RB.Generating Efficient Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetases for Structurally Diverse Non-Canonical Amino Acids.ACS Chem Biol.2022 Dec 16;17(12):3458-3469.doi:10.1021 / acschembio.2c00639.Epub 2022 Nov 16.PMID:36383641;PMCID:PMC9833845.
[0019] 6.Craig D,Howell MT,Gibbs CL,Hunt T,Jackson RJ.Plasmid cDNA-directed protein synthesis in a coupled eukaryotic in vitro transcription-translation system.Nucleic Acids Res.1992 Oct 11;20(19):4987-95.doi:10.1093 / nar / 20.19.4987.PMID:1383935;PMCID:PMC334274.
[0020] 7.Kim DM,Kigawa T,Choi CY,Yokoyama S.A highly efficient cell-free protein synthesis system from Escherichia coli.Eur J Biochem.1996 Aug 1;239(3):881-6.doi:10.1111 / j.1432-1033.1996.0881u.x.PMID:8774739.
[0021] 8.Jackson RJ,Hunt T.Preparation and use of nuclease-treated rabbit reticulocyte lysates for the translation of eukaryotic messenger RNA.Methods Enzymol.1983;96:50-74.doi:10.1016 / s0076-6879(83)96008-1.PMID:6656641.
[0022] 9.Erickson AH,Blobel G.Cell-free translation of messenger RNA in a wheat germ system.Methods Enzymol.1983;96:38-50.doi:10.1016 / s0076-6879(83)96007-x.PMID:6656637.
[0023] 10.Ezure T,Suzuki T,Higashide S,Shintani E,Endo K,Kobayashi S,Shikata M,Ito M,Tanimizu K,Nishimura O.Cell-free protein synthesis system prepared from insect cells by freeze-thawing.Biotechnol Prog.2006 Nov-Dec;22(6):1570-7.doi:10.1021 / bp060110v.PMID:17137303.
[0024] 11.Mikami S,Kobayashi T,Masutani M,Yokoyama S,Imataka H.A human cell-derived in vitro coupled transcription / translation system optimized for production of recombinant proteins.Protein Expr Purif.2008Dec;62(2):190-8.doi:10.1016 / j.pep.2008.09.002.Epub 2008Sep 11.PMID:18814849.
[0025] 12.Sander JD,Joung JK.CRISPR-Cas systems for editing,regulating and targeting genomes.Nat Biotechnol.2014Apr;32(4):347-55.doi:10.1038 / nbt.2842.Epub 2014Mar 2.PMID:24584096;PMCID:PMC4022601.
[0026] 13.Jinek M,Chylinski K,Fonfara I,Hauer M,Doudna JA,Charpentier E.A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.Science.2012Aug 17;337(6096):816-21.doi:10.1126 / science.1225829.Epub 2012Jun 28.PMID:22745249;PMCID:PMC6286148.
[0027] 14.Bibikova M,Carroll D,Segal DJ,Trautman JK,Smith J,Kim YG,Chandrasegaran S.Stimulation of homologous recombination through targeted cleavage by chimeric nucleases.Mol Cell Biol.2001Jan;21(1):289-97.doi:10.1128 / MCB.21.1.289-297.2001.PMID:11113203;PMCID:PMC88802.
[0028] 15.DiCarlo JE,Norville JE,Mali P,Rios X,Aach J,Church GM.Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems.Nucleic Acids Res.2013Apr;41(7):4336-43.doi:10.1093 / nar / gkt135.Epub 2013Mar 4.PMID:23460208;PMCID:PMC3627607.
[0029] 16.Bao Z,Xiao H,Liang J,Zhang L,Xiong X,Sun N,Si T,Zhao H.Homology-integrated CRISPR-Cas(HI-CRISPR)system for one-step multigene disruption in Saccharomyces cerevisiae.ACS Synth Biol.2015May 15;4(5):585-94.doi:10.1021 / sb500255k.Epub 2014Sep 19.PMID:25207793.
[0030] 17.Ryan OW,Skerker JM,Maurer MJ,Li X,Tsai JC,Poddar S,Lee ME,DeLoache W,Dueber JE,Arkin AP,Cate JH.Selection of chromosomal DNA libraries using a multiplex CRISPR system.Elife.2014Aug 19;3:e03703.doi:10.7554 / eLife.03703.PMID:25139909; PMCID:PMC4161972. Summary of the Invention
[0031] To overcome the shortcomings of existing technologies, this invention designs histidine tags at the N-terminus and C-terminus of the MaPylRS gene, respectively, as well as thrombin restriction sites for convenient tag excision. Results show that adding these sequences at the N-terminus or C-terminus not only promotes enzyme stability and facilitates purification, but also further enhances the enzyme's catalytic activity.
[0032] Furthermore, existing literature and commercial kits all involve manually adding the MaPylRS protein exogenously or introducing plasmids containing its expression structure through transformation / transfection. This negatively impacts the interpretation, complexity, and stability of experimental results, as well as the production cost of the target protein. This invention utilizes a CRISPR / Cas9 high-efficiency gene editing system to integrate the MaPylRS gene into the Kluyveromyces genome, achieving stable and efficient expression of the MaPylRS gene in yeast without the need for antibiotic selection and maintenance. Furthermore, it enables the insertion of non-natural amino acids into the target protein.
[0033] The first invention provides a recombinant aminoacyl-tRNA synthetase having the structure described in Formula I:
[0034] A1-A2-A3-A4(I):
[0035] In formula I,
[0036] A1 indicates that the tag is absent or is a histidine tag;
[0037] A2 is either absent or a thrombin cleavage site;
[0038] A3 is either absent or a tagged protein;
[0039] A4 is an aminoacyl-tRNA synthetase;
[0040] "-" indicates a bond or amino acid linkage sequence on its own.
[0041] And at least one of A1 to A3 exists.
[0042] The connection between A1 to A4 can be either from N to C or from C to N.
[0043] More preferably, the aminoacyl-tRNA synthetase is selected from natural or mutated Pyl-tRNA synthetase (PylRS), Leu-tRNA synthetase (LeuRS), Tyr-tRNA synthetase (TyrRS), Phe-tRNA synthetase (PheRS) or TrP-tRNA synthetase (TrpRS).
[0044] More preferably, the aminoacyl-tRNA synthetase is selected from natural or mutated MaPylRS, MmPylRS, MbPylRS, EcTyrRS, MjTyrRS, EcLeuRS, ScPheRS, ScTrpRS or BsTrpRS.
[0045] More preferably, the aminoacyl-tRNA synthetase is selected from natural or mutated MaPylRS.
[0046] More preferably, the sequence of the aminoacyl-tRNA synthetase is SEQ ID NO:60.
[0047] More preferably, the aminoacyl-tRNA synthetase comprises the sequence shown in SEQ ID NO:60 or its active fragment, or is a polypeptide having ≥85% homology (preferably ≥90% homology; equally preferably ≥95% homology; most preferably ≥97% homology, such as ≥98% or ≥99%) with the amino acid sequence shown in SEQ ID NO:60 and having the same activity as the sequence in SEQ ID NO:60.
[0048] More preferably, the histidine tag has an n×His structure, where 1≦n≦50; more preferably 2≦n≦30; more preferably 5≦n≦20; and more preferably 6≦n≦10.
[0049] More preferably, the tag protein is selected from T7 tag, CBP tag, CMyc tag, FLAG tag, Spot tag, C tag, Avi tag, Streg tag, SUMO tag, GST tag, MBP tag or a combination thereof; preferably T7 tag.
[0050] More preferably, the N-terminal or C-terminal of A4 is connected to A1-A2-A3:
[0051] More preferably, a his tag is attached to the N-terminus or C-terminus of the aminoacyl-tRNA synthetase.
[0052] More preferably, a thrombin cleavage site and a tag protein are connected to the N-terminus of the aminoacyl-tRNA synthetase in sequence from the N-terminus to the C-terminus.
[0053] More preferably, the N-terminus of the aminoacyl-tRNA synthetase is connected with a his tag, a thrombin cleavage site, and a tag protein in the order from the N-terminus to the C-terminus.
[0054] More preferably, the amino acid sequence of the recombinant aminoacyl-tRNA synthetase is selected from any one or more of the following: SEQ ID NO:1 to SEQ ID NO:3 or SEQ ID NO:64.
[0055] More preferably, the recombinant aminoacyl-tRNA synthetase comprises any one of the sequences described in SEQ ID NO:1-3 and SEQ ID NO:64 or its active fragment, or is a polypeptide that has ≥85% homology (preferably ≥90% homology; equally preferably ≥95% homology; most preferably ≥97% homology, such as 98% or more, 99% or more) with the amino acid sequences described in any one of SEQ ID NO:1-3 and SEQ ID NO:64 and has the same activity as the sequences described in any one of SEQ ID NO:1-3 and SEQ ID NO:64.
[0056] More preferably, the coding sequence of the recombinant aminoacyl-tRNA synthetase is selected from any one or more of the following: SEQ ID NO:4, SEQ ID NO:61, SEQ ID NO:62 or SEQ ID NO:63.
[0057] More preferably, the coding sequence of the recombinant aminoacyl-tRNA synthetase comprises any one of the sequences or active fragments of SEQ ID NO:4, SEQ ID NO:61, SEQ ID NO:62, and SEQ ID NO:63, or respectively, or each of the nucleotides having ≥85% homology (preferably ≥90% homology; equally preferably ≥95% homology; most preferably ≥97% homology, such as 98% or more, 99% or more) and each having the same activity as any one of the sequences of SEQ ID NO:4, SEQ ID NO:61, SEQ ID NO:62, or SEQ ID NO:63.
[0058] A second aspect of the present invention provides a nucleic acid construct encoding the recombinant aminoacyl-tRNA synthetase described in the first aspect of the present invention.
[0059] More preferably, the nucleic acid construct contains at least the nucleic acid sequence with the structure described in Formula II: Z1-Z2-Z3-Z4, where Z1 to Z4 are elements used to constitute the construct; "-" independently represents a bond or nucleotide linking sequence; Z1 is a coding sequence for a histidine tag that is absent or a thrombin cleavage site that is absent or a thrombin cleavage site that is absent or a tagged protein that is absent or a tagged protein that is absent or a tRNA synthetase sequence that is present; wherein at least one of Z1 to Z3 is present.
[0060] The structure of Formula II can be either from 5' to 3' or from 3' to 5':
[0061] That is, the encoding products of Z1-Z2-Z3 are connected to the N-terminus or C-terminus of the encoding product of Z4.
[0062] More preferably, the N-terminus of the encoded products of Z1-Z2-Z3 is connected to the N-terminus of the encoded product of Z4.
[0063] More preferably, the amino acid sequence encoded by Z1 is HHHHHH; more preferably, the Z1 encodes a sequence containing HHHHHH or its active fragment, or a nucleotide having ≥85% homology (preferably ≥90% homology; equally preferably ≥95% homology; most preferably ≥97% homology, such as 98% or more, 99% or more) and having the same activity as the above-described sequence.
[0064] More preferably, the amino acid sequence encoded by Z2 is LVPRGS; more preferably, the sequence encoded by Z2 contains LVPRGS or its active fragment, or is a nucleotide having ≥85% homology (preferably ≥90% homology; equally preferably ≥95% homology; most preferably ≥97% homology, such as 98% or more, 99% or more) and having the same activity as the sequence shown above.
[0065] More preferably, the amino acid sequence encoded by Z3 is SEQ ID NO:59; more preferably, the Z-encoded sequence comprises the sequence shown in SEQ ID NO:59 or its active fragment, or is a nucleotide having ≥85% homology (preferably ≥90% homology; equally preferably ≥95% homology; most preferably ≥97% homology, such as 98% or more, 99% or more) and having the same activity as the sequence shown above.
[0066] More preferably, the nucleic acid construct further includes a promoter.
[0067] More preferably, the nucleic acid structure contains the structure described in Formula III:
[0068] Z5-Z1-Z2-Z3-Z4.
[0069] More preferably, the promoter is selected from PGK1, GAP1, ADH1, HXK1, GAPDH1, TEF1 or TIF11.
[0070] A third aspect of the present invention provides a vector containing the nucleic acid construct provided in the second aspect of the present invention.
[0071] A fourth aspect of the present invention provides a genetically engineered strain, wherein one or more sites in the genome of the genetically engineered strain are integrated with the nucleic acid constructs described in the second aspect of the present invention.
[0072] More preferably, the site is selected from Lys1-5, glpA, or UPF1.
[0073] More preferably, the nucleic acid construct further includes a promoter and a terminator.
[0074] More preferably, the promoter is selected from PGK1, GAP1, ADH1, HXK1, GAPDH1, TEF1, TIF11, GAL1, GAL7 or GAL10.
[0075] More preferably, the terminator is selected from CYC1, GPM1, TDH2 or ACT1.
[0076] More preferably, the genetically engineered strain contains the recombinant aminoacyl-tRNA synthetase provided in the first aspect of the present invention.
[0077] More preferably, the genetically engineered strain contains the vector provided in the third aspect of the present invention.
[0078] More preferably, the strain is derived from one or any combination of mammalian cells, plant cells, yeast cells, insect cells, prokaryotic cells, or other similar sources.
[0079] The fifth aspect of the present invention provides a method for synthesizing proteins incorporating non-natural amino acids, using the recombinant aminoacyl-tRNA synthetase described in the first aspect of the present invention, or using the recombinant aminoacyl-tRNA synthetase provided by the genetically engineered strain described in the fourth aspect of the present invention.
[0080] The sixth aspect of the present invention provides a cell-free system for synthesizing proteins containing non-natural amino acids, characterized in that the cell-free system comprises at least: (a) a cell extract, and (b) one or more of the recombinant aminoacyl-tRNA synthetase provided in the first aspect of the present invention, the nucleic acid construct provided in the second aspect of the present invention, or the carrier provided in the third aspect of the present invention; wherein the cell extract is derived from one or any combination of mammalian cells, plant cells, yeast cells, insect cells, prokaryotic cells, or other similar cells.
[0081] The seventh aspect of the present invention provides a cell-free system for synthesizing proteins containing non-natural amino acids, characterized in that the cell-free system includes at least a cell extract derived from the genetically engineered strains provided in the fourth aspect of the present invention.
[0082] More preferably, the cell extract is selected from any one or a combination of the following sources: Escherichia coli, Kluyveromyces lactis, wheat germ cells, insect cells, rabbit reticulocytes, CHO cells, COS cells, VERO cells, BHK cells, human fibrosarcoma HT1080 cells, or a combination thereof.
[0083] More preferably, the cell extract is derived from yeast cells.
[0084] Furthermore, the yeast cells are selected from Pichia pastoris, Pichia finlandica, Pichia trehalophila, Pichia koclamae, Pichia membranaefaciens, Pichia minuta, Ogataeaminuta, Pichia lindneri, Pichia opuntiae, Pichia thermotolerans, Pichia salictaria, Pichia guercuum, Pichia pijperi, Pichiastiptis, Pichia methanolica, Pichia sp., and Saccharomyces cerevisiae. The yeasts include: cerevisiae, brewer's yeast, sugarcane molasses yeast, Saccharomyces sp., Hansenula polymorpha, Candida utilis, Kluyveromyces, or a combination thereof.
[0085] Further, and even more preferably, the Kluyveromyces further includes: Kluyveromyces lactis (K. lactis), Kluyveromyces marxianus, Kluyveromyces dobzhanskii, Kluyveromyces aestuarii, Kluyveromyces nonfermentans, Kluyveromyces wickerhamii, Kluyveromyces thermotolerans, Kluyveromyces fragilis, Kluyveromyces hubeiensis, Kluyveromyces polysporus, Kluyveromyces siamensis, and Kluyveromyces yaros. One or a combination of (yarrowii); preferably, the yeast cell is a Kluyveromyces cell, more preferably a Kluyveromyces lactis cell.
[0086] More preferably, the cell-free system further includes: non-natural amino acids, orthogonal tRNA, and a template containing the gene sequence of the target protein, wherein the codons encoding the amino acids in the gene sequence of the target protein are mutated.
[0087] More preferably, the non-natural amino acid has the structural formula of compound (2) or its salt form.
[0088] Where n is selected from natural numbers from 1 to 20, R1 is selected from substituted or unsubstituted C5-C60 aryl or heteroaryl, substituted or unsubstituted C1-C20 alkyl, substituted or unsubstituted C2-C20 alkenyl or substituted or unsubstituted C2-C20 alkynyl, and A is selected from O or -CH2-.
[0089] In another preferred embodiment, n is selected from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20.
[0090] In another preferred example, n is selected from natural numbers from 1 to 10.
[0091] In another preferred embodiment, n is selected from 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
[0092] In another preferred example, n is selected from natural numbers from 1 to 6.
[0093] In another preferred embodiment, n is selected from 1, 2, 3, 4, 5, or 6.
[0094] In another preferred embodiment, R1 is selected from substituted or unsubstituted C5-C30 aryl or heteroaryl groups.
[0095] In another preferred embodiment, R1 is selected from substituted or unsubstituted phenyl groups.
[0096] In another preferred embodiment, R1 is selected from substituted or unsubstituted C2-C20 alkenyl groups.
[0097] In another preferred embodiment, R1 is selected from substituted or unsubstituted C2-C10 alkenyl groups.
[0098] In another preferred embodiment, R1 is selected from substituted or unsubstituted C2-C6 alkenyl groups.
[0099] In another preferred embodiment, R1 is selected from substituted or unsubstituted C2-C20 alkynyl groups.
[0100] In another preferred embodiment, R1 is selected from substituted or unsubstituted C2-C10 alkynyl groups.
[0101] In another preferred embodiment, R1 is selected from substituted or unsubstituted C2-C6 alkynyl groups.
[0102] In another preferred embodiment, A is selected from O.
[0103] In another preferred embodiment, the A is selected from -CH2-.
[0104] In another preferred embodiment, the substituent is a common substituent group in the art, such as aryl, heteroaryl, alkyl, cycloalkyl, aryloxy, heteroaryloxy, alkyloxy, cycloalkyloxy, hydroxyl, mercapto, ester, carboxyl, cyano, halogen, nitro, sulfonic acid, azide, alkenyl, alkynyl, phosphate, etc.
[0105] In another preferred embodiment, the structural formula of the non-natural amino acid is selected from one or a combination of the following, or its salt form:
[0106] More preferably, the target protein is selected from: luciferin, luciferase (such as firefly luciferase), fluorescent protein (such as green fluorescent protein, yellow fluorescent protein), aminoacyl-tRNA synthetase, glyceraldehyde-3-phosphate dehydrogenase, catalase, actin, variable regions of antibodies, luciferase mutations, α-amylase, enterotoxin A, hepatitis C virus E2 glycoprotein, insulin precursor, interferon αA, cytokines, interferon α2b, interleukin-1β, lysozyme, serum albumin, single-chain antibody fragment (scFV), thyroxine transporter, tyrosinase, xylanase, or combinations thereof.
[0107] More preferably, the target protein includes a wild-type protein, a mutant protein, or a recombinant protein.
[0108] The eighth aspect of the present invention provides a method for preparing the genetically engineered strain described in the fourth aspect of the present invention, characterized in that the nucleic acid construct described in the second aspect of the present invention is transferred into or integrated into cells by transformation, transfection or gene editing technology.
[0109] More preferably, the nucleic acid constructs described in the second aspect of the present invention are integrated into the genome of a cell using gene editing technology.
[0110] More preferably, the nucleic acid construct is integrated into the cell's genome through an active site.
[0111] More preferably, the nucleic acid construct further includes a promoter and a terminator.
[0112] More preferably, the promoter is selected from PGK1, GAP1, ADH1, HXK1, GAPDH1, TEF1 or TIF11.
[0113] More preferably, the terminator is selected from CYC1, GPM1, TDH2 or ACT1.
[0114] More preferably, the site is selected from Lys1-5, glpA, or UPF1.
[0115] The ninth aspect of the present invention provides a reagent kit, characterized in that the reagent kit contains the reaction system described in the sixth or seventh aspect of the present invention.
[0116] The tenth aspect of the present invention provides a method for in vitro synthesis of proteins containing non-natural amino acids, characterized in that the preparation is carried out using the cell-free system described in the sixth or seventh aspect of the present invention or the kit described in the ninth aspect of the present invention.
[0117] The advantages of this invention are:
[0118] (1) Structures containing His tags were designed at the N-terminus and C-terminus of natural or modified MaPylRS synthase, which further promotes enzyme stability, makes it easier to purify, and improves enzyme catalytic activity, which is beneficial to improve the efficiency of non-natural amino acid introduction.
[0119] (2) Linking His tag, T7 tag and thrombin restriction site to the N-terminus of natural or modified MaPylRS synthase will further promote enzyme stability, make it easier to purify, and improve enzyme catalytic activity, which is beneficial to improve the efficiency of non-natural amino acid introduction.
[0120] (3) The N-terminal T7 tag and thrombin cleavage site of natural or modified MaPylRS synthase further promote enzyme stability, make it easier to purify, and improve enzyme catalytic activity, which is beneficial to improving the efficiency of non-natural amino acid introduction.
[0121] (4) This invention integrates the aforementioned designed MaPylRS recombinant expression structure into the cell genome using CRISPR / Cas9 and efficient transformation technology, thereby achieving the stable existence of MaPylRS in the cell genome and the continuous expression of MaPylRS protein.
[0122] (5) The Kluyveromyces strain with inserted recombinant MaPylRS was prepared into an in vitro expression system, which realized the site-directed insertion of non-natural amino acids into the exogenous target protein, greatly simplified the preparation steps, saved costs, and increased the stability of the synthesized protein with inserted non-natural amino acids. Attached Figure Description
[0123] Figure 1 shows the construction graph of long N-terminal mapylrs.
[0124] Figure 2 shows the N-his mapylrs construction diagram.
[0125] Figure 3 shows the construction diagram of C-his mapylrs.
[0126] Figure 4 shows the electrophoresis diagrams of the lysate supernatants of the three proteins after induction. 1 represents long N-terminal mapylrs, 2 represents N-his mapylrs, and 3 represents C-his mapylrs. Mapylrs with longer N-terminals exhibit higher protein expression levels and greater protein solubility. The highest protein concentration reached 120 mg / mL (3.4 mM).
[0127] Figure 5 shows a comparison of the activity of proteins expressed in three different forms with the addition of non-natural amino acids at the same concentration. All three forms of mapylrs can catalyze the expression of proteins containing non-natural amino acids. In the figure, the inverted triangles represent those without added non-natural amino acids, and the dots represent those with added non-natural amino acids. nhis represents N-his mapylrs, chis represents C-his mapylrs; N-terminal represents long N-terminal mapylrs.
[0128] Figure 6 shows a comparison of the activity of original mapylrs (referring to unrecombined mapylrs) and recombined mapylrs (long N-terminal mapylrs and N-his mapylrs) after the introduction of non-natural amino acids. In the figure, nhis represents N-his mapylrs, long terminal represents long N-terminal mapylrs, and the left side of each comparison bar represents the addition of non-natural amino acids, while the right side represents the addition of non-natural amino acids.
[0129] Figure 7 shows a comparison of RFP / GFP with original mapylrs and recombinant mapylrs (long N-terminal mapylrs and N-his mapylrs) containing non-natural amino acids. In the figure, nhis represents N-his mapylrs, long terminal represents long N-terminal mapylrs, and no tag represents original mapylrs.
[0130] Figure 8 shows a schematic diagram of the pKM-CAS1.0-KlLys1-5 plasmid structure.
[0131] Figure 9 shows a schematic diagram of the pKM-MaPylRS plasmid structure.
[0132] Figure 10 shows a schematic diagram of the pKM-CAS1.0-KlglpA plasmid structure.
[0133] Figure 11 shows a schematic diagram of the pKM-CAS1.0-KlUPF1 plasmid structure.
[0134] Figure 12 shows the RFP activity measured in Example 9.
[0135] Figure 13 shows the GFP activity measured in Example 9.
[0136] Figure 14 shows the RFP / GFP activity measured in Example 9.
[0137] Figure 15 shows the GFP activity measured in Example 10.
[0138] Figure 16 shows the RFP activity measured in Example 10.
[0139] Figure 17 shows the RFP / GFP activity measured in Example 10.
[0140] In the accompanying figures of this article, the presence of "non" or "control" indicates that non-natural amino acids were not added during the reaction. Detailed Implementation
[0141] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0142] The present invention will be further described below with reference to specific embodiments and examples. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments that do not specify specific conditions should preferably be performed according to the conditions indicated in the specific embodiments described above, and then may be performed under conventional conditions or as recommended by the manufacturer.
[0143] Unless otherwise stated, percentages and parts mentioned in this invention are weight percentages and weight parts.
[0144] Unless otherwise specified, all materials and reagents used in the embodiments of this invention are commercially available products.
[0145] Unless otherwise specified, all temperature units in this application are Celsius (°C).
[0146] Nouns and terms
[0147] The following are explanations or descriptions of the meanings of some relevant "nouns" and "terms" used in this invention to facilitate a better understanding of the invention. These explanations or descriptions apply to the entire text of this invention, both below and above. When references are made in this invention, the definitions of relevant terms, nouns, and phrases in the referenced documents are also cited; however, in case of conflict with the definitions in this invention, the definitions in this invention shall prevail. The conflict between the definitions in the referenced documents and the definitions in this invention does not affect the application of the cited components, substances, compositions, materials, systems, formulations, types, methods, equipment, etc., which shall be determined in the referenced documents.
[0148] In this invention, preferred embodiments such as “preferred,” “better,” “more preferred,” “even better,” “most preferred,” and “further preferred” do not constitute any limitation on the scope of the invention or its protection scope, and are not intended to limit the scope and implementation of the invention, but are only used to provide some embodiments as examples.
[0149] In the description of this invention, terms such as "one preferred method," "one preferred embodiment," "one preferred example," "preferred example," "in a preferred embodiment," "in some preferred examples," "in some preferred methods," "preferred," "preferred," "more preferred," "more preferably," "further preferred," and "most preferred," as well as illustrative enumerations such as "one embodiment," "one method," "example," "specific example," "for instance," "as an example," "for example," "like," etc., do not constitute any limitation on the scope or protection of the invention. The specific features described in each method are included in at least one specific embodiment of this invention. In this invention, the specific features described in each method can be combined in any suitable manner in one or more specific embodiments. In this invention, the technical features or technical solutions corresponding to each preferred method can also be combined in any suitable manner.
[0150] In this invention, "any combination thereof" means "greater than 1" in quantity and "a group consisting of the following situations in terms of scope: "any one of them, or a group consisting of at least two of them".
[0151] In this invention, the descriptions of "one or more", "one or more", etc., have the same meaning as "at least one", "at least one", "a combination thereof", "or a combination thereof", "and a combination thereof", "or any combination thereof", "and any combination thereof", etc., and can be used interchangeably to indicate a quantity equal to "1" or "greater than 1".
[0152] In this invention, "or / and" or "and / or" means "either one or a combination thereof", or at least one of them.
[0153] The term “about” can refer to a value or composition within an acceptable margin of error for a particular value or composition as determined by a person skilled in the art, depending in part on how the value or composition is measured or determined. For example, as used herein, the expression “about 100” includes all values between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).
[0154] Sequence identity (or homology) is determined by comparing two aligned sequences along a predetermined comparison window (which may be 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the length of a reference nucleotide sequence or protein) and determining the number of positions where identical residues occur. This is typically expressed as a percentage. The measurement of sequence identity of nucleotide sequences is a method well known to those skilled in the art.
[0155] The prior art methods described in this invention using terms such as "usually", "conventionally", "generally", "frequently", and "often" are also cited as references to the content of this invention. Unless otherwise specified, they can be regarded as one of the preferred embodiments of some technical features of this invention. It should be noted that this does not constitute any limitation on the scope of coverage and protection of the invention.
[0156] All documents mentioned in this invention, and those directly or indirectly cited by such documents, are incorporated herein by reference as if each document were cited individually.
[0157] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (including but not limited to embodiments) can be combined with each other to form new or preferred technical solutions, as long as they can be used to implement this invention. Due to space limitations, they will not be described in detail here.
[0158] In vitro protein synthesis refers to the synthesis of proteins in a cell-free in vitro synthesis system, including at least the translation process. This includes, but is not limited to, IVT (in vitro translation), IVTT (in vitro transcription-translation), and IVDTT (in vitro replication-transcription-translation). In this invention, the IVTT reaction is preferred. The IVTT reaction, corresponding to the IVTT system, is the process of transcribing and translating DNA into protein in vitro. Therefore, we also refer to this type of in vitro protein synthesis system as a D2P system, D-to-P system, or DNA-to-Protein system; and the corresponding in vitro protein synthesis methods are also referred to as D2P methods, D-to-P methods, or DNA-to-Protein methods.
[0159] "Cell-free system" refers to a method of in vitro protein synthesis that does not involve secretion and expression by intact cells. It should be noted that in the in vitro cell-free protein synthesis system of this invention, the addition of cellular components to promote the reaction is also permitted, but the added cells are not primarily intended for the secretion and expression of exogenous target proteins. Furthermore, in the D2P system constructed under the guidance of this invention, the intentional addition of a small number of intact cells (e.g., whose protein content does not exceed 30 wt% compared to the protein content provided by cell extracts) is also within the scope of protection of this invention.
[0160] Target protein: The target expression product of the in vitro protein synthesis system of this invention is not synthesized by host cell secretion, but is synthesized in vitro based on an exogenous nucleic acid template, and can also be called an exogenous protein. The exogenous protein can be a protein, a fusion protein, or a mixture containing protein molecules or fusion protein molecules; it also broadly includes polypeptides. The product obtained after the in vitro protein synthesis reaction based on the nucleic acid template encoding the target protein can be a single substance or a combination of two or more substances. "Exogenous protein," "target protein," "target protein," and "target translation product" have the same meaning and can be translated as "objective protein," "interested protein," "objective translated product," "interested protein product," etc., and can be used interchangeably in this invention.
[0161] D2P, DNA-to-Protein, refers to the process of converting a DNA template into a protein product. Examples include D2P technology, D2P systems, D2P methods, and D2P kits.
[0162] The terms "expression system of the present invention," "in vitro expression system of the present invention," "in vitro cell-free expression system," and "in vitro cell-free expression system" are used interchangeably and all refer to the in vitro protein expression system of the present invention. Other descriptive methods may also be used, such as: in vitro protein synthesis system, in vitro protein synthesis system, cell-free system, cell-free protein synthesis system, cell-free in vitro protein synthesis system, in vitro cell-free protein synthesis system, in vitro cell-free synthesis system, CFS system (cell-free system), CFPS system (cell-free protein synthesis system), etc. Depending on the reaction mechanism, it may include an in vitro translation system (which can be abbreviated as IVT system, a type of mR2P system), an in vitro transcription-translation system (which can be abbreviated as IVTT system, a type of D2P system), an in vitro replication-transcription-translation system (which can be abbreviated as IVDTT system, a type of D2P system), etc. In this invention, the IVTT system is preferred. We also refer to the in vitro protein synthesis system as a "protein factory" (or "protein factory"). The in vitro protein synthesis system provided by this invention uses an open-ended description of its components. The cell-free protein synthesis system of the present invention uses exogenous DNA, mRNA or a combination thereof as nucleic acid templates for protein synthesis, and achieves in vitro synthesis of target proteins by artificially controlling the addition of substrates and transcription and translation-related protein factors required for protein synthesis.
[0163] In this invention, "protein" and "protein protein" have the same meaning and are both translated as protein, and can be used interchangeably.
[0164] In this invention, "system" and "structure" are both translated as "system" and can be used interchangeably.
[0165] In this invention, "protein synthesis amount", "protein expression amount" and "protein expression yield" have the same meaning and can be used interchangeably.
[0166] In this invention, cell extract, cell extract solution, cell lysate, cell fragments, and cell lysate have the same meaning and can be used interchangeably. In English, they can be described as cell extract, cell lysate, etc.
[0167] In this invention, the terms "energy system," "energy system," and "energy supply system" have equivalent meanings and can be used interchangeably. Similarly, "energy regeneration system" and "energy regeneration system" have equivalent meanings and can be used interchangeably. An energy regeneration system is a preferred embodiment or component of an energy system.
[0168] Furthermore, the present invention provides a cell-free protein synthesis system, comprising at least cell extracts or cell lysates.
[0169] More preferably, the cell-free protein synthesis system further includes one or more components selected from the group consisting of: a substrate for RNA synthesis, a substrate for protein synthesis, polyethylene glycol or an analogue thereof, magnesium ions, potassium ions, a buffer, RNA polymerase, an energy regeneration system, dithiothreitol, and optionally an aqueous solvent.
[0170] More preferably, the substrate for the synthesized RNA includes one or a combination of nucleoside monophosphate, nucleoside triphosphate, or nucleoside triphosphate.
[0171] More preferably, the substrate for the synthesized protein includes 20 natural amino acids and non-natural amino acids.
[0172] More preferably, the magnesium ions are derived from a magnesium ion source, which is selected from the group consisting of magnesium acetate, magnesium glutamate, or a combination thereof.
[0173] More preferably, the potassium ions are derived from a potassium ion source, which is selected from the group consisting of potassium acetate, potassium glutamate, or a combination thereof.
[0174] More preferably, the energy regeneration system is selected from the group consisting of: creatine phosphate / creatine phosphate enzyme system, glycolysis pathway and its intermediate product energy system, or a combination thereof.
[0175] More preferably, the energy regeneration system includes a glucose / phosphate system, wherein the phosphate is selected from the group consisting of tripotassium phosphate, triammonium phosphate, trisodium phosphate, dipotassium hydrogen phosphate, diammonium hydrogen phosphate, disodium hydrogen phosphate, potassium dihydrogen phosphate, ammonium dihydrogen phosphate, sodium dihydrogen phosphate, or a combination thereof.
[0176] More preferably, the buffer is selected from the group consisting of 4-hydroxyethylpiperazine ethanesulfonic acid, tris(hydroxymethyl)aminomethane, or a combination thereof.
[0177] Further preferably, the in vitro protein synthesis system contains polyethylene glycol (PEG) or an analogue thereof. The concentration of PEG or an analogue thereof is not particularly limited, but typically, the concentration (w / v) of PEG or an analogue thereof is 0.1-8%, more preferably 0.5-4%, and even more preferably 1-2%, based on the total weight of the protein synthesis system. Representative PEGs are selected from the group consisting of PEG3000, PEG3350, PEG6000, PEG8000, or combinations thereof.
[0178] More preferably, the polyethylene glycol includes polyethylene glycol with a molecular weight (Da) of 200-10000, such as PEG200, 400, 1500, 2000, 4000, 6000, 8000, 10000, etc., and more preferably, polyethylene glycol with a molecular weight of 3000-10000.
[0179] In this invention, the RNA polymerase is not particularly limited and can be selected from one or more RNA polymerases, with T7 RNA polymerase being a typical RNA polymerase.
[0180] An optional approach is that the in vitro protein synthesis system provided by the present invention includes: cell extract, 4-hydroxyethylpiperazine ethanesulfonic acid, potassium acetate, magnesium acetate, adenine nucleoside triphosphate (ATP), guanine nucleoside triphosphate (GTP), cytosine nucleoside triphosphate (CTP), thymidine nucleoside triphosphate (TTP), a mixture of amino acids, creatine phosphate, dithiothreitol (DTT), creatine phosphate kinase, and RNA polymerase.
[0181] In this invention, the cell extract does not contain intact cells. Typical cell extracts include ribosomes for protein translation, aminoacyl-tRNA synthetase, initiation and elongation factors required for protein synthesis, and termination release factors. Furthermore, the cell extract also contains other proteins derived from the cytoplasm of cells, especially soluble proteins.
[0182] In this invention, the proportion of the cell extract in the in vitro cell-free protein synthesis system is not particularly limited. Typically, the cell extract accounts for 20-70% of the system, preferably 30-60%, and more preferably 40-50%.
[0183] In this invention, the protein content of the cell extract is 20-100 mg / mL, preferably 50-100 mg / mL. The method for determining the protein content is the Coomassie Brilliant Blue assay.
[0184] This invention also provides a vector or combination of vectors containing the nucleic acid constructs of this invention. Preferably, the vector is selected from bacterial plasmids, bacteriophages, yeast plasmids, animal cell vectors, and shuttle vectors; the vector is a transposon vector. Methods for preparing recombinant vectors are well known to those skilled in the art. Any plasmid and vector can be used as long as it can replicate and remain stable in the host.
[0185] Those skilled in the art can use well-known methods to construct expression vectors containing the promoter and / or target gene sequence described in this invention. These methods include in vitro recombinant DNA technology, DNA synthesis technology, in vivo recombination technology, etc.
[0186] Template DNA
[0187] Template DNA is a nucleotide sequence encoding any target protein to be synthesized. It can be a primitive sequence, an artificially synthesized sequence, or an artificially modified sequence. The corresponding RNA and / or protein can be synthesized using this template DNA.
[0188] In this invention, the preparation method of the cell extract is not limited, but a preferred preparation method is described below.
[0189] Includes the following steps:
[0190] (i) Provide cells;
[0191] (ii) The cells are washed to obtain washed cells;
[0192] (iii) The washed cells are subjected to cell-breaking treatment to obtain crude cell extract;
[0193] (iv) The crude cell extract is subjected to solid-liquid separation to obtain the liquid fraction, which is the cell extract.
[0194] In this invention, the solid-liquid separation method is not particularly limited, but centrifugation is a preferred method.
[0195] In a preferred embodiment, the centrifugation is performed in a liquid state.
[0196] In this invention, the centrifugation conditions are not particularly limited, but a preferred centrifugation condition is 5000-100000g, and more preferably, 8000-30000g.
[0197] In this invention, the centrifugation time is not particularly limited, but a preferred centrifugation time is 0.5 min to 2 h, and more preferably, 20 min to 50 min.
[0198] In this invention, the temperature of the centrifugation is not particularly limited. Preferably, the centrifugation is carried out at 1-10°C, and more preferably, at 2-6°C.
[0199] In this invention, the washing treatment method is not particularly limited. A preferred washing treatment method is to use a washing solution at a pH of 7-8 (preferably 7.4). The washing solution is not particularly limited, and a typical washing solution is selected from the group consisting of potassium 4-hydroxyethylpiperazine ethanesulfonate, potassium acetate, magnesium acetate, or a combination thereof.
[0200] In this invention, the method of cell disruption is not particularly limited, but a preferred method of cell disruption includes high-pressure disruption and freeze-thaw (e.g., liquid nitrogen cryogenic) disruption.
[0201] The nucleoside triphosphate mixture in the in vitro cell-free protein synthesis system comprises adenine nucleoside triphosphate, guanine nucleoside triphosphate, cytosine nucleoside triphosphate, and uracil nucleoside triphosphate. In this invention, the concentration of each mononucleotide is not particularly limited; typically, the concentration of each mononucleotide is 0.5-5 mM, preferably 1.0-2.0 mM.
[0202] The amino acid mixture in the in vitro cell-free protein synthesis system may include natural or non-natural amino acids, and may include D-type or L-type amino acids. Representative amino acids include (but are not limited to) 20 natural amino acids: glycine, alanine, valine, leucine, isoleucine, phenylalanine, proline, tryptophan, serine, tyrosine, cysteine, methionine, asparagine, glutamine, threonine, aspartic acid, glutamic acid, lysine, arginine, and histidine. The concentration of each amino acid is typically 0.01-0.5 mM, preferably 0.02-0.2 mM, such as 0.05, 0.06, 0.07, or 0.08 mM.
[0203] In a preferred embodiment, the in vitro cell-free protein synthesis system further contains polyethylene glycol or its analogues. The concentration of polyethylene glycol or its analogues is not particularly limited, but typically, the concentration (w / v) is 0.1-8%, more preferably 0.5-4%, and even more preferably 1-2%, based on the total weight of the biosynthetic system. Representative examples of PEG include (but are not limited to): PEG3000, PEG8000, PEG6000, and PEG3350. It should be understood that the system of the present invention may also include polyethylene glycols of various other molecular weights (such as PEG200, 400, 1500, 2000, 4000, 6000, 8000, 10000, etc.).
[0204] In a preferred embodiment, the in vitro cell-free protein synthesis system further contains sucrose. The concentration of sucrose is not particularly limited, but typically it is 0.03-40 wt%, more preferably 0.08-10 wt%, and even more preferably 0.1-5 wt%, based on the total weight of the protein synthesis system.
[0205] A particularly preferred in vitro cell-free protein synthesis system, in addition to yeast cell extract, contains the following components: 22 mM 4-hydroxyethylpiperazine ethanesulfonic acid at pH 7.4, 30-150 mM potassium acetate, 1.0-5.0 mM magnesium acetate, 1.5-4 mM nucleoside triphosphate mixture, 0.08-0.24 mM amino acid mixture, 25 mM creatine phosphate, 1.7 mM dithiothreitol, 0.27 mg / mL creatine phosphate kinase, 1%-4% polyethylene glycol, 0.5%-2% sucrose, and 0.027-0.054 mg / mL T7 RNA polymerase.
[0206] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only.
[0207] This is not intended to limit the scope of the invention. Experimental methods in the following examples, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sam Brook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight. This invention uses Kluyveromyces lactis (K. lactis or Kl) as an example, but the same design, analysis, and experimental methods are applicable to other yeasts, animal cells, eukaryotic cells, and prokaryotic cells.
[0208] This invention uses Kluyveromyces lactis (K. lactis) as an example, but the same design, analysis, and experimental methods are applicable to other yeasts and other lower eukaryotic cells as well as higher animal cells. The gene modification method used in this invention is CRISPR-Cas9 technology, but it is not limited to this and can be any known or existing gene modification method.
[0209] An in vitro protein synthesis reaction mixture system, also described as an in vitro protein synthesis reaction mixture, reaction mixture system, or reaction mixture, refers to a mixture system including an in vitro protein synthesis system and a nucleic acid template encoding the target protein; it can be homogeneous or heterogeneous, and is allowed to be a liquid system such as a solution, emulsion, or suspension.
[0210] Protein described in this invention The final concentrations of each component in the Factory are as follows: 80% (v / v) Kluyveromyces lactis extract, 15 mM glucose, 320 mM maltodextrin (molar concentration based on glucose monomer), 24 mM tripotassium phosphate, 1.8 mM nucleoside triphosphate mixture (a mixture of adenine, guanine, cytosine, and uracil triphosphates, each with a final concentration of 1.8 mM), 0.7 mM amino acid mixture (glycine, alanine, valine, leucine, isoleucine, phenylalanine, proline, tryptophan, serine, tyrosine, cysteine, methionine, asparagine, glutamine, threonine, aspartic acid, glutamic acid, lysine, arginine, and histidine, each with a final concentration of 0.7 mM), magnesium L-aspartate, 80 mM potassium acetate, 2% (w / v) polyethylene glycol 8000, and 9.78 mM pH 8.0. Tris·HCl buffer and 6% (w / v) trehalose were used. The preparation of *Kluyveromyces lactis* cell extract was carried out using conventional techniques, referring to the method described in CN109593656A. The preparation steps generally include: providing an appropriate amount of fermented *Kluyveromyces lactis* cells as raw material; flash-freezing the cells with liquid nitrogen; breaking the cells; centrifuging and collecting the supernatant to obtain the cell extract. The protein concentration in the obtained *Kluyveromyces lactis* cell extract was 20–40 mg / mL. In the following examples, [the following text is incomplete and requires further context to translate accurately]. (i.e., prock, N-E-propoxycarbonyl-L-lysine hydrochloride) is a representative of non-natural amino acids (abbreviated as ncaa), but it is not limited to ncaa in this application referring only to prock.
[0211] Example 1: Three construction methods for Mapylrs protein:
[0212] The first type, Long N-terminal mapylers:
[0213] The N-terminus of the Mapylrs protein contains a 31-amino acid leader peptide, which includes a 6*His affinity purification tag, a T7 tag, and a thrombin restriction site (Figure 1).
[0214] Long N-terminal mapylrs amino acid sequence:
[0215] Theoretical molecular weight: 34349.98 Da
[0216] The second type, N-his mapylrs:
[0217] The Mapylrs protein has only a 6*His affinity purification tag at its N-terminus and a flexible interface (Figure 2):
[0218] Theoretical molecular weight of N-his mapylrs: 32091.35 Da
[0219] The third type, C-his mapylrs:
[0220] The Mapylrs protein has one more Gly at the N-terminus than the reported sequence, and a 6*His affinity purification tag at the C-terminus (Figure 3):
[0221] C-his mapylrs theoretical molecular weight: 31685.98
[0222] The fourth type, compared to Long N-terminal mapylrs, removes the his tag:
[0223] Example 2: Construction of expression plasmids
[0224] Long-N-terminal mapylrs expression plasmid:
[0225] The complete mapylrs genome (WP_015505008) was synthesized by Sangon Biotech. Codons were optimized for expression in *E. coli*. The synthesized gene sequence is as follows:
[0226] Nde I
[0227] Xho I
[0228] Insert the pET28a vector through the NdeI / Xho I restriction site.
[0229] The N-his maplys and C-his maplys expression vectors were modified from pET28a-long-terminal mapylrs using PCR. The primers used were:
[0230] N-his mapylrs
[0231] C-his mapylrs:
[0232] C-his vector R:
[0233] The amplified product was digested with DnpI enzyme, ligated and transformed into DH5α competent cells, and the plasmid was extracted and sequenced.
[0234] Example 3
[0235] 1) Pick a single bacterial clone, inoculate it with 100ml LB (containing 100mg / L kanamycin), and incubate overnight at 37°C.
[0236] 2) On the second day, inoculate the culture medium at a ratio of 1:100 into fresh LB medium (containing 100 mg / L kanamycin), incubate at 37°C until OD600≈0.6, add IPTG to a final concentration of 0.1 mM, and induce expression at 16°C for 20 hours.
[0237] 3) Collect bacterial cells by centrifugation at 4℃ and 5000 rpm for 20 min. Resuspend the bacterial cells in 100 ml of lysis buffer (25 mM Tris-HCl, 500 mM NaCl, 25 mM imidazole, 5 mM β-mercaptoethanol, 1 mM PMSF, 0.1% Triton X-100) at a ratio of 20 g wet bacterial cells. Lyse the cells using a high-pressure homogenizer.
[0238] 4) Centrifuge twice at 20,000 rpm for 20 min at 4℃, repeating to remove bacterial debris. Use a HisTrap affinity chromatography column to separate and purify the target protein from the supernatant. Buffer A: 25 mM TrisHCl pH 7.8, 500 mM NaCl, 25 mM imidazole; Buffer B: 25 mM TrisHCl pH 7.8, 150 mM NaCl, 500 mM imidazole;
[0239] 5) After the cell lysis supernatant is passed through the affinity chromatography column, the affinity column is repeatedly washed with 10 cv buffer A, and then the target protein is eluted with 10 cv buffer B from 0% to 100%.
[0240] 6) Collect and combine the components containing the target protein, and concentrate the sample using ultrafiltration centrifuge tubes;
[0241] 7) After concentration, the sample was dialyzed thoroughly with 50% glycerol and 25 mM Hepes pH 7.5. Protein concentration was then determined, and the samples were aliquoted and stored at -80°C for at least one year. As shown in Figure 4, the supernatant of the first three forms of recombinant mapylrs was lysed after induced expression. Mapylrs with a longer N-terminal showed higher protein expression levels and higher protein solubility. The highest protein concentration reached 120 mg / mL (3.4 mM).
[0242] Example 4 compares the activity of proteins purified by three different construction methods.
[0243] The concentrations of the three proteins were determined using a nanodrop, and the concentrations of the three proteins were adjusted to be consistent using ddH2O that does not contain DNase Rnase.
[0244] Establishment of expression systems containing non-natural amino acids:
[0245] Protein Factory 100ul
[0246] Prock (non-natural amino acids) 500mM 1ul
[0247] matRNA CUA pyl 10ul of in vitro transcription product (unpurified)
[0248] 3ul of GFP-TAG-RFP dual fluorescent reporter gene PCR product
[0249] Different forms of Mapylrs purified protein to a final concentration of 5 μM
[0250] GFP fluorescence intensity was detected at Ex485nm / Em535nm, and RFP fluorescence intensity was detected at Ex535nm / Em595nm. The introduction efficiency of non-natural amino acids was determined based on the fluorescence intensity of RFP and the RFP / GFP ratio (see Figure 5 for details).
[0251] As shown in Figure 5, based on RFP / GFP, all three forms of mapylers can achieve efficient delivery of non-natural amino acids and exhibit high activity. The N-terminal tag, in particular, enhances activity by stabilizing the protein structure (long-N-terminal mapylers).
[0252] Example 5 examines the effect of purified products of raw Mapylrs (i.e., unrecombined Mapylrs) and N-terminal tagged Mapylrs on the activity of introducing non-natural amino acids.
[0253] The activity of the original Mapylrs was compared with that of Long-N-terminal Mapylrs and N-his Maplys. The activity assay method was the same as in Example 4, using the dual fluorescent protein expression method.
[0254] As shown in Figures 6 and 7, compared with the original Mapylrs (with no tag in the figure), both Long-N-terminal mapylrs and N-his maplys showed improved activity in introducing non-natural amino acids. This indicates that the recombinant Mapylrs obtained after the modification of this invention have significant progress in introducing non-natural amino acids compared with the original Mapylrs.
[0255] To overcome the shortcomings of existing non-natural amino acid insertion systems that require the manual addition of MaPylRS protein or the introduction of plasmids containing its expression structure through transformation / transfection, this invention further discloses a method for integrating MaPylRS protein into the cell genome using gene editing technology, creating a strain capable of stably and appropriately expressing MaPylRS protein, thereby forming a simple and efficient non-natural amino acid insertion system that does not require the addition of MaPylRS protein from the outside.
[0256] The following example uses Long-N-terminal maplyrs with the his tag removed (i.e., the fourth structure) to verify the integration of the MaPylRS protein into the cell genome. This does not limit the other MaPylRS in this invention.
[0257] Example 6: The MaPylRS expression cassette (i.e., the fourth type of expression cassette, hereinafter the same) was inserted near KlLys1-5 using CRISPR-Cas9 technology.
[0258] (1) KlLys1-5 sequence retrieval and CRISPR gRNA sequence determination
[0259] In order not to affect the normal expression of other genes in Kluyveromyces lactis, the present invention inserts the MaPylRS expression structure into the vicinity of KlLys1-5 of the Kluyveromyces lactis tDNA, and after expression, it binds to non-natural amino acids to perform site-directed insertion.
[0260] i. A search of http: / / gtrnadb.ucsc.edu / GtRNAdb2 / genomes / eukaryota / Kluy_lact_NRRL_Y_1140 / Kluy_lact_NRRL_Y_1140-gene-list.html yielded tRNA-Lys-CTT-1-5, providing the Lys1-5 gene sequence from K. lactis yeast. In this invention, this sequence is named KlLys1-5 (located on chromosome E, 856274-856346).
[0261] ii. Search for PAM sequences (NGG) within a 500bp range upstream and downstream of KlLys1-5, and finally select the PAM located downstream of the gene (positions 856876...856878 on chromosome E), and determine the KlLys1-5 gRNA sequence (GTTCCCATTGATCCCATATC (SEQ ID NO: 11), located at positions 856856...856875 on chromosome E).
[0262] (2) Construction of KlLys1-5 CRISPR-Cas9 plasmid
[0263] Based on the designed gRNA sequence, two 24nt primers were designed for vector construction. gRNA-F1: AATCGTTCCCATTGATCCCATATC (SEQ ID NO: 12); gRNA-R1: AAACGATATGGGATCAATGGGAAC (SEQ ID NO: 13). The primers were diluted to 10 μM, and 10 μL each of gRNA-F1 and gRNA-R2 were added to PCR tubes. After mixing, the mixture was centrifuged to the bottom of the tube, and annealing was performed according to the following procedure:
[0264] After 3 min at 95℃; 30 s at 72℃; 2 min at 65℃; 2 min at 60℃; 2 min at 55℃; 2 min at 50℃; and 2 min at 16℃, the ligation reaction was carried out.
[0265] A. Reaction system: 10×Buffer 1μL, plasmid 20-50ng, annealing product 1μL, enzyme 0.2μL, add water to 10μL.
[0266] B. Reaction procedure: 16℃, 60min.
[0267] Take 50 μL of commercially available E. coli DH5α competent cells, add all ligation products and mix well, then complete the transformation process according to the manufacturer's instructions. Screen on LB agar plates containing 50 mg / L kanamycin and culture overnight. Pick 5 single clones and culture with shaking in LB liquid medium. After sequencing confirmation of positive results, extract and preserve the plasmid, naming it pKM-CAS1.0-KlLys1-5 (Figure 8).
[0268] (3) Donor DNA construction and amplification
[0269] First, the donor DNA, namely the homologous recombination sequence at sites KlLys1-5 and the MaPylRS expression structure, was constructed. The KlPGK1 promoter and ScCYC1 terminator were used for MaPylRS. The donor DNA construction and transformation methods are as follows:
[0270] iii. Gene synthesis of a plasmid containing the MaPylRS expression cassette (named pKM-MaPylRS, see Figure 9), and using it as a template, PCR amplification was performed with primers PF1: AATTGTTCCCATTGATCCCATATCCTTCGAGCGTCCCAAAACC (SEQ ID NO: 14) and PR1: TTCAGTTCAAAAACGCCCCGTTCCTCATCACTAGAAG (SEQ ID NO: 15) to obtain the MaPylRS expression cassette fragment.
[0271] iv. Using Kluyveromyces lactis free plasmid as a template, PCR amplification was performed with primers PF2:GTTATTAATGTCGTGTGCCATAGGT (SEQ ID NO: 16) and PR2:AGGTTTTGGGACGCTCGAAGGATATGGGATCAATGGGAA (SEQ ID NO: 17) to obtain the homologous arm 1 fragment of the KlLys1-5 site; using Kluyveromyces lactis free plasmid as a template, PCR amplification was performed with primers PF3:TTCTAGTGATGAGGAACGGGGCGTTTTTGAACTGAATTTCG (SEQ ID NO: 18) and PR3:CAGCATAGCATTTGAGTATTGTG (SEQ ID NO: 19) to obtain the homologous arm 2 fragment of the KlLys1-5 site.
[0272] v. Mix the three PCR product fragments obtained above, dilute them 100 times as templates, and perform PCR amplification again with primers PF4:ATAGGTCAATTAATAATATGCCAGCAAT (SEQ ID NO: 20) and PR4:GGGGAGCATAGCATTCAAAAACTTC (SEQ ID NO: 21); the three fragments can be ligated together by overlap extension PCR to become a linear donor DNA. After sequencing confirmation, store it in a -20℃ freezer.
[0273] (4) Preparation and transformation of highly efficient Kluyveromyces lactis competent cells
[0274] Preparation of yeast competent cells
[0275] Streak Kluyveromyces lactis culture on YPD solid medium and pick single colonies, then culture overnight in 25 mL of 2×YPD liquid medium with shaking. Take 2 mL of the culture and culture in 50 mL of 2×YPD liquid medium with shaking for 2–8 h. Collect yeast cells by centrifugation at 3000g for 5 min at 20°C, resuspend in 500 μL of sterile water, and collect cells by centrifugation under the same conditions. Prepare competent cell solution (5% v / v glycerol, 10% v / v DMSO) and dissolve yeast cells in 500 μL of this solution. Aliquot 50 μL into 1.5 mL centrifuge tubes and store at -80°C.
[0276] Yeast DNA transformation
[0277] Thaw competent cells on ice for 30 seconds, add 200 ng of pKM-CAS1.0-KlLys1-5 plasmid and 2000 ng of donor DNA. Electrolyze at 1.5 kV for 5 ms, then immediately add 1 mL of YPD liquid medium and incubate for 2-3 hours. Spread 200 μL onto solid YPD (200 μg / mL G418) medium and incubate for 2-3 days until single colonies appear.
[0278] (5) Positive identification of gene editing
[0279] Pick 50-60 single clones from the Kluyveromyces lactis transformed plate, and place each single clone in 5 μL of yeast lysis buffer (Takara Mighty Prep Reagent for DNA). Using the cell lysis buffer as a template, primers Ma R1 (MaPylRS inner primer): ATCTCTTACTTGAACGGTGCTA (SEQ ID NO: 22); 1-5F1 (KlLys1-5 donor DNA 5' outer primer): GGTTATCCATTCAGGCAATGAAG (SEQ ID NO: 23) and primer Ma F1 (MaPylRS inner primer): CCATGTGAGAACCTCTTGG (SEQ ID NO: 24); 1-5R1 (KlLys1-5 donor DNA 5' outer primer): CAGCATAGCATTTGAGTATTGTG (SEQ ID NO: 23) were used. NO:25) PCR amplification was performed to detect the CRISPR insertion at the KlLys1-5 sites. A positive band indicates that the MaPylRS sequence was successfully inserted into the target site.
[0280] Example 7: The MaPylRS expression cassette was inserted near KlglpA using CRISPR-Cas9 technology.
[0281] (1) KlglpA sequence retrieval and CRISPR gRNA sequence determination
[0282] In order not to affect the normal expression of other genes in Kluyveromyces lactis, this invention inserts the MaPylRS expression structure near KlglpA in the Kluyveromyces lactis tDNA, and after expression, it binds to non-natural amino acids to perform site-directed insertion.
[0283] i. A search for glycerol-3-phosphate dehydrogenase at https: / / www.genome.jp / kegg / kegg2.html yielded the glpA gene sequence from K. lactis yeast. In this invention, this sequence is named KlglpA (located on chromosome A, 33084-35012).
[0284] ii. Search for PAM sequences (NGG) within 1000-2000 bp upstream of KlglpA, and finally select the PAM located downstream of the gene (positions 31956...31958 on chromosome A), and determine the KlglpA gRNA sequence (GAAGTAACTCTAGCCATCGG (SEQ ID NO: 26), located at positions 31936...31955 on chromosome A).
[0285] (2) Construction of KlglpA CRISPR-Cas9 plasmid
[0286] Based on the designed gRNA sequence, two 24nt primers were designed for vector construction. gRNA-F2: AATCGAAGTAACTCTAGCCATCGG (SEQ ID NO: 27); gRNA-R2: AAACCCGATGGCTAGAGTTACTTC (SEQ ID NO: 28). The primers were diluted to 10 μM, and 10 μL each of gRNA-F2 and gRNA-R2 were added to PCR tubes. After mixing, the mixture was centrifuged to the bottom of the tube, and annealing was performed according to the following procedure:
[0287] 95℃, 3min; 72℃, 30s; 65℃, 2min; 60℃, 2min; 55℃, 2min; 50℃, 2min; 16℃, 2min
[0288] Then the connection reaction occurs.
[0289] A. Reaction system: 10×Buffer 1μL, plasmid 20-50ng, annealing product 1μL, enzyme 0.2μL, add water to 10μL.
[0290] B. Reaction procedure: 16℃, 60min.
[0291] Take 50 μL of commercially available E. coli DH5α competent cells, add all ligation products and mix well, then complete the transformation process according to the manufacturer's instructions. Screen on LB agar plates containing 50 mg / L kanamycin and culture overnight. Pick 5 single clones and culture with shaking in LB liquid medium. After sequencing confirmation of positive results, extract and preserve the plasmid, naming it pKM-CAS1.0-KlglpA (Figure 10).
[0292] (3) Donor DNA construction and amplification
[0293] First, the donor DNA, namely the homologous recombination sequence at the KlglpA site, and the MaPylRS expression structure were constructed. The KlPGK1 promoter and ScCYC1 terminator were used for MaPylRS. The donor DNA construction and transformation methods are as follows.
[0294] iii. Gene synthesis of plasmids containing the MaPylRS expression cassette, and using the plasmids as templates, PCR amplification of the MaPylRS expression cassette fragments was performed with primers PF5: CCATCAGTTACGGTAGATTCTCCAGTGCCTACGTTCCTCATCACTAGAAG (SEQ ID NO: 29) and PR5: TGTTTTGCGCTTGGTTTTCTTTGTGGAGAAATTTCTTCGAGCGTCCCAAA (SEQ ID NO: 30).
[0295] iv. Using Kluyveromyces lactis free plasmid as a template, PCR amplification was performed with primers PF6:AAATTAAGGCAAACATACAGG (SEQ ID NO: 31) and PR6:CAACAGTTCGGCTTCTAGTGATGAGGAACGTAGGCACTGGAGAATCTACC (SEQ ID NO: 32) to obtain the homologous arm 1 fragment of the KlglpA site; using Kluyveromyces lactis free plasmid as a template, PCR amplification was performed with primers PF7:GCTTGAGAAGGTTTTGGGACGCTCGAAGAAATTTCTCCACAAAGAAAACC (SEQ ID NO: 33) and PR7:GACCTTTTATTTTGTCACCG (SEQ ID NO: 34) to obtain the homologous arm 2 fragment of the KlglpA site.
[0296] v. Mix the three PCR product fragments obtained above, dilute them 100 times as templates, and perform PCR amplification again with primers PF8:ATATCGGATGACATGCAGCAA (SEQ ID NO: 35) and PR8:TTGTGTACCAAAACTTTCACGG (SEQ ID NO: 36); the three fragments can be ligated together by overlap extension PCR to become a linear donor DNA. After sequencing confirmation, store it in a -20 degree Celsius freezer.
[0297] (4) Preparation and transformation of highly efficient Kluyveromyces lactis competent cells
[0298] Preparation of yeast competent cells
[0299] Streak Kluyveromyces lactis culture on YPD solid medium and pick single colonies, then culture overnight in 25 mL of 2×YPD liquid medium with shaking. Take 2 mL of the culture and culture in 50 mL of 2×YPD liquid medium with shaking for 2–8 h. Collect yeast cells by centrifugation at 3000g for 5 min at 20°C, resuspend in 500 μL of sterile water, and collect cells by centrifugation under the same conditions. Prepare competent cell solution (5% v / v glycerol, 10% v / v DMSO) and dissolve yeast cells in 500 μL of this solution. Aliquot 50 μL into 1.5 mL centrifuge tubes and store at -80°C.
[0300] Yeast DNA transformation
[0301] Thaw competent cells on ice for 30 seconds, add 200 ng of pKM-CAS1.0-KlglpA plasmid and 2000 ng of donor DNA. Electrolyze at 1.5 kV for 5 ms, then immediately add 1 mL of YPD liquid medium and incubate for 2-3 hours. Spread 200 μL onto solid YPD (200 μg / mL G418) medium and incubate for 2-3 days until single colonies appear.
[0302] (5) Positive identification of gene editing
[0303] Fifty to sixty single clones were picked from the Kluyveromyces lactis transformed plates, and each single clone was placed in 5 μL of yeast lysis buffer (Takara Mighty Prep Reagent for DNA). Using the cell lysis buffer as a template, PCR amplification was performed using primers Ma R2 (MaPylRS inner primer): CCATGTGAGAACCTCTTGG (SEQ ID NO: 37); glpA F1 (KlglpA donor DNA 5' outer primer): AAATTAAGGCAAACATACAGG (SEQ ID NO: 38) and primers Ma F2 (MaPylRS inner primer): ATCTCTTACTTGAACGGTGCTA (SEQ ID NO: 39); glpA R1 (KlglpA donor DNA 5' outer primer): GACCTTTTATTTTGTCACCG (SEQ ID NO: 40). CRISPR insertion at the KlglpA site was detected. A positive band indicated that the MaPylRS sequence was successfully inserted into the target site.
[0304] Example 8: KlUPF1 was knocked out and replaced with a MaPylRS expression cassette using CRISPR-Cas9 technology.
[0305] (1) KlUPF1 sequence retrieval and CRISPR gRNA sequence determination
[0306] According to the literature, the yeast UPF1 protein has been shown to be involved in mRNA degradation, and deleting this gene can slow down the degradation rate of immature mRNA transcripts. Therefore, this invention completely knocks out the KlUPF1 gene and replaces it with the MaPylRS sequence, causing non-natural amino acid transcripts to aggregate, thereby improving the efficiency of non-natural amino acid insertion.
[0307] i. Search for "UPF1" at http: / / www.yeastgenome.org / to obtain the ScUPF1 gene sequence from Saccharomyces cerevisiae. Perform BLAST alignment analysis on the UPF1 gene in the NCBI database to identify the UPF1 homologous gene sequence KlUPF1 in Kluyveromyces lactis (located on chromosome B at 567908...570817).
[0308] ii. Search for PAM (NGG) sequences at both ends of the KlUPF1 gene and determine the gRNA sequences. The principles for gRNA selection are: moderate GC content, with the standard of 40%-60% in this invention; and avoidance of polyT structures. Finally, the KlUPF1 gRNA1 sequence determined in this invention is TTGGCAAACGCATCGTCATA (SEQ ID NO: 41) and the gRNA1 sequence is CTTAAGGAAGTACAATGGAG (SEQ ID NO: 42).
[0309] (2) Construction of KlUPF1 CRISPR-Cas9 plasmid
[0310] Based on the designed gRNA1 and gRNA2 sequences, two 24nt primers were designed for vector construction. gRNA-F3: AATCTTGGCAAACGCATCGTCATA (SEQ ID NO: 43); gRNA-R3: AAACTATGACGATGCGTTTGCCAA (SEQ ID NO: 44); gRNA-F4: AATCTTAAGGAAGTACAATGGAG (SEQ ID NO: 45); gRNA-R4: AAACCTCCATTGTACTTCCTTAAG (SEQ ID NO: 46). The primers were diluted to 10 μM, and 10 μL each of gRNA-F and gRNA-R were added to PCR tubes. After mixing, the mixture was centrifuged to the bottom of the tube, and annealing was performed according to the following procedure:
[0311] 95℃, 3min; 72℃, 30s; 65℃, 2min; 60℃, 2min; 55℃, 2min; 50℃, 2min; 16℃, 2min
[0312] Then the connection reaction takes place.
[0313] A. Reaction system: 10×Buffer 1μL, plasmid 20-50ng, annealing product 1μL, enzyme 0.2μL, add water to 10μL.
[0314] B. Reaction procedure: 16℃, 60min.
[0315] Take 50 μL of commercially available E. coli DH5α competent cells, add all ligation products and mix well. Complete the transformation process according to the manufacturer's instructions. Select on LB agar plates containing 50 mg / L kanamycin and culture overnight. Pick 5 single clones and culture with shaking in LB liquid medium. After sequencing confirmation of positive results, extract and preserve the plasmid, naming it pKM-CAS1.0-KlUPF1 (Figure 11).
[0316] (3) Donor DNA construction and amplification
[0317] This invention first constructs a donor, Donor, by replacing the UPF1 coding sequence with the MaPylRS coding sequence. Specifically, the inserted MaPylRS uses the UPF1 promoter and terminator. The donor DNA construction and transformation methods are as follows.
[0318] iii. Gene synthesis of plasmids containing the MaPylRS expression cassette, and using plasmids as templates, PCR amplification of the MaPylRS coding sequence was performed with primers PF9: AGTACAATTAGAATCAAGTTTCCTTATGGGTTCTTCTTCTTCTGG (SEQ ID NO: 47) and PR9: TAATATTATTTAATTAATGGATTGATACGCGTTCATGTTTAGTTGATCTTAGCACCGTTC (SEQ ID NO: 48).
[0319] iv. Using Kluyveromyces lactis free plasmid as a template, PCR amplification was performed with primers PF10: CAATGGATACAGTTTCTCGCTA (SEQ ID NO: 49) and PR10: ACCAGAAGAAGAAGAACCCATAAGGAAACTTGATTCTAATTGT (SEQ ID NO: 50) to obtain the homologous arm 1 fragment of the KlUPF1 site; using Kluyveromyces lactis free plasmid as a template, PCR amplification was performed with primers PF11: TTGAACGGTGCTAAGATCAACTAAACATGAACGCGTATCAATC (SEQ ID NO: 51) and PR11: CTTCGAGACTTCCAATGATCTC (SEQ ID NO: 52) to obtain the homologous arm 2 fragment of the KlUPF1 site.
[0320] v. Mix the three PCR product fragments obtained above, dilute them 100 times as templates, and perform PCR amplification again with primers PF12: GATCGTCCATTAGCTTATCTACAAATGCC (SEQ ID NO: 53) and PR12: GTGAGAATGCCAGACGAT (SEQ ID NO: 54); the three fragments can be ligated together by overlap extension PCR to become a linear donor DNA. After sequencing confirmation, store it in a -20 degree Celsius freezer.
[0321] (4) Transformation and positive identification of Kluyveromyces lactis
[0322] Preparation of yeast competent cells
[0323] Streak Kluyveromyces lactis culture on YPD solid medium and pick single colonies, then culture overnight in 25 mL of 2×YPD liquid medium with shaking. Take 2 mL of the culture and culture in 50 mL of 2×YPD liquid medium with shaking for 2–8 h. Collect yeast cells by centrifugation at 3000g for 5 min at 20°C, resuspend in 500 μL of sterile water, and collect cells by centrifugation under the same conditions. Prepare competent cell solution (5% v / v glycerol, 10% v / v DMSO) and dissolve yeast cells in 500 μL of this solution. Aliquot 50 μL into 1.5 mL centrifuge tubes and store at -80°C.
[0324] Yeast DNA transformation
[0325] Thaw competent cells on ice for 30 seconds, add 200 ng of pKM-CAS1.0-KlUPF1 plasmid and 2000 ng of donor DNA. Electrolyze at 1.5 kV for 5 ms, then immediately add 1 mL of YPD liquid medium and incubate for 2-3 hours. Spread 200 μL onto solid YPD (200 μg / mL G418) medium and incubate for 2-3 days until single colonies appear.
[0326] (5) Positive identification of gene editing
[0327] Fifty to sixty single clones were picked from the Kluyveromyces lactis transformed plates, and each single clone was placed in 5 μL of yeast lysis buffer (Takara Mighty Prep Reagent for DNA). Using the cell lysis buffer as a template, PCR amplification was performed using primers Ma R3 (MaPylRS inner primer): GTCTCTAGAAGCCAAGTCTTC (SEQ ID NO: 55); UPF1 F1 (KlUPF1 donor DNA 5' outer primer): GAACTGCCACGGGCT (SEQ ID NO: 56) and primers Ma F3 (MaPylRS inner primer): GCTGCTCACGACGTTCA (SEQ ID NO: 57); UPF1 R1 (KlUPF1 donor DNA 5' outer primer): GCACTGTAATCAGGCAACT (SEQ ID NO: 58). CRISPR insertion at the KlUPF1 site was detected. A positive band indicated that the MaPylRS sequence was successfully inserted into the target site.
[0328] Example 9 Activity Assay
[0329] A genetically modified *Kluyveromyces lactis* strain was used to prepare an in vitro protein synthesis system (IVTT). A plasmid containing dual reporter genes (GFP and RFP) was added to determine the ability of the modified strain to insert non-natural amino acids at specific sites in specific proteins. GFP was used to detect overall protein expression, while RFP was used to detect whether a non-natural amino acid was successfully inserted at a specific location, namely the stop codon TAG of GFP. The presence of a red fluorescent signal indicated that the translation-related protein had read the stop codon TAG, meaning the non-natural amino acid was successfully inserted. The absence of a red fluorescent signal indicated that the translation-related protein stopped at the stop codon TAG, and the non-natural amino acid insertion failed.
[0330] Taking the Kluyveromyces lactis strain obtained in Example 6 as an example, and adjusting the promoter in the MaPylRS expression cassette, different Kluyveromyces lactis strains were obtained and prepared into Protein Factory.
[0331] Establishment of expression systems containing non-natural amino acids:
[0332] Protein Factory 100ul
[0333] Prock (non-natural amino acids) 500mM 1ul
[0334] matRNA CUA pyl In vitro transcription product (unpurified) 10ul
[0335] 3ul of GFP-TAG-RFP dual fluorescent reporter gene PCR product
[0336] The efficiency of non-natural amino acid introduction was determined by detecting the fluorescence intensity of RFP and the RFP / GFP ratio (see Figures 12-14 for details).
[0337] In Figures 12-14, DW14-1 and DW14-2 are two parallel experiments with TIF11 as the promoter; DW14-3 has TEF1 as the promoter; DW14-4 and DW14-5 are two parallel experiments with ADH1 as the promoter; DW14-6 and DW14-7 are two parallel experiments with GAP1 as the promoter; DW14-8 and DW14-9 are two parallel experiments with HXK4 as the promoter; and DW14-10 has PGK1 as the promoter.
[0338] As shown in Figure 12, all the modified strains successfully read the stop codon TAG, demonstrating that the modified strains of this invention can introduce non-natural amino acids. As shown in Figure 14, the modified strains of this invention all exhibit high efficiency in introducing non-natural amino acids, especially DW14-4 to DW14-7 and DW14-10, which have very high RFP / GFP values, indicating very high efficiency in introducing non-natural amino acids.
[0339] Example 10
[0340] The effects of the reaction systems obtained after integrating the modified MaPylRS (Example 6) and the original MaPylRS (Example 5) into yeast on the activity of ncaa introduction were compared. The specific measurement conditions are the same as in Example 9, and the specific results are shown in Figures 15-17. sl-3 represents the reaction system of yeast integrating the original MaPylRS, and sl-9 represents the reaction system of yeast integrating the modified MaPylRS of this application. Orthogonal tRNA and Prock were added to both systems, and a system without ncaa (i.e., non) was used as a control. As shown in Figures 15-17, the reaction systems obtained by integrating wild-type and modified MaPylRS into yeast cells can achieve ncaa introduction. However, the reaction system obtained after integrating the modified MaPylRS of this invention into yeast cells can significantly improve the ncaa introduction efficiency.
[0341] The sequences used in this invention are shown in Table 1 below.
[0342] Table 1
[0343] Based on the above-described preferred embodiments according to this application, and through the above description, those skilled in the art can make various changes and modifications without departing from the technical spirit of this application. The technical scope of this application is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A recombinant aminoacyl-tRNA synthetase having the structure described in Formula I: A1-A2-A3-A4(I): In Formula I, "-" independently represents a bond or amino acid linkage sequence. A1 is either absent or a histidine tag. A2 is either absent or a thrombin cleavage site. A3 is either absent or a tagged protein. A4 is an aminoacyl-tRNA synthetase, and at least one of A1 to A3 is present; the connection between A1 to A4 can be either from the N-terminus to the C-terminus or from the C-terminus to the N-terminus.
2. The recombinant aminoacyl-tRNA synthetase of claim 1, wherein: The aminoacyl-tRNA synthetase is selected from natural or mutated Pyl-tRNA synthetase (PylRS), Leu-tRNA synthetase (LeuRS), Tyr-tRNA synthetase (TyrRS), Phe-tRNA synthetase (PheRS) or TrP-tRNA synthetase (TrpRS).
3. The recombinant aminoacyl-tRNA synthetase of claim 1 or 2, characterized in that: The aminoacyl-tRNA synthetase is selected from natural or mutated MaPylRS, MmPylRS, MbPylRS, EcTyrRS, MjTyrRS, EcLeuRS, ScPheRS, ScTrpRS, or BsTrpRS; preferably, the aminoacyl-tRNA synthetase comprises the sequence shown in SEQ ID NO:60 or its active fragment, and more preferably, the sequence of the aminoacyl-tRNA synthetase is SEQ ID NO:62; or it is a polypeptide having ≥85%, ≥90%, ≥95%, ≥97%, ≥98%, or ≥99% homology with the amino acid sequence shown in SEQ ID NO:60 and having the same activity as the sequence in SEQ ID NO:
60.
4. The recombinant aminoacyl-tRNA synthetase of any one of claims 1-3, wherein: The histidine tag has an n×His structure, where 1≦n≦50; preferably 2≦n≦30; more preferably 5≦n≦20; and even more preferably 6≦n≦10.
5. The recombinant aminoacyl-tRNA synthetase of any of claims 1-4, wherein: The aminoacyl-tRNA synthetase has a his tag attached to its N-terminus or C-terminus; or has a thrombin cleavage site and a tag protein attached to its N-terminus in sequence from N-terminus to C-terminus; or has a his tag, a thrombin cleavage site, and a tag protein attached to its N-terminus in sequence from N-terminus to C-terminus.
6. The recombinant aminoacyl-tRNA synthetase of any of claims 1-5, wherein: The recombinant aminoacyl-tRNA synthetase comprises any one of the sequences described in SEQ ID NO:1-3 and SEQ ID NO:64, or its active fragment. More preferably, the amino acid sequence of the recombinant aminoacyl-tRNA synthetase is selected from any one or more of the following: SEQ ID NO:1-SEQ ID NO:3 and SEQ ID NO:64; or is a polypeptide having ≥85%, ≥90%, ≥95%, ≥97%, ≥98%, or ≥99% homology with any one of the amino acid sequences described in SEQ ID NO:1-3 and SEQ ID NO:64, and having the same activity as any one of the sequences corresponding to the homology in SEQ ID NO:1-3 and SEQ ID NO:64; or the coding sequence of the recombinant aminoacyl-tRNA synthetase is selected from any one or more of the following: SEQ ID NO:4, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63; or comprises SEQ ID NO:4, SEQ ID NO:61, SEQ ID NO:64 ... The sequence or its active fragment described in any one of SEQ ID NO:62 or SEQ ID NO:63, or a nucleotide having ≥85%, ≥90%, ≥95%, ≥97%, ≥98%, or ≥99% homology to any one of the sequences described in SEQ ID NO:4, SEQ ID NO:61, SEQ ID NO:62, or SEQ ID NO:63, and having the same activity as any one of the sequences described in SEQ ID NO:4, SEQ ID NO:61, SEQ ID NO:62, or SEQ ID NO:
63.
7. A nucleic acid construct encoding the recombinant aminoacyl-tRNA synthetase of any one of claims 1-6.
8. The nucleic acid construct of claim 7, wherein: The nucleic acid construct contains at least the structure described in Formula II: Z1-Z2-Z3-Z4, where Z1 to Z4 are elements used to constitute the construct; "-" independently represents a bond or nucleotide linking sequence; Z1 is a coding sequence for a histidine tag that is absent or a coding sequence for a thrombin cleavage site that is absent or a coding sequence for a tag protein that is absent or a coding sequence for an aminoacyl-tRNA synthetase; wherein at least one of Z1 to Z3 is present.
9. The nucleic acid construct of claim 8, wherein: The amino acid sequence encoded by Z1 is HHHHHH; the amino acid sequence encoded by Z2 is LVPRGS; the amino acid sequence encoded by Z3 is SEQ ID NO:59; or Z1, Z2 and Z3 respectively encode sequences containing HHHHHH, LVPRGS and SEQ ID NO:59 or their active fragments, or each of them has ≥85%, ≥90%, ≥95%, ≥97%, ≥98% or ≥99% homology with the nucleotide sequences corresponding to HHHHHH, LVPRGS and SEQ ID NO:59 and each of them has the same activity as the sequences encoding HHHHHH, LVPRGS and SEQ ID NO:
59.
10. The nucleic acid construct of claim 8 or 9, wherein It also includes a starter element. Preferably, the nucleic acid structure contains the structure described in Formula III: Z5-Z1-Z2-Z3-Z4, In the formula, Z5 is the starter element. More preferably, the promoter is selected from PGK1, GAP1, ADH1, HXK1, GAPDH1, TEF1 or TIF11.
11. A vector, characterized in that, The vector contains the nucleic acid construct according to any one of claims 7-10.
12. A genetically engineered bacterial strain, characterized in that, The genetically engineered strain has one or more sites in its genome that integrate the nucleic acid construct according to any one of claims 7-10, or the genetically engineered strain contains the recombinant aminoacyl-tRNA synthetase according to any one of claims 1-6, or the genetically engineered strain contains the vector according to claim 11.
13. The genetically engineered bacterial strain of claim 12, wherein, The strains described are derived from one or any combination of mammalian cells, plant cells, yeast cells, insect cells, and prokaryotic cells.
14. The genetically engineered bacterial strain of claim 12 or 13, wherein: The sites mentioned are selected from Lys1-5, glpA, or UPF1.
15. The genetically engineered bacterial strain according to any one of claims 12-14, characterized in that: The nucleic acid construct further includes a terminator; preferably, the terminator is selected from CYC1, GPM1, TDH2 or ACT1.
16. A method for synthesizing a protein incorporating a non-natural amino acid, comprising: The recombinant aminoacyl-tRNA synthetase is provided by any one of claims 1-6 or by any one of the genetically engineered strains of claims 12-15.
17. A cell-free system for synthesizing a protein containing a non-natural amino acid, comprising, The cell-free system comprises at least: (a) a cell extract, and (b) one or more of the recombinant aminoacyl-tRNA synthetase of any one of claims 1-6, the nucleic acid construct of any one of claims 7-10, or the vector of claim 11; wherein the cell extract is derived from one or any combination of mammalian cells, plant cells, yeast cells, insect cells, prokaryotic cells, or other types of cells.
18. A cell-free system for synthesizing a protein containing a non-natural amino acid, comprising, The cell-free system includes at least a cell extract derived from the genetically engineered strains described in any one of claims 12-15.
19. A cell-free system for synthesizing a protein containing unnatural amino acids according to claim 17 or 18, characterized in that, The cell-free system further includes: non-natural amino acids, orthogonal tRNA, and a template containing the gene sequence of the target protein, wherein the codons encoding the amino acids in the gene sequence of the target protein are mutated.
20. The method for preparing a genetically engineered bacterial strain according to any one of claims 12 to 15, characterized in that, The nucleic acid constructs described in any one of claims 7-10 are transferred into or integrated into cells using transformation, transfection, or gene editing techniques.
21. The method for preparing a genetically engineered strain according to claim 20, characterized in that, The nucleic acid construct according to any one of claims 7 to 10 is integrated into the genome of the cell by active site integration.
22. The method for preparing a genetically engineered strain according to claim 20 or 21, characterized in that, The nucleic acid construct further comprises a terminator; preferably, the terminator is selected from CYC1, GPM1, TDH2 or ACT1.
23. The method for preparing a genetically engineered bacterial strain according to claim 21 or 22, characterized in that, The site is selected from Lys1-5, glpA or UPF1.
24. A kit characterized in that, The kit comprises the reaction system according to any one of claims 17 to 19.
25. A method for in vitro synthesis of proteins containing non-natural amino acids, characterized in that, The preparation is performed using the cell-free system according to any one of claims 17 to 19 or the kit according to claim 24.