Method for directed evolution of proteins

By combining PCR and E. coli in vitro expression systems, mutant protein traits can be directly expressed and detected in vitro, solving the problems of long process and unstable results in protein directed evolution, and realizing rapid and stable protein directed evolution.

WO2025222647A1PCT designated stage Publication Date: 2025-10-30TIANJIN ASYMCHEM BIOTECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/105815
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-23
Filing Date
2024-07-16
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing protein directed evolution methods are lengthy, time-consuming, and prone to large fluctuations in results. Microbial manipulation poses risks of contamination and has high requirements for equipment and consumables.

Method used

By combining PCR into the protein mutation evolution process with the E. coli in vitro expression system, the characteristics of mutant proteins can be directly expressed and detected in vitro, simplifying the detection process to within 7 hours and avoiding traditional steps such as enzyme digestion, ligation, and transformation.

Benefits of technology

It greatly accelerates the directed evolution of proteins, improves the stability and efficiency of results, reduces operational complexity and contamination risk, and is suitable for the directed evolution of industrial enzymes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024105815_30102025_PF_FP_ABST
    Figure CN2024105815_30102025_PF_FP_ABST
Patent Text Reader

Abstract

A method for the directed evolution of proteins. The method comprises obtaining a PCR product with a gene mutation of a target protein by means of PCR amplification; placing the PCR product in an in-vitro protein expression system for gene expression to obtain a target protein with a mutation; and performing a characterization detection on the target protein with the mutation, wherein the target protein is a target enzyme, and the characterization detection comprises detecting the activity of the target enzyme and / or the corresponding enantiomeric excess. The method not only has high stability, but also shortens the time from a generally required 2-3 weeks to about 7 hours, which greatly accelerates the evolution speed.
Need to check novelty before this filing date? Find Prior Art

Description

Methods of directed protein evolution

[0001] This application is based on and claims priority to Chinese application CN application number 202410490345.3 filed on April 23, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] This invention relates to the field of directed protein evolution in protein engineering, and more specifically, to a method for directed protein evolution. Background Technology

[0003] Directed evolution of proteins aims to mimic natural evolution by repeatedly mutating, expressing, and screening target genes to achieve in a short time what would take thousands of years in nature, ultimately resulting in proteins with improved performance or new functions. Methods of directed evolution of proteins can be categorized into three strategies: irrational design, semi-rational design, and rational design.

[0004] Irrational design, or random evolution strategy, has the advantage of not requiring in-depth knowledge of protein sequences and structures; it simply simulates natural evolution through random mutations and fragment recombination. It mainly includes error-prone polymerase chain reaction (epPCR) and DNA recombination (DNA shuffling). DNA shuffling is primarily used for recombination of single or multiple genes. This technique uses DNase to cut a set of homologous genes with sense mutation sites into random fragments (usually 10-50 bp), and then uses PCR to extend and recombine them to obtain the full-length gene. Its advantages are simplicity, no need for protein structural information, and easy acquisition of sense mutations; its disadvantage is the requirement for at least 70% homology between gene sequences. Due to codon degeneracy, amino acid sequence variations are much smaller than base sequence variations. Therefore, 70% homology in gene sequences implies over 90% homology at the protein amino acid sequence level. This fatal flaw has prevented this technique from being widely adopted in the last 20 years. Error-prone PCR is relatively more widely used. Its basic principle is to increase the random mismatch rate of bases by changing the reaction conditions of the PCR system or using low-fidelity DNA polymerases, thereby causing multi-point mutations and generating a mutant library with sequence diversity. It is widely adopted by researchers because it does not require protein structural information and is simple to operate. However, the application of this technique is limited by the following factors: the base bias of the polymerase (usually AG>TC), low mutation efficiency (typically only one base is mutated per round), and the need for continuous accumulation. Typically, at least four consecutive rounds of epPCR are required to gradually accumulate positive mutations to obtain the target mutant with significantly improved protein performance. Limited by detection throughput, the library capacity of one round of epPCR is generally around 1000-2000.

[0005] Rational design is an intelligent modification method that relies on computer technology (in silico) to simulate the evolutionary trajectory of proteins in nature. Through virtual mutations, it screens for target mutants that can be predicted quickly and accurately. Using a series of algorithms and programs developed based on bioinformatics, it predicts protein active sites and examines the effects of mutations at specific sites on their stability, folding, and substrate binding, thus enabling targeted modification of proteins. Computer-aided design and large-scale molecular dynamics simulations can efficiently and quickly modify and screen biocatalysts, not only predicting protein structures with high accuracy but also designing novel proteins that do not exist in nature from scratch. Although new protein design has achieved some success, it still faces many challenges: first, its success rate is low; second, the computational workload is heavy, with a very high dependence on computer resources; third, the designed new proteins often have poor structures and stability, and their catalytic activity is often low. This is mainly because the understanding of the relationship between protein sequence / structure / function is still insufficient. Rational design generally introduces mutation sites through site-directed mutagenesis, with a library size ranging from tens to hundreds.

[0006] Semi-rational design mainly relies on bioinformatics methods. Based on homologous protein sequence alignment, three-dimensional structure, or existing knowledge, multiple amino acid residues are rationally selected as modification targets. Combined with the rational selection of effective codons, high-quality mutant libraries are constructed to specifically modify proteins. Mutations are generally introduced through degenerate primers, and the library capacity ranges from hundreds to thousands (Qu Ge, Zhao Jing, Zheng Ping, etc., Recent advances in directed evolution technology. Chinese Journal of Biotechnology, 2018, 34(1): 1-11).

[0007] In summary, site-directed mutagenesis and site-directed saturation mutagenesis are the most frequently used methods in the construction of mutant libraries in directed protein evolution. In addition, epPCR is also an effective method.

[0008] Currently, site-directed mutagenesis and site-directed saturation mutagenesis generally involve designing the mutation site as a target base or degenerate base, introducing it via PCR, constructing it onto a plasmid, transforming it into an expression host (often E. coli), culturing it, transferring it, inducing the expression of the target protein, then breaking it down to obtain the protein, and using crude protein extract or purified protein for reaction.

[0009] Error-prone PCR has a large library capacity, but due to limitations in screening throughput, it typically screens around 1000-2000 mutants. For industrial proteins, the number of amino acids is usually around 300. By designing 3-5 primers for each site and mutating that site with 3-5 amino acids representing different properties, error-prone PCR can also be solved using global PCR.

[0010] In summary, existing methods for directed protein evolution can generally be achieved using primers containing a single mutant. Due to continuous breakthroughs in gene synthesis technology in recent years, the cost of gene synthesis has been decreasing. Generally, a single primer costs around 10 yuan, and synthesizing a thousand primers would cost only around 10,000 yuan. Based on current trends, this cost is expected to decrease further in the future.

[0011] Traditional directed protein evolution involves more than 10 steps, from primer design to mutant phenotypic detection, including PCR, protein digestion, ligation, transformation, single clone selection, single clone culture, transfer, induction, expression, centrifugation, resuspension, and disruption, to obtain a crude cellular extract of the target protein. This process typically takes about 2-3 weeks. This series of operations is not only time-consuming and labor-intensive, but also carries the risk of contamination by other microbial bacteria or bacteriophages. It places high demands on equipment and consumables, as well as on personnel. Furthermore, due to the lengthy process, each step introduces a small amount of error, leading to significant fluctuations in the final results.

[0012] Summary of the Invention

[0013] The main objective of this invention is to provide a method for directed protein evolution, thereby solving the problem of long processes in the prior art for directed enzyme evolution.

[0014] To achieve the above objectives, according to one aspect of the present invention, a method for directed protein evolution is provided, the method comprising: obtaining a PCR product containing a gene mutation of a target protein by PCR amplification; placing the PCR product in an in vitro protein expression system for gene expression to obtain a target protein with mutation; and performing phenotypic testing on the target protein with mutation, the phenotypic testing including detecting the activity of a target enzyme and / or the excess percentage of the corresponding isoform (i.e., ee value); wherein the target protein is a target enzyme; the in vitro protein expression system is an *E. coli* in vitro expression system, the *E. coli* in vitro expression system comprising: basic components, energy-related components, additive components, cell extracts, and RNase inhibitors, wherein, in the *E. coli* in vitro expression system, the basic components include: 19 amino acids at a concentration of 2 mM each, 2 mM tyrosine, 14 mM magnesium acetate, 60 mM potassium acetate, and 7 mM DDT; the energy-related components in the *E. coli* in vitro expression system include: 1.2 mM AMP, 0.85 mM CMP, 0.85 mM GMP, and 0.85 mM... The in vitro expression system of *E. coli* contained the following components: UMP, 15–83 mM PEP, 0.4–0.6 mM NAD, 4 mM potassium oxalate, 90 mM potassium glutamate, and 2.5–10 mM magnesium glutamate. The added components included 1.5 mM spermidine and 157.33 mM HEPES. The concentration of the RNase inhibitor in the in vitro expression system was 150 U / 450 μL. The volume fraction of the cell extract in the in vitro expression system was 20–60%.

[0015] Further, using any one or more of the following methods, a PCR product containing the target protein gene mutation is obtained by PCR amplification: 1) Amplifying the PCR product containing the target protein gene mutation using a two-step PCR method; designing two primer pairs: 1) F1 and R1; 2) F2 and R2, and introducing the mutation sequence containing the mutation site into the two primer pairs; performing the first step of PCR using the two primer pairs to PCR out fragments L1 and L2 on both sides of the mutation site, wherein the overlapping region between fragments L1 and L2 is denoted as L, and the mutation site is located on L; then using the mixture of fragments L1 and L2 as a template, and using F1 and R2 as primers, performing the second step of PCR to obtain the full-length sequence, which is the PCR product containing the target protein gene mutation. CR products; or 2) Based on the principle of site-directed saturation mutagenesis, multiple PCR products containing gene mutations of the target protein are constructed by introducing mutations through PCR amplification, and the multiple PCR products containing gene mutations of the target protein are used to construct a saturation mutant library of the target protein gene; or 3) Mutations are introduced at specific sites through PCR amplification to obtain PCR products containing gene mutations of the target protein; or 4) Error-prone PCR is used to perform random mutations of the entire sequence to obtain multiple PCR products containing gene mutations of the target protein, and the multiple PCR products containing gene mutations of the target protein cover random mutations of the entire gene sequence of the target protein; or 5) Multiple mutation sites of the gene containing the target protein are obtained by using multi-point mutagenesis.

[0016] Furthermore, in the Escherichia coli in vitro expression system, the concentration of PEP is 30 mM; preferably, the concentration of NAD is 0.4 mM; preferably, the concentration of magnesium glutamate is 7.5 mM; preferably, the volume content of cell extract in the Escherichia coli in vitro expression system is 33.3%.

[0017] Furthermore, the target enzyme is selected from any of the following industrial proteases.

[0018] Furthermore, the industrial protease is the ester protein shown in SEQ ID NO: 1 or the transaminase TA-1 shown in SEQ ID NO: 2.

[0019] Furthermore, the phenotypic detection of target enzymes with gene mutations includes: using multiple target enzymes with different gene mutations to catalyze the same substrate reaction to produce the same product, detecting the conversion rate of the substrate and / or the excess percentage of the corresponding isomer of the product catalyzed by different target enzymes; using the conversion rate and / or excess percentage of the substrate catalyzed by the initial control enzyme as a reference, selecting target enzymes with improved conversion rate and / or excess percentage of the corresponding isomer from multiple target enzymes, and recording them as the initial +1 control enzyme.

[0020] Furthermore, after obtaining the initial +1 control enzyme, the method also includes: iterating the initial +1 control enzyme to the initial control enzyme, and then repeating steps S1 to S3, and so on, to obtain multiple target enzymes after directed evolution.

[0021] To achieve the above objectives, according to a second aspect of the present invention, an in vitro protein expression system is provided. This in vitro protein expression system is an *E. coli* in vitro expression system, comprising: basic components, energy-related components, additive components, cell extracts, and RNase inhibitors. The basic components of the *E. coli* in vitro expression system include: 19 amino acids at a concentration of 2 mM each; 2 mM tyrosine; 14 mM magnesium acetate; 60 mM potassium acetate; and 7 mM DDT. The energy-related components include: 1.2 mM AMP, 0.85 mM CMP, 0.85 mM GMP, 0.85 mM UMP, 15–83 mM PEP, 0.4–0.6 mM NAD, 4 mM potassium oxalate, 90 mM potassium glutamate, and 2.5–10 mM magnesium glutamate. The additive components include: 1.5 mM spermidine and 157.33 mM... HEPES; In the E. coli in vitro expression system, the concentration of RNase inhibitor was 150 U / 450 μL; The volume content of cell extract in the E. coli in vitro expression system was 20-60%.

[0022] Furthermore, in the Escherichia coli in vitro expression system, the concentration of PEP is 30 mM; preferably, the content of NAD is 0.4 mM.

[0023] Preferably, the content of magnesium glutamate is 7.5 mM; preferably, the volume content of the cell extract in the Escherichia coli in vitro expression system is 33.3%.

[0024] To achieve the above objectives, according to a third aspect of the present invention, a kit for directed protein evolution is provided. This kit includes an in vitro protein expression system, which is an *E. coli* in vitro expression system. The *E. coli* in vitro expression system includes: basic components, energy-related components, additive components, cell extracts, and RNase inhibitors. In the *E. coli* in vitro expression system, the basic components include: 19 amino acids at a concentration of 2 mM each; 2 mM tyrosine; 14 mM magnesium acetate; 60 mM potassium acetate; and 7 mM DDT. In the *E. coli* in vitro expression system, the energy-related components include: 1.2 mM AMP, 0.85 mM CMP, 0.85 mM GMP, 0.85 mM UMP, 15–83 mM PEP, and 0.4–0.6 mM... The ingredients included NAD, 4 mM potassium oxalate, 90 mM potassium glutamate, and 2.5–10 mM magnesium glutamate. In the *E. coli* in vitro expression system, the added components included 1.5 mM spermidine and 157.33 mM HEPES. The concentration of the RNase inhibitor in the *E. coli* in vitro expression system was 150 U / 450 μL. The volume percentage of cell extract in the *E. coli* in vitro expression system was 20–60%.

[0025] Furthermore, in the Escherichia coli in vitro expression system, the concentration of PEP is 30 mM; preferably, the content of NAD is 0.4 mM; preferably, the content of magnesium glutamate is 7.5 mM; preferably, the volume content of cell extract in the Escherichia coli in vitro expression system is 33.3%.

[0026] By applying the technical solution of this invention, the process of introducing PCR into protein mutation evolution is combined with in vitro protein expression. The improved in vitro expressed protein product is then directly used for performance verification and directed evolution screening, including enzyme activity and / or the percentage of corresponding isoform excess. A series of experimental verifications show that the results obtained by this method are consistent with those obtained by traditional methods, proving the feasibility and effectiveness of this improved directed protein evolution method. From the perspective of efficiency and result stability, it not only exhibits high stability but also reduces the time required from the usual 2-3 weeks to approximately 7 hours, significantly accelerating the evolution speed. Attached Figure Description

[0027] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0028] Figure 1 shows the standard curves plotted using different concentrations of sfGFP as the model protein in Example 1 of the present invention;

[0029] Figure 2 shows the Mg content in the in vitro protein expression system in Example 1 of the present invention. 2+ The results of concentration optimization are shown in the figure;

[0030] Figure 3 shows the results of optimizing the PEP concentration in the in vitro protein expression system in Example 1 of the present invention;

[0031] Figure 4 shows the results of optimizing the proportion of cell extract in the protein in vitro expression system in Example 1 of the present invention;

[0032] Figure 5 shows the results of optimizing the NAD concentration in the in vitro protein expression system in Example 1 of the present invention;

[0033] Figure 6 shows the results of optimizing the glutamate concentration in the in vitro protein expression system in Example 1 of the present invention;

[0034] Figure 7 shows the results of in vitro protein expression using different amounts of PCR products of the sfGFP gene in Example 2 of the present invention.

[0035] Figure 8 shows the results of in vitro protein expression of PCR products of sfGFP genes with different lengths upstream of the start codon in Example 2 of the present invention.

[0036] Figure 9 shows the SDS-PAGE electrophoresis results of eight randomly selected mutants in Example 3 of the present invention;

[0037] Figure 10 shows a comparison of the ee values ​​of mutants obtained by directed evolution using the method of this application in Embodiment 3 of the present invention with the ee values ​​of mutants obtained by directed evolution using conventional methods. Detailed Implementation

[0038] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the embodiments.

[0039] As mentioned in the background section, existing protein directed evolution methods are lengthy, time-consuming, and prone to significant fluctuations in results. To improve this situation, this application attempts to combine the protein mutation evolution process with in vitro protein expression and directly utilize the in vitro expressed protein product for performance verification. Through a series of experiments, it was found that the results obtained by this method are consistent with those obtained by traditional methods, proving the feasibility and effectiveness of this improved protein directed evolution method. From the perspective of efficiency and result stability, it not only exhibits high stability but also reduces the time required from the usual 2-3 weeks to approximately 7 hours, significantly accelerating the evolution speed.

[0040] Based on the above research results, the applicant has proposed a series of technical solutions for this application. In a first typical embodiment, a method for directed protein evolution is provided, comprising: obtaining a PCR product containing a gene mutation of a target protein through PCR amplification; expressing the PCR product in an in vitro protein expression system to obtain a target protein with mutation; and performing phenotypic analysis on the target protein with mutation, including detecting the activity of a target enzyme and / or the excess percentage of its corresponding isoform; wherein the target protein is a target enzyme; the in vitro protein expression system is an *E. coli* in vitro expression system, which includes: basic components, energy-related components, additive components, cell extracts, and RNase inhibitors; wherein, in the *E. coli* in vitro expression system, the basic components include: 19 amino acids at a concentration of 2 mM each, 2 mM tyrosine, 14 mM magnesium acetate, 60 mM potassium acetate, and 7 mM DDT; and the energy-related components include: 1.2 mM AMP, 0.85 mM CMP, 0.85 mM GMP, 0.85 mM UMP, and 15–83 mM... PEP, 0.4–0.6 mM NAD, 4 mM potassium oxalate, 90 mM potassium glutamate, 2.5–10 mM magnesium glutamate; In the *E. coli* in vitro expression system, the added components included: 1.5 mM spermidine and 157.33 mM HEPES; In the *E. coli* in vitro expression system, the concentration of the RNase inhibitor was 150 U / 450 μL; The volume content of cell extract in the *E. coli* in vitro expression system was 20–60%.

[0041] This application directly introduces mutations via PCR to obtain mutated PCR products. Then, using these mutated PCR products as templates, the improved in vitro protein expression system directly expresses the mutated target protein. Furthermore, the in vitro expressed protein products are used for screening and detection of desired enzyme traits, achieving directed evolution of the target enzyme. This method is not only simple, stable, and efficient, significantly accelerating the evolution process, but also low-cost, making it particularly suitable for the directed evolution of industrial enzymes.

[0042] It should be noted that the above-mentioned in vitro protein expression system can be an existing or commercially available system. In the preferred embodiments of this application, in order to further improve the in vitro expression level, an Escherichia coli in vitro expression system is preferably used, or it can be obtained by improving upon an existing Escherichia coli in vitro expression system.

[0043] This improved method for directed protein evolution, by introducing mutated PCR products, eliminates the need for more than 10 steps, including enzyme digestion, ligation, transformation, single-clone selection, single-clone culture, transfer, induction, expression, centrifugation, resuspension, and disruption. Instead, it allows for direct in vitro expression and synthesis, thus providing a simple and rapid way to obtain the expression product of the mutated target protein. This not only shortens the evolutionary process but also reduces the risk of contamination, simplifies operation, minimizes the introduction of errors, and ensures high stability of the final results.

[0044] In the steps of phenotyping target proteins with mutations, the traits tested will vary depending on the biological properties of the target protein. Besides enzyme activity and / or isomer excess percentage mentioned above, other factors may include substrate specificity, catalytic efficiency, catalytic reaction temperature, and reaction stability in different reaction solvents such as organic or aqueous phases. In actual production, a reasonable selection can be made based on the desired optimized enzyme performance.

[0045] In some embodiments, phenotypic detection of target enzymes with gene mutations includes: using multiple target enzymes with different gene mutations to catalyze the same substrate reaction to produce the same product, detecting the conversion rate of the substrate and / or the excess percentage of the corresponding isomer of the product catalyzed by different target enzymes; using the conversion rate and / or the excess percentage of the isomer of the substrate catalyzed by the initial control enzyme as a reference, screening from multiple target enzymes to obtain target enzymes with improved conversion rate and / or excess percentage of the corresponding isomer, and recording them as the initial +1 control enzyme.

[0046] In other embodiments, after obtaining the initial +1 control enzyme, the method for directed enzyme evolution further includes: iterating the initial +1 control enzyme to the initial control enzyme, and then repeating steps S1 to S3, and so on, to obtain multiple directed evolution target enzymes.

[0047] The method described above for obtaining PCR products containing gene mutations of the target protein via PCR amplification can employ any known method that can achieve mutation via PCR. In some preferred embodiments of this application, the product is obtained using any one or more of the following methods:

[0048] 1) Amplify the PCR product containing the gene mutation of the target protein using a two-step PCR method;

[0049] Two primer pairs were designed: 1) F1 and R1; 2) F2 and R2. A mutated sequence containing the mutation site was introduced into each primer pair. The first step of PCR was performed using these two primer pairs, PCR-generating fragments L1 and L2 flanking the mutation site. The overlapping region between fragments L1 and L2 is denoted as L, and the mutation site is located on L. Then, using a mixture of fragments L1 and L2 as a template, and F1 and R2 as primers, the second step of PCR was performed to obtain the full-length sequence. This full-length sequence is the PCR product containing the gene mutation of the target protein.

[0050] 2) Based on the principle of site-directed saturation mutagenesis, multiple PCR products containing gene mutations of the target protein are constructed by introducing mutations through PCR amplification. These multiple PCR products containing gene mutations of the target protein are then used to construct a saturation mutant library of the target protein gene; or

[0051] 3) Mutations are introduced at specific sites using PCR amplification to obtain PCR products containing the target protein gene mutation; or

[0052] 4) Utilize error-prone PCR methods to perform random full-sequence mutations, thereby obtaining multiple PCR products containing gene mutations of the target protein. These multiple PCR products covering random mutations of the target protein's entire gene sequence are obtained; or

[0053] 5) Using the multi-point mutagenesis method, PCR products containing multiple mutation sites of the gene carrying the target protein are obtained.

[0054] The above-mentioned methods for introducing mutations are not particularly improved in this application; the specific operation can be carried out by referring to existing methods.

[0055] The preferred in vitro protein expression system described in this application is optimized and can improve protein expression levels compared to existing in vitro protein expression systems. The cell extract mentioned above refers to Escherichia coli cell extract, which mainly includes ribosomes, RNA polymerase, transcription and translation proteins, as well as enzymes and cofactors used for energy metabolism.

[0056] In some preferred embodiments, the concentration of PEP in the E. coli in vitro expression system is 30 mM; preferably, the content of NAD is 0.4 mM; preferably, the content of magnesium glutamate is 7.5 mM; preferably, the volume content of cell extract in the E. coli in vitro expression system is 33.3%. These preferred conditions result in relatively higher protein expression levels.

[0057] The target protein in this application may be different proteins depending on the actual research purpose. This application preferably uses industrial proteins, especially industrial proteases. In some preferred embodiments, the industrial protease is selected from any of the following proteins: ester protein (amino acid sequence as shown in SEQ ID NO: 1, nucleotide sequence as shown in SEQ ID NO: 3) or transaminase TA-1 (amino acid sequence as shown in SEQ ID NO: 2, nucleotide sequence as shown in SEQ ID NO: 4).

[0058] SEQ ID NO: 1 (Amino acid sequence---264aa):

[0059] SEQ ID NO: 3 (nucleotide sequence --- 792bp):

[0060] SEQ ID NO: 2 (Amino acid sequence---341aa):

[0061] SEQ ID NO: 4 (nucleotide sequence --- 1035bp):

[0062] In a second typical embodiment of this application, a protein in vitro expression system is provided. This protein in vitro expression system is an *E. coli* in vitro expression system, comprising: basic components, energy-related components, additive components, cell extracts, and RNase inhibitors. The basic components of the *E. coli* in vitro expression system include: 19 amino acids at a concentration of 2 mM each; 2 mM tyrosine; 14 mM magnesium acetate; 60 mM potassium acetate; and 7 mM DDT. The energy-related components include: 1.2 mM AMP, 0.85 mM CMP, 0.85 mM GMP, 0.85 mM UMP, 15–83 mM PEP, 0.4–0.6 mM NAD, 4 mM potassium oxalate, 90 mM potassium glutamate, and 2.5–10 mM magnesium glutamate. The additive components include: 1.5 mM spermidine and 157.33 mM... HEPES; In the E. coli in vitro expression system, the concentration of RNase inhibitor was 150 U / 450 μL; The volume content of cell extract in the E. coli in vitro expression system was 20-60%.

[0063] In some preferred embodiments, the concentration of PEP in the above-mentioned Escherichia coli in vitro expression system is 30 mM; preferably, the content of NAD is 0.4 mM; preferably, the content of magnesium glutamate is 7.5 mM; preferably, the volume content of cell extract in the Escherichia coli in vitro expression system is 33.3%.

[0064] The preferred in vitro protein expression system described in this application is optimized and can improve protein expression levels compared to existing in vitro protein expression systems. The cell extract mentioned above refers to Escherichia coli cell extract, which mainly includes ribosomes, RNA polymerase, transcription and translation proteins, as well as enzymes and cofactors used for energy metabolism.

[0065] In existing technologies, some in vitro protein expression systems have been reported in the literature, but their protein expression levels are generally low. Furthermore, aside from model proteins such as green fluorescent protein (GFP) or its variants, reports on other biocatalytic enzymes are even fewer. Currently, commercially available in vitro expression kits produce protein yields at the tens of mg / mL level, but these proteins are mostly used in proteomics research and rarely for industrial enzyme catalytic applications. The in vitro protein expression system optimized in this application can achieve yields of 1-2 mg / mL for many enzyme proteins without specific optimization. Further optimization is expected to yield even higher protein expression levels, thus meeting the requirements for high-throughput enzyme screening.

[0066] In a third typical embodiment of this application, a kit for directed protein evolution is provided. This kit includes an in vitro protein expression system, which is an *E. coli* in vitro expression system. The *E. coli* in vitro expression system includes: basic components, energy-related components, additive components, cell extracts, and RNase inhibitors. Specifically, the basic components in the *E. coli* in vitro expression system include: 19 amino acids at a concentration of 2 mM each; 2 mM tyrosine; 14 mM magnesium acetate; 60 mM potassium acetate; and 7 mM DDT. The energy-related components in the *E. coli* in vitro expression system include: 1.2 mM AMP, 0.85 mM CMP, 0.85 mM GMP, 0.85 mM UMP, 15–83 mM PEP, 0.4–0.6 mM NAD, 4 mM potassium oxalate, 90 mM potassium glutamate, and 2.5–10 mM magnesium glutamate. The additive components in the *E. coli* in vitro expression system include: 1.5 mM spermidine and 157.33 mM... HEPES; In the E. coli in vitro expression system, the concentration of RNase inhibitor was 150 U / 450 μL; The volume content of cell extract in the E. coli in vitro expression system was 20-60%.

[0067] In some preferred embodiments, the concentration of PEP in the above-mentioned Escherichia coli in vitro expression system is 30 mM; preferably, the content of NAD is 0.4 mM; preferably, the content of magnesium glutamate is 7.5 mM; preferably, the volume content of cell extract in the Escherichia coli in vitro expression system is 33.3%.

[0068] The beneficial effects of this application will be further illustrated below with reference to specific embodiments.

[0069] It should be noted that, unless otherwise specified, the in vitro protein expression system used in the following examples is the in vitro protein expression system of Escherichia coli, and unless otherwise specified, the experiments are all conducted using sfGFP (superfolder Green fluorescent protein) as an example.

[0070] Example 1:

[0071] Establishment and optimization of a cell-free protein synthesis system (also referred to as an in vitro protein expression system in this application) (1 mL system)

[0072] Table 1:

[0073] The final concentrations of Solution A in Table 1 are: 1.2 mM ATP, 0.85 mM GMP, 0.85 mM UMP, 0.85 mM CMP, 31.50 ug / mL leucovorin, 170.60 ug / mL tRNA, 0.40 mM NAD, 0.27 mM Coprotein A (CoA), 4 mM oxalic acid, 1 mM diammonium succinate, 1.50 mM spermidine, and 57.33 mM HEPES buffer.

[0074] The final concentration of Solution B is: 10mM Mg(Glu)2, 10mM NH4(Glu), 130mM K(Glu), 2mM 20 amino acids, and 0.03M phosphoenolpyruvate (PEP).

[0075] The system was reacted at 30℃ and 220 rpm for 16 h.

[0076] Using sfGFP as the model protein, the expression systems in Table 1 were optimized according to the existing publicly available in vitro expression systems (systems 1 to 3) in Table 2. Specifically, this included optimizing the expression systems for Mg. 2+ The optimization of the concentration of PEP, the proportion of cell extract in the whole reaction system, the amount of NAD, and the concentration of glutamate were all optimized.

[0077] Fluorescence intensity detection: excitation light 485nm, emission light 525nm, 50μL system detected in 96-well plate. The standard curve is shown in Figure 1. The detection results under different optimized conditions are shown in Figures 2 to 6.

[0078] 1) The system Mg 2+ The optimized concentration results of Mg are shown in Figure 2. 2+ sfGFP can be produced in in vitro protein expression systems at concentrations ranging from 2.5 mM to 19.5 mM. 2+ The effect is better at concentrations of 2.5 mM to 10 mM, and optimal at around 7.5 mM.

[0079] 2) The optimization results of PEP concentration in the system are shown in Figure 3. The protein expression system can produce sfGFP when the PEP concentration is between 5mM and 83mM. The protein synthesis is better when the PEP concentration is between 15 and 83mM, and the sfGFP protein synthesis is the highest when the PEP concentration is 30mM.

[0080] 3) The optimized proportion of cell extract in the whole reaction system is shown in Figure 4. The target protein can be generated when the proportion of cell extract in the whole system is between 20% and 60%, and better results can be obtained when it exceeds 33%.

[0081] 4) The optimization results of NAD dosage are shown in Figure 5. NAD plays an important role in energy cycling. The expression of NAD is best at a concentration of 0.6 mM, and there is no significant difference between 0.4 mM and NAD at a concentration of 0.6 mM.

[0082] 5) The results of the optimized concentration of glutamate are shown in Figure 6. The experimental results show that glutamate does not play any role in the whole experiment and the effect is better when it is not added.

[0083] Therefore, after optimizing the above parameters, the in vitro expression system of this application, as shown in the last column of Table 2, was obtained. The pH of this system is 7-8, typically pH 7.5. The reaction time is 3-16 h, typically 4 h.

[0084] Table 2:

[0085] The cell extracts in the table above were obtained using the following methods:

[0086] Activated strain BL21 Star(DE3) was streaked to produce single colonies. After activation, BL21 Star(DE3) was inoculated into 50 ml of LB liquid medium and cultured overnight at 37°C and 200 rpm. The overnight cultured BL21 Star(DE3) was then transferred to 400 ml of 2×YT medium to achieve an initial OD600 of 0.1. IPTG was added to a final concentration of 0.5 mM, and the culture was incubated at 37°C until OD600 reached 3.8–4.0. The bacterial sludge was collected: centrifuged at 5000 g at 10°C for 10 min. The supernatant was slowly poured off, and the sludge was transferred to a pre-chilled 50 ml centrifuge tube. 30 ml of S30 buffer was added to the 50 ml centrifuge tube to resuspend the cells. The tube was centrifuged at 5000 g at 10°C for 10 min, the supernatant was discarded, and the water in the centrifuge tube was removed using clean filter paper. 1 ml of pre-chilled S30 buffer was added for every 0.6 g of bacterial sludge. Resuspend cells, sonicate to disrupt, and add 65 μl of 1M DTT to 5 ml of cell lysis buffer. Centrifuge at 12000 rpm, 4°C for 10 min. Aliquot and store at -80°C for use.

[0087] Example 2

[0088] The mutation is introduced using a two-step PCR method, which directly yields the gene fragment with the mutated site, allowing for direct in vitro protein expression. This eliminates a series of intermediate steps, including PCR product recovery, protein digestion, ligation, transformation, single-clone culture, sequencing, shake-flask or plate culture, induction, centrifugation for bacterial collection, disruption, and centrifugation for supernatant collection. The mutant protease can be obtained directly from the mutated PCR product.

[0089] The mutation point is introduced by a two-step PCR method. Primers are designed near the mutation site, and the mutation sequence containing the mutation site is introduced into the primer sequence. In the first step of PCR, the fragments on both sides of the mutation site are PCRed out. Then, in the second step of PCR, the two products of the first step are mixed as a template, and primers at both ends are added. The full-length sequence is obtained by PCR and can be used as a template for protein expression in vitro.

[0090] In a 450 μL in vitro protein expression system, using the PCR product of the sfGFP gene as a DNA template, different amounts of PCR product were added for in vitro protein expression. Figure 7 shows that from when the PCR product volume exceeded 22.5 μL until it reached 90 μL, the amount of protein produced by the in vitro expression system remained relatively uniform, indicating that fluctuations in the amount of PCR product added within this range resulted in relatively parallel production of the target protein.

[0091] Furthermore, using the PCR product of the sfGFP gene as a DNA template, different PCR products included different lengths upstream of the start codon, namely 0bp, 50bp, 100bp, 115bp, 130bp, and 140bp. When these PCR products were used as DNA templates in the reaction system, the effect of the length upstream of the start codon on the sfGFP expression results is shown in Figure 8. The length upstream of 50bp to 140bp had little effect on the expression level of sfGFP protein. Among them, when the length including 50-100bp upstream of the start codon was included, the protein expression level was the highest.

[0092] Example 3

[0093] Site-directed saturation mutation is a common method for constructing mutants in directed protein evolution. It is used in various mutation methods, including semi-rational design, random mutation, even rational mutation, or simplified codon mutation.

[0094] In this embodiment, a site-directed saturation mutagenesis was constructed using an in vitro protein expression system. The ester protein Asym-503029, with the amino acid sequence shown in SEQ ID NO: 1, underwent a saturation mutagenesis, and the proteolytic reaction it catalyzed is shown in the reaction formula above:

[0095] Asymchem-503029 is active against the target substrate, but its stereoselectivity is not good, with an ee of approximately 61%. Based on computer structural simulation results, a saturation mutation was performed at the G19 site. Nineteen primers for the G19 site were synthesized, and the mutation was introduced using PCR. The PCR product was directly used in an in vitro protein expression system for protein synthesis, and the in vitro expression system was then used to validate the protein-catalyzed reaction.

[0096] Figure 9 shows the electrophoresis results of eight randomly selected mutants (from left to right: molecular marker, G19D, G19A, G19Y, G19H, G19N, G19M, G19F, and G19S). The SDS-PAGE in Figure 9 shows that the different mutants produced a good amount of protein. Using the Bio-Rad gel imaging system, the protein concentration was calculated to be within 1.5 ± 0.1 mg / mL for all mutants.

[0097] Table 3: Results of the response of the in vitro protein expression system to the saturation mutation at the G19 site.

[0098] As shown in the table above, the G19S mutant significantly increased the ee value to 75.73%, and its transformation rate also increased by about 100% to 33.96%. We also performed a control reaction using the traditional PCR-protein digestion-ligation-transformation-clone picking-shake flask culture-ultrasonic disruption method. Figure 10 shows that the reaction results of the in vitro protein expression system of this application have a good correlation with the shake flask reaction results, indicating that this method is feasible for directed protein evolution.

[0099] However, the two methods differ significantly in efficiency. PCR took 3 hours, in vitro protein expression took 3 hours, and the protein catalytic reaction took 1 hour, totaling 7 hours to complete the entire process from gene to protein to performance detection. This process typically takes 2-3 weeks using traditional methods. Therefore, the protease evolution method proposed in this application can significantly improve the efficiency of evolutionary screening.

[0100] Example 4:

[0101] This embodiment uses an in vitro protein expression system to construct a site-directed mutagenesis system. Site-directed mutagenesis is the most commonly used method in directed protein evolution, used for rational design, semi-rational design, and the stacking of mutation sites, etc.

[0102] As shown in Example 3, Asymchem-503029 is active against the target substrate; however, its stereoselectivity is not good enough, with an ee of approximately 61%. A site-directed mutagenesis was performed on the G19S site of the ester protein Asym-503029 shown in SEQ ID NO: 1.

[0103] G19S primers were synthesized, and then the mutation was introduced using PCR. At the same time, the parent fragment of Asymchem-503029 was amplified by PCR using conventional primers (primers without the G19S mutation) as a control. The products of both PCRs were directly used in the in vitro protein expression system for protein synthesis. The in vitro protein expression system after protein synthesis was directly used to verify the protein catalytic reaction.

[0104] The results were the same as in Example 3. The parent parent had an ee value of 61.1% and a transformation rate of 15.4%. The mutant G19S had an ee value of 75.7% and a transformation rate of 33.9%.

[0105] In terms of efficiency, in this embodiment, PCR took 3 hours, in vitro protein expression took 3 hours, and protein catalytic reaction took 1 hour. The complete process from gene to protein to trait was achieved in a total of 7 hours, while site-directed mutagenesis generally takes 1-2 weeks using traditional methods.

[0106] Example 5:

[0107] Full-sequence random mutation and multi-point mutation were performed using an in vitro protein expression system.

[0108] For random mutations, the most frequently used method is error-prone PCR. Error-prone PCR has a large library capacity, but it is limited by the screening throughput. Generally, error-prone PCR can screen about 1,000-2,000 mutants. Due to the probability distribution, many of these mutants are repetitive mutations. At the same time, due to the base preference of PCR, the location of the introduced mutation and the mutated amino acid are not evenly distributed.

[0109] For industrial proteins, the number of amino acids is typically around 300. Five primers are designed for each site, each mutating the site to one of five amino acids representing different properties: alanine (A, or G if the site is originally A), serine (S), lysine (K), aspartic acid (D), and phenylalanine (F). These represent five types of amino acids: less sterically hindered, polar, positively charged, negatively charged, and aromatic. In this way, error-prone PCR can be addressed using global PCR.

[0110] The transaminase protein TA-1 shown in SEQ ID NO: 2 has high selectivity for the substrate (see reaction formula II), but its activity is relatively poor and a large amount of protein is required.

[0111] Using global PCR instead of error-prone PCR, three mutants, L76A, S125A, and A226G, were obtained after activity testing, and their protein activity was improved (see Table 4).

[0112] Then, a 3-point mutant was directly constructed using a multi-point mutagenesis method. Four PCR product fragments were obtained by PCR: T7 to L76A, L76A to S125A, S125A to A226G, and A226G to T7 terminal. These four PCR product fragments were used for the second step of over-lap PCR. The products of the second step PCR were directly used in the in vitro protein expression system to obtain the target 3-point mutant L76A+S125A+A226G for protein activity assay.

[0113] In this embodiment, the first round of irrational evolution plus the second round of mutation site combination takes a total of 1 week, while the traditional method takes 1.5 to 2 months.

[0114] Table 4:

[0115] As can be seen from the above description, the embodiments of the present invention achieve the following technical effects: This application utilizes an in vitro protein synthesis system to accelerate directed protein evolution, introduces mutations using PCR, and combines this with an in vitro protein expression system for directed protein evolution. This method can be used for all current protein evolution techniques, including irrational design, rational design, and semi-rational design. Furthermore, the process is simple, requiring only two steps: PCR and directly using the PCR product for in vitro protein expression. The target mutant can be obtained within 6-8 hours. Moreover, it does not involve microbial manipulation, greatly reducing operational difficulty and risk. The experimental results also exhibit high parallelism and good robustness.

[0116] Compared with existing directed evolution methods, the method of this invention has the following advantages:

[0117] 1) The protein expression system synthesizes a high amount of protein, which can be used for protein evolution.

[0118] 2) The protein evolution effect data of the present invention are consistent with the effect data obtained by traditional methods.

[0119] 3) The protein evolution method provided by this invention saves more than 80% of the time compared with traditional experiments, greatly accelerating the speed of protein evolution.

[0120] 4) The evolution method provided by this invention has simple steps and obtains better data parallelism.

[0121] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for directed evolution of proteins, characterized in that, The method includes: S1, PCR products containing gene mutations of the target protein are obtained through PCR amplification; S2, The PCR product is placed in an in vitro protein expression system for gene expression to obtain the target protein with the gene mutation; S3, perform phenotypic testing on the target protein with the gene mutation, the phenotypic testing including testing the activity of the target enzyme and / or the excess percentage of the corresponding isoform; The target protein is a target enzyme; The protein in vitro expression system is an *E. coli* in vitro expression system, which includes: basic components, energy-related components, additives, cell extracts, and RNase inhibitors. In the Escherichia coli in vitro expression system, the basic components include: 19 amino acids at a concentration of 2 mM each, 2 mM tyrosine, 14 mM magnesium acetate, 60 mM potassium acetate, and 7 mM DDT. In the Escherichia coli in vitro expression system, the energy-related components include: 1.2 mM AMP, 0.85 mM CMP, 0.85 mM GMP, 0.85 mM UMP, 15–83 mM PEP, 0.4–0.6 mM NAD, 4 mM potassium oxalate, 90 mM potassium glutamate, and 2.5–10 mM magnesium glutamate; In the Escherichia coli in vitro expression system, the added components include: 1.5 mM spermidine and 157.33 mM HEPES; In the in vitro expression system of Escherichia coli, the concentration of the RNase inhibitor is 150 U / 450 μL; The cell extract contained 20-60% by volume in the Escherichia coli in vitro expression system.

2. The method according to claim 1, characterized in that, Use any one or more of the following methods to obtain PCR products containing gene mutations of the target protein via PCR amplification: 1) Amplify the PCR product containing the gene mutation of the target protein using a two-step PCR method; Two primer pairs were designed: F1 and R1; and F2 and R2. A mutated sequence containing the mutation site was introduced into each primer pair. A first-step PCR was performed using the primer pairs to PCR-generate fragments L1 and L2 flanking the mutation site. The overlapping region between fragments L1 and L2 is denoted as L, and the mutation site is located on fragment L. Then, using a mixture of fragments L1 and L2 as a template, a second-step PCR was performed using primers F1 and R2 to obtain the full-length sequence. This full-length sequence is the PCR product of the gene mutation containing the target protein. 2) Based on the principle of site-directed saturation mutagenesis, multiple PCR products containing gene mutations of the target protein are constructed by introducing mutations through PCR amplification. These multiple PCR products containing gene mutations of the target protein are then used to construct a saturation mutant library of the target protein gene; or 3) Mutations are introduced at specific sites using PCR amplification to obtain PCR products containing the target protein gene mutation; or 4) Using error-prone PCR, random mutations are performed on the entire gene sequence to obtain multiple PCR products containing the target protein gene mutations, wherein the multiple PCR products containing the target protein gene mutations cover the random mutations of the entire gene sequence of the target protein; or 5) Using a multi-point mutagenesis method, obtain PCR products containing multiple mutation sites of the gene carrying the target protein.

3. The method according to claim 1, characterized in that, In the in vitro expression system of *E. coli*, the concentration of PEP is 30 mM; The concentration of NAD was 0.4 mM; The concentration of magnesium glutamate was 7.5 mM; The cell extract contained 33.3% by volume in the Escherichia coli in vitro expression system.

4. The method according to claim 1, characterized in that, The target protein is selected from industrial proteases.

5. The method according to claim 4, characterized in that, The industrial protease is the ester protein shown in SEQ ID NO: 1 or the transaminase TA-1 shown in SEQ ID NO:

2.

6. The method according to claim 1, characterized in that, The phenotypic detection of the target enzyme carrying the gene mutation includes: Multiple target enzymes with different gene mutations were used to catalyze the same substrate reaction to produce the same product, and the conversion rate of the substrate catalyzed by different target enzymes and / or the excess percentage of the corresponding isomers of the product were detected. Using the conversion rate and / or isomer excess percentage of the substrate catalyzed by the initial control enzyme as a reference, the target enzymes that improve the conversion rate and / or corresponding isomer excess percentage are screened from a plurality of target enzymes and are designated as the initial +1 control enzyme.

7. The method according to claim 6, characterized in that, After obtaining the initial +1 control enzyme, the method further includes: iterating the initial +1 control enzyme back to the initial control enzyme, and then repeating steps S1 to S3, and so on, to obtain multiple target enzymes after directed evolution.

Citation Information

Patent Citations

  • New gene mutation recombination method and application thereof

    CN101532043A

  • New method for in-situ evolution of target protein in escherichia coli cells

    CN106011162A

  • Engineered transaminase polypeptide for preparing sitagliptin

    CN112048485A

  • In-vitro cell-free protein synthesis system (D2P system) and kit and use thereof

    CN113215005A

  • Efficient cell-free in-vitro protein expression system

    CN116103210A