A computational method for designing protease mutants with non-native substrate activity

CN115295077BActive Publication Date: 2026-08-11ENZYMASTER NINGBO BIO ENG CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-03
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但是由于这些算法本身都属于经验数学计算形式,计算公式中包含一些物理原理的能量计算项,也包含对现有数据库统计获得的统计能量项,到目前为止计算机辅助突变设计方法都有其各自的局限性,而通过计算预测得出的突变体的性能也很难与实验验证结果有可靠的吻合

Benefits of technology

[0020]Based on the analysis of the parameters of the energy function in the current algorithm, the calculations in the current algorithm design actually calculate the stability of the enzyme mutant protein structure, without involving the determination of the activity of the enzyme mutant in the specific reaction process. The overall structural stability of the enzyme mutant is one aspect of enzyme engineering modification and a prerequisite for catalytic activity. However, the discovery and improvement of enzyme catalytic activity (especially the creation of activity for non-natural substrates) is what the industry is more concerned with, because using highly active enzyme mutants to improve the catalytic efficiency for substrates can achieve industrial practicality. However, existing calculation methods have failed to determine the catalytic activity after stability screening. Therefore, this invention, based on the optimization of the stability calculation method, further applies the reaction energy barrier calculation to determine the activity of the enzyme mutant in the specific reaction process, taking into account multiple factors such as substrate molecules, intermediates, and reaction transition states, and realizing the determination of the level of catalytic activity of the predicted stable mutant. The computational design method disclosed in this invention can obtain a more concise mutant library for experimental verification, which greatly improves the accuracy and R&D efficiency of virtual screening of protease mutations; the calculation of reaction barriers effectively screens active enzyme mutants, improving the efficiency and effect of enzyme engineering modification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115295077B_ABST
    Figure CN115295077B_ABST
Patent Text Reader

Abstract

This disclosure provides a computational method for designing protease mutants with non-natural substrate activity, enabling the design of non-natural substrate-specific mutations in proteases. This invention constructs a unique method for processing stability calculation results and creatively adds a calculation process for reaction energy barriers, which improves the accuracy of virtual screening for protease mutations. This not only significantly reduces the number of mutations required for screening, saving manpower and resources, but also unexpectedly achieves the effect of engineering enzyme modification that is impossible with traditional enzyme-directed evolution methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to computer-aided design and virtual screening of protease mutants, and more specifically, to a computational method that combines virtual screening of stable protease mutants and virtual screening of catalytically active mutants to design protease mutants with non-natural substrate activity. Background Technology

[0002] Protease catalysts play a crucial role in modern synthetic industries. With the ever-expanding applications of enzyme catalysis, the catalytic performance of naturally occurring enzymes can no longer meet the requirements of enzymological research and industrial applications. Directed evolution, a key technique for modifying proteases, is a protein engineering strategy that more closely approximates natural evolution. Also known as laboratory evolution, it is a method for rapidly modifying proteins in vitro, mimicking the natural evolutionary process. Even without knowing the three-dimensional structure and mechanism of action of the target protein, it can complete an evolutionary process that would normally take millions of years in nature in a very short time, thus obtaining enzymes with the desired functions. In recent years, directed evolution technology has been widely applied in the development of enzyme catalysts required in pharmaceuticals, food, and chemical engineering, triggering another revolution in the field of biocatalysis and greatly expanding the research and application scope of protein engineering. The applicant has been dedicated to the research and application of directed evolution technology for enzymes and has successfully developed a large number of protease catalysts for the production of pharmaceuticals and fine chemicals. However, laboratory directed evolution typically requires screening a large number of enzyme mutant libraries, and the construction of these libraries and subsequent screening are arduous tasks for researchers. Based on extensive prior research into directed evolution samples, the applicant, combining computational biology and bioinformatics techniques, has developed a feasible computational method capable of effectively screening envisioned enzyme mutants virtually and reliably predicting enzyme mutants with desired performance. This significantly reduces the scope of mutant libraries that need to be constructed and screened in laboratory directed evolution processes. The computational method disclosed in this application not only powerfully complements laboratory directed evolution processes and, when combined with laboratory directed evolution techniques, overcomes the limitations of the libraries that can be constructed and screened in laboratories. It can reduce R&D costs, improve R&D efficiency, and more effectively obtain enzyme mutants with desired performance.

[0003] Many computational methods have been published for protein research. For example, there are molecular docking algorithms for observing the binding of small molecule substrates to proteins, homology modeling and de novo modeling algorithms for constructing three-dimensional protein models, and tools such as Foldx, I-mutant, and Rosetta for calculating protein structural stability. These methods have been widely applied in protein mutation design. However, because these algorithms are all empirical mathematical calculations, their formulas include energy calculation terms based on physical principles and statistical energy terms obtained from existing databases. Therefore, computer-aided mutation design methods currently have their limitations, and the performance of mutants predicted through computation is difficult to reliably match experimental results. Currently, no computational design procedure for protease mutants has been provided to reliably predict enzyme mutants with the desired performance. Summary of the Invention

[0004] This invention specifically develops a computational method for effective computer-based virtual screening of enzyme mutants and reliable prediction of enzyme mutants with desired performance. This invention constructs a unique method for processing stability calculation results and creatively adds a calculation process for reaction energy barriers, which improves the accuracy of virtual screening of protease mutants. This not only significantly reduces the number of mutants required for screening, saving manpower and resources, but also unexpectedly achieves the effect of engineering modification of enzymes that is impossible with traditional directed enzyme evolution methods.

[0005] The calculation method of the present invention is as follows: Figure 1 As shown, it includes the following four specific steps:

[0006] (1) Protein structure model acquisition: A three-dimensional structural model of the protein is obtained based on the amino acid sequence of the target protease. This structural model can be an experimentally obtained structural model recorded in the PDB (protein data bank) database, or a structural model virtually constructed based on the protein sequence using homology modeling or de novo modeling methods. Generally, homology modeling with higher homology is more accurate than de novo modeling. For protein structures from different sources, they must be in the catalytic conformation, i.e., the substrate molecule (product or transition state)-protein complex.

[0007] (2) Substrate docking analysis: The amino acid sites on the target protease's three-dimensional structure that bind to the substrate molecule are determined. Then, the enzyme's native or target substrate (including non-natural substrates) is docked with the protein. Suitable mutation sites within the protein's active site are selected based on the different conformations of the docking results. Many computational software programs can perform molecular docking, such as Discovery Studio, Schrodinger, and Yasara (including Autodock and Autodockvina plugins). Amino acid sites requiring mutation are screened by comparing the docking conformations.

[0008] (3) Mutant stability calculation: The stability of single-site mutants and / or multi-site combination mutants is calculated based on the sites selected by docking results. First, a Python script is used to generate a batch of mutant sets to be virtually screened. Then, software such as Yasara or Rosetta is used to generate the structural model of each mutant, and algorithms such as ddg_monomer, Cartesian_ddg, FoldX, Provean, ELASPIC, or Amber TI are used to calculate the structural stability of each mutant. Finally, a Python analysis script is used to calculate the free energy difference (ΔΔG) between the structure of each mutant and the wild-type enzyme.

[0009] There are two strategies for processing the results of the stability calculations mentioned above: simple sorting method [3a] and statistical method [3b].

[0010] [3a] Simple sorting method: that is, simply sort the ΔΔG results of the mutants from low to high, and select the mutants with the highest ranking as stable mutants obtained by computer virtual screening prediction.

[0011] [3b] Statistical method: The mutants are sorted from low to high ΔΔG results, and the mutants at the top and bottom of the sort are selected for frequency analysis of mutated amino acid residues. For a specific amino acid site, the amino acid residue types that appear more frequently in the stable mutants are subtracted from the amino acid residue types that appear more frequently in the unstable mutants to obtain the set of amino acid residue types that can be mutated at that site. Finally, the mutable amino acid residues at each site are arranged and combined as stable mutants predicted by computer virtual screening.

[0012] The criteria for evaluating mutant stability calculation results are as follows: ΔΔG ≤ -1 kcal / mol indicates a stable mutant; ΔΔG ≥ 1 kcal / mol indicates an unstable mutant; and -1 kcal / mol < ΔΔG < 1 kcal / mol indicates an invalid mutant. This criterion also applies to the evaluation of stability results obtained using other calculation methods.

[0013] There are already many methods available for calculating protein structural stability, so this process is not limited to using the Rosetta program. The inventive contribution of this invention in the stability calculation process lies in the conception and use of a statistical method [3b] to process the calculation results. Statistical analysis is used to screen for amino acid residue mutations at each site, and then the screened mutations are combined as stable mutants predicted by computer virtual screening.

[0014] Current computer-based virtual screening methods generally stop there. After obtaining the predicted stable mutants using a simple ranking method [3a], these mutants are then validated in the laboratory using specific experimental protocols, without further computational methods to evaluate their catalytic activity. Furthermore, due to the limitations of stability calculations, the stable mutants identified by the simple ranking method [3a] may miss some mutants that actually exhibit high stability and activity in experimental verification. This invention presents a statistical method [3b] for screening stable mutants, which can obtain stable mutants that current algorithms cannot predict. Moreover, while current common computational methods eliminate a large number of unstable mutants through virtual screening, reducing the number of mutants requiring laboratory verification, they cannot further determine the catalytic activity of the predicted stable mutants through computation, nor can they assess the activity of the predicted stable mutants on non-natural substrates.

[0015] Based on this invention, the protease mutant design method further adds a method and process for calculating the reaction energy barrier of each mutant in the catalytic chemical reaction (natural or non-natural substrate), realizing the judgment of the catalytic activity level of the predicted stable mutant, and being able to evaluate whether the predicted stable mutant has activity on non-natural substrate.

[0016] (4) Reaction Barrier Calculation: This method is based on the force field description of different substrate reaction states, and also provides a quantum mechanical description of chemical reactions within the framework of valence bond theory. This allows the reaction barrier calculation to utilize the computational speed of classical force field-based methods while carrying a large amount of chemical and thermodynamic information, thus providing a meaningful physical description of bonding and breaking processes. Preparations before the reaction barrier calculation include selecting the simulation force field, determining the rate-limiting step of the target enzyme reaction, the reaction transition state form of the substrate small molecule compound binding, etc. After determining the calculation parameters, the reaction barrier calculation will be performed in the cadee flow. This invention uses Qtools to analyze the calculation results. In the cadee calculation flow, for example, the default simulation calculation time is set to 12.6 ns. First, the system is gradually heated from 0.01 K to 300 K within a 90 ps simulation time. For example, 200 kcal mol is applied to all protein atoms in the simulation. -1 Harmonic suppression, applying 20 kcal mol to all water atoms -1 Harmonic suppression was applied initially, then gradually decreased as the temperature increased. A Berendsen thermostat was used to regulate the temperature, with a time step of 1 fs. For all simulations, the reaction coordinates were set to λ = 0.5 to begin subsequent reaction barrier calculations for reaction steps approaching the transition state. Molecular dynamics (MD) simulations were performed for each of the four parallel calculations at 8 ns intervals, and the results were used as a starting point for empirical valence bond simulations. In the 8ns MD simulation, a snapshot was taken every 1ns to obtain the structure close to the transition state. Based on this structure, an empirical valence bond simulation was performed at 520ps, distributed in 26 FEP / US windows at 20ps each (λ = 0, 0.05, 0.075, 0.1, 0.125, 0.15, 0.2, 0.25, 0.30, 0.35, 0.40, 0.425, 0.45, 0.55, 0.575, 0.6, 0.65, 0.70, 0.75, 0.80, 0.85, 0.875, 0.90, 0.925, 0.95, 1). Using the default Cadee computation flow would be a very lengthy process. Therefore, this invention reduces the MD simulation time from 8 ns to 4 ns, and the number of repeated computation cycles in each parallel computation copy is reduced to 4. This reduces the computation time to less than 24 hours, without significantly altering the results. Cadee can perform multi-task parallel computation on multi-core computers; therefore, on high-performance computers, Cadee can screen the catalytic activity of large-scale mutants.

[0017] The reaction energy barrier is the minimum energy required for reactant molecules to reach the activated molecule; the magnitude of the barrier reflects the ease with which a reaction occurs. For example... Figure 2 As shown, the difference between the potential energy of the enzyme and substrate in their free state and the potential energy of the activated molecule formed by their combination is the reaction energy barrier, that is, the difference between the lowest energy point (the optimal conformation for enzyme-substrate binding) and the highest energy point (the optimal conformation for enzyme-substrate transition state binding). In our calculation process, this is the energy difference between the lowest energy point of the initial state and the highest energy point of the activated state.

[0018] Experimental verification of the catalytic activity of the mutant: Figure 1 The calculated catalytically active mutants were constructed and expressed in the laboratory, and their catalytic activity against the target substrate was then tested to verify whether the mutants predicted by the computational method had the desired catalytic performance.

[0019] In traditional directed evolution, there are usually no clear mutation sites for the proteases being evolved, and it is impossible to predict which amino acid residues will be beneficial mutations. This often results in the construction of large, diverse libraries, and the screening of these libraries consumes significant time and resources. However, after implementing the computational design process of this invention—that is, calculating enzyme structural stability based on the docking of the protease and substrate molecules, and performing virtual screening using the stability result processing method disclosed in this invention—a large number of unstable and invalid mutations can be filtered out, greatly reducing the number of mutant libraries required for laboratory screening. Regarding current enzyme stability calculations, on the one hand, the calculations themselves are empirical algorithms, which cannot guarantee complete accuracy. Furthermore, for the sake of computational feasibility, many factors, such as the flexibility of the amino acid backbone, are not considered variables. In addition, small molecules in substrate docking are often treated as rigid bodies, and surrounding environmental factors are also simplified. Therefore, in practical applications of computational methods, the success rate of predicting mutants in laboratory verification is not high, and many good mutants are often missed. Furthermore, current computational methods typically process stability calculation results by directly sorting the energy levels to identify mutants with the lowest possible energy, assuming these are the most likely mutant sequences to be active against the target response. In this invention, the applicant, drawing on years of practical experience, has further optimized the stability result processing method. This invention specifically proposes mutation frequency analysis, which involves subtracting the most frequent amino acid residue mutation types from those in unstable mutations to obtain the theoretically mutable amino acid residue types at that site. Finally, the mutable amino acid residues at each site are combined to obtain stable mutants. This approach ensures, to a certain extent, that effective mutations are not easily excluded, and these types of mutants are often overlooked in the virtual screening process of currently known algorithms.

[0020] Based on the analysis of the parameters of the energy function in the current algorithm, the calculations in the current algorithm design actually calculate the stability of the enzyme mutant protein structure, without involving the determination of the activity of the enzyme mutant in the specific reaction process. The overall structural stability of the enzyme mutant is one aspect of enzyme engineering modification and a prerequisite for catalytic activity. However, the discovery and improvement of enzyme catalytic activity (especially the creation of activity for non-natural substrates) is what the industry is more concerned with, because using highly active enzyme mutants to improve the catalytic efficiency for substrates can achieve industrial practicality. However, existing calculation methods have failed to determine the catalytic activity after stability screening. Therefore, this invention, based on the optimization of the stability calculation method, further applies the reaction energy barrier calculation to determine the activity of the enzyme mutant in the specific reaction process, taking into account multiple factors such as substrate molecules, intermediates, and reaction transition states, and realizing the determination of the level of catalytic activity of the predicted stable mutant. The computational design method disclosed in this invention can obtain a more concise mutant library for experimental verification, which greatly improves the accuracy and R&D efficiency of virtual screening of protease mutations; the calculation of reaction barriers effectively screens active enzyme mutants, improving the efficiency and effect of enzyme engineering modification. Attached Figure Description

[0021] Figure 1 Computational design flow for protease mutants

[0022] Figure 2 Schematic diagram of reaction energy barrier

[0023] Figure 3 This is the reaction catalyzed by ketone reductase to produce 4-hydroxy-2-butanone from 1,3-butanediol.

[0024] Figure 4 Flowchart for calculating reaction energy barrier

[0025] Figure 5 The atomic number of 1,3-butanediol

[0026] Figure 6 This is the reaction catalyzed by ketone reductase to produce 1,3-butanediol from 4-hydroxy-2-butanone.

[0027] Figure 7 Ketoreductase catalyzes the production of 4-hydroxy-2-butanone from 1,3-butanediol.

[0028] Figure 8 GC spectra of 4-hydroxy-2-butanone and 1,3-butanediol

[0029] Figure 9 GC spectra of (R)-(-)-1,3-butanediol and (S)-(-)-1,3-butanediol

[0030] Figure 10The reaction process of catalyzing the synthesis of β-alanine from acrylic acid by an amino-lysin mutant.

[0031] Figure 11 Procedure for calculating the reaction energy barrier of mutants

[0032] Figure 12 Acrylic acid atomic number

[0033] Figure 13 Spectra for the detection of acrylic acid and β-alanine Detailed Implementation

[0034] The following examples further illustrate the present invention, providing a clear and complete description of its technical solutions. However, the present invention is not limited thereto. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Furthermore, unless otherwise specified, the equipment and reagents involved in the following embodiments are commercially available.

[0035] Example 1: Designing a mutant ketone reductase active against a non-natural substrate (an enantiomer of the natural substrate).

[0036] The ketone reductase disclosed in patent CN111321129A can asymmetricly convert 4-hydroxy-2-butanone into the more expensive (R)-(-)-1,3-butanediol, which is of great significance and industrial value. Simultaneously, this enzyme can also catalyze the conversion of alcohols to ketones, but due to substrate specificity, it can only catalyze the substrate of (R)-(-)-1,3-butanediol. If racemic 1,3-butanediol is used as the substrate (i.e., a 1:1 mixture of (R)-(-)-1,3-butanediol and (S)-(-)-1,3-butanediol), converting (R)-(-)-1,3-butanediol to 4-hydroxy-2-butanone while retaining (S)-(-)-1,3-butanediol unchanged results in a very low overall conversion rate. In actual industrial production, racemic 1,3-butanediol is readily available and its price is far lower than that of chiral pure (R)-(-)-1,3-butanediol. Developing a ketone reductase capable of converting all racemic 1,3-butanediol to 4-hydroxy-2-butanone would have significant industrial application value. This example uses the ketone reductase disclosed in CN 111321129A, which has a selectivity of >99% for (R)-(-)-1,3-butanediol, as its starting point. Its amino acid sequence is shown in SEQ ID NO: 2, and its DNA sequence is shown in SEQ ID NO: 1. SEQ ID NO: 2 shows no activity against (S)-(-)-1,3-butanediol, meaning that (S)-(-)-1,3-butanediol is a non-natural substrate for SEQ ID NO: 2. Using the computational design algorithm disclosed in this invention, 10 ketone reductase mutants with simultaneous activity for (R)-(-)-1,3-butanediol and (S)-(-)-1,3-butanediol were predicted through virtual screening. After experimental verification of these 10 predicted mutants, a ketone reductase mutant that maintains high catalytic activity and can simultaneously convert (R)-(-)-1,3-butanediol and (S)-(-)-1,3-butanediol to 4-hydroxy-2-butanone was successfully found.

[0037] Calculation design process for Example 1:

[0038] (1) Structure acquisition: The structural model of the target protein was obtained by homology modeling of SEQ ID NO:2 using the YASARA software package.

[0039] (2) Substrate docking analysis: (R)-(-)-1,3-butanediol and (S)-(-)-1,3-butanediol were docked with the target protein, respectively, using the Yasara software package. The docking results were visualized and analyzed in Yasara. The docking results of (S)-(-)-1,3-butanediol with the target protein showed that the amino acid residue side chains at sites I144, H145, Q150, and Y188 had excessive steric hindrance. Therefore, these sites were selected for the next step of virtual screening in this embodiment.

[0040] (3) Mutation stability calculation: Select appropriate amino acid types based on the mutated amino acids, as shown in Table 1. Then, use a Python script to generate the input file of mutation combinations required by Rosetta. There are 400 possible mutation combinations. Use the Cartesian_ddg algorithm to calculate the free energy difference (ΔΔG) between the structure of each mutant and the wild-type enzyme structure.

[0041] Table 1 Types of amino acid mutations

[0042]

[0043] The calculation results were then analyzed using the statistical analysis method described in [3b] above, and the results are shown in Table 2.

[0044] Table 2 Calculation and Analysis Results

[0045]

[0046] (4) Reaction barrier calculation: The mutations obtained in the previous step were combined to form 36 mutants. The reaction barrier was calculated using the Cadee calculation process, as follows: Figure 4 As shown, the atomic numbering of 1,3-butanediol is as follows: Figure 5 As shown.

[0047] Simulation setup: First, the simulation system is dissolved in a solution with (S)-(-)-1,3-butanediol as the substrate, centered at the C3 atom, with a radius of [missing information]. The TIP3P model of water molecules in a spherical water droplet, in which all atoms are in The simulation center is fully mobile and uses 10 kcal / mol. -1 The harmonic suppressor is constrained in the simulation center located at 17 and All atoms between, and constrained using a harmonic force constant of 200 kcal. Atoms other than those in the solvent. The SHAKE algorithm is used to confine H atoms within the solvent. The cutoff value is used to calculate the nonbonding interactions between all atoms except those in the empirical valence bond region. For these atoms, all interactions are explicitly calculated up to the cutoff value. All long-distance static electricity exceeding this critical value is handled using the Local Reactive Field (LRF) method.

[0048] From the calculated reaction barrier results, the 10 mutants with the lowest reaction barrier were selected, as shown in Table 3.

[0049] Table 3. The 10 ketoreductase mutants with the lowest catalytic energy barrier for (S)-(-)-1,3-butanediol.

[0050]

[0051]

[0052] (5) Experimental verification:

[0053] (5.1) with Figure 6 The reaction was performed to verify the altered chiral selectivity of the 10 mutants for 1,3-butanediol shown in Table 3. 0.1 g of 4-hydroxy-2-butanone, 0.5 mL of isopropanol, 0.1 g of wet bacterial cells expressing ketoreductase (recombinant expression process according to the method disclosed in patent CN111321129A), and 0.005 g of cofactor NAD+ were added to the reaction flask. The final reaction volume was brought up to 5 mL with 0.1 M PBS (pH 7). The reaction was carried out in a water bath at 40°C with a stirring speed of 400 rpm. After 1 hour, samples were taken and analyzed by HPLC. The chiral values ​​(ee%) of the product 1,3-butanediol were calculated, as shown in Table 4. The formula for calculating ee% is ee%=([R]-[S]) / ([R]+[S]), where [R] represents the concentration of (R)-(-)-1,3-butanediol in the sample, and [S] represents the concentration of (S)-(-)-1,3-butanediol in the sample.

[0054] Table 4

[0055]

[0056]

[0057] Note 1: (S)-(-)-1,3-butanediol was not detected.

[0058] The ee% results in Table 4 indicate that the 10 ketoreductase mutants (i.e., SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20 and 22) catalyze Figure 6During the reaction, not only (R)-(-)-1,3-butanediol was generated, but also (S)-(-)-1,3-butanediol was generated, indicating that the 10 ketone reductase mutants predicted by calculation possessed activity for (S)-(-)-1,3-butanediol.

[0059] (5.2) For the racemic substrates containing both (R)-(-)-1,3-butanediol and (S)-(-)-1,3-butanediol, which are readily available industrially, the applicant experimentally verified the catalytic performance of the above 10 ketone reductase mutants. The reaction is as follows: Figure 7 As shown, the specific experimental procedure is as follows:

[0060] Add 0.1 g of racemic 1,3-butanediol (i.e., a 1:1 mixture of (R)-(-)-1,3-butanediol and (S)-(-)-1,3-butanediol), 0.5 mL of acetone, 0.1 g of wet bacterial cells expressing ketoreductase (the recombinant expression process is based on the method disclosed in patent CN111321129A), and 0.005 g of cofactor NAD+ to the reaction flask, and then make up the final reaction volume to 5 mL with 0.1 M PBS (pH 7). Maintain the reaction at 40°C using a water bath and a stirring speed of 400 rpm. After 1 h, sample was taken for HPLC analysis, and the calculated molar conversion rates are shown in Table 5.

[0061] Table 5

[0062]

[0063]

[0064] Since SEQ ID NO: 2 is only active for (R)-(-)-1,3-butanediol and not for (S)-(-)-1,3-butanediol, its catalytic activity... Figure 7 The theoretically highest conversion rate achievable for the substrate (i.e., racemic 1,3-butanediol) in the reaction shown is 50%. The conversion data in Table 5 indicate that the 10 designed ketoreductase mutants (i.e., SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20 and 22) can achieve a conversion rate >50%, demonstrating that these mutants can simultaneously convert (R)-(-)-1,3-butanediol and (S)-(-)-1,3-butanediol to 4-hydroxy-2-butanone.

[0065] Analytical detection methods for compounds

[0066] GC conversion analysis method: The chromatographic column was a DB-WAX 15m*0.25mm*0.25μm, the carrier gas was N2, the detector was FID, the injection port temperature was 250℃, the split ratio was 28:1, the detector temperature was 300℃, the injection volume was 1μL, the column temperature was 130℃, and the temperature was increased to 150℃ at 10℃ / min, then increased to 160℃ at 20℃ / min. The retention time of 4-hydroxy-2-butanone was 1.5 min, and the retention time of 1,3-butanediol was 2.3 min. Figure 8 .

[0067] GC chiral analysis method: Sample pretreatment involved adding 200 μL of inactivation solution, 50 μL of MSTFA, and 30 μL of anhydrous pyridine to a 1.5 mL centrifuge tube, mixing thoroughly, and shaking for 30 min. The chromatographic column was a CP-Chirasil Dex CB (CP7502) 25 m * 0.25 mm * 0.25 μm, carrier gas was N2, detector was FID, injection port temperature was 250℃, split ratio was 28:1, detector temperature was 300℃, injection volume was 1 μL, column temperature was 105℃, stop time was 9 min, retention time of (R)-(-)-1,3-butanediol was 6.4 min, and retention time of (S)-(-)-1,3-butanediol was 6.6 min. Figure 9 .

[0068] Example 2: Design of mutants of aspartate amino lyase using acrylic acid as a substrate

[0069] Patent CN109385415A discloses a mutant of aspartate amino lyase used to catalyze the synthesis of β-alanine from acrylic acid as a substrate (e.g. Figure 10 (As shown). The wild-type aspartate amino lyase is derived from Bacillus sp. YM55-1, and its amino acid sequence is shown in SEQ ID NO: 24. The recombinant DNA sequence of the aspartate amino lyase is shown in SEQ ID NO: 23. The natural substrate of SEQ ID NO: 24 is aspartic acid, and acrylic acid is a non-natural substrate for SEQ ID NO: 24.

[0070] To further verify the effectiveness of the computational design method disclosed in this invention, in this embodiment, starting with SEQ ID NO: 24, the computational design algorithm disclosed in this invention was used to virtually screen and predict 10 pairs. Figure 10 The mutants shown exhibit catalytic activity in the reaction; further experimental verification revealed two previously undisclosed effective mutations, N326T and N326V, and mutants containing these two effective mutations, SEQ ID NO: 26 and SEQ ID NO: 28, were obtained. Figure 10 The reaction shown has a very good catalytic effect.

[0071] Calculation and design process of Example 2:

[0072] (1) Structure acquisition: The three-dimensional structure of the aspartate amino lyase with ID 3R6V was retrieved from the PDB database.

[0073] (2) Substrate docking analysis: The molecular structures of aspartic acid and β-alanine were docked with 3R6V, respectively, using the Yasara software package. The docking results were visualized and analyzed in Yasara. The results showed that sites Q142, T187, H188, M321, K324, N326, and L358 in 3R6V had a significant impact on the binding of β-alanine to the active site. Therefore, these amino acids were selected as mutation sites.

[0074] (3) Mutation stability calculation: Select appropriate amino acid types based on the mutated amino acids, as shown in Table 6. Then, use a Python script to generate the input file of the mutation combinations required by Rosetta, which has 16384 possible mutations. Use the Cartesian_ddg algorithm to calculate the free energy difference (ΔΔG) between each mutant structure and the wild-type enzyme structure.

[0075] Table 6 Types of Amino Acid Mutations

[0076]

[0077] The calculation results were then analyzed using the statistical analysis method described in [3b] above, and the results are shown in Table 7.

[0078] Table 7 Calculation and Analysis Results

[0079]

[0080] (4) Reaction barrier calculation: A total of 432 variants were obtained by combining the mutations from the previous step. The Cadee calculation process was used to perform the calculations as follows: Figure 11 As shown, the atomic number of acrylic acid is as follows: Figure 12 As shown.

[0081] Simulation setup: First, dissolve the simulation system in a solution centered on the C1 atoms of acrylic acid substrate, with a radius of [missing information]. The TIP3P model of water molecules in a spherical water droplet, where all atoms are in The simulation center is fully mobile, using 10 kcal / mol. -1 The harmonic suppressor is constrained in the simulation center located at 17 and All atoms between, and constrained using a harmonic force constant of 200 kcal. Atoms other than those in the solvent. The SHAKE algorithm is used to confine H atoms within the solvent. The cutoff value is used to calculate the nonbonding interactions between all atoms except those in the empirical valence bond region. For these atoms, all interactions are explicitly calculated up to the cutoff value. (That is, no cutoff value was applied). All long-distance static electricity exceeding this critical value was handled using the Localized Reactive Field (LRF) method.

[0082] From the above reaction barrier calculation results, the 10 mutants with the lowest reaction barrier were selected, as shown in Table 8.

[0083] Table 8

[0084] Calculate the predicted mutants Amino acid residue mutation relative to SEQ ID NO: 24 1 T187I; M321I; K324V; N326T 2 T187C; M321I; K324L; N326V 3 T187C; K324I; N326V 4 T187I; M321I; K324L; N326C 5 T187C; M321I; K324V; N326V 6 T187I; M321I; K324V; N326T; L358A 7 T187I; H188L; M321I; K324V; N326T 8 T187V; K324V; N326T; L358A 9 T187I; K324V; N326V 10 T187I; K324V; N326T

[0085] PDBID:3R6V aspartate amino acid lyase is one of the most widely studied enzymes in the industry. In this application, the applicant used the calculation method provided by this invention to predict its effect on... Figure 10 The results show that the computational method of the present invention not only obtained some published effective mutations (such as T187I and N326C disclosed in CN109385415A), but also obtained mutations (i.e. N326T and N326V) that have not been reported by other computer-aided mutation design methods or laboratory directed evolution techniques.

[0086] (5) Experimental verification:

[0087] (5.1) Two mutants containing N326T or N326V predicted in Table 8 were selected, and their amino acid and DNA sequence numbers are shown in Table 9. They were recombinantly expressed (refer to ChemCatChem 2014, 6, 965–968. doi:10.1002 / cctc.201300986), and their activity was verified using the following method: Wet cells expressing SEQ ID NO: 26 or SEQ ID NO: 28 were added to a reaction flask at a final concentration of 10 g / L, along with 300 g / L acrylic acid (adjusted to pH 9 with ammonia). The mixture was heated in a water bath at 40℃ with a stirring speed of 400 rpm for 24 h. The conversion rate was measured by HPLC and is shown in Table 9. The results showed that the mutants containing N326T or N326V, SEQ ID NO: 26 or SEQ ID NO: 28, were effective against recombinant bacterial infections. Figure 10 The catalytic activity of the reaction shown is also significantly improved compared to the disclosed enzymes, demonstrating that the calculation method provided by this invention is very effective.

[0088] Table 9

[0089]

[0090] The HPLC instrument used to detect the above reaction was a commercially available Agilent 1100 chromatograph equipped with an Agilent ZORBAX-NH2 column (4.6*150mm, 5μm). The parameters for detection were as follows: mobile phase: 50% potassium dihydrogen phosphate and 50% acetonitrile; flow rate: 1 mL / min; detection wavelength: 205 nm; column temperature: 40℃. The obtained chromatogram is shown below. Figure 13 As shown (the peak time for acrylic acid is 3.589 min, and the peak time for β-alanine is 4.051 min).

[0091] (5.2) SEQ ID NO: 26 The reaction of acrylic acid to β-alanine under different temperature and pH conditions

[0092] Reaction 5.2.1

[0093] Preparation of bacterial culture: Weigh 2.5 g of wet bacterial cells expressing SEQ ID NO: 26 into a 1 L three-necked flask, add 100 mL of pure water, and stir in a 60 °C water bath. Preparation of substrate solution: Add 25 g of acrylic acid to the flask, adjust the pH to 10 with ammonia, and dilute to 400 mL with water at 60 °C. Add the substrate solution to the reaction flask and mix with the bacterial culture. The reaction volume is 500 mL, the acrylic acid concentration is 50 g / L, the bacterial cell concentration is 5 g / L, the reaction temperature is set at 60 °C, and the mechanical stirring is 200 rpm. Samples are taken during the transformation process to monitor the reaction progress.

[0094] Reaction 5.2.2

[0095] Preparation of bacterial culture: Weigh 2.5 g of wet bacterial cells expressing SEQ ID NO: 26 into a 1 L three-necked flask, add 100 mL of pure water, and stir in a 50 °C water bath. Preparation of substrate solution: Add 25 g of acrylic acid to the flask, adjust the pH to 7 with ammonia, and dilute to 400 mL with water at 50 °C. Add the substrate solution to the reaction flask and mix with the bacterial culture. The reaction volume is 500 mL, the acrylic acid concentration is 50 g / L, the bacterial cell concentration is 5 g / L, the reaction temperature is set at 50 °C, and the mechanical stirring is 200 rpm. Samples are taken during the transformation process to monitor the reaction progress.

[0096] Reaction 5.2.3

[0097] Preparation of bacterial culture: Weigh 2.5 g of wet bacterial cells expressing SEQ ID NO: 26 into a 1 L three-necked flask, add 100 mL of pure water, and stir in a 30 °C water bath. Preparation of substrate solution: Add 25 g of acrylic acid to the flask, adjust the pH to 9 with ammonia, and dilute to 400 mL with water at 30 °C. Add the substrate solution to the reaction flask and mix with the bacterial culture. The reaction volume is 500 mL, the acrylic acid concentration is 50 g / L, the bacterial cell concentration is 5 g / L, the reaction temperature is set at 30 °C, and the mechanical stirring is 200 rpm. Samples are taken during the transformation process to monitor the reaction progress.

[0098] Reaction 5.2.4

[0099] Preparation of bacterial culture: Weigh 2.5 g of wet bacterial cells expressing SEQ ID NO: 26 into a 1 L three-necked flask, add 100 mL of pure water, and stir in a 45 °C water bath. Preparation of substrate solution: Add 25 g of acrylic acid to the flask, adjust the pH to 11 with ammonia, and dilute to 400 mL with water at 45 °C. Add the substrate solution to the reaction flask and mix with the bacterial culture. The reaction volume is 500 mL, the acrylic acid concentration is 50 g / L, the bacterial cell concentration is 5 g / L, the reaction temperature is set at 45 °C, and the mechanical stirring is 200 rpm. Samples are taken during the transformation process to monitor the reaction progress.

[0100] In the above experimental examples 5.2.1 to 5.2.4, samples were taken at different reaction time points (6 hours and 24 hours) during the reaction process, and the conversion rate at each reaction time point was detected by HPLC. The detection results are shown in Table 10.

[0101] Table 10

[0102] reaction time 6h 24h Conversion rate of reaction 5.2.1 in Experimental Example 70.5% 99.9% Conversion rate of reaction 5.2.2 in Experimental Example 64.3% 98.9% Conversion rate of reaction 5.2.3 in Experimental Example 69.1% 99.6% The conversion rate of reaction 5.2.4 in Experimental Example 67.9% 99.4%

[0103] The enzyme mutant provided by this invention not only improves the catalytic activity of ammonia addition to acrylic acid, but also extends the optimal temperature for catalytic substrate from 45℃-50℃ to 30℃-60℃ and the optimal pH range from 8.5-9.0 to 7-11.

[0104] (5.3) SEQ ID NO: 26 The reaction of acrylic acid to β-alanine under different enzyme dosages and / or substrate loadings.

[0105] Reaction 5.3.1

[0106] Preparation of bacterial culture: Weigh 7.5 g of wet bacterial cells expressing SEQ ID NO: 26 into a 1 L three-necked flask, add 100 mL of pure water, and stir in a 30 °C water bath. Preparation of substrate solution: Add 150 g of acrylic acid to the flask, adjust the pH to 9 with ammonia, and dilute to 400 mL with water at 30 °C. Add the substrate solution to the reaction flask and mix with the bacterial culture. The reaction volume is 500 mL, the acrylic acid concentration is 300 g / L, the bacterial cell concentration is 15 g / L, the reaction temperature is set at 30 °C, and the mechanical stirring is 200 rpm. Samples were taken and tested during the transformation process, and the results are shown in Table 11.

[0107] Table 11

[0108] reaction time 2h 4h 8h 20h 24h reaction conversion rate 44.0% 82.5% 99.1% 99.9% 99.9%

[0109] Reaction 5.3.2

[0110] Preparation of bacterial culture: Weigh 10g of wet bacterial cells expressing SEQ ID NO: 26 into a 1L three-necked flask, add 100mL of pure water, and stir in a 60℃ water bath. Preparation of substrate solution: Add 200g of acrylic acid to the flask, adjust the pH to 9 with ammonia, and dilute to 400mL with water at 60℃. Add the substrate solution to the reaction flask and mix with the bacterial culture. The reaction volume is 500mL, the acrylic acid concentration is 400g / L, the bacterial cell concentration is 20g / L, the reaction temperature is set at 60℃, and the mechanical stirring is 200rpm. Samples were taken and tested during the transformation process, and the results are shown in Table 12.

[0111] Table 12

[0112] reaction time 2h 4h 8h 20h 24h reaction conversion rate 31.0% 65.3% 88.6% 99.5% 99.9%

[0113] Reaction 5.3.3

[0114] Preparation of bacterial culture: Weigh 5g of wet bacterial cells expressing SEQ ID NO: 26 into a 1L three-necked flask, add 100mL of pure water, and stir in a 30℃ water bath. Preparation of substrate solution: Add 150g of acrylic acid to the flask, adjust the pH to 9 with ammonia, and dilute to 400mL with water at 30℃. Add the substrate solution to the reaction flask and mix with the bacterial culture. The reaction volume is 500mL, the acrylic acid concentration is 300g / L, the bacterial cell concentration is 10g / L, the reaction temperature is set at 30℃, and the mechanical stirring is 200rpm. Samples were taken and analyzed during the transformation process, and the results are shown in Table 13.

[0115] Table 13

[0116] reaction time 2h 4h 8h 20h 24h reaction conversion rate 39.5% 60.1% 87.6% 96.7% 99.9%

[0117] SEQ ID NO: 26 can tolerate high substrate concentrations up to 400 g / L, has a wide temperature and pH tolerance range, and high catalytic efficiency.

[0118] The reaction system is simple and suitable for industrial production.

[0119] It should be understood that after reading the above description of the present invention, those skilled in the art can make various alterations or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims. sequence list <110> Ningbo Enzyme Biotechnology Co., Ltd. <120> A computational method for designing protease mutants with non-natural substrate activity <130> CAD <160> 28 <170> SIPOSequenceListing 1.0 <210> 1 <211> 747 <212> DNA <213> Artificial Sequence <400> 1 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taaagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagca ttcatggtat cgtgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg ctacattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 2 <211> 249 <212> PRT <213> Artificial Sequence <400> 2 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Ile 130 135 140 His Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Tyr Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 3 <211> 747 <212> DNA <213> Artificial Sequence <400> 3 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagcg gccatggtat cgtgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg ctacattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 4 <211> 249 <212> PRT <213> Artificial Sequence <400> 4 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Gly 130 135 140 His Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Tyr Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 5 <211> 747 <212> DNA <213> Artificial Sequence <400> 5 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagca ttcatggtat cgtgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg cgcgattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 6 <211> 249 <212> PRT <213> Artificial Sequence <400> 6 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Ile 130 135 140 His Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Ala Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 7 <211> 747 <212> DNA <213> Artificial Sequence <400> 7 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taaagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagca ttcatggtat cgttgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg cggcattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 8 <211> 249 <212> PRT <213> Artificial Sequence <400> 8 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Ile 130 135 140 His Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Gly Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 9 <211> 747 <212> DNA <213> Artificial Sequence <400> 9 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taaagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagca tttgcggtat cgttgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg ctacattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 10 <211> 249 <212> PRT <213> Artificial Sequence <400> 10 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Ile 130 135 140 Cys Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Tyr Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 11 <211> 747 <212> DNA <213> Artificial Sequence <400> 11 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taaagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagca tttgcggtat cgttgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg cgcgattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 12 <211> 249 <212> PRT <213> Artificial Sequence <400> 12 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Ile 130 135 140 Cys Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Ala Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 13 <211> 747 <212> DNA <213> Artificial Sequence <400> 13 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taaagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagca tttgcggtat cgttgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg cggcattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 14 <211> 249 <212> PRT <213> Artificial Sequence <400> 14 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Ile 130 135 140 Cys Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Gly Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu With Lys Ser 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Glu Lys Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 15 <211> 747 <212> DNA <213> Artificial Sequence <400> 15 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc tAAagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gataagggta aaaaccgt tgagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagcg gctgcggtat cgtgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg ctacattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 16 <211> 249 <212> PRT <213> Artificial Sequence <400> 16 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Gly 130 135 140 Cys Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Tyr Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 17 <211> 747 <212> DNA <213> Artificial Sequence <400> 17 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagcg gctgcggtat cgtgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg cgcgattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 18 <211> 249 <212> PRT <213> Artificial Sequence <400> 18 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Gly 130 135 140 Cys Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Ala Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 19 <211> 747 <212> DNA <213> Artificial Sequence <400> 19 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagcg gctgcggtat cgtgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg cggcattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 20 <211> 249 <212> PRT <213> Artificial Sequence <400> 20 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Gly 130 135 140 Cys Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Gly Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 21 <211> 747 <212> DNA <213> Artificial Sequence <400> 21 atgggtatcc tggacaacaa agtcgcactg gttacgggcg ctggttcggg catcggtctg 60 gcggtggcac actcctacgc taagaaggc gctaaagtca ttgtgtcaga tatcaacgaa 120 gacaagggta ataaaaccgt tgaagatatt aaagcacagg gcggtgaagc tagttttgtg 180 aaagcggaca ccagcaaccc ggaagaagtg gaagccctgg ttaaacgtac ggtcgaaatt 240 tatggtcgcc tggatgtggc atgcaacaat gctggcattg cgggtgaaaa ggcactggct 300 ggtgattacg gcctggacag ctggcgtaaa gttctgtctg tgaatctgga cggtgtcttc 360 tatggctgta aatacgaact ggaacaaatg gagaaaaacg gcggtggcgt tatcgtcaat 420 atggccagcg gccatggtat cgtgcgcag ccgctgaact ctgcatatac ctctgcgaaa 480 cacgccgtgg ttggcctgac gaaaaacatt ggtgctgatt atggccagaa aaacatccgt 540 tgcaatgcgg tgtgcccggg cgcgattgaa accccgctgc tggaatcact gacgaaagaa 600 atgaaagaag ccctgatctc gaaacccg atgggtcgcc tgggcaaacc ggaagaagtg 660 gcagaactgg ttctgtttct gagttccgaa aaatcatcgt tcatgaccgg tggctattac 720 ctggtcgatg gtggctacac ggcagtg 747 <210> 22 <211> 249 <212> PRT <213> Artificial Sequence <400> 22 Met Gly Ile Leu Asp Asn Lys Val Ala Leu Val Thr Gly Ala Gly Ser 1 5 10 15 Gly Ile Gly Leu Ala Val Ala His Ser Tyr Ala Lys Glu Gly Ala Lys 20 25 30 Val Ile Val Ser Asp Ile Asn Glu Asp Lys Gly Asn Lys Thr Val Glu 35 40 45 Asp Ile Lys Ala Gln Gly Gly Glu Ala Ser Phe Val Lys Ala Asp Thr 50 55 60 Ser Asn Pro Glu Glu Val Glu Ala Leu Val Lys Arg Thr Val Glu Ile 65 70 75 80 Tyr Gly Arg Leu Asp Val Ala Cys Asn Asn Ala Gly Ile Ala Gly Glu 85 90 95 Lys Ala Leu Ala Gly Asp Tyr Gly Leu Asp Ser Trp Arg Lys Val Leu 100 105 110 Ser Val Asn Leu Asp Gly Val Phe Tyr Gly Cys Lys Tyr Glu Leu Glu 115 120 125 Gln Met Glu Lys Asn Gly Gly Gly Val Ile Val Asn Met Ala Ser Gly 130 135 140 His Gly Ile Val Ala Gln Pro Leu Asn Ser Ala Tyr Thr Ser Ala Lys 145 150 155 160 His Ala Val Val Gly Leu Thr Lys Asn Ile Gly Ala Asp Tyr Gly Gln 165 170 175 Lys Asn Ile Arg Cys Asn Ala Val Cys Pro Gly Ala Ile Glu Thr Pro 180 185 190 Leu Leu Glu Ser Leu Thr Lys Glu Met Lys Glu Ala Leu Ile Ser Lys 195 200 205 His Pro Met Gly Arg Leu Gly Lys Pro Glu Glu Val Ala Glu Leu Val 210 215 220 Leu Phe Leu Ser Ser Glu Lys Ser Ser Phe Met Thr Gly Gly Tyr Tyr 225 230 235 240 Leu Val Asp Gly Gly Tyr Thr Ala Val 245 <210> 23 <211> 1404 <212> DNA <213> Bacillus sp. YM55-1 <400> 23 atgaataccg atgttcgtat tgagaaagac ttttaggag aaaaggagat tccgaaagac 60 gcttattg gcgtacaaac aattcgggca acggaaaatt ttccaattac aggttatcgt 120 attcatccag aattaattaa atcactaggg attgtaaaaa aatcagccgc attagcaaac 180 atggaagttg gcttactcga taaagaagtt gggcaatata tcgtaaaagc tgctgacgaa 240 gtgattgaag gaaaatggaa tgatcaattt attgttgacc caattcaagg cggggcagga 300 acttccatta atatgaatgc aaatgaagtg attgctaacc gcgcattaga attaatggga 360 gaggaaaaag gaaactattc aaaatagt caaactccc atgtaaatat gtctcaatca 420 acaaacgatg cttccctac tgcaacgcat attgctgtgt taagtttatt aaatcaatta 480 attgaaacta caaaatacat gcaacaagaa ttcatgaaaa aagcagatga attcgctggc 540 gttattaaaa tgggaagaac gcacttgcaa gacgctgttc ctattttatt aggacaagag 600 tttgaagcat atgctcgtgt aattgcccgc gatattgaac gtattgccaa tacgagaaac 660 aatttatacg acatcaacat gggtgcaaca gcagtcggca ctggcttaaa tgcagatcct 720 gaatatataa gcatcgtaac agaacattta gcaaaattca gcggacatcc attaagaagt 780 gcacaacatt tagtggacgc aactcaaaat acagactgct atacagaagt ttcttctgca 840 ttaaaagttt gcatgatcaa catgtctaaa attgccaatg atttacgctt aatggcatct 900 ggaccacgcg caggcttatc agaaatcgtt cttcctgctc gacaacctgg atcttctatc 960 atgcctggta aagtgaatcc tgttatgcca gaagtgatga accaagtggc attccaagtg 1020 ttcggtaatg atttaacaat tacatctgct tctgaagcag gccaatttga attaaatgtg 1080 atggaacctg tgttattctt caatttaatt caatcgattt cgattatgac taatgtcttt 1140 aaatccttta cagaaaactg cttaaaaggt attaaggcaa atgaagaacg catgaaagaa 1200 tatgttgaga aaagcattgg atcattact gcaattacc cacatgtagg ctgaaca 1260 gctgcaaaat tagcacgtga agcatatctt acagggaat ccatccgtga actttgcatt 1320 aagtatggcg tattacaga agaacagtta atgaatct taaatccata tgaatgaca 1380 catccgggaa ttgctggaag aaaa 1404 <210> 24 <211> 468 <212> PRT <213> Bacillus sp. YM55-1 <400> 24 Met Asn Thr Asp Val Arg With Glu Lys Asp Phe Leu Gly Glu Lys Glu 1 5 10 15 Pro Lys Asp with Tyr Tyr Gly Val Gln and Arg with Thr Glu 20 25 30 Asn Phe Pro Ile Thr Gly Tyr Arg Ile His Pro Glu Leu Ile Lys Ser 35 40 45 Leu Gly Ile Val Lys Ser Ala Ala Leu Ala Asn Met Glu Val Gly 50 55 60 Leu Leu Asp Lys Glu Val Gly Gln Tyr Ile Val Lys Ala Ala Ala Asp Glu 65 70 75 80 Val Ile Glu Gly Lys Trp Asn Asp Gln Phe Ile Val Asp Pro Ile Gln 85 90 95 Gly Gly Ala Gly Thr Ser Ile Asn Met Asn Ala Asn Glu Val Ile Ala 100 105 110 Asn Arg Ala Leu Glu Leu Met Gly Glu Glu Lys Gly Asn Tyr Ser Lys 115 120 125 Ile Ser Pro Asn Ser His Val Asn Met Ser Gln Ser Thr Asn Asp Ala 130 135 140 Phe Pro Thr Ala Thr His Ile Ala Val Leu Ser Leu Leu Asn Gln Leu 145 150 155 160 Ile Glu Thr Thr Lys Tyr Met Gln Gln Glu Phe Met Lys Lys Ala Asp 165 170 175 Glu Phe Ala Gly Val Ile Lys Met Gly Arg Thr His Leu Gln Asp Ala 180 185 190 Val Pro Ile Leu Leu Gly Gln Glu Phe Glu Ala Tyr Ala Arg Val Ile 195 200 205 Ala Arg Asp Ile Glu Arg Ile Ala Asn Thr Arg Asn Asn Leu Tyr Asp 210 215 220 Ile Asn Met Gly Ala Thr Ala Val Gly Thr Gly Leu Asn Ala Asp Pro 225 230 235 240 Glu Tyr Ile Ser Ile Val Thr Glu His Leu Ala Lys Phe Ser Gly His 245 250 255 Pro Leu Arg Ser Ala Gln His Leu Val Asp Ala Thr Gln Asn Thr Asp 260 265 270 Cys Tyr Thr Glu Val Ser Ser Ala Leu Lys Val Cys Met Ile Asn Met 275 280 285 Ser Lys Ile Ala Asn Asp Leu Arg Leu Met Ala Ser Gly Pro Arg Ala 290 295 300 Gly Leu Ser Glu Ile Val Leu Pro Ala Arg Gln Pro Gly Ser Ser Ile 305 310 315 320 Met Pro Gly Lys Val Asn Pro Val Met Pro Glu Val Met Asn Gln Val 325 330 335 Ala Phe Gln Val Phe Gly Asn Asp Leu Thr Ile Thr Ser Ala Ser Glu 340 345 350 Ala Gly Gln Phe Glu Leu Asn Val Met Glu Pro Val Leu Phe Phe Asn 355 360 365 Leu Ile Gln Ser Ile Ser Ile Met Thr Asn Val Phe Lys Ser Phe Thr 370 375 380 Glu Asn Cys Leu Lys Gly Ile Lys Ala Asn Glu Glu Arg Met Lys Glu 385 390 395 400 Tyr Val Glu Lys Ser Ile Gly Ile Ile Thr Ala Ile Asn Pro His Val 405 410 415 Gly Tyr Glu Thr Ala Ala Lys Leu Ala Arg Glu Ala Tyr Leu Thr Gly 420 425 430 Glu Ser Ile Arg Glu Leu Cys Ile Lys Tyr Gly Val Leu Thr Glu Glu 435 440 445 Gln Leu Asn Glu Ile Leu Asn Pro Tyr Glu Met Thr His Pro Gly Ile 450 455 460 Ala Gly Arg Lys 465 <210> 25 <211> 1404 <212> DNA <213> Artificial Sequence <400> 25 atgaacaccg acgttcgtat cgaaaaggac ttcctgggcg aaaaagagat cccgaaagat 60 gcctattacg gcgtgcagac aatccgtgca accgagaact ttccgattac cggttatcgt 120 atccacccgg aactgattaa gagcctgggt attgtgaaga agagcgcagc actggcaaat 180 atggaagttg gcctgctgga taaggaggtt ggtcaatata tcgtgaaggc cgccgatgaa 240 gttatcgaag gtaaatggaa cgaccagttc atcgtggatc cgattcaggg tggtgcaggt 300 accagcatca atatgaatgc caacgaggtt attgcaaacc gcgcattaga gctgatgggc 360 gaggagaagg gtaattagat taagatcagc cctaacagcc acgtgaatat gagtcagagc 420 accaatgacg cattcccgac agcaacacat atcgctgttc tgagcctgct gaatcagctg 480 attgaaacca ccaagtacat gcagcaggag tttatgaaga aggcggatga attcgcgggc 540 gttattaaga tgggtcgtat ccatctgcag gatgctgttc ctattctgct gggtcaggaa 600 tttgaggcct atgctcgtgt gattgcccgt gacattgaac gtattgcaaa cacccgtaac 660 aacctgtacg acattaacat gggcgccacc gcagttggta caggtctgaa cgcagatcct 720 gaatatatca gcatcgtgac cgagcatctg gcaaaatttt ccggtcatcc gctgcgtagc 780 gctcagcacc tggtggacgc aacccagaac accgactgct acaccgaggt gagcagcgca 840 ctgaaggtgt gcatgattaa tatgagcaag atcgcaaacg acctgcgtct gatggcatca 900 ggtcctcgtg caggtttaag cgaaattgtt ctgccggccc gtcagcctgg tagcagcatc 960 atccctggtg tggtgacccc tgttatgccg gaagttatga atcaggtggc attccaggtg 1020 ttcggcaatg atctgaccat taccagtgca agcgaggctg gtcaatttga actgaatgtg 1080 atggagccgg tgctgttttt caacctgatt cagagcatca gcatcatgac caacgtgttt 1140 aagagcttca ccgagaactg cctgaagggt attaaggcaa acgaggaacg tatgaaggag 1200 tacgttgaga agagcatcgg tatcatcacc gcaattaacc cgcatgttgg ctatgaaacc 1260 gccgcaaaac tggctagaga agcatatctg accggtgaaa gtatccgtga actgtgcatt 1320 aagtacggcg tgctgaccga agaacagtta aatgaaatcc tgaaccctta cgagatgatc 1380 caccgggta ttgcaggtag aaaa 1404 <210> 26 <211> 468 <212> PRT <213> Artificial Sequence <400> 26 Met Asn Thr Asp Val Arg Ile Glu Lys Asp Phe Leu Gly Glu Lys Glu 1 5 10 15 Ile Pro Lys Asp Ala Tyr Tyr Gly Val Gln Thr Ile Arg Ala Thr Glu 20 25 30 Asn Phe Pro Ile Thr Gly Tyr Arg Ile His Pro Glu Leu Ile Lys Ser 35 40 45 Leu Gly Ile Val Lys Lys Ser Ala Ala Leu Ala Asn Met Glu Val Gly 50 55 60 Leu Leu Asp Lys Glu Val Gly Gln Tyr Ile Val Lys Ala Ala Asp Glu 65 70 75 80 Val Ile Glu Gly Lys Trp Asn Asp Gln Phe Ile Val Asp Pro Ile Gln 85 90 95 Gly Gly Ala Gly Thr Ser Ile Asn Met Asn Ala Asn Glu Val Ile Ala 100 105 110 Asn Arg Ala Leu Glu Leu Met Gly Glu Glu Lys Gly Asn Tyr Ser Lys 115 120 125 Ile Ser Pro Asn Ser His Val Asn Met Ser Gln Ser Thr Asn Asp Ala 130 135 140 Phe Pro Thr Ala Thr His Ile Ala Val Leu Ser Leu Leu Asn Gln Leu 145 150 155 160 Ile Glu Thr Thr Lys Tyr Met Gln Gln Glu Phe Met Lys Lys Ala Asp 165 170 175 Glu Phe Ala Gly Val Ile Lys Met Gly Arg Ile His Leu Gln Asp Ala 180 185 190 Val Pro Ile Leu Leu Gly Gln Glu Phe Glu Ala Tyr Ala Arg Val Ile 195 200 205 Ala Arg Asp Ile Glu Arg Ile Ala Asn Thr Arg Asn Asn Leu Tyr Asp 210 215 220 Ile Asn Met Gly Ala Thr Ala Val Gly Thr Gly Leu Asn Ala Asp Pro 225 230 235 240 Glu Tyr Ile Ser Ile Val Thr Glu His Leu Ala Lys Phe Ser Gly His 245 250 255 Pro Leu Arg Ser Ala Gln His Leu Val Asp Ala Thr Gln Asn Thr Asp 260 265 270 Cys Tyr Thr Glu Val Ser Ser Ala Leu Lys Val Cys Met Ile Asn Met 275 280 285 Ser Lys Ile Ala Asn Asp Leu Arg Leu Met Ala Ser Gly Pro Arg Ala 290 295 300 Gly Leu Ser Glu Ile Val Leu Pro Ala Arg Gln Pro Gly Ser Ser Ile 305 310 315 320 Ile Pro Gly Val Val Thr Pro Val Met Pro Glu Val Met Asn Gln Val 325 330 335 Ala Phe Gln Val Phe Gly Asn Asp Leu Thr Ile Thr Ser Ala Ser Glu 340 345 350 Ala Gly Gln Phe Glu Leu Asn Val Met Glu Pro Val Leu Phe Phe Asn 355 360 365 Leu Ile Gln Ser Ile Ser Ile Met Thr Asn Val Phe Lys Ser Phe Thr 370 375 380 Glu Asn Cys Leu Lys Gly Ile Lys Ala Asn Glu Glu Arg Met Lys Glu 385,390,395,400 Tyr Val Glu Lys Ser Ile Gly Ile Ile Thr Ala Ile Asn Pro His Val 405 410 415 Gly Tyr Glu Thr Ala Ala Lys Leu Ala Arg Glu Ala Tyr Leu Thr Gly 420 425 430 Glu Ser Ile Arg Glu Leu Cys Ile Lys Tyr Gly Val Leu Thr Glu Glu 435 440 445 Gln Leu Asn Glu Ile Leu Asn Pro Tyr Glu Met Ile His Pro Gly Ile 450 455 460 Ala Gly Arg Lys 465 <210> 27 <211> 1404 <212> DNA <213> Artificial Sequence <400> 27 atgaacaccg acgttcgtat cgaaaaggac ttcctgggcg aaaaagagat cccgaaagat 60 gcctattacg gcgtgcagac aatccgtgca accgagaact ttccgattac cggttatcgt 120 atccacccgg aactgattaa gagcctgggt attgtgaaga agagcgcagc actggcaaat 180 atggaagttg gcctgctgga taaggaggtt ggtcaatata tcgtgaaggc cgccgatgaa 240 gttatcgaag gtaaatggaa cgaccagttc atcgtggatc cgattcaggg tggtgcaggt 300 accagcatca atatgaatgc caacgaggtt attgcaaacc gcgcattaga gctgatgggc 360 gaggagaagg gtaattagat taagatcagc cctaacagcc acgtgaatat gagtcagagc 420 accaatgacg cattcccgac agcaacacat atcgctgttc tgagcctgct gaatcagctg 480 attgaaacca ccaagtacat gcagcaggag tttatgaaga aggcggatga attcgcgggc 540 gttattaaga tgggtcgttg ccatctgcag gatgctgttc ctattctgct gggtcaggaa 600 tttgaggcct atgctcgtgt gattgcccgt gacattgaac gtattgcaaa cacccgtaac 660 aacctgtacg acattaacat gggcgccacc gcagttggta caggtctgaa cgcagatcct 720 gaatatatca gcatcgtgac cgagcatctg gcaaaatttt ccggtcatcc gctgcgtagc 780 gctcagcacc tggtggacgc aacccagaac accgactgct acaccgaggt gagcagcgca 840 ctgaaggtgt gcatgattaa tatgagcaag atcgcaaacg acctgcgtct gatggcatca 900 ggtcctcgtg caggtttaag cgaaattgtt ctgccggccc gtcagcctgg tagcagcatc 960 atccctggtt tagtggtgcc tgttatgccg gaagttatga atcaggtggc attccaggtg 1020 ttcggcaatg atctgaccat taccagtgca agcgaggctg gtcaatttga actgaatgtg 1080 atggagccgg tgctgttttt caacctgatt cagagcatca gcatcatgac caacgtgttt 1140 aagagcttca ccgagaactg cctgaagggt attaaggcaa acgaggaacg tatgaaggag 1200 tacgttgaga agagcatcgg tatcatcacc gcaattaacc cgcatgttgg ctatgaaacc 1260 gccgcaaaac tggctagaga agcatatctg accggtgaaa gtatccgtga actgtgcatt 1320 aagtacggcg tgctgaccga agaacagtta aatgaaatcc tgaaccctta cgagatgatc 1380 caccgggta ttgcaggtag aaaa 1404 <210> 28 <211> 468 <212> PRT <213> Artificial Sequence <400> 28 Met Asn Thr Asp Val Arg Ile Glu Lys Asp Phe Leu Gly Glu Lys Glu 1 5 10 15 Ile Pro Lys Asp Ala Tyr Tyr Gly Val Gln Thr Ile Arg Ala Thr Glu 20 25 30 Asn Phe Pro Ile Thr Gly Tyr Arg Ile His Pro Glu Leu Ile Lys Ser 35 40 45 Leu Gly Ile Val Lys Lys Ser Ala Ala Leu Ala Asn Met Glu Val Gly 50 55 60 Leu Leu Asp Lys Glu Val Gly Gln Tyr Ile Val Lys Ala Ala Asp Glu 65 70 75 80 Val Ile Glu Gly Lys Trp Asn Asp Gln Phe Ile Val Asp Pro Ile Gln 85 90 95 Gly Gly Ala Gly Thr Ser Ile Asn Met Asn Ala Asn Glu Val Ile Ala 100 105 110 Asn Arg Ala Leu Glu Leu Met Gly Glu Glu Lys Gly Asn Tyr Ser Lys 115 120 125 Ile Ser Pro Asn Ser His Val Asn Met Ser Gln Ser Thr Asn Asp Ala 130 135 140 Phe Pro Thr Ala Thr His Ile Ala Val Leu Ser Leu Leu Asn Gln Leu 145 150 155 160 Ile Glu Thr Thr Lys Tyr Met Gln Gln Glu Phe Met Lys Lys Ala Asp 165 170 175 Glu Phe Ala Gly Val Ile Lys Met Gly Arg Cys His Leu Gln Asp Ala 180 185 190 Val Pro Ile Leu Leu Gly Gln Glu Phe Glu Ala Tyr Ala Arg Val Ile 195 200 205 Ala Arg Asp Ile Glu Arg Ile Ala Asn Thr Arg Asn Asn Leu Tyr Asp 210 215 220 Ile Asn Met Gly Ala Thr Ala Val Gly Thr Gly Leu Asn Ala Asp Pro 225 230 235 240 Glu Tyr Ile Ser Ile Val Thr Glu His Leu Ala Lys Phe Ser Gly His 245 250 255 Pro Leu Arg Ser Ala Gln His Leu Val Asp Ala Thr Gln Asn Thr Asp 260 265 270 Cys Tyr Thr Glu Val Ser Ser Ala Leu Lys Val Cys Met Ile Asn Met 275 280 285 Ser Lys Ile Ala Asn Asp Leu Arg Leu Met Ala Ser Gly Pro Arg Ala 290 295 300 Gly Leu Ser Glu Ile Val Leu Pro Ala Arg Gln Pro Gly Ser Ser Ile 305 310 315 320 Ile Pro Gly Leu Val Val Pro Val Met Pro Glu Val Met Asn Gln Val 325 330 335 Ala Phe Gln Val Phe Gly Asn Asp Leu Thr Ile Thr Ser Ala Ser Glu 340 345 350 Ala Gly Gln Phe Glu Leu Asn Val Met Glu Pro Val Leu Phe Phe Asn 355 360 365 Leu Ile Gln Ser Ile Ser Ile Met Thr Asn Val Phe Lys Ser Phe Thr 370 375 380 Glu Asn Cys Leu Lys Gly Ile Lys Ala Asn Glu Glu Arg Met Lys Glu 385 390 395 400 Tyr Val Glu Lys Ser Ile Gly Ile Ile Thr Ala Ile Asn Pro His Val 405 410 415 Gly Tyr Glu Thr Ala Ala Lys Leu Ala Arg Glu Ala Tyr Leu Thr Gly 420 425 430 Glu Ser Ile Arg Glu Leu Cys Ile Lys Tyr Gly Val Leu Thr Glu Glu 435 440 445 Gln Leu Asn Glu Ile Leu Asn Pro Tyr Glu Met Ile His Pro Gly Ile 450 455 460 Ala Gly Arg Lys 465

Claims

1. A computational method for designing protease mutants with non-natural substrate activity, the design process of which is as follows: First, obtain a three-dimensional structural model of the target protease; then, perform docking and analysis of the protein structure and substrate molecule to determine the amino acid sites to be mutated; for each amino acid site to be mutated, specify replaceable amino acid residue mutations, exhaustively list all mutation combinations for all sites to be mutated, and calculate the stability of each mutant using a protein stability calculation method; process the stability calculation results and select stable amino acid residue mutation combinations to obtain the predicted stable mutants; finally, perform reaction barrier calculations on the obtained stable mutants for the target reaction, select mutants with lower reaction barriers, and obtain the predicted catalytically active mutants; The method for determining the amino acid site to be mutated is as follows: select a suitable mutation site from the active sites of the protein structure based on the different conformations of the docking results of the natural substrate or the target substrate; The results of the stability calculations were processed using statistical methods to obtain the predicted stable mutants: The free energy difference ΔΔG calculated for stability is sorted from low to high, and the mutants at the top and bottom of the sort are selected for mutation frequency analysis. For a specific amino acid site, the amino acid residue types that appear more frequently in the stable mutants are subtracted from the amino acid residue types that appear more frequently in the unstable mutants to obtain the set of amino acid residue types that can be mutated at that site. Finally, the mutable amino acid residues at each site are arranged and combined to obtain the stable mutants predicted by computer virtual screening.

2. The calculation method according to claim 1, characterized in that, When obtaining the three-dimensional structure of a protease, the three-dimensional structure of the protease is obtained from a structural model in a protein data bank database or a protein structural model obtained using computer modeling methods.

3. The calculation method according to claim 1, characterized in that, The software tools used for docking substrate molecules and proteases are Yasara, Discovery Studio, or Rosetta.

4. The calculation method according to claim 1, characterized in that, The algorithms used to calculate the structural stability of protein mutants are ddg_monomer, Cartesian_ddg, FoldX, Provean, ELASPIC, or Amber TI.

5. The calculation method according to claim 1, characterized in that, In "reaction barrier calculation", "reaction barrier" is defined as the difference between the potential energy of the enzyme and substrate in their free state and the potential energy of the activated molecule formed by their combination.

6. The amino-lysin mutant polypeptide obtained by the calculation method according to claim 1, wherein the polypeptide, under suitable reaction conditions, is capable of catalyzing the synthesis of β-alanine from acrylic acid substrate with better activity and / or stability than SEQ ID NO: 24, wherein the amino acid sequence of the polypeptide contains residue differences X326T or X326V compared to the sequence of SEQ ID NO:

24.

7. The polypeptide according to claim 6, wherein the amino acid sequence of the polypeptide further comprises residue differences X187I or X187C, X321I, X324L or X324V compared with the sequence of SEQ ID NO:

24.

8. The polypeptide according to claim 6 or 7, wherein the amino acid sequence of the polypeptide is SEQ ID NO: 26 and SEQ ID NO:

28.

9. The polypeptide according to claim 6 or 7, wherein the temperature range for catalyzing the synthesis of β-alanine from acrylic acid substrate is 30-60°C, and the pH range is 7-11.

Citation Information

Patent Citations

  • Aspartase variant, preparation method and applications thereof

    CN109385415A

  • Engineered ketoreductase polypeptide and application thereof

    CN111321129A

  • Molecular modification method of efficient aromatic hydrocarbon dioxygenase

    CN112391361A

  • Method for improving hydrolase robustness in combination with high-pressure molecular dynamics simulation and free energy calculation

    CN112582031A