A method for designing and modifying new thermostable enzymes based on a deep learning model
By constructing a protein sequence generation model and a data cycle feedback optimization of the thermal stability protein discrimination model, the problem of low thermal stability transformation efficiency in the existing technology is solved, and efficient generation and screening of enzyme sequences with thermal stability is achieved, and the enzyme transformation efficiency is improved.
Patent Information
- Application Number
- CN202310369852.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-07
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-04-07
AI Technical Summary
The prior art has problems of low screening efficiency and high workload in enhancing enzyme thermal stability. Traditional methods require the construction of a large number of mutants and the accuracy is not ideal, making it difficult to efficiently generate enzyme sequences with specified properties through data-driven methods.
A protein sequence generation model and thermally stable protein discriminant model were constructed, and the enzyme sequence with thermal stability was directly generated through data loop feedback optimization, combining deep learning methods to expand the sequence space of enzymes and reduce the experimental screening workload.
Efficient generation and screening of enzyme sequences with thermally stable properties is achieved, significantly reducing the number of experimentally detected candidate sequences and improving the efficiency of enzyme transformation.
Smart Images

Figure CN118782150B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of molecular biology and protein engineering, and specifically relates to a method for designing and modifying new thermostable enzymes based on a deep learning model for thermostability discrimination and protein sequence generation. The method extracts thermostable protein features through a deep learning model and directly uses the protein generation model to generate a new enzyme amino acid sequence with thermostable properties. Background Art
[0002] Enzymes, key to modern bioengineering, are widely used in various fields, including food, medicine, and chemistry. However, most natural enzymes are easily inactivated at high temperatures, limiting their use in industrial applications. Therefore, enhancing enzyme thermostability has become a common goal in enzyme engineering. Currently, directed evolution remains the most common approach to improving enzyme thermostability. While this approach has achieved numerous successes over the years, it still suffers from low screening efficiency and a high workload, often requiring the construction of thousands of mutants to identify a suitable sequence. In recent years, various rational design approaches have been developed to enhance enzyme thermostability. These methods include adjusting amino acid preferences, truncating flexible regions of the target protein, and increasing intra-enzyme interactions such as disulfide bonds and salt bridges. Most of these approaches require a deeper understanding of the target enzyme's structure and catalytic mechanism to identify key sites, which limits their application. Computational design can facilitate the prediction of mutations that improve thermostability, but the accuracy of currently used computational tools remains suboptimal.
[0003] With the increase in the amount of biological data, machine learning methods for improving protein thermal stability have developed rapidly. These data-driven methods aim to establish a connection between protein features (such as sequence and / or structure) and stability-related measurements (such as Tm), thereby improving the thermal stability of enzymes. For example, the DeepDDG method uses a neural network to predict the effect of point mutations on protein stability. However, this method does not consider the wider range of interactions between amino acids and has the problem of less training data, which limits their application in enzyme engineering. The DeepET method uses large-scale genomic data to predict the optimal temperature of proteins, but is limited by the accuracy of the data itself, making it difficult to achieve high-accuracy protein property predictions and not easy to apply to enzyme engineering.
[0004] In recent years, protein language models (such as ESM) have developed rapidly, becoming an effective means of representing protein sequences. Protein sequence generation models (such as ProteinGAN) have also rapidly developed, helping to expand the protein sequence space and overcome the difficulty of combining point mutations. However, current protein sequence generation models primarily focus on generating functional protein sequences and are unable to generate sequences with specified properties (such as heat resistance or acid resistance).
[0005] Therefore, there is an urgent need in this field to develop a deep learning-based method that uses a protein language model to extract thermal stability information from a large number of protein sequences, and directly generates thermally stable enzyme sequences based on a protein sequence generation model, thereby expanding the sequence space of enzymes and reducing the workload of experimental screening. Summary of the Invention
[0006] In order to solve the problems existing in the above-mentioned prior art, the present invention provides a method for designing and modifying new thermostable enzymes based on a deep learning model to make up for the shortcomings of existing directed evolution and rational design methods. The present invention constructs a sequence generation model and a thermostable protein discrimination model, and performs data loop feedback between the two, and finally optimizes to obtain a deep learning model that can directly generate protein sequences with thermostable properties. The present invention utilizes a protein sequence generation model to overcome the limitation that traditional machine learning methods can only predict the effects of single-point mutations on protein stability, and utilizes a thermostable protein discrimination model for data optimization to overcome the limitation that traditional sequence generation models cannot generate sequences according to specified properties. At the same time, the present invention can efficiently generate and screen protein sequences with thermostable properties, significantly reducing the number of candidate sequences for experimental detection compared to the directed evolution method, and improving the efficiency of enzyme modification.
[0007] To achieve the above objectives, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a method for designing and modifying a new thermostable enzyme based on a deep learning model, comprising the following steps:
[0009] 1) Collect marker-related sequences;
[0010] 2) Construct a protein sequence generation model;
[0011] 3) Construct a thermostable protein discrimination model;
[0012] 4) Using the thermal stability protein discrimination model to tune the protein sequence generation model;
[0013] 5) Computational screening to generate sequences;
[0014] 6) Experimental verification of the designed sequences that were identified as thermostable.
[0015] Preferably, step 1) specifically comprises: collecting the optimal growth temperature information OGT of different species in a database, and defining the organisms as high-temperature species HTO, medium-temperature species MTO, and low-temperature species LTO according to their optimal growth temperatures; extracting the gene sequences corresponding to the species in the gene dataset according to the previously marked species information, and marking them as high-temperature sequences, medium-temperature sequences, and low-temperature sequences, respectively; collecting enzyme sequences with desired functions, and obtaining potential sequences of enzymes with desired functions for constructing a protein sequence generation model;
[0016] Preferably, in step 1), organisms with OGT>=50°C, 30°C≤OGT<50°C and OGT<30°C are defined as high-temperature species HTO, medium-temperature species MTO and low-temperature species LTO respectively;
[0017] Further preferably, the step 1) collects the optimal growth temperature information OGT of different species from the TEMPURA, ExProtDB, NCBI, and BacDive databases;
[0018] More preferably, in step 1), the gene sequence corresponding to the species is extracted from the unipro and uniref gene datasets according to the species information marked above.
[0019] Preferably, step 2) specifically comprises: filtering the relevant sequences of potential enzymes with desired functions obtained in step 1) according to sequence length, building a generative adversarial network or diffusion network model using a deep learning framework, expanding the sequence space of the target enzyme, and evaluating the quality of the generative model. When the model is stable, the generative network can generate new protein sequences with a distribution similar to that of natural sequences;
[0020] Further preferably, in step 2), a generative adversarial network or diffusion network model is constructed using a deep learning framework such as tensorflow or pytorch, and the quality of the generative model is evaluated using methods such as tSNE and Shannon entropy calculation.
[0021] Preferably, the step 3) specifically comprises: building a classification model using the high temperature sequence, medium temperature sequence and low temperature sequence obtained in step 1), and performing model training to determine whether the input sequence has thermal stability characteristics;
[0022] Further preferably, in step 3), the high-temperature sequence, medium-temperature sequence and low-temperature sequence obtained in step 1) are encoded using a protein language model, randomly divided into a training set and a test set, a classification model is built using a deep learning framework tensorflow or pytorch, and the model is trained using the training set to determine whether the input sequence has a thermostable feature. When the model is stable, the model quality is judged using the accuracy of the model on the test set and the recall rate of the high-temperature sequence.
[0023] Preferably, step 4) specifically comprises: encoding the new protein sequence generated in step 2) with a protein language model and inputting the encoded data into the thermostable protein discrimination model constructed in step 3); screening the new protein sequences identified as high-temperature sequences as new training sequences, and optimizing the protein sequence generation model obtained in step 2) until the optimized protein sequence generation model can directly generate a large proportion of new sequences that can be identified as thermostable sequences.
[0024] Preferably, step 5) is specifically as follows: transferring the sequence generated by the final protein sequence generation model into the thermal stability protein discrimination model for thermal stability evaluation, performing similarity comparison and important functional domain comparison with the natural sequence to evaluate the foldability of the sequence, and using alphafold for structural modeling to evaluate the quality of the protein generation sequence, and screening an appropriate number of sequences for experimental verification according to the throughput of the experimental detection.
[0025] Preferably, step 6) specifically comprises: synthesizing the screened designed sequence using DNA synthesis technology, transferring it into cells and measuring the expression of the designed sequence; purifying the designed protein and measuring the enzyme activity at different temperatures to verify whether the designed sequence has thermal stability.
[0026] Further preferably, the enzyme with the desired function in step 1) is glyceraldehyde-3-phosphate dehydrogenase.
[0027] More preferably, the amino acid sequence of the thermostable new enzyme obtained by verification in step 6) is as shown in any one of SEQ ID NOs. 1-6, and the thermostability refers to the enzyme activity at 65°C.
[0028] In a second aspect, the present invention provides a glyceraldehyde-3-phosphate dehydrogenase with thermal stability, wherein the amino acid sequence of the glyceraldehyde-3-phosphate dehydrogenase is shown in any one of SEQ ID NOs. 1-6, and the thermal stability refers to the enzyme activity at 65°C.
[0029] Beneficial effects of the present invention:
[0030] The present invention constructs a protein sequence generation model and a thermostable protein discrimination model, with data loop feedback between the two, enabling the direct generation of new thermostable enzyme sequences. The method for designing and modifying new thermostable enzymes provided by this invention expands the enzyme sequence space and is unbiased in terms of training data. Results demonstrate that thermostable enzyme variants can be effectively obtained, demonstrating promising application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Flowchart of the present invention;
[0032] Figure 2 Generate sequence quality assessment schematics for the present invention;
[0033] Among them: orange points represent the distribution of natural sequences in the tSNE two-dimensional projection space, and purple points represent the distribution of generated sequences in the same space;
[0034] Figure 3 Activity assay data of the new thermostable enzyme designed for the present invention.
[0035] Wherein: G3PDH is the natural G3PDH enzyme in Escherichia coli, B1, C1, E, F, M, and D1 are the G3PDH enzymes with thermostable properties designed in Example 1 of the present invention, G1, G2, and G3 are the G3PDH enzymes designed in the comparative example, B1C1E_nat is the natural G3PDH enzyme closest to B1, C1, and E, and F_nat, M_nat, D1_nat, G1_nat, G2_nat, and G3_nat are the natural G3PDH enzymes closest to F, M, D1, G1, G2, and G3, respectively. Blue indicates enzyme activity at 30 degrees Celsius, and red indicates enzyme activity at 65 degrees Celsius. DETAILED DESCRIPTION
[0036] The technical solutions of the present invention will be described in further detail below with reference to specific embodiments. It should be understood that the following embodiments are merely illustrative and explanations of the present invention and should not be construed as limiting the scope of protection of the present invention. All technologies implemented based on the above content of the present invention are encompassed within the scope of protection that the present invention is intended to protect.
[0037] Unless defined otherwise or clearly indicated by the context, all technical and scientific terms in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0038] Unless otherwise specified, the materials and reagents in the following examples or comparative examples are commercially available or can be prepared by known methods. Experimental methods in the following examples, where specific conditions are not specified, are generally performed according to the conditions described in the General Conditions or the conditions recommended by the manufacturer.
[0039] The present invention uses functional sequences with similar catalytic domains to construct a protein sequence generation model to expand the sequence space; by collecting the optimal growth temperatures of different species from different databases and further extracting the genomic sequences of the corresponding species, a thermostable protein sequence discrimination model is constructed; the sequence information feedback loop between the two is used to directly generate new thermostable enzyme sequences, and through computational filtering and screening, a designed sequence that can be verified experimentally is obtained. Finally, new enzymes with thermostable properties are detected through experimental means. In the comparative example, we only constructed a protein sequence generation model to generate new functional enzyme sequences, but because it was not optimized by the thermostable protein discrimination model, the generated sequence had poor activity at high temperatures, demonstrating that the present invention has the ability to generate thermostable protein sequences through data cycle optimization of the two models.
[0040] Example 1: Design of a G3PDH enzyme sequence with thermostable properties based on a deep learning model
[0041] Step 1: Collect marker-related sequences:
[0042] (1) Collecting sequences tagged with temperature information: The present invention collected the optimal growth temperature information of different species from four different sources. The first source is the TEMPURA database, which collects the growth temperatures of common and rare prokaryotes. The present invention obtained about 8,000 organisms and their optimal growth temperatures (OGT) from it; the second source is the ExProtDB database, which collects extremophilic proteins and their host organisms. The present invention collected about 300 thermophilic organisms from it; the third source is the NCBI database. The names of all microorganisms with sequenced genomes were searched from the NCBI database, and then these names were searched in NCBI and Wikipedia. If the webpage contains keywords such as "extreme thermophile", "thermophile", "high temperature", "thermoacidic" or "multiple extremophiles", the microorganisms were checked to see if they are thermophiles. In this way, about 500 thermophilic organisms were collected. The fourth source is the BacDive database, which covers relevant information on bacterial and archaeal biodiversity. About 5,000 microorganisms with growth temperature information were collected from this database. After removing duplicate species, the present invention collected a total of 10,189 species, some of which did not have OGT and were searched separately on the website. 805 organisms with OGT ≥ 50°C, 5122 organisms with 30°C ≤ OGT < 50°C, and 4262 organisms with OGT < 30°C were defined as high-temperature species (HTO), medium-temperature species (MTO), and low-temperature species (LTO), respectively. Subsequently, the corresponding genes were obtained from three downloaded gene sets (UniProt reference proteome, UniRef90, and UniRef50). The obtained genes were further divided into HTO, MTO, and LTO genes. In the UniProt reference proteome, a total of 25,724,264 genes were obtained, including 1,393,345 HTO genes, 12,317,734 MTO genes, and 12,013,185 LTO genes. A total of 15,901,817 genes were obtained from UniRef90, including 973,655 HTO genes, 7,941,331 MTO genes, and 6,986,831 LTO genes. A total of 2,199,998 genes were obtained from UniRef50, including 165,625 HTO genes, 1,120,580 MTO genes, and 913,793 LTO genes. These genes were considered as potential training sets for constructing thermostable protein discrimination models.
[0043] (2) Collecting enzyme / protein sequences with potential specific functions: In this example, the glyceraldehyde-3-phosphate dehydrogenase (G3PDH) gene was collected as an example of functional enzyme generation. The non-redundant database was searched using the Pfam domain ID PF02800, and 67,493 genes with this domain were obtained. Secondly, all G3PDH genes were retrieved from the KEGG database. Then, by using the G3PDH gene in NCBI as the query sequence, a local blastp search was performed in KEGG to screen out potential G3PDH genes with the best match, a similarity greater than 40%, and an alignment length greater than 200, and a total of 54,896 were screened out. After filtering out genes that were too long and too short, 40,000 potential G3PDH genes were finally selected to construct a protein sequence generation model.
[0044] Step 2: Constructing a protein sequence generation model: In this example, a generative adversarial network (GAN) was used to construct a protein sequence generation model for the G3PDH gene. Specifically, the 40,000 potential G3PDH genes obtained in Step 1 were length-filtered and similarity-screened. Sequences with a length of less than 400 and a pairwise similarity between 30% and 90% were selected, resulting in 15,454 genes for constructing the generative model. These 15,454 gene sequences were then randomly partitioned into training and test sets, and a GAN model was constructed based on the ProteinGAN model. The model used ResNet blocks as the discriminator and generator networks. The random vector input to the generator was 128 dimensions, and the output was a 512x21 matrix corresponding to a one-hot encoded sequence of length 512, containing 20 amino acids and a gap at the beginning or end of the sequence. During training, the generator generated genes in groups of 64 sequences, mixed them with 64 natural gene sequences from the training set using a weighted sampling method, and then passed them to the discriminator for discrimination. In particular, when there is a target sequence, the weights of sequences similar to the target sequence can be increased to make the generated sequence closer to the target sequence. The non-saturated loss function and R1 regularization are used for optimization, and the Adam algorithm is selected for network optimization. Finally, after 150,000 training steps, the model converged. There is no significant difference in the discriminator scores between the generated sequence and the test set sequence. At this time, tSNE analysis is performed on the generated sequence, as shown in the figure below. Figure 2 As shown, the generated sequences and natural sequences have similar distributions, and the generated sequences expand the distribution space of natural sequences.
[0045] Step 3: Construct a thermostable protein discrimination model. The HTO and LTO genes obtained in Step 1 were converted into protein sequences. Sequences were filtered by length, and 30,968 high-temperature protein sequences and 162,891 low-temperature protein sequences with lengths between 300 and 600 amino acids were selected as the dataset. These sequences were then divided into training and test sets. Next, the training data was preprocessed using the protein language model ESM 1b, encoding each sequence into a 1280-dimensional vector. These encoded vectors were then used to train a three-layer neural network classifier using the PyTorch framework with a dimensionality of 1280:64:16, using a binary cross-entropy loss function. Training was performed for 1,000 epochs to stabilize the loss. After the model stabilized, it achieved an overall accuracy of 95% on the test set, with a recall rate of 78.0% for the high-temperature sequences.
[0046] Step 4: Use the thermostable protein discrimination model to tune the sequence generation model. This embodiment combines the GAN model with the thermostable discrimination model and uses a feedback data loop method to enhance sampling in the G3PDH sequence space with thermostable characteristics. In this embodiment, 100,000 G3PDH sequences are first generated, and the top 20,000 sequences scored by the GAN discriminator are selected for G3PDH conserved domain screening. 18,238 sequences with functionally conserved residues are selected for thermostability classification, and 1,354 high-temperature stable protein sequences are screened and added to the training set of the GAN model for fine-tuning. After retraining, the proportion of protein generated sequences that can be classified as thermostable protein sequences increased from 7.4% to 14.9%. At the same time, the overall distribution of the newly generated sequences is similar to that of the original generated sequences.
[0047] Step 5: Screening Protein Generator Sequences. In this example, 100,000 new G3PDH sequences were generated using the optimized protein sequence generation model. The top 20,000 sequences were screened for G3PDH conserved domains using the GAN discriminator, resulting in approximately 16,000 designed sequences. After filtering using the thermostable protein discrimination model, 30 designed sequences were selected based on their similarity to the closest native sequence and their structure prediction scores using AlphaFold2. These sequences had similarities between 70% and 95% to the closest native sequence and were used for experimental validation.
[0048] Step 6: Experimental verification of the designed highly thermostable sequence. The gene sequence encoding the protein was synthesized and cloned into the pET28a expression vector. After sequence verification, the constructed vector was transformed into BL21 (DE3) Escherichia coli. The aforementioned E. coli was inoculated into 2YT medium containing 50 μg / mL kanamycin at a ratio of 1:160 and grown at 37°C and 220 rpm. When the OD600 of the cells reached 0.4-0.6, IPTG was added to a final concentration of 0.5 mM to induce expression. The strain was cultured overnight at 16°C and 220 rpm and then harvested by centrifugation. The cells were resuspended in 50 mM Tris-HCl (pH 6.8) lysis buffer and lysed using a high-pressure homogenizer for 2-3 cycles at 1200-1500 bar. Cell debris was removed by centrifugation (10,000 g, 40 min), and the Ni-NTA agarose column was equilibrated with ddH2O and lysis buffer for 2 column volumes. The supernatant was applied to the column and the protein was eluted using a gradient elution buffer (50 mM Tris-HCl containing 10 mM, 50 mM, and 200 mM imidazole, respectively). The separated protein was analyzed by SDS-PAGE and then concentrated by centrifugation (4,000 g, 30 min) in a 10 kDa ultrafiltration tube (Centriplus YM series, Millipore), and finally quickly frozen in liquid nitrogen and stored at -80 ° C.
[0049] The activity of G3PDH was detected by measuring the formation of NADH. Three parallel samples of the purified protein were mixed with 10mM NAD in 993μL reaction solution (40mM triethanolamine, 50mM Na2HPO4, 5mM EDTA, 0.1mM DTT, pH8.6). 7μL of DL-G3PDH solution (Sigma) was added to the above system, and A340 was immediately measured. The reaction system was incubated at 30°C for 10 minutes, and then A340 was measured again. The activity of G3PDH was calculated as unit / mg / min = ΔA340 x VT (volume of the tube) / (6.22 x concentration (mg) x time (seconds)). For the determination of thermal stability, 100μL of the reaction system was placed in a 96-well plate. The plate was incubated at the design temperature of the thermostatic microwell reader for 30 minutes, and A340 was continuously read. Of the 30 designed G3PDHs tested in this example, 23 proteins were correctly expressed and purified. According to the results of the G3PDH activity test, 17 of the proteins showed normal G3PDH catalytic activity at 30°C. Next, the thermal stability of these 17 proteins at 65°C was measured. Among them, five proteins, B1, C1, E, F, and M, showed significantly increased enzyme activity at high temperatures compared to their closest natural sequence proteins, while one protein, D1, had a slight increase in enzyme activity compared to the natural sequence protein ( Figure 3 These novel enzymes differ from the natural sequences by approximately 20-50 amino acid residues (Table 1). Thus, this invention demonstrates that novel enzymes with high thermostability can be effectively obtained.
[0050] Table 1
[0051]
[0052] Comparative Example 1: Generate G3PDH enzyme sequence using conventional protein sequence generation model
[0053] Step 1: Collect enzyme / protein sequences with potential specific functions: In this comparative example, the glyceraldehyde-3-phosphate dehydrogenase (G3PDH) gene was collected as an example of functional enzyme generation. According to (2) in step 1 of Example 1, the Pfam domain IDPF02800 was used to search the non-redundant database, and 67,493 genes with this domain were obtained. Secondly, all G3PDH genes were retrieved from the KEGG database. Then, by using the G3PDH gene in NCBI as the query sequence, a local blastp search was performed in KEGG to screen out potential G3PDH genes with the best match, a similarity greater than 40% and an alignment length greater than 200, and a total of 54,896 were screened out. After filtering out genes that were too long and too short, 40,000 potential G3PDH genes were finally selected to construct a generation model.
[0054] Step 2: Constructing a protein sequence generation model. In this comparative example, a generative model for the G3PDH gene was constructed based on a generative adversarial network (GAN). Specifically, the 40,000 potential G3PDH sequences obtained in Step 1 were length- and similarity-filtered. Sequences with lengths less than 400 and pairwise similarities between 30% and 90% were selected, resulting in 15,454 genes for constructing the generative model. These 15,454 gene sequences were then randomly partitioned into training and test sets. A GAN model was then constructed using the ProteinGAN model. The model used ResNet blocks as the discriminator and generator networks. The generator input was a 128-dimensional random vector, and the output was a 512 x 21 matrix corresponding to a one-hot encoded sequence of length 512, containing 20 amino acids and a gap at the beginning or end of the sequence. During training, the generator generated genes in groups of 64 sequences. These were mixed with 64 natural gene sequences from the training set using a weighted sampling method, and then passed to the discriminator for discrimination. In particular, when a target sequence is present, the weights of sequences similar to the target sequence can be increased to make the generated sequence closer to the target sequence. A non-saturating loss function and R1 regularization were used for optimization, and the Adam algorithm was selected for network optimization. The model converged after 150,000 training steps. There was no significant difference in the discriminator scores between the generated and test set sequences.
[0055] Step 3: Screening and Generating Sequences. In this comparative example, 100,000 new G3PDH sequences were generated using the protein sequence generation model. The top 20,000 sequences were screened for conserved G3PDH domains using the GAN discriminator, resulting in approximately 16,000 designed sequences. Based on their similarity to the closest native sequence and their structure prediction scores using AlphaFold2, six designed sequences were selected for experimental validation. These sequences had similarities between 70% and 95% with the closest native sequence.
[0056] Step 4:
[0057] The designed highly thermostable sequence was experimentally verified. The protein-encoding DNA gene sequence was synthesized and cloned into the pET28a expression vector. After sequence verification, the constructed vector was transformed into BL21 (DE3) Escherichia coli. The aforementioned E. coli cells were inoculated into 2YT medium containing 50 μg / mL kanamycin at a ratio of 1:160 and grown at 37°C and 220 rpm. When the OD600 of the cells reached 0.4-0.6, IPTG was added to a final concentration of 0.5 mM to induce expression. The strain was cultured overnight at 16°C and 220 rpm and then harvested by centrifugation. The cells were resuspended in 50 mM Tris-HCl (pH 6.8) lysis buffer and lysed using a high-pressure homogenizer for 2-3 times at 1200-1500 bar. Cell debris was removed by centrifugation (10,000 × g, 40 min), and the Ni-NTA agarose column was equilibrated with ddH2O and lysis buffer for 2 column volumes. The supernatant was applied to the column and the protein was eluted using a gradient elution buffer (50 mM Tris-HCl containing 10 mM, 50 mM, and 200 mM imidazole, respectively). The separated protein was analyzed by SDS-PAGE and then concentrated by centrifugation (4,000 g, 30 min) in a 10 kDa ultrafiltration tube (Centriplus YM series, Millipore), and finally quickly frozen in liquid nitrogen and stored at -80 ° C.
[0058] G3PDH activity was determined by measuring NADH formation. Three replicates of the purified protein were mixed with 10 mM NAD in 993 μL of a reaction solution (40 mM triethanolamine, 50 mM Na2HPO4, 5 mM EDTA, 0.1 mM DTT, pH 8.6). 7 μL of DL-G3PDH solution (Sigma) was added to the reaction mixture, and the A340 reading was immediately measured. The reaction mixture was incubated at 30°C for 10 minutes, and the A340 reading was repeated. G3PDH activity was calculated as units / mg / min = ΔA340 x VT (tube volume) / (6.22 x concentration (mg) x time (seconds)). For thermal stability measurements, 100 μL of the reaction mixture was placed in a 96-well plate. The plate was incubated at the design temperature of a thermostatted microplate reader for 30 minutes, and the A340 reading was continuously taken. In this control example, among the 6 designed G3PDHs tested, 3 proteins were correctly expressed and purified, namely G1, G2, and G3, and their amino acid sequences are shown in SEQ ID NO.7-9, respectively. According to the results of the G3PDH activity test, these 3 proteins showed normal G3PDH catalytic activity. Next, the thermal stability of these 3 proteins at 65°C was measured, among which G2 and G3 did not show activity at high temperature, and only G1 showed weak high temperature activity. This comparative example shows that it is difficult to obtain a new enzyme with thermostable properties only through the existing protein sequence generation model. The thermostable protein discrimination model in the present invention and the data loop feedback optimization between the two models are crucial to the generation of new enzymes with thermostable properties.
[0059] The above describes the embodiments of the present invention. However, the present invention is not limited to the above embodiments. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A method for designing and modifying a new thermostable enzyme based on a deep learning model, characterized in that: The following steps are involved: 1) Collecting Marker-Related Sequences: Collect the optimal growth temperature information (OGT) of different species from the database and define organisms as high-temperature species (HTO), medium-temperature species (MTO), and low-temperature species (LTO) according to their optimal growth temperatures. Extract the gene sequences corresponding to the species in the gene dataset according to the previously marked species information and mark them as high-temperature sequences, medium-temperature sequences, and low-temperature sequences, respectively. Collect enzyme sequences with the desired function to obtain potential sequences of enzymes with the desired function for constructing protein sequence generation models. 2) Constructing a generative model for protein sequences: The potential enzyme sequences with the desired function obtained in step 1) are filtered by sequence length. A generative model is constructed using a deep learning framework to expand the sequence space of the target enzyme. The quality of the generative model is evaluated using relevant algorithms. Once the model is stable, it can generate new protein sequences with a distribution similar to that of natural sequences. 3) Constructing a thermostable protein discrimination model: Using the high-temperature, medium-temperature, and low-temperature sequences obtained in step 1), a classification model is constructed and trained to determine whether the input sequence has thermostable characteristics. Once the model is stable, its quality is judged by its accuracy on the test set and its recall rate on the high-temperature sequence. 4) Optimizing the protein sequence generation model using the thermostable protein discrimination model: Inputting the new protein sequence generated in step 2) into the thermostable protein discrimination model constructed in step 3), screening new protein sequences identified as high-temperature sequences as new training sequences, and optimizing the protein sequence generation model obtained in step 2) until the optimized protein sequence generation model can directly generate a large proportion of new sequences that can be identified as thermostable sequences; 5) Screening generated sequences: The sequences generated by the final protein sequence generation model are passed to the thermal stability protein discrimination model for thermal stability evaluation. The quality of the sequences is evaluated by computational methods, and an appropriate number of sequences are screened for experimental verification based on the throughput of the experimental test. 6) Experimental verification of the designed sequences that are identified as having thermal stability: Use experimental methods to obtain the screened designed sequences, measure the expression of the designed sequences, and then experimentally verify whether the designed sequences are thermally stable.
2. The method according to claim 1, characterized in that In the step 1), the organisms with OGT>=50°C, 30°C≤OGT<50°C and OGT<30°C are defined as high-temperature species HTO, medium-temperature species MTO and low-temperature species LTO respectively.
3. The method according to claim 1, characterized in that The step 1) collects the optimal growth temperature information OGT of different species in the TEMPURA, ExProtDB, NCBI, and BacDive databases.
4. The method according to claim 1, wherein In the step 1), the gene sequence corresponding to the species is extracted from the unipro and uniref gene data sets according to the species information marked above.
5. The method according to claim 1, wherein In step 2), a generative adversarial network or diffusion network model is constructed using a deep learning framework such as tensorflow or pytorch.
6. The method according to claim 1, wherein The step 2) uses tSNE or Shannon entropy calculation method to evaluate the quality of the generated model.
7. The method according to claim 1, characterized in that In the step 3), the high-temperature sequence, the medium-temperature sequence, and the low-temperature sequence obtained in step 1) are encoded using the protein language model ESM 1b, randomly divided into a training set and a test set, and a classification model is built using the deep learning framework tensorflow or pytorch. The model is trained using the training set to determine whether the input sequence has a thermostable feature.
8. The method according to claim 1, characterized in that In step 5), alphafold is used to perform structural modeling to evaluate the quality of the protein generated sequence.
9. The method according to any one of claims 1 to 8, characterized in that The enzyme with the desired function in step 1) is glyceraldehyde-3-phosphate dehydrogenase.
10. The method according to claim 9, characterized in that Step 6) Verify that the amino acid sequence of the obtained thermostable new enzyme is as shown in any one of SEQ ID NOs. 1-6, wherein the thermostability refers to the enzyme activity at 65°C.
Citation Information
Patent Citations
Method for identifying thermophilic protein based on machine learning
CN110517730A
Method and device for generating protein sequence of high-thermal-stability enzyme, medium and equipment
CN113539374A
Optimizing Proteins Using Model Based Optimizations
US20230054552A1