EnzyForge AI deep learning screening method for novel enzyme mining
The EnzyForge AI deep learning screening method combines sequence and structure screening to identify key catalytic domains, solving the problem of low homology enzymes being difficult to discover in traditional methods. This enables efficient and accurate discovery of novel lipases with significantly improved catalytic activity.
Patent Information
- Application Number
- CN202511884439.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-03
AI Technical Summary
When existing technologies rely on sequence homology to discover lipases, it is difficult to find enzymes with low homology but similar functions, resulting in high blindness and low success rate in experimental verification, which cannot meet industrial needs.
Using the EnzyForge AI deep learning screening method, combined with sequence screening and structural screening, key catalytic domains were identified through molecular docking and kinetic simulation, highly active lipases were screened, a dedicated lipase resource library was constructed, and heterologous expression was performed.
A novel lipase was discovered efficiently and precisely, with catalytic activity twice that of commercially available enzymes, breaking the limitations of traditional sequence dependence, reducing costs and improving efficiency.
Smart Images

Figure CN121601035A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of enzyme engineering, specifically involving a novel EnzyForge AI deep learning screening method for enzyme mining, thereby efficiently discovering novel lipases. Background Technology
[0002] Lipases are in high demand due to their highly efficient catalytic role in green industries such as biodiesel, food, pharmaceuticals, and detergents. The market urgently needs novel lipases with higher activity, stability, and specific selectivity. 99% of microorganisms in nature have not yet been cultured in the laboratory, resulting in a large amount of potential high-quality enzyme resources remaining undiscovered and their functions unknown. Enzyme mining technology, through bioinformatics methods, directly and directionally discovers new enzyme genes from environmental metagenomics or databases, becoming a key approach to breaking through the performance bottlenecks of existing lipases. Current mainstream methods rely on sequence homology or conserved motifs for mining, but this has significant limitations: it is highly dependent on known sequences, making it difficult to discover lipases with large evolutionary distances; sequence similarity cannot reliably predict key industrial properties of enzymes (such as substrate specificity, stereoselectivity, and stability), because enzyme function is ultimately determined by three-dimensional structure, and relying on sequence similarity mining often leads to high uncertainty and low success rate in subsequent experimental verification.
[0003] To overcome the limitations of sequence mining and accurately acquire high-performance lipases, structure-based mining strategies exhibit core advantages: 1) Overcoming sequence limitations: Utilizing structural conservation (folding patterns, catalytic centers), distant homologs with extremely low sequence similarity but similar functions can be identified, greatly expanding the pool of mineable resources; 2) Precise prediction capabilities: By analyzing the geometry of active sites, combined with molecular docking and molecular dynamics simulations, binding affinity to substrates, substrate preference, stereoselectivity, and stability can be directly assessed, significantly improving targeting and prediction reliability. However, relying solely on structure prediction presents challenges for massive amounts of unknown sequences.
[0004] Integrating a synergistic strategy of sequence screening and structural depth screening, taking into account both breadth and depth of discovery, is the inevitable development direction for achieving efficient and accurate discovery of novel lipases that meet demanding industrial requirements. Summary of the Invention
[0005] The purpose of this invention is to provide a novel EnzyForge AI deep learning screening method for enzyme discovery, addressing the aforementioned problems and further identifying novel lipases. This invention primarily develops an EnzyForge AI deep learning screening method for novel enzyme discovery. Addressing the technical bottleneck of traditional enzyme screening relying on sequence similarity and failing to effectively discover enzymes with low homology but similar functions, this method proposes an innovative system integrating multimodal features and multi-level screening strategies, enabling high-accuracy and high-efficiency lipase resource discovery. The method first screens highly catalytically active lipases as seed proteins through literature review and previous laboratory results. It then searches for proteins with sequence and structural similarity to the seed enzymes to construct a dedicated lipase resource library. The seed enzymes are experimentally verified and subjected to molecular dynamics simulations to screen seed proteins. Molecular docking and dynamics simulations are used to analyze the key catalytic domains of the seed proteins. By comparing the local structures with the localized dedicated lipase resource library, potential proteins with similar binding regions to the seed proteins are obtained. Candidate proteins are screened by EC numbering and integrated with catalytic activity prediction tools and soluble expression analysis to obtain preferred proteins. Finally, the preferred proteins are heterologously expressed using *E. coli* BL21(DE3). This invention integrates a synergistic strategy of sequence screening and structural depth screening, taking into account both breadth and depth of discovery, and can achieve efficient and accurate discovery of novel lipases that meet demanding industrial requirements.
[0006] To solve the technical problem of this invention, the technical solution proposed by this invention includes the following steps:
[0007] 1) Based on previous laboratory results, lipases with high catalytic activity were screened out; based on the compatibility requirements of industrial production with microbial expression systems, 12 microbial lipases were finally screened and retained as search seeds for subsequent sequence structure comparison, and the crystal structure and amino acid sequence of the lipases were obtained from the Protein Database (PDB).
[0008] 2) For the seed enzyme obtained in step 1), use Protein Cartography software to search for protein sequences with similar sequences and structures; obtain other proteins belonging to the same cluster as the seed enzyme through the AlphaFold Clusters website, and integrate the two to build a dedicated lipase resource library.
[0009] 3) Experimental verification and molecular dynamics simulation analysis were performed on the seed enzymes obtained in step 1). CRL, CALB and CALA were selected as seed proteins. Molecular docking and molecular dynamics simulation were used to analyze the key structural regions of catalytic binding.
[0010] 4) Based on the simulation results obtained in step 3), the residues within a 6 Å radius of the ligand binding center are defined as the catalytic binding region. The local structure of this region is then compared with that of the localized dedicated lipase resource library using Pyscomotif and Folddisco to obtain potential proteins with binding regions similar to the seed protein.
[0011] 5) For the potential proteins obtained in step 4), candidate proteins of the 3.1.1.3 classification enzyme system are predicted and screened by EC numbering; the candidate proteins are screened and retained by substrate docking; four catalytic activity prediction tools (UniKP, Turnup, CataPro, CatPred) are integrated and soluble expression prediction is performed to finally obtain the preferred proteins;
[0012] 6) The preferred protein obtained in step 5) was heterologously expressed in Escherichia coli BL21(DE3).
[0013] The dedicated lipase resource library is based on twelve highly efficient microbial lipases as seeds. Protein cartography is used to obtain protein sequences with similar sequences and structures to the seed enzymes. Proteins with low overall structural similarity to the seed enzymes but belonging to the same structural cluster are obtained through the AlphaFoldClusters website, thus constructing a dedicated lipase resource library.
[0014] Preferably, the seed enzymes finally screened in step (1) include yarrow lipolytic bacteria lipase, cottony thermophilic bacteria lipase, Candida antarcticis lipase A, Candida antarcticis lipase B, Rhizopus miltiorrhiza lipase, Candida pleurisy lipase, cholesterol esterase, Aspergillus fumigatus lipase, Burkholderia cepacia lipase, Burkholderia gladiolus lipase, Pseudomonas fluorescens lipase, and Pseudomonas aeruginosa lipase.
[0015] Preferably, in step (2), the method for constructing a dedicated lipase resource library involves using Protein Cartography to obtain proteins with similar sequences and structures to the seed enzyme; to expand the depth of the database and obtain proteins that have low structural similarity to the seed enzyme but belong to the same structural cluster, the structure of other proteins in the same cluster of the seed enzyme is obtained through the AlphaFold Clusters website; after integrating the database and removing redundancy, a dedicated lipase resource library is constructed.
[0016] Preferably, the molecular docking and molecular dynamics simulation method for analyzing key catalytic domains described in step (3) involves performing 1000 semi-flexible molecular dockings for each seed protein. After conformational filtering to eliminate docking postures with spatial conflicts, the K-means clustering algorithm is used to select the conformation with the highest frequency as the representative binding posture for subsequent 100 ns all-atom molecular dynamics simulations.
[0017] In step (4), after the simulation, in order to obtain the key catalytic domain, the performance of fpocket, PrankWeb and dogsite3 protein binding pocket prediction tools was compared. Finally, the residues within a 6 Å radius of the ligand binding center were determined as the definition standard for the catalytic binding region. In the local structure comparison stage, the computational efficiency of Pyscomotif and Folddisco was compared, and a hierarchical screening strategy was adopted: first, Folddisco was used to perform preliminary overall structure matching of the catalytic binding region, and then Pyscomotif was used to perform fine comparison of key residues for proteins with a matching score higher than 2000.
[0018] Preferably, the method for screening preferred proteins described in step (4) uses the EC number prediction Clean tool to exclude enzyme systems that are not classified in 3.1.1.3, leaving 312 candidate proteins. After screening with substrates through 100 repeated semi-flexible docking, 104 sequences are retained. Finally, the preferred proteins are obtained by integrating the results of four catalytic activity prediction tools, UniKP, Turnup, CataPro, and CatPred, with the soluble expression prediction results.
[0019] The protein screened by the above method yielded the preferred protein H7, whose amino acid sequence is shown in SEQ ID NO.1 and its nucleotide sequence is shown in SEQ ID NO.2.
[0020] In the application of the protein, the hydrolytic activity of the preferred protein H7 for p-nitrophenol palmitate was measured after expression in E. coli. It was found that the activity of the screened and expressed protein H7 was about twice that of the commercially available lipase CALB.
[0021] This invention breaks through the limitations of traditional sequence-based searches, and for the first time utilizes microbial lipases obtained through experimental verification and molecular dynamics simulation screening as "seed enzymes." Combining their multidimensional sequence-structure characteristics, a dual-channel extended search is performed in UniProt and the AlphaFold Structure Database to construct a localized lipase resource library. Based on the hypothesis that "the substrate-binding domain of lipases is a key factor in their catalytic ability," this invention further performs multiple rounds of molecular docking between the seed enzyme and the target substrate, selecting the complex conformation with the optimal docking score and performing molecular dynamics simulations. Through multiple rounds of comparative analysis, the set of residues within a 6 Å radius of the ligand binding center is defined as the key catalytic binding region. Subsequently, using local structure alignment tools such as Pyscomotif and FoldDisco, the aforementioned catalytic binding region is used as a structural probe to perform a local structure search in the constructed local resource library, obtaining candidate proteins with highly similar binding regions. After obtaining preliminary candidates, functional prediction is performed using the Clean protein language model, eliminating sequences with predicted EC numbers other than 3.1.1.3, thereby excluding non-lipase proteins. For the remaining candidate sequences, a comprehensive evaluation and further screening were conducted using substrate molecule docking calculations, combined with four enzyme catalytic activity prediction tools (UniKP, Turnup, CataPro, and CatPred) and soluble expression prediction tools (DeepSoluE, SoluProt, and TISIGNER).
[0022] The selected proteins were cloned into the pET-28a vector and heterologously expressed in E. coli BL21(DE3), forming a closed-loop process from computational prediction to experimental verification, and realizing the systematic and large-scale intelligent mining of lipases.
[0023] Beneficial effects:
[0024] This invention develops an EnzyForge AI deep learning screening method for novel enzyme discovery. Addressing the technical bottleneck of traditional enzyme screening, which relies on sequence similarity and cannot effectively discover enzymes with low homology but similar functions, this method proposes an innovative system integrating multimodal features and a multi-level screening strategy, enabling high-accuracy and high-efficiency discovery of lipase resources. This invention innovatively combines sequence homology search with structural similarity search, using local catalytic key domains as the core of functional determination. Based on this, it further performs local residue alignment and combines a protein big language model to conduct multi-dimensional evaluation of the activity, catalytic potential, and stability of candidate enzymes, achieving deep screening and ranking of candidate enzymes. This breaks through the limitations of traditional methods based solely on sequence search and can identify enzymes with extremely low overall sequence similarity but completely similar catalytic core conformations. The three candidate proteins finally screened all exhibited significant enzyme activity. Among them, H7 showed the highest catalytic activity, with a catalytic efficiency twice that of commercially available CALB. It is worth emphasizing that the gene fragment corresponding to H7 has never been reported to possess lipase activity in previous literature and databases. This invention is the first to successfully mine this functional enzyme molecule from the natural sequence space, representing a novel, previously undisclosed lipase resource. Compared to traditional enzyme modification strategies that rely on random mutation or directed evolution, the mining method constructed in this invention can directly obtain high-performance candidate enzymes from the natural sequence system without extensive experimental screening, offering significant advantages such as low cost, high efficiency, and strong scalability. The successful acquisition of H7 not only verifies the unique advantage of this method in discovering enzyme molecules with low homology but high functional potential, but also fully demonstrates the reliability and innovation of the enzyme mining strategy based on sequence and structural features proposed in this invention.
[0025] First, to obtain template enzymes for subsequent screening, 14 highly efficient lipases were screened based on previous laboratory results. To adapt to industrial microbial expression systems, 12 microbial lipases were retained as templates, and their crystal structures and sequences were obtained from the PDB database. Second, a database was constructed based on proteins in the obtained template enzyme search database. Protein cartography was used to search for proteins with similar sequences and structures to the template enzymes, and AlphaFold Clusters was used to obtain clusters of proteins with the same structure as the template enzymes, integrating them to construct a dedicated lipase resource library. Subsequently, suitable enzymes were screened based on their catalytic binding regions. First, the key catalytic domains of the enzymes were analyzed, and the template enzymes were experimentally verified and subjected to molecular dynamics simulations. CRL, CALB, and CALA were selected as seed proteins. Molecular docking and dynamics simulations were used to analyze the key catalytic domains. The catalytic binding region was defined by residues within a 6 Å radius of the ligand binding center. Pyscomotif and Folddisco were used for local structural comparison to screen for potential proteins with similar binding regions. Preferred proteins were screened based on predicted enzyme type, catalytic activity, and soluble expression. Candidate proteins were further screened using EC number 3.1.1.3 classification. Substrate and candidate proteins were docked to select target proteins with correct binding postures and strong affinity. Four catalytic activity prediction tools (UniKP, Turnup, CataPro, and CatPred) were integrated, and soluble expression analysis was performed to obtain the preferred proteins. Finally, the preferred proteins were heterologously expressed using the *E. coli* BL21(DE3) system, and their hydrolytic activity was verified.
[0026] The analysis of the key catalytic domain involved performing 1000 semi-flexible molecular dockings on seed proteins with excellent catalytic efficiency and substrate binding energy. After eliminating spatially conflicting docking postures, the K-means clustering algorithm was used to select the most frequent conformation as the representative binding posture for subsequent 100 ns all-atom molecular dynamics simulations. Residues within a 6 Å radius of the ligand binding center were used as the definition standard for the catalytic binding region. Folddisco was used for preliminary overall structural matching of the catalytic binding region, and then Pyscomotif was used to perform fine-grained comparison of key residues for proteins with matching scores higher than 2000. Attached Figure Description
[0027] Figure 1 The workflow of EnzyForge AI deep learning screening methods;
[0028] Figure 2 Seed enzyme database;
[0029] Figure 3 : The source composition and distribution of structural cluster resource library;
[0030] Figure 4 Extraction of key residues from the catalytic binding region;
[0031] Figure 5 : Phylogenetic analysis and docking binding energy screening;
[0032] Figure 6 The heterologous expression bands of H7 protein are shown in the following order from left to right: Marker, empty supernatant after 4 hours of induction, empty precipitate after 4 hours of induction, empty supernatant after 24 hours of induction, empty precipitate after 24 hours of induction, H7 supernatant after 4 hours of induction, H7 precipitate after 4 hours of induction, H7 supernatant after 24 hours of induction, and H7 precipitate after 24 hours of induction. Detailed Implementation
[0033] This invention constructs and proposes an EnzyForge AI deep learning screening method for novel enzyme mining, such as... Figure 1 As shown. The present invention will be further described below with reference to specific embodiments, but is not limited thereto. Unless otherwise specified, the reagent raw materials described in the following embodiments are all commercially available common raw materials, and the reagents are prepared using conventional methods. Methods not detailed in the embodiments are all conventional operations in the art.
[0034] Example 1: Construction of a dedicated lipase resource library.
[0035] (1) Based on literature review and previous laboratory results, lipases with high catalytic activity were screened. Based on the compatibility requirements of industrial production with microbial expression systems, microbial lipases were ultimately selected and retained as search seeds for subsequent structure comparisons. These included *Achillea millefolium* lipase, *Thermophilus sparsely cottony* lipase, *Candida antarcticus* lipase A, *Candida antarcticus* lipase B, *Rhizopus oryzae* lipase, *Candida globosum* lipase, cholesterol esterase, *Aspergillus fumigatus* lipase, *Burkholderia cepacia* lipase, *Burkholderia gladiolus* lipase, *Pseudomonas fluorescens* lipase, and *Pseudomonas aeruginosa* lipase. The crystal structures and amino acid sequences of the lipases to be modified were obtained from the Protein Database (PDB). The three-dimensional structures of the seed proteins are shown below. Figure 2 As shown.
[0036] (2) An initial database was constructed by searching for enzymes with similar sequences and structures to these seed proteins. Protein cartography yielded 28,438 protein sequences with sequences and structures similar to the seed enzymes. To expand the database's depth, proteins with lower overall structural similarity to the seed enzymes but belonging to the same structural cluster were identified. The AlphaFold Clusters website was used to obtain 11,991 other protein structures from the seed enzyme cluster. After integrating and removing redundancy, a dedicated lipase resource library with 32,635 proteins was finally established. To verify the database's comprehensiveness, structural clustering analysis was performed on the database proteins, such as... Figure 3 As shown.
[0037] Example 2: Structure and sequence-based lipase mining
[0038] (1) Based on experimental verification data and molecular dynamics simulation analysis, CRL, CALB, and CALA were selected as seed proteins. The three proteins showed significantly better catalytic efficiency and substrate binding energy than other candidate enzymes. To elucidate the key catalytic domains, 1000 semi-flexible molecular dockings were performed between each seed protein and the substrate. After conformational filtering to eliminate spatially conflicting docking postures, the K-means clustering algorithm was used to select the most frequent conformation as the representative binding posture. The final clustering method was used for subsequent 100 ns all-atom molecular dynamics simulations.
[0039] (2) To obtain the key catalytic domains after the simulation, the performance of fpocket, PrankWeb, and dogsite3 protein binding pocket prediction tools was compared. It was found that none of these tools could adequately represent the entire catalytic binding region between the protein and ligand. Therefore, the residues within a 6 Å radius of the ligand binding center were defined as the catalytic binding region. During the local structure alignment stage, the computational efficiency of Pyscomotif and Folddisco was compared. It was found that Pyscomotif significantly increased the time required for global alignment. Therefore, a hierarchical screening strategy was adopted: first, Folddisco was used for preliminary overall structure matching of the catalytic binding region; then, Pyscomotif was used for fine alignment of key residues for proteins with matching scores higher than 2000, such as... Figure 4 As shown.
[0040] (3) After the above screening process, more than 500 potential proteins were obtained. Enzymes not classified as 3.1.1.3 (lipases) were excluded using EC number prediction (Clean algorithm), leaving 312 candidate proteins. After 100 repeated docking with the substrate, 104 protein sequences with correct binding postures were retained. Four catalytic activity prediction tools (UniKP, Turnup, CataPro, CatPred) and soluble expression prediction were integrated to screen for more than twenty lipases with higher activity and soluble expression prediction results than the seed protein. These proteins were compared with the seed protein CRL using TM-align. Proteins with a TM-score below 0.8 were excluded, and the top three proteins were finally selected for experimental verification. Proteins H7, P8, and U6 were obtained. The amino acid sequence of H7 is shown in SEQ ID NO.1, and its nucleotide sequence is shown in SEQ ID NO.2. The amino acid sequence of P8 is shown in SEQ ID NO.3, and its nucleotide sequence is shown in SEQ ID NO.4. The amino acid sequence of U6 is shown in SEQ ID NO.5, and its nucleotide sequence is shown in SEQ ID NO.7. Figure 5 As shown, the free energy calculation results indicate that all three candidate proteins can form stable complexes with the target substrate, exhibiting a certain degree of binding affinity. Based on the above energy assessment results and the sequence and structural feature screening strategy, H7 from *Verticillium longisporum*, P8 from *Athelia psychrophila*, and U6 from *Colletotrichum nymphaeae* were selected as preferred candidates, and they were constructed into the pET-28a vector for heterologous expression in *E. coli* BL21(DE3) to further verify their catalytic function.
[0041] Heterologous expression of proteins H7, P8, and U6 in *E. coli* BL21(DE3) all exhibited significant hydrolytic activity against p-nitrophenol palmitate, validating the effectiveness of the constructed computational screening system. Among them, H7 showed the most outstanding catalytic activity, with an efficiency twice that of commercially available CALB and comparable to commercially available CRL, demonstrating extremely high application potential. It is worth emphasizing that the gene fragment corresponding to H7 had not been previously reported functionally. This invention is the first to resolve and verify its lipase activity, indicating that the proposed enzyme discovery method can overcome the limitations of traditional strategies relying on sequence homology or site-directed mutagenesis, achieving the effective discovery of previously unrecognized potentially highly active enzymes in nature.
[0042] The amino acid sequence of H7 (SEQ ID NO.1):
[0043] MLFKLPILLGLLGTVAAQSDATPALDERAAGTATVVLPLATVLGNVMNKVESFGGIPFAEPPVGRLRLKPPQRIARNLGTFDATGPAAACPQMVSSSESDNFLFDVLGEIANLPFVQKPPVGRLRLKPPQRIARNLGTFDATGPAAACPQMVSSSESENFLFNLLGDIANLPFVQKVTGQTEDCLSITVARPEGTKADAKLPVLYWIFGGGFELGWSSMHDGTGLIKHGVDLKKPFIFVAVNYRVAGFGFMPGKEILADGSSNLGLLDQRMGIEWVADNIASFGGDPSKVTIWGESAGAISVFDQMALYDGDNTYKGKPLFRGAIMNSGSMVPADPVDCPKGQAVYDLVVKEAGCAGQADTLNCLRDLPYQTFLKAVTAPPGILSYNSVALSYLPRPDGKVLTQSPDVLAATGKYAAVPMIIGNQEDEGTLFGLFQPNLTTTDRFVDYLQQLFFNSATKAQLTTLVNTYDNGVAAVLAGSPHRTALLNEIFPGFKRRAAVLGDLVFTLTRRAFLSLTKAAHPDVPAWSYLATYNYGTPILGTFHGSDLLQIFPGFKRRAAVLGDLVFTLTRRAFLSLTKAAHPDVPAWSYLATYNYGTPILGTFHGSDILQVFYGVLPNYASRQIRTYYTNFVHDLDPNVGAAAQYPSWPRWDQGKKLINFLANRAGGLLDDNFRSDSYDFIASNVGAFYI
[0044] Nucleotide sequence of H7 (SEQ ID NO. 2):
[0045]
[0046] Amino acid sequence of P8 (SEQ ID NO. 3):
[0047] MDILLSEREGLRWIQKYIHSFGGDPSKVTLFGVSSGGISTALHMLLNDGNTEGLFRAAFSQSGAPIPVGSYTHGQKWYDGAVQAANCTAAKDTLECLRGADVEVLQAYFSTTPSKLSYQALPSAWLPRVDGKYLKDDPQQLILKGSVAKIPFISGNDDDEGTLFSLYTTNITTDADFREWVQSDYFPNATSAEIDLILKLYPSDPTVGSPFDTGINNTLYPQYKRISAFQGDVVFQGPRRFYLQHRASLQKAWSYLDKRGKTSSSLGSQHGIDMLSMYGTGNGTELQDYVINFAYRLDPNGPTVPAWPQYTLVSPKLLTLVDGAIPVTITEDTYRVAPMEGVTALGLKYPLLE
[0048] Nucleotide sequence of P8 (SEQ ID NO. 4):
[0049]
[0050] Amino acid sequence of U6 (SEQ ID NO.5):
[0051] VSDHVQPLEDRAVANATVVLKTATVVGNVMNNVESFGGIPYAKPPTGQLRLKPPVRLTDNIGTFDATGPAAACPQMVSSSDSDNILFNLLGDIANLPFVQKATGQTEDCLTITVARPQGTTADAKLPVLYWIFGGGFELGWSSMYDGTGLVQHGVDISKPFIFVAVNYRVAGFGFMPGKEILADGSANLGLLDQRMGLEWVADNIAAFGGDPDKVTIWGESAGAISVFDQMALYNGNNKYNGKALFRGAIMNSGSIVPTDPVDCPKGQAVYDKVVSEAGCAGKADTLACLRGVDYTTFLNAVTSVPGILSYNSLALSYLPRPDGKTLTASPDVLAKNGQYAAVPMIIGDQEDEGTLFGIFQPNLTTTEKLVTYLKNYYFATATTAQITAYVATYDDGVTAVINGSPHRTGLLNEIFPGFKRRSAVLGDLVFTLTRRVFLTIANSVQPTVPSWSYLSSYDYGTPILGTFHGSDLLQVFYGIKDNYAARSIRTYYTNFVYASDPTVGLNGAYPTWPQWSQGQNLMQFFADKASTLKDDFRKSSSDWILNNAGSLYFLEHHHHHH
[0052] Nucleotide sequence of U6 (SEQ ID NO.6):
[0053]
[0054] Example 3: Expression of a novel lipase
[0055] (1) Using the three protein genes selected as search seeds, whose amino acid sequences are shown in SEQ ID NO:1, 3 and 5, they were inserted into the pET21b(+) vector with EcoRI-NcoI as the cloning site, and the wild recombinant expression plasmid pET21b-protein was synthesized.
[0056] (2) The above plasmid was transformed into BL21(DE3) competent cells by heat shock. 500 µL of LB medium without any antibiotics was added and incubated at 37°C for 60 min. After a brief centrifugation at 3000 rpm, the remaining 100 µL of supernatant was mixed with the cells. An appropriate amount was spread on an LB plate containing 100 mg / L ampicillin and cultured overnight at 37°C until a clear single colony grew.
[0057] (3) Pick a single colony from the overnight plate into 5 mL of LB liquid medium containing 100 mg / L ampicillin resistance, and culture it in a shaker at 37°C for 12 h to preserve the bacteria and sequence them.
[0058] (4) Select the successfully verified single clones and culture them in LB liquid medium containing 100 mg / L ampicillin resistance for 12-16 h at 37℃ and 220 rpm. Transfer 1-2% of the clones to 50 mL of TB medium containing 100 mg / L ampicillin resistance and culture them at 37℃ and 220 rpm for 2-4 h until the OD reaches 0.6-0.8. Add IPTG to a final concentration of 0.1 mM and induce for 24 h at 18℃ and 220 rpm.
[0059] (5) Centrifuge the fermentation broth at 4℃ and 6000 rpm for 20 min to collect the cells. Add 10 mL of PBS buffer to each g of cells to resuspend the cells. Perform sonication at 3 s pause and 5 s sonication at 30% power for 20 min per 10 mL. Centrifuge the broken bacterial solution at 4℃ and 6000 rpm for 20 min to remove cell debris and collect the supernatant.
[0060] (6) The supernatant was subjected to 10% SDS-PAGE electrophoresis to identify the expression of the target protein. The expression and purification process of protein H7 is as follows: Figure 6 As shown.
[0061] (7) Determination of lipase activity: Solution A: 15 mg p-nitrophenyl palmitate (pNPP) dissolved in 10 mL isopropanol, stored in a brown bottle at 4℃. Solution B: 1.51 g tris(hydroxymethyl)aminomethane (Tris) dissolved in 200 mL water, 1 mL Triton-X100, 0.25 g gum arabic, and 0.72 mL concentrated hydrochloric acid were added, stirred evenly, pH adjusted to 7.5, and the volume was brought to 250 mL. Stored at 4℃. Mixed solution A and B: A and B solutions were mixed in a volume ratio of 1:9 and incubated at 40℃ for later use. 4.5 mL of substrate solution was added to 0.1 mL of the obtained wild-type lipase crude enzyme solution, and 4.5 mL of substrate solution was added to 0.1 mL of blank culture medium as a control. The reaction was incubated in a water bath at 40℃ for 5 min, and 1 mL of acetone was added to terminate the reaction. Lipase hydrolyzes pNPB to release p-nitrophenol pNP. p-nitrophenol solution is yellow-green under alkaline conditions. Enzyme activity is defined as follows: one enzyme activity unit (U) is the amount of pNPB released by 1 μmol of p-nitrophenol per 1 min of hydrolysis by 1 mL of lipase, expressed in μmol / mL. -1 min -1 .
[0062] Table 1: Comparison of hydrolysis activity of commercially available Novozym CALB (purchased from Beijing Gaoruisen Technology Co., Ltd.) with H7, P8, and U6.
[0063]
[0064] The activity assay results showed that the successfully mined and expressed H7, P8, and U6 enzymes all exhibited significant hydrolytic activity. Among them, H7 showed the highest enzyme activity, with a catalytic efficiency twice that of commercially available CALB. These results fully validate the effectiveness and reliability of the enzyme mining method proposed in this invention.
Claims
1. An EnzyForge AI deep learning screening method for novel enzyme mining, characterized in that, Includes the following steps: (1) Based on the compatibility requirements of industrial production for microbial expression systems, microbial lipases were screened and retained as search seeds for subsequent sequence and structure comparisons. The crystal structure and amino acid sequence of the lipases were obtained from the protein database PDB. (2) For the seed enzyme obtained in step (1), use Protein Cartography software to search for proteins with similar sequences and structures; obtain the structures of other proteins belonging to the same cluster as the seed enzyme through the AlphaFold Clusters website, and construct a lipase resource library; (3) Experimental verification and molecular dynamics simulation analysis were performed on the seed enzymes obtained in step (1). Candida rhombifolia lipase CRL, Candida antarcticis lipase B (CALB) and Candida antarcticis lipase A (CALA) were selected as seed proteins. Molecular docking and molecular dynamics simulation were used to analyze the key catalytic binding domains. (4) For the simulation results obtained in step (3), the residues within a radius of 6 Å of the ligand binding center are defined as the catalytic binding region. The local structure of this region is then compared with that of the localized dedicated lipase resource library using Pyscomotif and Folddisco to obtain potential proteins whose binding region is highly similar to that of the seed protein. (5) For the potential proteins obtained in step (4), candidate proteins of the lipase enzyme system in the 3.1.1.3 classification enzyme system are predicted and screened by EC number; the candidate proteins are screened and retained by docking with the substrate; four catalytic activity prediction tools UniKP, Turnup, CataPro and CatPred are integrated and soluble expression prediction is performed to finally obtain the preferred protein; (6) The preferred protein obtained in step (5) was heterologously expressed in Escherichia coli BL21(DE3) and its hydrolytic activity was determined.
2. The novel enzyme mining EnzyForge AI deep learning screening method according to claim 1, characterized in that, The seed enzymes finally screened in step (1) include yarrow lipase, cottony thermophilic bacteria lipase, Candida antarcticis lipase A, Candida antarcticis lipase B, Rhizopus miltiorrhiza lipase, Candida pleurisy lipase, cholesterol esterase, Aspergillus fumigatus lipase, Burkholderia cepacia lipase, Burkholderia gladiolus lipase, Pseudomonas fluorescens lipase, and Pseudomonas aeruginosa lipase.
3. The enzyme molecule mining method based on sequence and structural features according to claim 1, characterized in that, In step (2), the method for constructing a dedicated lipase resource library involves using Protein Cartography to obtain proteins with similar sequences and structures to the seed enzyme. To expand the depth of the database and obtain proteins that have low structural similarity to the seed enzyme but belong to the same structural cluster, the structure of other proteins in the same cluster of the seed enzyme is obtained through the AlphaFold Clusters website. After integrating the database and removing redundancy, a dedicated lipase resource library is constructed.
4. The novel enzyme mining EnzyForge AI deep learning screening method according to claim 1, characterized in that, The molecular docking and molecular dynamics simulation method for analyzing key catalytic domains described in step (3) involves performing 1000 semi-flexible molecular dockings for each seed protein and filtering the obtained docking conformations to eliminate binding postures with spatial conflicts. Subsequently, the K-means clustering algorithm was used to perform cluster analysis on the remaining docking conformations, and the representative binding conformation with the highest frequency was selected as the input structure for subsequent simulations; and 100 ns all-atom molecular dynamics simulations were carried out based on the representative conformation to resolve the key catalytic structural domains. In step (4), after the simulation, in order to obtain the key catalytic domain, the performance of fpocket, PrankWeb and dogsite3 protein binding pocket prediction tools was compared. Finally, the residues within a 6 Å radius of the ligand binding center were determined as the definition standard for the catalytic binding region. In the local structure comparison stage, the computational efficiency of Pyscomotif and Folddisco was compared, and a hierarchical screening strategy was adopted: first, Folddisco was used to perform preliminary overall structure matching of the catalytic binding region, and then Pyscomotif was used to perform fine comparison of key residues for proteins with a matching score higher than 2000.
5. The novel enzyme mining EnzyForge AI deep learning screening method according to claim 1, characterized in that, The method for screening and selecting preferred proteins described in step (4) involves using the EC number prediction model Clean tool to exclude enzyme systems not classified as 3.1.1.3, leaving 312 candidate proteins. After screening with substrates through 100 repeated semi-flexible docking, 104 sequences are retained. Finally, the results of four catalytic activity prediction tools (UniKP, Turnup, CataPro, and CatPred) and soluble expression prediction are integrated to make a multi-dimensional comprehensive judgment on the candidate proteins. The top three proteins with the highest scores are then selected for experimental verification.
6. The protein screened by the method according to claims 1-5, characterized in that, Proteins H7, P8, and U6 were obtained. The amino acid sequence of H7 is shown in SEQ ID NO.1, and its nucleotide sequence is shown in SEQ ID NO.2; the amino acid sequence of P8 is shown in SEQ ID NO.3, and its nucleotide sequence is shown in SEQ ID NO.4; the amino acid sequence of U6 is shown in SEQ ID NO.5, and its nucleotide sequence is shown in SEQ ID NO.
7.
7. The application of the protein according to claim 6, characterized in that, Heterologous expression of proteins H7, P8, and U6 in Escherichia coli BL21(DE3) all showed significant hydrolytic activity against p-nitrophenol palmitate, validating the effectiveness of the constructed computational screening system. Among them, H7 exhibited the most outstanding catalytic activity, with a catalytic efficiency twice that of commercially available CALB, demonstrating extremely high application potential.