Protocatechuic acid O-methyltransferase based on artificial intelligence design and design method thereof
By constructing a highly selective and active protocatechuic acid O-methyltransferase using artificial intelligence design methods, the problem of lack of site specificity of OMT catalysts in the process of PCA to vanillin was solved, achieving the goal of efficient biosynthesis of vanillin and improving catalytic efficiency and product purity.
Patent Information
- Application Number
- CN202511597503.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-02-10
AI Technical Summary
Existing OMT catalysts lack site specificity in the process of catalyzing PCA to vanillin, and are prone to simultaneous methylation of 3-OH and 4-OH, leading to the formation of byproducts and increasing the pressure of subsequent separation and cost. Current technology lacks a systematic method for designing novel OMT frameworks with high selectivity and high activity.
Using an artificial intelligence design approach, the protein backbone structure was generated through the RF diffusion AA model. The amino acid sequence was optimized by combining the ProteinMPNN and AlphaFold3 models. Hydrogen bond donors and hydrophobic residues were introduced to form a stable octahedral coordination configuration, thus preparing a highly selective and highly active protocatechuic acid O-methyltransferase.
It significantly improves the catalytic efficiency and target product purity of OMT, produces almost no byproducts, and provides a highly efficient tool for the biosynthesis of vanillin, with broad application prospects and economic value.
Smart Images

Figure CN121495892A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of technology, specifically to a protocatechuic acid O-methyltransferase designed based on artificial intelligence and its design method. Background Technology
[0002] O-methyltransferases (OMTs) are important biocatalysts that play a crucial role in the secondary metabolic networks of plants and microorganisms. They can perform O-methylation modification of polyphenolic substrates such as protocatechuic acid (PCA), thereby regulating their biological activity, chemical stability, and metabolic flux. In the natural metabolic pathway, PCA is catalyzed by OMTs to produce vanillic acid, an intermediate product of vanillin. This process is one of the core steps in many current microbial routes for vanillin synthesis. Vanillin, as an important natural flavor compound, is widely used in the food, pharmaceutical, and cosmetic industries, and its demand continues to grow, driving research into efficient biosynthetic strategies for it.
[0003] In industrial production, to improve the yield of target products and simplify the purification process, there is an urgent need for highly selective methylation of specific sites (such as 3-OH of PCA) by OMT. However, natural OMT usually lacks site specificity and tends to methylate both 3-OH and 4-OH simultaneously, leading to the formation of byproducts (such as isopranic acid), which increases the pressure on subsequent separation and costs.
[0004] Current optimization strategies for OMT selectivity mainly include site-directed mutagenesis, directed evolution, and rational modification assisted by structure analysis. Furthermore, recent research has largely relied on local mutation optimization using natural proteins as templates, making it difficult to systematically model complex catalytic reactions involving multiple electron transfers, metal ions, and cofactor synergistic regulation. Therefore, existing technologies lack a design method that can systematically construct novel OMT frameworks with high selectivity and activity from the core catalytic structure to meet the industrial application needs of synthesizing high-value products such as vanillin. Summary of the Invention
[0005] To address the aforementioned problems, this invention provides a protocatechuic acid O-methyltransferase designed based on artificial intelligence.
[0006] An artificial intelligence-based protocatechuic acid O-methyltransferase, wherein the amino acid sequence of the protocatechuic acid O-methyltransferase is SEQ ID NO: 1 or SEQ ID NO: 2.
[0007] Note: The protocatechuic acid O-methyltransferase (OMT) with the above-mentioned sequence obtained in this invention exhibits approximately 15-fold increased activity compared to the wild type, significantly improving catalytic efficiency. Furthermore, it produces almost no byproducts during the catalytic reaction and demonstrates excellent regioselectivity control. This AI-designed protocatechuic acid O-methyltransferase not only effectively enhances enzyme activity but also significantly improves the purity of the target product, providing a powerful tool for the efficient biosynthesis of high-value-added products such as vanillin. It has broad application prospects and significant economic value in the field of biomanufacturing.
[0008] Furthermore, the distance between the methyl carbon atom of S-adenosylmethionine and the 3-hydroxy oxygen atom of protocatechuic acid in the protocatechuic acid O-methyltransferase is 2.8~3.0 Å.
[0009] Note: The above precisely reveals the spatial distance relationships between key atoms in the reaction at the atomic level. This precise distance definition helps to deepen the understanding of the microscopic mechanism of the enzyme-catalyzed reaction, providing important structural basis data for further research on how enzymes accurately recognize substrates and catalyze specific chemical reactions. It also helps to conduct related research such as rational drug design and targeted modification of enzymes based on this spatial structure information, so as to achieve artificial regulation and optimized utilization of enzyme function.
[0010] Furthermore, the active site of the protocatechuic acid O-methyltransferase comprises Mg 2+ An octahedral coordination configuration is formed by coordination with six coordinating atoms; the six coordinating atoms are the second carboxyl oxygen atom of the two aspartic acid side chains in the amino acid sequence, the first carbonyl oxygen atom of the asparagine in the amino acid sequence, the 3-hydroxy oxygen atom of the protocatechuic acid, the 4-hydroxy oxygen atom of the protocatechuic acid, and a water molecule.
[0011] Note: The above-mentioned stable octahedral coordination configuration of the active site can assist in substrate conformational locking, improve reaction selectivity, and use Mg 2+ As a structural hub, it integrates catalysis (the 3-hydroxy oxygen atom of protocatechuic acid), orientation control (the 4-hydroxy oxygen atom of protocatechuic acid), and protein anchoring (Asp140 / 168 / Asn169) to achieve quantum-precision pre-organization of substrate conformation. The carboxyl oxygen of the aforementioned aspartic acid Asp140 / Asp168 provides strong electrostatic binding, and the carbonyl oxygen of asparagine forms a weak coordination buffer layer with water molecules, jointly resisting conformational drift caused by substrate vibration or thermal perturbation, thus thoroughly suppressing substrate rotational freedom and improving selectivity.
[0012] Furthermore, in the octahedral coordination configuration, the bond length of the coordination bond formed by the coordination center and each coordination atom is 1.95~2.10 Å, and the coordination angle is 90~180°.
[0013] Note: The above settings clearly indicate that the bond length of the coordinate bond formed between the coordination center and each coordinating atom in the octahedral coordination configuration is in the range of 1.95~2.10 Å and the coordination angle is in the range of 90~180°, providing key quantitative basis for improving stability and spatial structure.
[0014] Furthermore, the glutamate residue side chain in the protocatechuic acid O-methyltransferase forms a hydrogen bond with the 4-hydroxy oxygen atom of protocatechuic acid, and the O1–CE–S angle in the protocatechuic acid O-methyltransferase is 168~172°, wherein O1 is the 3-hydroxy oxygen atom of protocatechuic acid, CE is the methyl carbon atom of S-adenosylmethionine, and S is the sulfur atom of S-adenosylmethionine.
[0015] Explanation: The above method achieves a breakthrough improvement in catalytic efficiency and selectivity through dual geometric locking. The hydrogen bond between Glu198 and 4-OH isolates the competing group from the reaction center and inhibits the formation of byproducts. At the same time, the forced constraint of the O1–CE–S angle improves the catalytic efficiency.
[0016] An artificial intelligence-based enzyme design method for designing the above-mentioned protocatechuic acid O-methyltransferase includes the following steps: S1. Generate protein backbone structures using the RF diffusion AA model; the input of the RF diffusion AA model is the fragment structure of the active site, as well as the fragment structure of the methyl carbon atom of S-adenosylmethionine and the 3-hydroxy oxygen atom of protocatechuic acid; the output is 100~1000 protein backbone structures. S2. Using the ProteinMPNN model, generate 10 amino acid sequences on each protein backbone structure obtained in S1; S3. Using the AlphaFold3 model, input the amino acid sequence generated on the protein backbone structure of S2 to obtain the three-dimensional structure of the protein. S4. By optimizing the binding sites in the three-dimensional structure of the protein, protocatechuic acid O-methyltransferase was designed.
[0017] Note: The RFdiffusion_AA model in the above method supports small molecules (such as SAM, PCA, Mg). 2+The protein motif is input together with the protein and participates in backbone generation; the basic RFdiffusion model does not support heterologous ligands and is difficult to use for designing fine catalytic pockets to achieve active site-dominated design based on ternary complexes (substrate + coenzyme + metal ion); ProteinMPNN fixes the catalytic residue sequence, and AlphaFold3 double-round screening ensures folding stability and avoids active site distortion; through a closed-loop design of structure generation → sequence fitting → dynamic verification, the limitations of natural enzyme catalytic efficiency and purity are finally overcome.
[0018] Furthermore, in S2, after generating the 10 amino acid sequences, hydrogen bond donor residues are introduced into the amino acid sequences near the 3-hydroxy oxygen atom position of protocatechuic acid, and hydrophobic residues are introduced into the amino acid sequences near the 4-hydroxy oxygen atom position of protocatechuic acid, using the LigandMPNN model.
[0019] Note: The above method uses LigandMPNN to specifically optimize the substrate binding region by introducing residues to strengthen 3-OH hydrogen bonds, and combines FastRelax to fine-tune the hydrogen bond network, which greatly improves the substrate recognition specificity.
[0020] Furthermore, the hydrogen bond donor residues include one or more of serine, threonine, and tyrosine; the hydrophobic residues include phenylalanine and / or leucine.
[0021] Explanation: By precisely arranging hydrogen bond donors and hydrophobic residues in a spatial manner, the performance of the enzyme is optimized at the atomic scale. This not only significantly improves regioselectivity and greatly reduces byproduct formation through bidirectional manipulation of polarity and hydrophobicity, but also greatly improves catalytic efficiency by breaking through the catalytic energy barrier in both directions, while enhancing structural stability.
[0022] Furthermore, in S4, the binding sites in the optimized protein three-dimensional structure are determined using any one of LigandMPNN, FastRelax, or FastDesign.
[0023] Note: In the above methods, LigandMPNN improves specificity, FastRelax ensures metal coordination stability, and FastDesign solves stubborn sites, forming a global optimization closed loop covering "electron interaction → spatial matching → long-range stability".
[0024] Furthermore, the protocatechuic acid O-methyltransferase was prepared by recombinant DNA technology combined with Escherichia coli heterologous expression.
[0025] Note: The above methods are commonly used methods for the preparation and synthesis of enzymes. Recombinant DNA technology can precisely manipulate and modify the target gene to achieve targeted regulation and optimization of the enzyme gene. Escherichia coli, as a commonly used heterologous expression host, has advantages such as rapid reproduction, low culture cost, and ease of large-scale fermentation production. It can efficiently obtain protocatechuic acid O-methyltransferase with specific functions and activities, while reducing production costs and improving production efficiency.
[0026] The beneficial effects of this invention are: This invention combines AI with structural biology to construct a bottom-up de novo protein design process (protein backbone design, sequence generation, multi-model activity prediction, and reaction geometry constraints), balancing ligand binding geometry, site selectivity, stability, and activity to obtain OMTs with clear substrate selectivity and catalytic efficiency. The resulting protocatechuic acid O-methyltransferase (OMT) exhibits significantly enhanced activity and catalytic efficiency compared to the wild type; it produces almost no byproducts during the catalytic reaction and demonstrates excellent control over regioselectivity. The AI-designed protocatechuic acid O-methyltransferase not only effectively enhances enzyme activity but also significantly improves the purity of the target product, providing a powerful tool for the efficient biosynthesis of high-value-added products such as vanillin, and possesses broad application prospects and significant economic value in the field of biomanufacturing. Attached Figure Description
[0027] Figure 1 This is one of the template diagrams of the active site of protocatechuic acid O-methyltransferase in the embodiments of the present invention.
[0028] Figure 2 This is the second template diagram of the active site of protocatechuic acid O-methyltransferase in the embodiments of the present invention.
[0029] Figure 3 This is a flowchart illustrating the design of protocatechuic acid O-methyltransferase according to an embodiment of the present invention. Detailed Implementation
[0030] To further illustrate the methods and effects of this invention, the technical solution of this invention will be clearly and completely described below in conjunction with experiments.
[0031] As the background information indicates, current protocatechuic acid O-methyltransfer (OMT) enzymes exhibit poor site selectivity during catalysis, easily leading to multi-site methylation byproducts. Their catalytic efficiency is also unstable due to limitations in the substrate adaptability and conformational flexibility of natural enzymes. Existing methods for OMT modification primarily rely on empirical mutation or site-specific saturation screening. There is a lack of automated processes that systematically integrate AI design with reaction geometry constraints. Furthermore, current AI-driven design methods mainly focus on binding pocket construction or single substrate identification, lacking effective modeling capabilities for complex catalytic reactions (such as those involving multiple electron transfers, synergistic regulation by metal ions and cofactors). Most successful cases are limited to static substrate-binding enzymes or simple hydrolases, and there is no systematic achievement of de novo construction of peroxidase enzymes, multi-step catalytic reactions, or complex transition state-regulated enzymes. Design strategies are often immature in handling situations where metal coordination and cofactors coexist, especially lacking a universal design paradigm in terms of geometric conformation coupling, electronic state regulation, and ligand-co-conformational stability.
[0032] In this invention, AI design can simulate complex environments and predict the interactions between proteins and substrates and cofactors through deep learning models and molecular dynamics simulations, thereby optimizing the geometric conformation and chemical environment of active sites. Through computational tools (such as Rosetta and LigandMPNN), specific amino acid residues or chemical modifications can be introduced to enhance the specific recognition of 3-OH sites and screen enzyme designs that perform best on multiple targets.
[0033] Example 1: Protocatechuic acid O-methyltransferase designed based on artificial intelligence, wherein the amino acid sequence of the protocatechuic acid O-methyltransferase is SEQ ID NO: 1 or SEQ ID NO: 2.
[0034] SEQ ID NO: 1: MEPSVEEAVLAVIERVAEPNNPWAALKAAEALRQIPQIVNFVGPERLDILVTAIKQQKPKKAVVFGGGVGLSSVALAALIGESEKMATIEILEEAAKVVEKIVKLGGV EDKVIPIVGKSQEIVGKLAEELGFDQLDFILFDHWKEEYLEIIENLIKNGNLKKGTTILCDNCLYPGAPKLIDYVQNSKLFTGETIPSKAEWTDTPGGAYVGQFDHN; SEQ ID NO: 2: MGDTKEQRILNYVKETANNKTPESVLKAIDEYEKNWEGGLSMNVGDKKGEVVRAEAARLQERGAPALLELGAYCGYSAVLMADAAGDAKLITIEINPDCAAITRETCKNAGM LDKGRVEVLVGASQDIIRNPPAEWLGLEIDLVFLDHWKDRYLPDTRLLHAMQPRIAVRGILADNVICPGAPGLREGMRELGWRFGEPIPVFLEYRGEDGLEKAITPAAVAR; The distance between the methyl carbon atom of S-adenosylmethionine and the 3-hydroxy oxygen atom of protocatechuic acid in the protocatechuic acid O-methyltransferase is 2.8~3.0 Å. like Figure 1 and Figure 2 As shown, the methyl carbon atom of S-adenosylmethionine is the methyl donor carbon, which is the positive electron center of the nucleophilic substitution reaction; the 3-hydroxy oxygen atom of protocatechuic acid is the nucleophilic attacking atom, which accepts the methyl group from the methyl carbon atom of S-adenosylmethionine to form a transition state C–O bond. The NZ atom at the end of the side chain of the lysine residue (Lys143) acts as a catalytic base, forming a transition state with O1 via hydrogen bonding to assist in deprotonation, with a hydrogen bond distance of approximately 2.8 Å. The active site of protocatechuic acid O-methyltransferase includes Mg 2+ It has an octahedral coordination configuration formed by coordination with six coordinating atoms, a SAM binding site (methyl donor), a catalytic residue (lysine), and an auxiliary residue (glutamic acid). The six coordinating atoms mentioned above are the second carboxyl oxygen atom of the two aspartic acid side chains in the amino acid sequence, the first carbonyl oxygen atom of the asparagine in the amino acid sequence, the 3-hydroxy oxygen atom of the protocatechuic acid, the 4-hydroxy oxygen atom of the protocatechuic acid, and a water molecule. In the above octahedral coordination configuration, the bond length of the coordination bond formed by the coordination center and each coordinating atom is 1.95~2.10 Å, and the coordination angle is 90~180°. The glutamate residue side chain in the above protocatechuic acid O-methyltransferase is connected to the 4-hydroxy oxygen atom of protocatechuic acid by a hydrogen bond, and the O1–CE–S angle in the protocatechuic acid O-methyltransferase is 168~172°, where O1 is the 3-hydroxy oxygen atom of protocatechuic acid, CE is the methyl carbon atom of S-adenosylmethionine, and S is the sulfur atom of S-adenosylmethionine.
[0035] Combination Figure 3 As shown, this invention provides a method for designing enzymes based on artificial intelligence, comprising the following steps: S1. Generate protein backbone structures using the RF diffusion AA model; the input of the RF diffusion AA model is the fragment structure of the active site, as well as the fragment structure of the methyl carbon atom of S-adenosylmethionine and the 3-hydroxy oxygen atom of protocatechuic acid; the output is 100~1000 protein backbone structures. Specifically, first, input a motif containing the aforementioned active sites, including the geometry of the substrates SAM and PCA, and including the following residues: (1) Deprotonated residues: lysine (Lys143); (2) Metal ion coordination residues: aspartic acid (Asp140, Asp168) and asparagine (Asn169). (3) Anchoring substrate localization residues: glutamic acid (Glu198); Except for the residues and their three-dimensional conformations mentioned above, all other amino acids are replaced with alanine to ensure that the RFdiffusion AA model focuses on structural construction around the active site; Parameter settings: 50–80 residues are designed on each side, with a total output structure length of approximately 280 amino acids; Output: Generate 100–1000 backbone structures, requiring that the RMSD deviation between their active sites and the backbone of the motif is less than 1.5 Å.
[0036] In this embodiment of the invention, the RFdiffusion_AA model supports small molecules (such as SAM, PCA, Mg). 2+ The RFdiffusion model is input together with the protein motif and participates in backbone generation. In contrast, the basic version of the RFdiffusion model does not support heterologous ligands and is difficult to use for designing fine catalytic pockets. Therefore, the RFdiffusion_AA version is specially selected in this embodiment of the invention to achieve active center-dominated design based on ternary complex (substrate + coenzyme + metal ion). S2. Using the ProteinMPNN model, generate 10 amino acid sequences for each protein backbone structure obtained in S1. The selection criteria for the above 10 amino acid sequences are: the structure prediction (ESMfold) quality score (pLDDT) is not less than 0.8, ensuring that the protein can fold correctly. Specifically, the residue type probability distribution based on the ProteinMPNN model samples 10 sequences for each backbone; ESMfold rapid folding validation: Using a protein language model, the three-dimensional structure of each sequence is predicted within seconds, and the global pLDDT (predicted Local Distance Difference Test) value is calculated. This value reflects the prediction confidence of each residue. pLDDT ≥ 0.8 indicates that the prediction error of more than 80% of the Cα atoms (referring to the central carbon atom in amino acids that is directly connected to the carbonyl carbon of the main chain, amino nitrogen, and side chain groups) is ≤ 1.5 Å, in order to meet the requirements of high-precision folding. After generating the 10 amino acid sequences, hydrogen bond donor residues were introduced into the amino acid sequences near the 3-hydroxy oxygen atom position of protocatechuic acid, and hydrophobic residues were introduced into the amino acid sequences near the 4-hydroxy oxygen atom position of protocatechuic acid, using the LigandMPNN model. The hydrogen bond donor residues included serine, threonine, and tyrosine; the hydrophobic residues included phenylalanine and leucine. S3. Using the AlphaFold3 model, input the amino acid sequence generated on the protein backbone structure of S2 to obtain the three-dimensional structure of the protein. Specifically, AlphaFold3 is used for structure prediction to evaluate whether the overall protein structure is stable and reasonable, and whether the active site maintains the expected conformation. This is used to eliminate problems such as protein structure collapse and pocket closure, and to ensure that the design results have a structural basis for further verification of reactivity. Protein sequences that meet the criteria are selected. The selection criteria are: the overall structure score (pLDDT) should not be lower than 90, and the key pocket region should also be higher than 80 to ensure structural stability and no distortion of active sites.
[0037] S4. Optimize the binding sites in the three-dimensional structure of the protein to design protocatechuic acid O-methyltransferase; the optimization of the binding sites in the three-dimensional structure of the protein was accomplished using LigandMPNN.
[0038] It should be understood that the binding sites in the three-dimensional structure of the aforementioned protein refer to regions close to the substrate (e.g., A138–A145, A165–A170, A197–A200), which enhance substrate binding efficiency and specificity. Among them, LigandMPNN is suitable for situations where a reasonable backbone has been constructed and a preliminary substrate binding model has been established. The user inputs the ligand (such as SAM, PCA) and the protein structure, and the system will automatically optimize the amino acid sequence in the defined binding site residues to improve the binding affinity and spatial complementarity with the substrate.
[0039] The binding energy between the ligand and the enzyme (≤–10 kcal / mol) was calculated using molecular docking tools (such as AutoDockVina); molecular dynamics simulations were performed on the preferred structure to confirm structural stability and that the binding pocket did not collapse; based on indicators such as binding energy, number of hydrogen bonds, spatial configuration, and structure prediction score, highly active candidate enzymes were finally screened, resulting in the following (SEQ ID NO: 1-SEQ ID NO: 9, excluding SEQ ID NO: 3 which is a natural enzyme) highly active candidate enzymes designed. SEQ ID NO: 1 (Design-03) MEPSVEEAVLAVIERVAEPNNPWAALKAAEALRQIPQIVNFVGPERLDILVTAIKQQKPKKAVVFGGGVGLSSVALAALIGESEKMATIEILEEAAKVVEKIVKLGGV EDKVIPIVGKSQEIVGKLAEELGFDQLDFILFDHWKEEYLEIIENLIKNGNLKKGTTILCDNCLYPGAPKLIDYVQNSKLFTGETIPSKAEWTDTPGGAYVGQFDHN; SEQ ID NO: 2 (Design-06) MGDTKEQRILNYVKETANNKTPESVLKAIDEYEKNWEGGLSMNVGDKKGEVVRAEAARLQERGAPALLELGAYCGYSAVLMADAAGDAKLITIEINPDCAAITRETCKNAGM LDKGRVEVLVGASQDIIRNPPAEWLGLEIDLVFLDHWKDRYLPDTRLLHAMQPRIAVRGILADNVICPGAPGLREGMRELGWRFGEPIPVFLEYRGEDGLEKAITPAAVAR; SEQ ID NO: 3 (HsOMT_WT) MGDTKEQRILNHVLQHAEPGNAQSVLEAIDTYCEQKEWAMNVGDKKGKIVDAVIQEHQPSVLLELGAYCGYSAVRMARLLSPGARLITIEINPDCAAITQRMVDFAGVKDK VTLVVGASQDIIPQLKKKYDVDTLDMVFLDHWKDRYLPDTLLLEECGLLRKGTVLLADNVICPGAPDFLAHVRGSSCFECTHYQSFLEYREVVDGLEKAIYKGPGSEAGP; SEQ ID NO: 4 (Design-01) MPPTLAEKVLAVLREVGPRGDPRAVLEAAAALREAPDPVPVVGPGILAILERVVKESKPKQGLVFGGGVGVSPVLLAALLDEDQQMASVEIDPEAAAVVREAAVAAGGAEGKIVIVGNSQEVVKKLKEELGFDRLDFTYFDHWKEIYREVLEGLIKAGNLAEGTRLLADNALYPGSPELVDYAQESPLFTSEVIPAAEWTDAPGGAFVAVYKGE; SEQ ID NO: 5 (Design-02) MLPTVYKKIVEILKKASKPNDPKQIGEAYKLLKELPEKVPRVYPGITEIIRRMVEENKPKTCVVFGGISPIELAEILGEDERMASIEINPEAAEMSKEVVKLGNAEDKIKFVVGKSQEVVFTLKKDLGIDKVDFYFDHWKDDYKNVIEKLIHEGCLSEGTRIVADNALKPGAPEFVEYAKESPLFTAEEMKAKRERTKEEGGAFTAIAKGA; SEQ ID NO: 6 (Design-04) MEPTRYQALLALVAEAAAPGDPWDVLAAAEELAKTPYPIVVVGPGILEILEKVKENKPKKGVVFGGGAGISPVRLAAILGKDEKMATIEIDPEAAAVLEEAVRLGDAEDKVIPVVGRSQEIVKKLEEELGFKEIDFAFFDHWKEEYLKIIENLIKEGNLKEGTRILADNALEPGAPELIDYMKNSSLFESEVIPSKKEYTDTPGGAVTAIYKGE; SEQ ID NO: 7 (Design-05) MGDTKEQRILDRVKALPAGAGPDSVLAAIDAYGKEHKDTMNVGDKKGEVVREWVAEWQDDLLLELGAYCGYSAVRMAQLLGEGQRLITIEINPDCAAITREMLKLAGME DKVEVLVGASQDIIPKMKWLAKQKSVGVFLDHWKDRYLPDTDLILKACKGCERVTFLADNVICPGAAGFAEHMRRMGCRVESLGELFLEYRRDGLEKAICAMETASKA; SEQ ID NO: 8 (Design-07) MGDTKEQRILAHAREGAAKEGGPESVLKAIDEIDKEVGIMNVGDKKGKILEEWVKESDADVLLELGAYCGYSAVRMAQQLKEGQRLITIEINPDCAAITRETLELAGMLG TAKGVEVLVGASQDIIPKLAERFKAEGWNVEVFLDHWKDRYLPDTKALLELLKDQKEVSFLADNVICPGAQEFLDFAKKLEGAEVELIETFLEYRGKDGLEKAIFKDVTS; SEQ ID NO: 9 (Design-08) MGDTKEQRILATIGGMSPESVLEAIDAAVKELGYAMNVGDKKGKVLAEWAAKSDSKYLLELGAYCGYSAVRMAANLKEGSKLITIEINPDCAAITRQTVEAAGMAD KVTVRVGASQDIIPQLAAEVSISEVFLDHWKDRYLPDTILILNLAKEGQVITFLADNVICPGAEELLQFFRELPLLSRELIKTFLEYREDGLEKAIVLVTPESKL.
[0040] In this embodiment of the invention, protocatechuic acid O-methyltransferase was prepared by recombinant DNA technology combined with Escherichia coli heterologous expression; the preparation method includes steps (1) to (2): Step 1) Based on the amino acid sequence of the designed optimal enzyme (i.e., SEQ ID NO: 1-SEQ ID NO: 9 above), the codon-optimized gene (suitable for E. coli, CAI>0.8) was synthesized, cloned into the pRSF expression vector (containing C-terminal His tag and kanamycin resistance), and transformed into E. coli BL21. After positive clones were verified by sequencing, they were inoculated into 3-M9Y medium containing 50 μg / mL kanamycin (yeast extract 5 g / L, KH2PO4 3 g / L, MgSO4·7H2O 1 mM 0.25 g / L, sodium chloride 0.5 g / L, anhydrous calcium chloride 0.1 mM 0.012 g / L, Na2HPO4·12H2O 15.14 g / L, MOPS 1 g / L, NH4Cl 1 g / L). The medium was incubated at 37℃ and 200 rpm until OD6 = 1.6–1.7. IPTG was then added to a final concentration of 0.2 mM (1 mM stock solution plus 25 μL / 125 mL culture medium), and the medium was simultaneously cooled to 30℃ for induction for 72 hours. The cells were harvested by centrifugation, and the fermentation supernatant was filtered through a 0.22 μm PES membrane for sterilization before being used for product analysis. Step 2) Purify the enzyme obtained above: ProteinIsoRNi-IDAResin column purification step (for His-tag proteins). The target proteins to be purified all contain His-tags, therefore ProteinIsoRNi-IDAResin is used for purification. The purification process mainly consists of five steps: column packing, equilibration, sample loading, washing, and elution; the specific steps are as follows: Ⅰ. Column packing: Pipette 4 mL of resin (loading capacity: 10-20 mg protein / mL wetgel) into the column, let stand, and allow the liquid to flow out completely naturally.
[0041] II. Equilibration: Slowly add 5-10 times (7 times in this example) of Ni column equilibration buffer along the column wall to equilibrate the chromatography column. Let it stand until the liquid flows out completely naturally to complete the equilibration.
[0042] III. Sample loading: Filter the crude enzyme solution obtained from cell disruption through a 0.22 mm filter membrane and slowly add it to the column. Let it stand until the liquid flows out completely naturally.
[0043] IV. Washing: After sample loading, wash the chromatography column with 5-10 times (7 times in this example) of Ni column equilibration buffer.
[0044] V. Elution: Add 2 mL of imidazole solution (column equilibration buffer) to a final concentration of 250 mM to elute the His-tag protein. Collect the eluent, determine the pure enzyme concentration using NanoDrop, and quickly freeze it with liquid nitrogen. Store it in a -80°C freezer and subsequently verify the success of enzyme extraction using SDS-PAGE.
[0045] In step II, the Ni column equilibration buffer consists of: 300mM NaCl, 50mM NaH2PO4, 10mM Mimidazole, 10mM Trisbase, pH=7.0 – soluble protein; in step IV, the Ni column equilibration buffer consists of: 6M GuHCl, 100mM NaH2PO4, 10mM Trisbase, pH 8.0 or 8M Murea, 100mM NaH2PO4, 10mM Trisbase, pH 8.0 – inclusion body protein.
[0046] The enzyme activity assay used a standard reaction system: 20 μg purified enzyme, 2 mM MSAM, 1 mM protocatechuic acid, 20 mM Tris-HCl pH 7.5, reacted at 30℃ for 30 min, and the reaction was terminated by boiling water bath. Vanillic acid was then quantitatively detected by HPLC (C18 column, methanol:water:0.5% phosphoric acid water:acetonitrile = 25:40:25:10, 1 mL / min, UV 260 nm).
[0047] At the structural verification level, the AlphaFold3 predicted structure of the enzyme designed in step S3 showed an overall pLDDT value higher than 85, with no backbone collapse or abnormal active pocket structure. The distance between the C5 methyl donor atom of SAM and the 3-OH atom of PCA was controlled at 2.8–3.0 Å, with an angle close to 180°, forming a geometric conformation consistent with SN2 nucleophilic substitution reaction. Mg 2+ The coordination bond lengths with the six coordinating atoms are 2.0–2.3 Å, maintaining a stable octahedral structure. In the designed structure, the 3-OH region of PCA forms multiple hydrogen bond donor residues (Ser, Thr, Tyr, etc.), while the 4-OH region is sterically repelled, ensuring site selectivity. In summary, this method can reliably construct artificial OMT enzyme candidates with catalytic geometry, structural stability, and ligand synergistic binding, demonstrating good potential for industrial and scientific research applications.
[0048] The specific properties of the obtained protocatechuic acid O-methyltransferase are shown in Table 1. Table 1 shows the performance of the designed protocatechuic acid O-methyltransferase.
[0049] Referring to Table 1, in in vitro enzymatic experiments, this invention designed and expressed eight OMT candidate enzymes (SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4~SEQ ID NO: 9), of which Design-02, Design-03, and Design-06 showed detectable catalytic activity. Among the active candidate enzymes, Design-03 and Design-06 showed significantly enhanced enzyme activity, with the optimal candidate enzyme Design-06 achieving a specific activity of 1.38 U / mg, approximately 15-fold higher than wild-type HsOMT (0.092 U / mg), and the other candidate enzyme Design-03 showing approximately 6-fold increased specific activity. These results demonstrate that the design strategy of this invention can yield variants with high catalytic efficiency. In fermentation validation, the optimal candidate enzyme, under 72-hour fermentation conditions, selectively converted 0.8 g / L protocatechuic acid (PCA) to vanillic acid (VA), achieving a maximum yield of 842.76 mg / L, significantly higher than the known wild-type HsOMT-WT's 445.38 mg / L. Simultaneously, this candidate enzyme produced almost no byproducts such as isopranic acid, demonstrating excellent regioselectivity. Overall, the results indicate that the OMT design method proposed in this invention can significantly improve enzyme activity and target product purity, making it suitable for the efficient biosynthesis of high-value-added products such as vanillin.
[0050] Example 2: The difference from Example 1 is that the optimized protein three-dimensional structure in S4 uses the FastRelax model; the FastRelax model is used to make energy-minimizing adjustments to the protein structure without changing the amino acid sequence, only fine-tuning the backbone and side chain conformation, especially the shape of the binding pocket, hydrophobic arrangement and hydrogen bond network, which can improve the stability of the binding site.
[0051] Example 3: The difference from Example 1 is that the optimized protein three-dimensional structure in S4 adopts the FastDesign model. The FastDesign model is used for more extensive structure and sequence co-optimization. It is suitable for situations where the binding pocket conformation is not ideal or there is a lack of clear ligand sites in the initial model. It can simultaneously sample side chain type, conformation and backbone fine-tuning in multiple iteration rounds to find a better binding site configuration.
Claims
1. A protocatechuic acid O-methyltransferase designed based on artificial intelligence, characterized in that, The amino acid sequence of the protocatechuic acid O-methyltransferase is SEQ ID NO:1 or SEQ ID NO:
2.
2. The protocatechuic acid O-methyltransferase based on artificial intelligence design as described in claim 1, characterized in that, The distance between the methyl carbon atom of S-adenosylmethionine in the protocatechuic acid O-methyltransferase and the 3-hydroxy oxygen atom of protocatechuic acid is 2.8~3.0 Å.
3. The protocatechuic acid O-methyltransferase based on artificial intelligence design as described in claim 1, characterized in that, The active site of the protocatechuic acid O-methyltransferase includes Mg 2+ An octahedral coordination configuration is formed by coordination with six coordinating atoms; the six coordinating atoms are the second carboxyl oxygen atom of the two aspartic acid side chains in the amino acid sequence, the first carbonyl oxygen atom of the asparagine in the amino acid sequence, the 3-hydroxy oxygen atom of the protocatechuic acid, the 4-hydroxy oxygen atom of the protocatechuic acid, and a water molecule.
4. The protocatechuic acid O-methyltransferase based on artificial intelligence design as described in claim 3, characterized in that, In the octahedral coordination configuration, the bond length of the coordinate bond formed by the coordination center and each coordinating atom is 1.95~2.10 Å, and the coordination angle is 90~180°.
5. The protocatechuic acid O-methyltransferase based on artificial intelligence design as described in claim 3, characterized in that, In the protocatechuic acid O-methyltransferase, the carboxyl oxygen atom of the glutamate residue side chain forms a hydrogen bond with the 4-hydroxy oxygen atom of protocatechuic acid, and the O1–CE–S angle in the protocatechuic acid O-methyltransferase is 168~172°, where O1 is the 3-hydroxy oxygen atom of protocatechuic acid, CE is the methyl carbon atom of S-adenosylmethionine, and S is the sulfur atom of S-adenosylmethionine.
6. A method for designing enzymes based on artificial intelligence, used to design protocatechuic acid O-methyltransferase as described in any one of claims 3 to 5, characterized in that, Includes the following steps: S1. Generate protein backbone structures using the RF diffusion AA model; the input of the RF diffusion AA model is the fragment structure of the active site, as well as the fragment structure of the methyl carbon atom of S-adenosylmethionine and the 3-hydroxy oxygen atom of protocatechuic acid; the output is 100~1000 protein backbone structures. S2. Using the ProteinMPNN model, generate 10 amino acid sequences on each protein backbone structure obtained in S1; S3. Using the AlphaFold3 model, input the amino acid sequence generated on the protein backbone structure of S2 to obtain the three-dimensional structure of the protein. S4. By optimizing the binding sites in the three-dimensional structure of the protein, protocatechuic acid O-methyltransferase was designed.
7. The method for designing enzymes based on artificial intelligence as described in claim 6, characterized in that, In S2, after generating 10 amino acid sequences, hydrogen bond donor residues are introduced into the amino acid sequences near the 3-hydroxy oxygen atom position of protocatechuic acid, and hydrophobic residues are introduced into the amino acid sequences near the 4-hydroxy oxygen atom position of protocatechuic acid, using the LigandMPNN model.
8. The method for designing enzymes based on artificial intelligence as described in claim 7, characterized in that, The hydrogen bond donor residues include one or more of serine, threonine, and tyrosine; the hydrophobic residues include phenylalanine and / or leucine.
9. The method for designing enzymes based on artificial intelligence as described in claim 6, characterized in that, In S4, the binding sites in the optimized protein three-dimensional structure are determined using any one of LigandMPNN, FastRelax, or FastDesign.
10. The protocatechuic acid O-methyltransferase according to any one of claims 1 to 5, characterized in that, The protocatechuic acid O-methyltransferase was prepared by recombinant DNA technology combined with Escherichia coli heterologous expression.