A method for accurate prediction of substrate protein hydrolysis sites in microbial fermentation processes

By combining molecular docking and molecular dynamics simulations with whole-genome data and modeling software, the problem of inaccurate prediction of substrate protein hydrolysis sites during microbial fermentation was solved, enabling accurate prediction and efficient screening of enzyme-substrate binding sites.

CN120126556BActive Publication Date: 2025-10-21OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510298179.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-10-21
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately predict the hydrolysis sites of substrate proteins during microbial fermentation, resulting in inaccurate results in subsequent experimental analysis.

Method used

By targeting and screening proteases in the microbial fermentation process, molecular docking and molecular dynamics simulations, combined with whole genome data and modeling software, were used to establish three-dimensional structural models of the proteases and perform molecular dynamics simulations to predict the binding sites between the enzyme and the substrate protein.

Benefits of technology

It enables accurate prediction of enzyme and substrate protein binding sites during microbial fermentation, improves the efficiency and accuracy of protease mining in fermentation agents, and targets the screening of microbial strains that meet actual production needs, thus narrowing the scope of experimental screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126556B_ABST
    Figure CN120126556B_ABST
Patent Text Reader

Abstract

The application provides a method for accurately predicting a hydrolysis site of a substrate protein in a microbial fermentation process, and belongs to the technical field of molecular biology. The method comprises the following steps: screening a protease secreted by a target microorganism based on whole genome data and using bioinformatics means; establishing a three-dimensional structure model of the protease secreted by the target microorganism; performing molecular docking of the protease secreted by the target microorganism and a substrate protein to obtain an optimal docking conformation; studying the interaction relationship between the protease secreted by the target microorganism and the substrate protein under fermentation conditions by using molecular dynamics simulation; and analyzing the molecular docking and molecular dynamics simulation results to predict the hydrolysis site of the substrate protein in the microbial fermentation process. The method can accurately predict the site of the enzyme and the substrate protein in the microbial fermentation process by using molecular docking and molecular dynamics simulation, and can analyze the polypeptide produced by the substrate protein after fermentation, so that the microbial protease hydrolyzing specific polypeptides can be targeted and positioned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and in particular relates to a method for accurately predicting substrate protein hydrolysis sites in a microbial fermentation process. Background Art

[0002] Fermentation is the process of transforming organic matter into products through the metabolic activity of microorganisms using biocatalysts (microbial cells or enzymes). Microbial fermentation is a complex process, primarily relying on the enzymes carried by the microorganisms to hydrolyze proteins, fats, carbohydrates, and other substances into small molecules, producing unique flavors. Proteases are the most notable of these substances. Fermentation is a traditional method of food processing and preservation. Fermentation not only aids in the digestion and absorption of food but also, to a certain extent, breaks down antigenic food proteins or denatures them into small peptides and amino acids, thereby destroying certain antigenic epitopes and reducing allergenicity. Fermentation also has the advantages of improving product texture, flavor, nutritional value, and stability.

[0003] However, due to the complexity of microbial fermentation, the multiple enzymes it produces act in combination on the same protein, which makes it difficult for us to accurately discover the fermentation site of the substrate protein. We can only roughly judge the fermentation site by the composition of the polypeptide after protein fermentation. Therefore, this method will make the site results we obtain for the fermentation of its substrate protein inaccurate, which will have a relatively large impact on subsequent experimental analysis.

[0004] Traditional molecular biology experimental methods, such as X-ray crystallography and nuclear magnetic resonance, provide powerful support for obtaining microscopic details of complex protein structures. However, many protein functions are regulated by specific biomolecular dynamic processes, such as protein conformational changes, protein folding / unfolding, and protein activation, and traditional biological methods are unable to accurately describe the dynamic behaviors of these proteins. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method for accurately predicting the hydrolysis sites of substrate proteins during microbial fermentation. By targeted screening of proteases in the microbial fermentation process and combining them with target protein substrates using molecular docking and molecular dynamics simulation, accurate prediction of the binding sites between enzymes and substrate proteins during microbial fermentation can be achieved.

[0006] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0007] The present invention provides a method for accurately predicting substrate protein hydrolysis sites during microbial fermentation, comprising the following steps:

[0008] 1) The whole genome data of the target microorganism is transcribed and translated into a protein amino acid sequence by bioinformatics means; the protein amino acid sequence is compared with known homologous microbial proteases to screen for proteases secreted by the target microorganism;

[0009] 2) using modeling software to establish a three-dimensional structural model of the protease secreted by the target microorganism;

[0010] 3) Based on the structure of the substrate protein, amino acid residues that may be binding sites are selected, and the protease secreted by the target microorganism is molecularly docked with the substrate protein to obtain the optimal docking conformation;

[0011] 4) Molecular dynamics simulations were used to study the interactions between proteases secreted by target microorganisms and substrate proteins under fermentation conditions;

[0012] 5) Analyze the results of molecular docking and molecular dynamics simulation to predict the hydrolysis sites of substrate proteins during microbial fermentation.

[0013] Preferably, the whole genome data in step 1) is derived from whole genome sequencing or the NCBI database.

[0014] Preferably, the modeling software in step 2) includes AlphaFold 2; before establishing the three-dimensional structural model of the protease secreted by the target microorganism, signal peptide prediction and structural domain division are performed on the protease secreted by the target microorganism.

[0015] Preferably, the signal peptide prediction tool includes SignalP v4.0; the domain division tool includes NCBI Conserved Domain Search.

[0016] Preferably, the molecular docking tool in step 3) includes Discovery Studio; the ZDock module in the Discovery Studio performs simulated docking of the catalytic domain of the protease secreted by the target microorganism and the substrate protein, and the RDock module in the Discovery Studio optimizes the docking configuration.

[0017] Preferably, the program settings of the ZDock module include: defining the PR domain of the protease secreted by the target microorganism as a receptor, defining the substrate protein as a ligand, setting the amino acid residues that may be binding sites as Binding Site Residues, and setting the Angular Step Size to 6°;

[0018] The configuration force field of the RDock module includes the CHARMM force field, and the RDock module extracts toprefinedpose as the final docking configuration after optimization.

[0019] Preferably, the molecular dynamics simulation in step 4) comprises the following steps:

[0020] A: The catalytic domain-substrate protein complex of the protease secreted by the target microorganism is generated using a topology parameter generation program, and the force field is configured to generate the topology parameters of the catalytic domain-substrate protein complex of the protease secreted by the target microorganism;

[0021] B: Based on the topological parameters of the catalytic PR domain-substrate protein complex of the protease secreted by the target microorganism, a simulation box is set up, a solvent model is added, and energy minimization is performed after adding counterions. Then, ensemble equilibrium is performed to obtain unified environmental conditions; the simulation time is set, and molecular dynamics simulation is performed under the said unified environmental conditions.

[0022] Preferably, the topological parameter generation program includes GROMACS; the force field includes AMBER14SB force field; the type of the simulation box includes a cubic box, and the distance between the simulation box and the solute is 0.5 to 1.5 nm; the solvent model includes the TIP3P water model; the counter ion includes Na + or Cl - ; The energy minimization method includes the steepest descent method; the ensemble equilibrium includes NPT and NVT ensemble equilibrium; the simulation time is 90 to 110 ns.

[0023] Preferably, the tool for obtaining the pH conditions of the molecular dynamics simulation includes a PDB2PQR server, and the temperature conditions of the molecular dynamics simulation are derived from the temperature required for the actual fermentation of the target microorganism.

[0024] Preferably, the analysis indicators in step 5) include root mean square deviation, root mean square fluctuation, radius of gyration, accessible surface area, number of hydrogen bonds and binding free energy.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] The present invention utilizes molecular docking and molecular dynamics simulation to accurately predict the binding sites between enzymes and substrate proteins during microbial fermentation. At the same time, it analyzes the production of polypeptides from protein substrates after fermentation, thereby enabling the targeted positioning of microbial proteases that hydrolyze specific polypeptides.

[0027] The present invention can also analyze the polypeptides hydrolyzed by the screened microbial proteases in combination with different protein substrates to achieve a dual screening purpose, that is, it can screen out the required microorganisms through the protein substrates, and can also screen out the required protein substrates through the microorganisms.

[0028] The present invention uses whole genome analysis in conjunction with databases such as NCBI to achieve preliminary screening of microorganisms at the molecular level. It can quickly and accurately analyze the protease resources of fermentation agents, improve the efficiency and accuracy of mining high-quality fermentation agents and their proteases, and target the screening of microbial strains that meet actual production needs, reducing the scope and workload of experimental screening, laying the foundation for subsequent related research. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a three-dimensional structural model of Lactobacillus helveticus Lh191404 cell wall protease and parvalbumin (wherein, A is Lactobacillus helveticus Lh191404 cell wall protease; B is parvalbumin);

[0030] Figure 2 is the RMSD value of the catalytic (PR) domain-parvalbumin complex of CEP during the 100 ns molecular dynamics simulation;

[0031] Figure 3 is the total binding free energy of the CEP catalytic (PR) domain-parvalbumin complex during a 100 ns molecular dynamics simulation and the main forces involved in its binding (where A is the total binding free energy and B is the binding free energy of each force);

[0032] Figure 4 Figure 3 is a graph showing the change in the number of hydrogen bonds in the CEP catalytic (PR) domain-parvalbumin complex during a 100 ns molecular dynamics simulation. DETAILED DESCRIPTION

[0033] The present invention provides a method for accurately predicting substrate protein hydrolysis sites during microbial fermentation, comprising the following steps:

[0034] 1) The whole genome data of the target microorganism is transcribed and translated into a protein amino acid sequence by bioinformatics means; the protein amino acid sequence is compared with known homologous microbial proteases to screen for proteases secreted by the target microorganism;

[0035] 2) using modeling software to establish a three-dimensional structural model of the protease secreted by the target microorganism;

[0036] 3) Based on the structure of the substrate protein, amino acid residues that may be binding sites are selected, and the protease secreted by the target microorganism is molecularly docked with the substrate protein to obtain the optimal docking conformation;

[0037] 4) Molecular dynamics simulations were used to study the interactions between proteases secreted by target microorganisms and substrate proteins under fermentation conditions;

[0038] 5) Analyze the results of molecular docking and molecular dynamics simulation to predict the hydrolysis sites of substrate proteins during microbial fermentation.

[0039] In this study, the whole genome data of the target microorganism was transcribed and translated into protein amino acid sequences using bioinformatics methods. These protein amino acid sequences were then compared with known homologous microbial proteases to screen for proteases secreted by the target microorganism. The whole genome data was obtained from whole genome sequencing or the NCBI database (https: / / www.ncbi.nlm.nih.gov / ).

[0040] In the present invention, a three-dimensional structural model of the protease secreted by the target microorganism is established using modeling software. Before establishing the three-dimensional structural model of the protease secreted by the target microorganism, signal peptide prediction and domain segmentation are performed on the protease secreted by the target microorganism; the signal peptide prediction tool includes SignalP v4.0; the domain segmentation tool includes NCBI Conserved Domain Search; and the modeling software includes AlphaFold 2.

[0041] In the present invention, based on the structure of the substrate protein, amino acid residues that may be binding sites are selected, and molecular docking is performed between the protease secreted by the target microorganism and the substrate protein to obtain the optimal docking conformation. The amino acid sequence of the substrate protein is obtained from the Uniprot database; the three-dimensional structure of the substrate protein is modeled by AlphaFold 2; the molecular docking tool includes Discovery Studio; the ZDock module in Discovery Studio performs a simulated docking between the catalytic (PR) domain of the protease secreted by the target microorganism and the substrate protein. The program settings of the ZDock module include: defining the catalytic domain of the protease secreted by the target microorganism as a receptor, defining the substrate protein as a ligand, setting the amino acid residues that may be binding sites as Binding Site Residues, and setting the Angular StepSize to 6°; the RDock module in Discovery Studio optimizes the docking conformation, and the configuration force field of the RDock module includes the CHARMM force field. After optimization, the RDock module extracts the top refined pose as the final docking conformation.

[0042] In the present invention, molecular dynamics simulation is used to study the interaction between the protease secreted by the target microorganism and the substrate protein under fermentation conditions. The molecular dynamics simulation includes the following steps:

[0043] A: The catalytic domain-substrate protein complex of the protease secreted by the target microorganism is generated using a topology parameter generation program, and the force field is configured to generate the topology parameters of the catalytic domain-substrate protein complex of the protease secreted by the target microorganism;

[0044] B: Based on the topological parameters of the catalytic domain-substrate protein complex of the protease secreted by the target microorganism, a simulation box is set up, a solvent model is added, and energy minimization is performed after adding counterions. Then, ensemble equilibrium is performed to obtain unified environmental conditions; the simulation time is set, and molecular dynamics simulation is performed under the said unified environmental conditions.

[0045] In the present invention, the topology parameter generation program includes GROMACS, and the PDB file of the catalytic domain-substrate protein complex of the protease secreted by the target microorganism is imported into GROMACS; the force field includes the AMBER14SB force field; the type of the simulation box includes a cubic box, and the size of the simulation box is defined by editconf; the distance between the simulation box and the solute is preferably 0.5 to 1.5 nm, more preferably 1 nm; the solvent model includes the TIP3P water model; the counterion includes Na + or Cl - , the counter ions are added through the genion module, and the counter ions are added to make the system electrically neutral; the energy minimization method includes the steepest descent method; the ensemble equilibrium includes NPT and NVT ensemble equilibrium, and the ensemble equilibrium is run through the mdrun module; the duration of the NPT ensemble equilibrium is preferably 0.5-1.5ns, and more preferably 1ns; the temperature of the NPT ensemble equilibrium is preferably 310-320K, and more preferably 313.15K; the duration of the NVT ensemble equilibrium is preferably 2-4ns, and more preferably 3ns; the pressure of the NVT ensemble equilibrium is preferably 0.5-1.5bar, and more preferably 1bar; the simulation duration is preferably 90-110ns, and more preferably 80ns. The tool for obtaining the pH conditions of the molecular dynamics simulation includes the PDB2PQR server (http: / / agave.wustl.edu / pdb2pqr / ), which assigns the protonation state of a single charged residue of the catalytic domain-substrate protein complex of the protease secreted by the target microorganism. The temperature conditions of the molecular dynamics simulation are derived from the temperature required for the actual fermentation of the target microorganism.

[0046] In this study, molecular docking and molecular dynamics simulation results were analyzed to predict the hydrolysis sites of substrate proteins during microbial fermentation. The analysis metrics included root mean square deviation (RMS), RMS fluctuation, gyration radius, accessible surface area, number of hydrogen bonds, and binding free energy; these RMS, RMS fluctuation, gyration radius, accessible surface area, and number of hydrogen bonds were calculated based on the simulation trajectories; the binding free energy was analyzed using the MMPBSA method developed by Valdés-Tresanco et al.

[0047] The technical solutions provided by the present invention are described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.

[0048] Example 1

[0049] 1. The whole genome data of Lactobacillus helveticus Lh191404 was transcribed and translated into protein amino acid sequences using bioinformatics methods; the protein amino acid sequences were compared with known homologous microbial proteases to screen for cell wall proteases (CEPs) that can target and hydrolyze Atlantic cod parvalbumin.

[0050] 2. CEP belongs to the PrtH1 variant (PrtH1 V) genotype. The three-dimensional structure of CEP was characterized and the three-dimensional structure of the protein was determined by dividing the domains. The domain composition distribution of PrtH1_V-191404 was determined by combining the prediction results of the signal peptide of SignalP v4.0 and NCBI Conserved Domain Search. Finally, AlphaFold 2 was used to model and analyze each domain, and the three-dimensional structural model of Lactobacillus helveticus Lh191404 cell wall protease was obtained (such as Figure 1 (as shown in A in the figure).

[0051] After obtaining the three-dimensional structure of the cell wall proteinase, the amino acid sequence of the Atlantic cod parvalbumin allergen Gad m 1 (SEQ ID: Q90YK9) was obtained using the Uniprot database. The three-dimensional structure of the Atlantic cod parvalbumin allergen Gad m 1 (PDB ID 2MBX) was obtained using the PDB database obtained by AlphaFold2 modeling. Figure 1 As shown in B.

[0052] 3. The complete sequence of parvalbumin (PV) was analyzed using four immunoinformatics-based computational methods: BepiPred v2.0 (http: / / www.cbs.dtu.dk / services / BepiPred / ), IEDBElipro (http: / / tools.iedb.org / ellipro), ABCPred (http: / / www.imtech.res.in / raghava / abcpred / dataset.html), and DNASTAR (DNAS-TAR, Madison, WI). BepiPred 2.0 comprehensively assesses the likelihood of each amino acid as a B cell epitope by considering factors such as its physicochemical properties, hydrophilicity, and charge. IEDB Elipro and DNASTAR analyze the complete amino acid sequence to predict antigenic epitopes based on hydrophilicity, surface accessibility, and antigenicity index. Regions with the highest overall scores were selected as candidate B cell epitopes. In DiscoTope v2.0, the number of antigen-antibody contacts and amino acid tendencies are calculated to derive potential conformational epitopes of allergens. A novel spatial accessibility algorithm is used to estimate the possibility of surface exposure, thereby increasing the accuracy of the prediction. In ABCPred, advanced machine learning algorithms are utilized to conduct in-depth analysis of the amino acid sequences of allergens to identify amino acid fragments that may become B cell epitopes. The screening overlap rate is ≥75%, that is, at least 3 software predict consistent amino acid sequences as the final B cell linear epitope sequence of Gad m 1. It can be concluded that the B cell epitope of parvalbumin is 5 GILNDADIT 13 、 19 CKAEGSFD 26 、 52 DQDKSDFVEED 62 、 72 SAGA 75 、 79 SDA 81 、 90 DSDGDGKIGVDEFGAMIKA 109 .

[0053] Random coil regions on the protein surface have good hydrophilicity, flexibility, and antigenicity index, which can be used as reference indicators for B cell epitope prediction. Therefore, we selected five highly flexible regions of parvalbumin, located at AA1-8, AA19-26, AA52-60, AA72-79, and AA91-99.

[0054] In step 2, the structural models of each domain of CEP have been characterized by AlphaFold 2. The ZDock module in Discovery Studio v2019 was used to perform molecular docking of the catalytic (PR) domain of CEP and parvalbumin. The PR domain of CEP was defined as a receptor, parvalbumin was defined as a ligand, and the five peptide segments (AA1-8, AA19-26, AA52-60, AA72-79, and AA91-99) were set as Binding Site Residues. In addition, to ensure the comprehensiveness of the docking, we set these five regions as a whole as Binding Site Residues, set the Angular Step Size to 6°, and set the other options to default. Molecular simulation docking was performed one by one. After the docking was completed, the RDock module was used to optimize the docking configuration based on the CHARMM force field, and the top refined pose was extracted as the final docking configuration.

[0055] 4. After obtaining the optimal docking conformation, molecular dynamics simulations were performed at 40°C at pH 6.5 using the Gromacs 2022.5 software package. In order to obtain the required pH conditions, the protonation states of the single charged residues of the catalytic (PR) domain-parvalbumin complex of CEP were assigned according to the PDB2PQR server (http: / / agave.wustl.edu / pdb2pqr / ), and the temperature conditions were derived from the temperature required for the actual target microbial fermentation. The PDB file of the PR domain-parvalbumin complex of CEP was imported into GROMACS, and the AMBER14SB force field was used to generate the topological parameters of the PR domain-parvalbumin complex of CEP. The editconf module was used to define the size of the box, the distance between the solute and the box was 1.0 nm, and the box type was a cubic box. The solvate module was used to add solvent to the box, and the solvent model adopted the TIP3P water model. The genion module was used to add counterions (Na + or Cl - ) to neutralize the system. The grompp module was used to minimize the system, specifying the steepest descent method and 50,000 iterations. A 1-ns canonical ensemble pre-equilibration (NVT) run was run using the mdrun module to stabilize the system's temperature to 300 K. A 3-ns isothermal and isobaric ensemble pre-equilibration (NPT) run was run using the mdrun module to stabilize the system's pressure to 1 bar. Finally, a 100-ns molecular dynamics simulation of the entire system was performed using the mdrun module.

[0056] 5. Based on the simulation trajectories, we calculated the root mean square deviation (RMSD), root mean square fluctuation (RMSF), radius of gyration (Rg), accessible surface area (SASA), binding free energy, and number of hydrogen bonds. The binding free energy of the interaction between the CEP catalytic domain and parvalbumin was analyzed using the MMPBSA method developed by Valdés-Tresanco et al.

[0057] Experimental results: Figure 2 、 Figure 3 、 Figure 4 and as shown in Table 1.

[0058] The root mean square deviation (RMSD) value can reflect the stability of the system during the MD simulation process and is used to evaluate the structural stability of the complex between the catalytic domain of the cell wall protease and the parvalbumin. Figure 2 As can be seen, the RMSD of the complexes all reached equilibrium after 100 ns of MD simulation. The complex with designated docking sites AA72-79 experienced fluctuations at 53 ns but quickly stabilized at around 0.55 nm. The complex with designated docking sites encompassing all five regions reached equilibrium around 40 ns, with the smallest RMSD value of approximately 0.32 nm. The remaining RMSD values ​​stabilized around 0.46 nm, 0.40 nm, 0.35 nm, and 0.33 nm, respectively. These complexes all experienced significant fluctuations in the early stages of the simulation, likely due to conformational changes between the catalytic domain and parvalbumin to achieve optimal binding.

[0059] The MMPBSA method was used to analyze the binding free energy and the main forces involved in the interaction between the cell wall protease catalytic domain and parvalbumin, such as Figure 3 The average binding free energies of these six complexes with different docking sites during MD simulations were -49.59 kcal / mol, -19.59 kcal / mol, -42.14 kcal / mol, -42.36 kcal / mol, -51.08 kcal / mol, and -41.18 kcal / mol, respectively. The complex with docking site AA91-99 had the lowest overall binding free energy. Lower binding free energies indicate more stable binding between the muramase and parvalbumin.

[0060] Figure 3Graph B shows the contribution of each energy term to the total binding free energy. Polar solvation energy (Egb) negatively impacts the stability of the complex, while van der Waals interaction energy (Evdw), electrostatic interaction energy (Eele), and nonpolar solvation energy (Esurf) contribute positively to the stabilization of the complex, with electrostatic interaction energy making the largest contribution. Overall, the energy change in the gas phase (ΔGgas) contributes positively to the total binding free energy, while solvation free energy (ΔGsolv) has a negative effect.

[0061] Hydrogen bonds play an important role in intermolecular interactions. Generally speaking, the specific recognition of hydrogen bonds plays a key role in the binding process between proteases and substrates. Figure 4 The results revealed changes in the number of hydrogen bonds between the muramase catalytic domain and parvalbumin. The number of hydrogen bonds in these complexes with different docking conformations stabilized in the later stages of the simulation. Notably, the number of hydrogen bonds in the complex with the entire flexible region as the docking site fluctuated in the early stages of the simulation, presumably due to conformational changes in the complex, leading to the breaking or forming of hydrogen bonds.

[0062] Table 1 Hydrogen bonding between microbial proteases and parvalbumin simulated by molecular dynamics

[0063]

[0064]

[0065]

[0066]

[0067] The number and distance of hydrogen bonds between the complexes and whether they exist in the cleavage site of fermentation or enzymatic hydrolysis are analyzed in Table 1. In the complex with docking site AA1-8, the catalytic domain formed 11 hydrogen bonds with 8 amino acid residues of parvalbumin (Ala4, Glu82, Ala2, Lys65, Gln69, Ala73, Arg76 and Ala77), with hydrogen bond distances ranging from The average hydrogen bond distance is The complexes with docking sites of AA19-26 and AA52-60 formed 7 and 5 hydrogen bonds, respectively, with average hydrogen bond distances of and In the complex with docking site AA72-79, the catalytic domain is in close contact with nine residues of parvalbumin (Ala4, Phe3, Asp80, Gln53, Glu49, Asp52, Asp43, Lys39, and Gln69) through 11 hydrogen bonds, with hydrogen bond distances ranging from The average hydrogen bond distance is The complex docking site is the catalytic domain of AA91-99, which forms 12 hydrogen bonds with 8 amino acid residues of parvalbumin (Ala42, Glh49, Gln53, Lys46, Ash57, Val59, Asp52 and Lys55). The hydrogen bond distance range is The average hydrogen bond distance is In the complex where the docking site is located in the entire flexible region, the protease binds to eight residues of parvalbumin, namely Ala4, Lys55, Ash95, Ash54, Glh60, Gly74, Ala77, and Arg76, through 11 hydrogen bonds with hydrogen bond distances ranging from The average hydrogen bond distance is For short distances (less than ) plays a major role in stabilizing the complex structure. The interactions among these complexes in the present invention are mostly short-range hydrogen bonds, and the complexes with docking sites AA1-8, AA72-79, AA91-99, and the entire flexible region have more hydrogen bonds than the AA19-26 and AA52-60 complexes, indicating that the binding of these four complexes is more stable.

[0068] The MD simulation results above reveal differences in structural stability, solvent-accessible surface area, and hydrogen bonding among complexes with different docking sites. However, it is clear that complexes with these five highly flexible regions as docking sites exhibit a more stable overall conformation, and the amino acid residues involved in hydrogen bonding are closer to the actual cleavage site. Therefore, using all flexible regions of a protein as docking sites for molecular dynamics simulations is a reasonable and reliable approach.

[0069] As can be seen from the above examples, the present invention utilizes molecular docking and molecular dynamics simulation to accurately predict the binding sites between enzymes and substrate proteins during microbial fermentation, and simultaneously analyzes the production of polypeptides from protein substrates after fermentation, thereby enabling the targeted location of microbial proteases that hydrolyze specific polypeptides.

[0070] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for accurately predicting substrate protein hydrolysis sites during microbial fermentation, characterized in that: The following steps are involved: 1) The whole genome data of the target microorganism is transcribed and translated into protein amino acid sequences using bioinformatics methods; the protein amino acid sequences are compared with known homologous microbial proteases to screen for proteases secreted by the target microorganism; 2) using modeling software to establish a three-dimensional structural model of the protease secreted by the target microorganism; 3) Based on the structure of the substrate protein, amino acid residues that may be binding sites are selected, and the protease secreted by the target microorganism is molecularly docked with the substrate protein to obtain the optimal docking conformation; 4) Molecular dynamics simulations were used to investigate the interactions between proteases secreted by the target microorganisms and substrate proteins under fermentation conditions; 5) Analyze the results of molecular docking and molecular dynamics simulation to predict the hydrolysis sites of substrate proteins during microbial fermentation; The target microorganism is Lactobacillus helveticus Lh191404, the protease is cell wall protease, and the substrate protein is Atlantic cod parvalbumin allergen Gad m 1; The PR domain of the cell wall protease was defined as the receptor, the Atlantic cod parvalbumin allergen Gad m 1 was defined as the ligand, and the five peptide segments AA1-8, AA19-26, AA52-60, AA72-79, and AA91-99 were set as Binding Site Residues, respectively, and the Angular Step Size was set to 6°.

2. The method according to claim 1, characterized in that Step 1) The whole genome data is derived from whole genome sequencing or NCBI database.

3. The method according to claim 1, characterized in that Step 2) The modeling software includes AlphaFold 2; before establishing the three-dimensional structural model of the protease secreted by the target microorganism, signal peptide prediction and structural domain division are performed on the protease secreted by the target microorganism.

4. The method according to claim 3, characterized in that The signal peptide prediction tool includes SignalP v4.0; the domain division tool includes NCBI Conserved Domain Search.

5. The method according to claim 1, wherein Step 3) The molecular docking tool includes Discovery Studio; the ZDock module in the Discovery Studio performs simulated docking between the catalytic domain of the protease secreted by the target microorganism and the substrate protein, and the RDock module in the Discovery Studio optimizes the docking configuration.

6. The method according to claim 5, characterized in that The program setting of the ZDock module includes: the configuration force field of the RDock module includes the CHARMM force field, and the top refined pose is extracted as the final docking configuration after the RDock module is optimized.

7. The method according to claim 1, characterized in that Step 4) The molecular dynamics simulation includes the following steps: A: The catalytic domain-substrate protein complex of the protease secreted by the target microorganism is generated using a topology parameter generation program, and the force field is configured to generate the topology parameters of the catalytic domain-substrate protein complex of the protease secreted by the target microorganism; B: Based on the topological parameters of the catalytic domain-substrate protein complex of the protease secreted by the target microorganism, a simulation box is set up, a solvent model is added, and energy minimization is performed after adding counterions. Then, ensemble equilibrium is performed to obtain unified environmental conditions; the simulation time is set, and molecular dynamics simulation is performed under the said unified environmental conditions.

8. The method according to claim 7, characterized in that The topological parameter generation program includes GROMACS; the force field includes AMBER14SB force field; the type of the simulation box includes a cubic box, and the distance between the simulation box and the solute is 0.5~1.5 nm; the solvent model includes the TIP3P water model; the counter ion includes Na + or Cl - ; The energy minimization method includes the steepest descent method; the ensemble equilibrium includes NPT and NVT ensemble equilibrium; the simulation time is 90~110ns.

9. The method according to claim 1 or 7, characterized in that The tool for obtaining the pH conditions of the molecular dynamics simulation includes a PDB2PQR server, and the temperature conditions of the molecular dynamics simulation are derived from the temperature required for the actual fermentation of the target microorganism.

10. The method according to claim 1, characterized in that Step 5) The analysis indicators include root mean square deviation, root mean square fluctuation, radius of gyration, accessible surface area, number of hydrogen bonds and binding free energy.

Citation Information

Patent Citations

  • Method for reducing sensitization of beta-lactoglobulin by combining targeted positioning with enzyme hydrolysis

    CN110531060A

  • Method for improving hydrolase robustness in combination with high-pressure molecular dynamics simulation and free energy calculation

    CN112582031A