Enzyme activity detection-based tannase catalytic efficiency evaluation method

By optimizing the enzyme active site through molecular dynamics simulations and nonlinear regression models, the problem of difficulty in assessing the impact of substrate structure changes on tannin catalytic behavior in existing technologies has been solved, enabling dynamic adjustment of the enzyme active site and improvement of catalytic efficiency.

CN120877863AInactive Publication Date: 2025-10-31GUANGZHOU VOCATIONAL & TECH COLLEGE OF HEALTH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510978824.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing methods are insufficient to fully reveal the profound impact of substrate structure changes on the catalytic behavior of tanninases. In particular, the fluctuations in spatial configuration adaptability, the number of hydrogen bonds formed, and the strength of hydrophobic interactions caused by changes in substrate side chain length result in insufficient dynamic adjustment ability of enzyme active sites, affecting the accuracy and adaptability of catalytic efficiency assessment.

Method used

The interaction trajectories between the enzyme active site and gallic esters with different side chain lengths were obtained by molecular dynamics simulation. The conformational matching degree, hydrogen bond dynamics and hydrophobic interaction strength were calculated. A nonlinear regression model was constructed to optimize the enzyme active site structure and establish a catalytic efficiency evaluation model.

Benefits of technology

This improved the binding efficiency of gallic esters to tanninases, optimized the substrate design and selection for enzyme-catalyzed reactions, and enhanced the accuracy and sensitivity of catalytic efficiency assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877863A_ABST
    Figure CN120877863A_ABST
Patent Text Reader

Abstract

The invention provides a tannase catalytic efficiency evaluation method based on enzyme activity detection, which comprises the following steps: analyzing the influence of side chain length distribution on hydrogen bond formation quantity and hydrogen bond fracture frequency according to a conformation matching degree parameter, extracting a conformation energy difference and a dynamic free energy change, and determining a hydrogen bond dynamic characteristic parameter; if the hydrogen bond fracture frequency in the hydrogen bond dynamic characteristic parameters exceeds a preset threshold value, hydrophobic interaction fluctuation is analyzed based on a molecular dynamics trajectory, and hydrophobic interaction intensity distribution is determined; analyzing the Spearman correlation between the dynamic adaptation capability parameter and the binding efficiency predicted value, generating the weight distribution of the side chain length distribution to the binding rate constant, and determining a catalytic efficiency evaluation model; the method comprises the following steps: predicting binding rate constants of various gallate structures, and sequencing side chain length distribution and conformation matching degrees according to the predicted binding rate constants to obtain a substrate selection scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method for evaluating the catalytic efficiency of tannins based on enzyme activity detection. Background Technology

[0002] Tanninases, as an important research subject in the field of biocatalysis, have key applications in food, pharmaceutical, and industrial production, and their catalytic efficiency directly affects the performance and output of related processes. Tanninases can decompose complex tannin compounds and regulate substrate degradation processes, thereby achieving efficient biotransformation. However, existing methods for evaluating the catalytic efficiency of tanninases based on enzyme activity detection have significant limitations, failing to fully reveal the profound impact of substrate structural changes on enzyme catalytic behavior. Current methods mostly focus on activity testing of single substrate types, neglecting the diversity of substrate molecular structures, especially the dynamic adaptation mechanism of gallic esters with different side chain lengths to enzyme binding sites, resulting in insufficient sensitivity and accuracy of evaluation results. In addition, traditional methods often lack a comprehensive consideration of changes in spatial configuration adaptability, the number of hydrogen bonds, and hydrophobic interactions when analyzing conformational adjustments of enzyme active sites, making it difficult to capture the subtle dynamics of enzyme-substrate interactions. The core challenge in this field lies in how to accurately characterize the impact of substrate structural differences on the catalytic activity of tanninases, particularly the decrease in spatial configuration adaptability, the reduction in the number of hydrogen bonds formed, and the fluctuations in the strength of hydrophobic interactions caused by changes in substrate side chain length. These factors collectively constrain the dynamic adjustment capability of the enzyme's active site, leading to biases in catalytic efficiency assessments when dealing with complex substrates. Insufficient substrate-enzyme active site fit can result in decreased binding efficiency, while changes in hydrogen bonding and hydrophobic interactions further affect catalytic rate and stability. These unresolved technical challenges make existing methods ill-suited for detecting diverse substrates, limiting the optimization and widespread application of tanninases in practical applications. Therefore, systematically analyzing the impact of substrate side chain length variations on the spatial conformational adjustment of the tanninase active site, and revealing its comprehensive regulatory mechanism on conformational fit, hydrogen bond formation, and hydrophobic interactions, thereby improving the accuracy of catalytic efficiency assessments, has become a critical issue that urgently needs to be addressed. Summary of the Invention

[0003] This invention provides a method for evaluating the catalytic efficiency of tanninases based on enzyme activity detection, mainly comprising:

[0004] The interaction trajectories between gallic acid esters with different side chain lengths and the active site of tanninase were obtained by molecular dynamics simulation. The side chain length distribution and substrate displacement distance were calculated to obtain conformational matching parameters.

[0005] Based on conformational matching parameters, the influence of side chain length distribution on the number of hydrogen bonds formed and the frequency of hydrogen bond breaking is analyzed. Conformational energy difference and dynamic free energy change are extracted to determine the dynamic characteristic parameters of hydrogen bonds.

[0006] If the hydrogen bond breaking frequency in the hydrogen bond dynamic characteristic parameters exceeds the preset threshold, the hydrophobic interaction fluctuation is analyzed based on molecular dynamics trajectory analysis to determine the distribution of hydrophobic interaction intensity.

[0007] Feature fusion was performed on conformational matching parameters, hydrogen bond dynamic characteristic parameters, and hydrophobic interaction intensity distribution to construct a nonlinear regression mapping model between side chain length distribution and binding rate constant, and the predicted value of binding efficiency was obtained.

[0008] Based on the predicted binding efficiency, the binding pocket volume and residue side chain angle of the enzyme active center are optimized using the Monte Carlo simulation method to determine the configuration parameters of the optimized enzyme active center. If the pocket deformation rate in the configuration parameters matches the preset adaptation time scale, the substrate rotational degrees of freedom and binding rate constant of the optimized configuration are verified to determine the improvement in binding efficiency.

[0009] Based on the improvement in binding efficiency, principal component analysis was used to extract the side chain swing amplitude and hydrophobic interaction fluctuation characteristics of the enzyme active center, and the distribution of dynamic free energy change was calculated to obtain the dynamic adaptability parameters.

[0010] The Spearman correlation between dynamic adaptability parameters and predicted binding efficiency was analyzed, and the weighted distribution of side chain length distribution on binding rate constant was generated to determine the catalytic efficiency evaluation model.

[0011] The binding rate constants of various gallic acid ester structures were predicted. Based on the predicted binding rate constants, the side chain length distribution and conformational matching degree were ranked to obtain the substrate selection scheme.

[0012] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0013] This invention discloses a method for evaluating the catalytic efficiency of tanninases based on enzyme activity detection. By analyzing the interaction between gallic esters with different side chain lengths and the enzyme active site, parameters such as conformational matching, hydrogen bond dynamics, and hydrophobic interaction strength are calculated to construct a nonlinear regression model of side chain length distribution and binding rate constant. The structure of the enzyme active site is further optimized to improve binding efficiency, and a catalytic efficiency evaluation model is established. Finally, by predicting the binding rate constants of various gallic ester structures, the suitability ranking of side chain length distribution and conformational matching is obtained, achieving optimal selection of substrate structures. This invention can effectively improve the binding efficiency of gallic esters to tanninases, providing theoretical guidance for substrate design and optimization of related enzyme-catalyzed reactions. Attached Figure Description

[0014] Figure 1 This is a flowchart of a method for evaluating the catalytic efficiency of tanninase based on enzyme activity detection according to the present invention. Detailed Implementation

[0015] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0016] like Figure 1 This embodiment of a method for evaluating the catalytic efficiency of tanninases based on enzyme activity detection may specifically include:

[0017] S101. The interaction trajectories between gallic acid esters with different side chain lengths and the active center of tanninase were obtained by molecular dynamics simulation. The side chain length distribution and substrate displacement distance were calculated to obtain conformational matching parameters.

[0018] A three-dimensional spatial model of the gallate molecule side chain and the tanninase active center was constructed using a molecular force field, and equilibrium conformational data was obtained using a temperature-pressure coupling algorithm. Based on the equilibrium conformational data, the total intermolecular potential energy was calculated using the group contribution method, and the spatial distribution probability density of the gallate molecule side chain around the tanninase active center was obtained from the total intermolecular potential energy. For the spatial distribution probability density, Monte Carlo sampling was performed on the side chain conformation to obtain the conformational ensemble, and the ratio of the cavity volume to the steric hindrance parameter of the enzyme active center was calculated from the conformational ensemble to obtain the conformational matching degree function value. Based on the conformational matching degree function value, a Gaussian distribution function was established to fit the displacement distance distribution curve, and the conformational free energy change value was calculated using the parameters of the distribution curve. Based on the conformational matching degree function value and the free energy change value, the conformational matching degree parameter was obtained.

[0019] Specifically, a three-dimensional spatial model of the side chain of gallic ester molecules and the active site of tanninase is constructed using a molecular force field. Hydrogen bonding potential, van der Waals force, and electrostatic interaction potential energy functions are defined within the force field. Periodic boundary conditions are applied to the enzyme-substrate complex for kinetic optimization, and a temperature-pressure coupling algorithm is used to obtain equilibrium conformational data. Based on this equilibrium conformational data, the total intermolecular interaction potential energy, including hydrogen bonding potential, van der Waals force, and electrostatic interaction potential, is calculated using the group contribution method. The spatial distribution probability density of the gallic ester molecule's side chain around the active site of tanninase is obtained from the potential energy data.

[0020] Based on the spatial distribution probability density, Monte Carlo sampling is performed on the side chain conformations to obtain the conformational ensemble. The conformational matching function value is obtained by calculating the ratio of the cavity volume of the enzyme active center to the steric hindrance parameter. Brownian dynamics is used to track the trajectory of the gallate molecule near the enzyme active center. The change curve of the substrate molecule's centroid position over time is extracted from the trajectory coordinate data, and the displacement distance is calculated by integrating the curve. For the displacement distance, a Gaussian distribution function with three degrees of freedom is established to fit the displacement distance distribution curve, and the expected value and variance of the displacement distance are calculated using the curve fitting parameters. Based on the spatial structure of the binding site of the enzyme active center, constraints such as hydrogen bond distance and dihedral angle are set, and the conformational free energy change value is calculated using the thermodynamic perturbation method under the constraints.

[0021] Based on the conformational matching function value and the free energy change value, a scoring function is constructed using a linear weighting method, and the optimal conformation is determined by the numerical value of the scoring function. In molecular dynamics simulations, when the interaction between gallate molecules and the active site of tanninase is calculated using the Amber force field, the potential energy function includes bond length, bond angle, dihedral angle, and non-bonded interaction terms, among which non-bonded interactions include van der Waals forces and electrostatic interactions.

[0022] For the enzyme-substrate complex system, periodic boundary conditions were set, with a box size of 10 nm × 10 nm × 10 nm. Kinetic optimization was performed using the NPT ensemble at 300 K and 1 atmosphere, with a time step of 2 femtoseconds and a total simulation duration of 100 nanoseconds. The conformation was saved every 10 picoseconds. In the kinetic trajectory analysis, intermolecular interaction potential energy was calculated based on the group contribution method. Different potential energy parameters were assigned to the functional groups (hydroxyl, carbonyl, benzene ring, etc.) of the gallic ester molecule. Statistical analysis of the interaction potential energy yielded a spatial probability density map of the side chain near the active site, with high probability density regions corresponding to stable conformations.

[0023] For regions with high spatial probability density, Monte Carlo sampling was employed for conformational sampling at a sampling interval of 100 picoseconds and a sampling temperature of 300 K. Conformational fit was determined by calculating the ratio of the enzyme active site cavity volume to the substrate molecule van der Waals volume. The cavity volume was calculated using a three-dimensional grid scanning method with a grid spacing of 0.2 Å. A point was considered to be within the cavity if its distance from a grid point to a protein atom was greater than the sum of the van der Waals radii. In Brownian dynamics simulations, the substrate molecule's center of mass was tracked at a sampling interval of 1 picosecond, recording the changes in its three-dimensional coordinates over time. The Euclidean distance between the centers of mass of two adjacent conformations was calculated to obtain the displacement distance over time. Gaussian distribution fitting was applied to the displacement distance data, yielding a distribution function with a mean of 0.5 nm and a standard deviation of 0.2 nm. In constraint optimization, a hydrogen bond distance threshold of 0.35 nm and a dihedral angle constraint range of ±30 degrees were set.

[0024] The free energy changes of different conformations were calculated using a thermodynamic perturbation method, with a simulation time of 2 nanoseconds for each perturbation window. The conformational fit and free energy change were weighted using a ratio of 0.6:0.4 to construct a scoring function to determine the optimal conformation. A higher scoring function value indicates a more stable conformation, which is more conducive to the enzyme-catalyzed reaction.

[0025] S102. Based on the conformational matching parameters, analyze the influence of side chain length distribution on the number of hydrogen bonds formed and the frequency of hydrogen bond breaking, extract the conformational energy difference and dynamic free energy change, and determine the dynamic characteristic parameters of hydrogen bonds.

[0026] The intermolecular hydrogen bond formation process is extracted from the trajectory data based on conformational matching parameters. Hydrogen bond spatial distribution characteristic data is obtained by performing a double Gaussian fit on the distance between hydrogen bond donor and acceptor atoms and the bond angle. The number of hydrogen bonds is counted based on this spatial distribution characteristic data, and the number of hydrogen bond breaks per unit time is calculated by setting a threshold for hydrogen bond length variation to obtain breakage frequency data. A curve relating side chain length to hydrogen bond breakage frequency is established based on the breakage frequency data, and the cumulative breakage probability distribution is obtained through curve integration. The free energy of the cumulative breakage probability distribution is calculated using the Boltzmann weighted calculation method, and free energy barrier data is obtained by setting the perturbation step size and controlling the equilibrium time. The free energy barrier data is optimized by iteratively calculating the minimum free energy path, and the hydrogen bond strength, breakage time, and bond length distribution data are linearly normalized within the 0-1 interval to obtain the dynamic characteristic parameters of the hydrogen bonds.

[0027] Specifically, the intermolecular hydrogen bond formation process is extracted from the trajectory data based on the conformational matching degree parameter. The distance and bond angle between hydrogen bond donor and acceptor atoms are statistically analyzed, and the spatial distribution characteristic data of hydrogen bonds are obtained by performing double Gaussian fitting on the distance and angle distribution curves.

[0028] Based on the spatial distribution characteristics of hydrogen bonds, a quantitative method was used to count the number of hydrogen bonds formed between side chains of different carbon chain lengths and amino acid residues in the active center. A hydrogen bond length change exceeding 0.35 nm was set as the breakage criterion, and the number of hydrogen bond breaks per unit time was calculated to obtain breakage frequency data. Time series analysis was performed on the breakage frequency data to establish a correlation curve between side chain length and hydrogen bond breakage frequency. The cumulative breakage probability distribution under different side chain lengths was calculated by curve integration. Based on the cumulative breakage probability distribution, the contribution of each hydrogen bond in the hydrogen bond network to conformational stability was quantified. The conformational potential difference was obtained by integrating the curve of the potential energy change of the side chain-active center complex over time using the Boltzmann weighted calculation method.

[0029] The free energy of the conformational potential energy difference was calculated using a thermodynamic perturbation method. The perturbation step size was set to 0.02, and equilibration was performed for 50 picoseconds for each perturbation window to obtain the free energy barrier data during the conformational transformation process.

[0030] The free energy barrier data were optimized using the steepest descent method, and the minimum free energy path was iteratively calculated. Hydrogen bond strength, breakage time, and bond length distribution data were linearly normalized within the 0-1 range to obtain dynamic hydrogen bond characteristic parameters. In molecular dynamics simulations, hydrogen bonds play a crucial role in the binding of substrates to enzyme active sites; analyzing the dynamic characteristics of hydrogen bonds can reveal the substrate recognition mechanism.

[0031] To determine the interaction between gallic ester molecules and tanninases, it is necessary to first identify the geometric characteristics of the hydrogen bonds, including the distance between the hydrogen bond donor and acceptor atoms and the hydrogen bond angle. In practical calculations, a donor-hydrogen-acceptor angle greater than 150 degrees and a donor-acceptor distance less than 0.35 nm were selected as criteria for hydrogen bond formation. Analysis of the trajectory data revealed a bimodal distribution of the donor-acceptor distance, with peaks at 0.28 nm and 0.32 nm, corresponding to strong and weak hydrogen bonds, respectively. Analysis of the hydrogen bond characteristics of side chains with different carbon chain lengths showed a significant difference in the distribution of hydrogen bonds formed between side chains with carbon chain lengths of 4 to 8 carbon atoms and the active center. Among these, a carbon chain length of 6 carbon atoms resulted in an average of 3.5 hydrogen bonds formed, and the hydrogen bond breaking frequency was the lowest, at 0.85 times per nanosecond.

[0032] As the carbon chain length increases to 8 carbon atoms, the average number of hydrogen bonds decreases to 2.8 due to enhanced steric hindrance, while the breakage frequency rises to 1.2 times per nanosecond. Time-series analysis of the hydrogen bond breakage frequency was used to construct a curve relating side chain length to breakage frequency. The cumulative breakage probability obtained by curve integration shows that the breakage probability is lowest at 0.35 when the carbon chain length is 6 carbon atoms, while the breakage probabilities are 0.48 and 0.52 for carbon chain lengths of 4 and 8 carbon atoms, respectively. This difference reflects the influence of side chain length on the stability of the hydrogen bond network.

[0033] In the conformational potential energy analysis, the Boltzmann weighted calculation method was used to integrate the potential energy curves to obtain the conformational potential energy difference for different side chain lengths. The results show that the conformational potential energy difference is smallest (12.5 kJ / mol) when the carbon chain length is 6 carbon atoms, indicating the most stable conformation at this length. Thermodynamic perturbation calculations revealed that the free energy barrier for conformational transformation exhibits a parabolic distribution with side chain length, reaching a minimum of 15.8 kJ / mol when the carbon chain length is 6 carbon atoms. When normalizing the hydrogen bond characteristic parameters, the hydrogen bond strength range was selected to be 5–15 kJ / mol, the breakage time range to be 0.5–2.0 nanoseconds, and the bond length range to be 0.25–0.35 nanometers. The normalized parameters show that all indicators are optimal when the carbon chain length is 6 carbon atoms, with a comprehensive score of 0.85, while the comprehensive scores for carbon chain lengths of 4 and 8 carbon atoms are 0.65 and 0.58, respectively.

[0034] S103. If the hydrogen bond breaking frequency in the hydrogen bond dynamic characteristic parameters exceeds the preset threshold, the hydrophobic interaction fluctuation is analyzed based on molecular dynamics trajectory analysis to determine the distribution of hydrophobic interaction intensity.

[0035] Conformational coordinate data were extracted from the binding process of gallate esters to the enzyme active site in the molecular dynamics trajectory, and the binding free energy was calculated using the temperature integral method. Based on the conformational coordinate data, the molecular surface area was calculated using the spherical surface integral method, and a hydrophobic surface distribution map was obtained by scanning the solvent-accessible surface area. Based on the hydrophobic surface distribution map, the interaction energy between hydrophobic residues on the molecular surface was quantified, and the energy fluctuation amplitude was obtained by performing a discrete Fourier transform on the curve of hydrophobic interaction over time. Based on the energy fluctuation amplitude, the hydrophobic interaction regions on the molecular surface were divided using the K-means clustering method, and the hydrophobic interaction intensity distribution was obtained by integrating the interaction energy of hydrophobic residues in each region using the Monte Carlo method.

[0036] Specifically, the hydrogen bond breaking frequency in the dynamic characteristic parameters of hydrogen bonds is compared with a preset threshold of 1.0 times per nanosecond. Conformational coordinate data of the binding process between gallate ester and enzyme active site is extracted from the molecular dynamics trajectory, and the binding free energy is calculated using the temperature integral method. For the conformational coordinate data, the molecular surface area is calculated using the spherical surface integral method. A probe radius of 0.14 nm is set to scan the solvent-accessible surface area to obtain a hydrophobic surface distribution map. The contact region is divided into grids with a spacing of 0.1 nm to obtain the hydrophobic contact area value. Based on the hydrophobic contact area value, the interaction energy between hydrophobic residues on the molecular surface is quantified. The energy fluctuation amplitude is obtained by performing a discrete Fourier transform on the curve of hydrophobic interaction over time in the dynamic trajectory.

[0037] Based on the energy fluctuation amplitude, the K-means clustering method was used to divide the hydrophobic interaction regions on the molecular surface, with a cluster size of 3. The interaction energies of the hydrophobic residue pairs within each region were accumulated. The Monte Carlo method was used to integrate the interaction energies of the hydrophobic residue pairs, and the hydrophobic interaction intensity distribution curve was obtained by normalizing the integral values ​​of each hydrophobic interaction region. For the hydrophobic interaction intensity distribution curve, an intensity threshold range of 5 to 15 kJ / mol was set, and the curve was piecewise integrated to obtain the proportion of interaction intensity in each range.

[0038] In the analysis of molecular binding processes, hydrophobic interactions and hydrogen bonds jointly determine the binding characteristics between the substrate and the enzyme active site. For gallic ester molecules, when the hydrogen bond breaking frequency exceeds 1.0 times per nanosecond, it indicates a decrease in the stability of the hydrogen bond network, requiring further investigation into the contribution of hydrophobic interactions. The binding free energy calculated by temperature integration reflects the thermodynamic driving force of molecular binding, with typical binding free energies ranging from -20 to -40 kJ / mol. When calculating the hydrophobic surface area, the spherical surface integration method was used to scan the molecular surface, with the probe sphere radius set to 0.14 nm, which approximates the van der Waals radius of a water molecule. The resulting hydrophobic surface distribution map shows that the benzene ring and alkyl side chain regions of the gallic ester molecule have a large hydrophobic surface area, forming complementary contact with the hydrophobic pocket of the enzyme active site.

[0039] A grid with a spacing of 0.1 nanometers was used to calculate a hydrophobic contact area of ​​approximately 2.5 square nanometers. In analyzing the dynamic characteristics of hydrophobic interactions, hydrophobic residue pairs on the molecular surface were identified and quantified. Hydrophobic residues, represented by aromatic amino acid residues such as tryptophan and phenylalanine, typically exhibit interaction energies between -8 and -15 kilojoules per mole. Fourier transforms of the hydrophobic interaction intensity in the kinetic trajectory yielded an energy fluctuation spectrum showing that the main fluctuation period was between 5 and 20 picoseconds. K-means clustering was used to divide the molecular surface into three hydrophobic interaction regions, corresponding to high, medium, and low intensity hydrophobic interactions, respectively.

[0040] High-intensity regions are mainly distributed in the benzene ring of the substrate molecule, with an average interaction energy of -12 kJ / mol; medium-intensity regions are located in the alkyl side chain, with an average interaction energy of -8 kJ / mol; and low-intensity regions are distributed in other parts of the molecule, with an average interaction energy of less than -5 kJ / mol. The hydrophobic interactions in each region were quantitatively analyzed using the Monte Carlo integration method, with 1000 sampling points, and the normalized distribution of hydrophobic interaction strengths was calculated. The results show that within the intensity range of 5 to 15 kJ / mol, high-intensity regions account for 35%, medium-intensity regions account for 45%, and low-intensity regions account for 20%. This non-uniform distribution reflects the spatial heterogeneity of hydrophobic interactions on the molecular surface, which has a significant impact on molecular recognition and binding. At a temperature of 300 K, the dynamic fluctuations of hydrophobic interactions exhibit a significant time correlation. Analysis revealed that the fluctuations in the intensity of hydrophobic interactions are directly related to the recombination of solvent molecules. When hydrophobic surfaces come into contact, water molecules at the interface are expelled, leading to an increase in system entropy, thereby promoting molecular binding. This entropy-driven binding process plays a crucial role in the recognition of many biomolecules.

[0041] S104. Feature fusion is performed on conformational matching parameters, hydrogen bond dynamic characteristic parameters, and hydrophobic interaction intensity distribution to construct a nonlinear regression mapping model of side chain length distribution and binding rate constant, and the predicted value of binding efficiency is obtained.

[0042] A feature matrix is ​​constructed using conformational matching parameters, hydrogen bond dynamic characteristic parameters, and hydrophobic interaction intensity distribution. Dimensionality-reduced feature data is obtained using principal component analysis. Random sampling is performed based on this dimensionality-reduced feature data, and a random forest training set is constructed using the minimum Gini coefficient as the node splitting criterion. Voting weights are calculated for the random forest training set, and a weighted average method is used to obtain the feature importance sequence and regression prediction values. Based on the feature importance sequence and regression prediction values, the tree depth is set to a range of five to fifteen. An optimal parameter combination is obtained using a grid search method, and the random forest is retrained to obtain the predicted combination efficiency.

[0043] Specifically, a feature matrix is ​​constructed based on conformational matching parameters, hydrogen bond dynamic characteristic parameters, and hydrophobic interaction intensity distribution. The feature data is standardized using mean normalization, with a variance contribution rate threshold set to 0.95. Dimensionality-reduced feature data is obtained through principal component analysis. Random sampling is performed on this dimensionality-reduced feature data, with a sampling ratio of 0.8. The minimum Gini coefficient is used as the node splitting criterion to construct a random forest training set containing 500 decision trees. Based on this training set, the maximum depth of the decision trees is set to 10 layers. For each node, the number of features randomly selected is the square root of the total number of features, and a feature importance sequence is obtained by calculating voting weights. Based on the feature importance sequence, a weighted average method is used to calculate the regression prediction value for each sample. The nonlinear mapping function between the sidechain length distribution and the binding rate constant is fitted using the least squares method.

[0044] The performance of the nonlinear mapping function was evaluated using a five-fold cross-validation method, and the prediction results were quantitatively evaluated by calculating the root mean square error and coefficient of determination. Based on the evaluation results, the parameters of the random forest were optimized using a grid search method, setting the tree depth to between 5 and 15 and the minimum number of leaf node samples to between 5 and 20, to obtain the optimal parameter combination. The random forest was retrained based on the optimal parameter combination, and the predicted binding efficiency and confidence interval were obtained by predicting the validation set samples. In the enzyme-substrate binding characteristic analysis, multiple physicochemical parameters are involved, including conformational matching, hydrogen bond dynamics, and hydrophobic interactions, which together determine the kinetic characteristics of the binding process. By constructing a feature matrix, parameters at different scales were integrated. Conformational matching reflects spatial geometric complementarity, with values ​​ranging from 0 to 1; hydrogen bond dynamics include bond length, bond angle, and breakage frequency information, with typical hydrogen bond lengths between 0.25 and 0.35 nanometers; and the hydrophobic interaction intensity distribution reflects the spatial distribution characteristics of nonpolar interactions, with intensity values ​​typically ranging from 5 to 15 kilojoules per mole. When preprocessing feature data, mean normalization is used to eliminate dimensional differences. Each feature value is subtracted from the mean and then divided by the standard deviation, so that the mean of the processed data is 0 and the variance is 1.

[0045] Principal component analysis (PCA) was used for dimensionality reduction, selecting principal components with a cumulative variance contribution rate of 0.95. This typically reduced the original 10+ features to 3-5 principal components, effectively reducing data redundancy. In the random forest model construction, 500 decision trees were used to form an ensemble learner, with a maximum depth of 10 layers per tree. Node splitting was performed using the minimum Gini coefficient criterion. For a training set containing 1000 samples, sampling with replacement was performed using a sampling ratio of 0.8 to ensure sufficient independence for each decision tree. During feature selection, if the original number of features was 9, 3 features were randomly selected for optimal splitting at each node. Feature importance analysis showed that conformational matching contributed the most to the prediction results, with an importance score of 0.45; followed by hydrogen bond breaking frequency and hydrophobic interaction strength, with importance scores of 0.30 and 0.25, respectively. This quantitative assessment of feature importance helps to understand the relative importance of different physicochemical parameters in the molecular recognition process. During the model validation phase, five-fold cross-validation was used to evaluate the predictive performance. The calculated root mean square error was 0.15 and the coefficient of determination was 0.85, indicating that the model has good predictive ability.

[0046] Optimizing hyperparameters through grid search revealed that the model performed optimally when the decision tree depth was 12 and the minimum number of leaf node samples was 10. Prediction results on the validation set showed that the relative error between the predicted and experimentally determined values ​​of the binding rate constant was within 15%, and the confidence interval was approximately 20% of the predicted value. In practical applications, this prediction model can effectively capture the binding characteristics of substrates with different side chain lengths to the enzyme's active site.

[0047] For example, for a substrate with a carbon chain length of 6, the predicted binding rate constant is 5 × 10^6 moles per second, which agrees well with the experimentally determined value of 4.8 × 10^6 moles per second. Furthermore, the model reflects the bell-shaped relationship between carbon chain length and binding affinity, indicating that the binding rate reaches its maximum at the optimal carbon chain length.

[0048] S105. Based on the predicted binding efficiency, optimize the binding pocket volume and residue side chain angles of the enzyme active center using Monte Carlo simulation to determine the optimized conformational parameters of the enzyme active center. If the pocket deformation rate in the conformational parameters matches the preset adaptation timescale, verify the substrate rotational degrees of freedom and binding rate constant of the optimized conformation to determine the improvement in binding efficiency.

[0049] The Monte Carlo method was used to randomly perturb the dihedral angles of the amino acid residue side chains within the binding pocket within a torsion angle range. An initial configuration set was obtained by sampling the binding pocket volume through internal coordinate optimization. Based on each structure in the initial configuration set, dynamic simulations were performed under temperature conditions. Carbon atom coordinate matrices were extracted from the trajectory data, and the principal component vectors of the pocket deformation motion were obtained by diagonalizing the covariance matrix. Time series analysis was performed on the principal component vectors, and pocket deformation rate data were obtained by statistically analyzing the time evolution of each vibrational mode. Configurations were screened based on the pocket deformation rate data. For the screened configurations, a spherical coordinate grid method was used to scan the rotation angles of the substrate molecule within the optimized configuration. The binding rate constant was obtained by Boltzmann weighted averaging, and the improvement in binding efficiency was obtained by comparing the values ​​before and after optimization.

[0050] Specifically, the structure of the enzyme active site is parameterized based on the predicted binding efficiency. The Monte Carlo method is used to randomly perturb the dihedral angles of the amino acid residue side chains within the binding pocket within a range of ±30 degrees and the torsion angles within a range of ±60 degrees. An initial configuration set is generated by sampling the configuration of the binding pocket volume through internal coordinate optimization. Dynamic simulations are performed on each structure in the initial configuration set at 300 K, with a sampling time interval of 1 picosecond. Carbon atom coordinate matrices are extracted from the trajectory data, and the principal component vector of the pocket deformation motion is obtained by diagonalizing the covariance matrix.

[0051] Time series analysis was performed on the principal component vectors to calculate the frequency of deformation amplitude changes, and pocket deformation rate data was obtained by statistically analyzing the time evolution of each vibration mode. Based on the pocket deformation rate data, the configuration optimization parameters were compared with a preset time scale of 5 picoseconds, and an adaptive annealing algorithm with an initial temperature of 500K was used to screen the optimized configurations. For the screened optimized configurations, the root mean square fluctuations of the carbon atom coordinates in the pocket residues over time were calculated, and the flexibility index of the active center was obtained by normalizing the fluctuation values.

[0052] Based on the aforementioned flexibility index of the active site, a spherical coordinate grid method was used to scan the rotation angles of the substrate molecule in the optimized configuration in 360 degrees, with an angle interval of 5 degrees. Based on the rotation angle scan data, the conformational transition energy barrier for each orientation was calculated, and the binding rate constant was obtained through Boltzmann weighted averaging. The improvement in binding efficiency was obtained by comparing the values ​​before and after optimization. In the optimization of enzyme active site structure, the flexibility and dynamic characteristics of the configuration have a significant impact on substrate binding. In the initial configuration set generation stage, the side chains of amino acid residues within the binding pocket were randomly perturbed. For example, perturbing the dihedral angle of the tryptophan side chain within a range of ±30 degrees can simulate its degrees of freedom of motion under physiological conditions, while the change in the torsion angle within a range of ±60 degrees reflects the diversity of side chain conformations. Through this perturbation sampling, typically 1000 different initial configurations can be obtained. In the kinetic simulation stage, a temperature of 300K was used to simulate physiological conditions, and each configuration was simulated for 100 nanoseconds. The configuration was saved every picosecond, generating 100,000 frames of trajectory data. The coordinate matrix of carbon atoms in the pocket is extracted from the trajectory, and the motion of each atom can be decomposed into multiple vibrational modes.

[0053] By diagonalizing the covariance matrix, the obtained eigenvectors reflect the main motion directions of pocket deformation, with the first three principal components typically accounting for over 85% of the total fluctuations. The pocket deformation rate is calculated based on the temporal evolution of the principal components; a Fourier transform of each vibrational mode yields the frequency distribution spectrum. The main deformation frequencies are concentrated in the 0.1 to 1.0 picosecond timescale, and configurations matching the preset 5 picosecond substrate binding timescale are retained for further optimization. An adaptive annealing algorithm cools from 500K, sampling the configuration every 50K until reaching 300K. The root mean square fluctuations of carbon atoms are calculated for the selected configurations, typically ranging from 0.5 to 2.0 Å. Regions with larger fluctuation values ​​correspond to the flexible segments of the active center, which play a "induced adaptation" role during substrate binding. Rotational degrees of freedom analysis of the substrate molecule is performed using 5-degree angular intervals, generating a total of 72 × 72 different orientations. The conformational transformation energy barrier is calculated for each orientation, and the barrier height is typically between 2 and 8 kilojoules per mole.

[0054] The optimized active site configuration exhibits better dynamic adaptability, mainly in three aspects: enhanced synergistic pocket deformation, optimized flexibility of key residues, and reduced substrate rotational barrier. Taking the binding rate constant as an example, the optimized configuration improves by 2 to 5 times compared to the initial configuration. This improvement stems from the dynamic optimization of the active site configuration. Specifically, the pocket volume fluctuation range decreases from the initial 15% to 10%, the deformation frequency is more closely matched to the substrate binding timescale, the flexibility index of the residue side chains increases from 1.5 to 1.8, and the average rotational barrier of the substrate molecule at the binding site decreases from 6 kJ / mol to 4 kJ / mol.

[0055] S106. Based on the increase in binding efficiency, principal component analysis is used to extract the side chain swing amplitude and hydrophobic interaction fluctuation characteristics of the enzyme active center, and the distribution of dynamic free energy change is calculated to obtain the dynamic adaptability parameters.

[0056] A three-dimensional spatial coordinate matrix is ​​obtained by periodic sampling of the enzyme active site conformation. The main motion direction is obtained by calculating the side chain atom covariance based on the three-dimensional spatial coordinate matrix. The displacement amplitude of the side chain atoms is calculated by statistical analysis method based on the main motion direction. Significant motion sites and key hydrophobic interaction residue pairs are identified by the displacement amplitude. A fast Fourier transform of van der Waals interaction energy is performed on the key hydrophobic interaction residue pairs to obtain the hydrophobic interaction characteristic spectrum. A conformational transition simulation is performed using the free energy perturbation method based on the hydrophobic interaction characteristic spectrum. The dynamic free energy distribution curve is calculated based on the conformational transition simulation. The conformational transition rate is obtained from the dynamic free energy distribution curve. The dynamic adaptability parameter is obtained by normalizing the rate data.

[0057] Specifically, the conformation of the enzyme active site is periodically sampled based on the improvement in binding efficiency, with a time interval of 1 picosecond. Principal component analysis is used to calculate the covariance of the three-dimensional spatial coordinate matrix of the side chain atoms, and the eigenvectors with a cumulative contribution rate of 90% are selected as the main motion directions. For the main motion directions, the displacement amplitude of the side chain atoms in each direction is calculated, and the oscillation amplitude distribution is obtained by statistically analyzing the displacement data. Displacements exceeding two standard deviations are considered significant motions.

[0058] Based on the significant motion sites, key residue pairs involved in hydrophobic interactions were identified. A fast Fourier transform was performed on the van der Waals interaction energies between residue pairs using a 10-picosecond time window to obtain spectral data. Peak values ​​were extracted from the spectral data, and a signal-to-noise ratio threshold of 3 was set to screen the main hydrophobic interaction fluctuation frequencies. The hydrophobic interaction characteristic spectrum was obtained through statistical analysis of the fluctuation periods. For the hydrophobic interaction characteristic spectrum, the conformational transition process was simulated using a free energy perturbation method, with a perturbation step size of 0.1 and a 50-picosecond equilibration period for each perturbation window. Based on the equilibrium trajectory, the free energy change value within each perturbation window was calculated, and the dynamic free energy distribution curve was obtained by thermodynamic averaging of the conformational ensemble.

[0059] Based on the dynamic free energy distribution curve, the Boltzmann weighted method was used to calculate the conformational transition rates at different time scales, and the dynamic adaptability parameters were obtained by normalizing the rate data. In molecular dynamics, the dynamic characteristics of the enzyme active site play a crucial role in substrate binding, among which the mobility of side chains and the dynamic changes in hydrophobic interactions are particularly important. By sampling the conformation of the active site every 1 picosecond, the complete conformational evolution trajectory can be obtained. Principal component analysis shows that typically the first 3 to 4 eigenvectors can explain more than 90% of the conformational fluctuations, and these eigenvectors represent the main modes of side chain motion.

[0060] In practical calculations, eigenvalue analysis of the covariance matrix of an active center containing 300 atoms shows that the first principal component contributes 45%, corresponding to large swings in the side chains; the second principal component accounts for 25%, reflecting local torsional motion; and the third principal component accounts for 20%, describing overall respiratory motion. The displacement amplitude distribution of the side chains exhibits significant regional differences, with atomic displacements reaching 0.8 nm in the flexible ring region, while remaining below 0.3 nm in the core region.

[0061] In hydrophobic interaction analysis, key residue pairs typically include hydrophobic amino acids such as phenylalanine, leucine, and isoleucine. The van der Waals interactions between these residue pairs exhibit complex fluctuations over time. Fourier transform analysis with a 10-picosecond time window reveals multiple characteristic frequencies. The main fluctuation periods are distributed between 0.5 and 5 picoseconds, with the most significant fluctuations around 1 picosecond, resulting in a signal-to-noise ratio exceeding 3. Free energy perturbation calculations use a step size of 0.1 for conformational transition sampling, with each perturbation window balancing for 50 picoseconds. During conformational transitions, the hydrophobic interaction energy varies between 2 and 8 kilojoules per mole. Statistical analysis of 100 perturbation windows reveals a distinct multi-peaked free energy distribution curve, reflecting multiple stable states of conformational transitions. Boltzmann weighted analysis considers multiple time scales from picoseconds to nanoseconds. The calculation results show that the conformational transition rate reaches its maximum value of about 10^9 per second on a timescale of 1 to 10 picoseconds; while on a longer timescale, the rate gradually decreases to 10^7 per second.

[0062] This time-scale dependence reflects the multi-level dynamic adaptation characteristics of the active center. The normalized dynamic adaptation capability parameter integrates both conformational flexibility and energy accessibility. The parameter values ​​are distributed between 0 and 1, with values ​​above 0.7 indicating strong dynamic adaptation capability, values ​​between 0.3 and 0.7 indicating a moderate level, and values ​​below 0.3 indicating weak adaptation capability. Statistical analysis shows that approximately 60% of the regions in the optimized active center configuration achieve strong dynamic adaptation capability, which is significantly correlated with the improved binding efficiency.

[0063] S107. Analyze the Spearman correlation between dynamic adaptability parameters and predicted binding efficiency, generate the weighted distribution of side chain length distribution on binding rate constant, and determine the catalytic efficiency evaluation model.

[0064] Data pairing sequences are obtained based on dynamic adaptability parameters and predicted binding efficiency values. A correlation matrix is ​​calculated using the Spearman rank correlation coefficient. The correlation matrix is ​​then fitted using the least squares method to obtain initial parameter weights. These initial weights are updated to a convergence threshold using an iterative optimization method, yielding convergent weights. The binding rate constant is integrated piecewise using these convergent weights, and the integration result is smoothed using a Gaussian kernel function to obtain a weight distribution curve. A partial least squares regression model is established based on this weight distribution curve. The model is then subjected to a residual distribution test, and the significance level is determined using an F-test. A catalytic efficiency evaluation index system is established, comprising three parameters: binding rate constant, substrate conversion efficiency, and product formation rate. A weighted combination of these parameters yields the catalytic efficiency evaluation model.

[0065] Specifically, a data pairing sequence is established based on the dynamic adaptability parameters and the predicted binding efficiency. The parameter sequence is sorted by size to obtain rank data, and the correlation matrix is ​​obtained using the Spearman rank correlation coefficient formula. For the correlation matrix, the adaptability parameters are normalized according to a range of 4 to 8 carbon atoms. The cross-validation factor is set to 5, and the initial weight values ​​of the parameters are obtained through least squares fitting.

[0066] Based on the initial weight values, an iterative optimization method is used to update the weights, setting a convergence threshold of 0.001. Residual calculations are performed on the results of each iteration until convergence. According to the convergent weight values, the binding rate constants for different sidechain length intervals are piecewise integrated, and the integral curve is smoothed using a Gaussian kernel function to obtain the weight distribution curve. For the weight distribution curve, partial least squares regression is used to model the binding rate, setting the number of principal components to 3, and selecting the optimal number of principal components through cross-validation. The residual distribution of the partial least squares regression model is tested, the root mean square error between the predicted and measured values ​​is calculated, and the significance level of the model is determined using an F-test.

[0067] Based on the significance test results, a catalytic efficiency evaluation index system was established, including three parameters: binding rate constant, substrate conversion efficiency, and product formation rate. A weighted combination was used to obtain the final evaluation model. In enzyme catalytic efficiency evaluation, the correlation analysis between dynamic adaptability and binding efficiency is a crucial step. Taking the gallate-tannin enzyme system as an example, the dynamic adaptability parameter covered the range of 0.3 to 0.9, while the predicted binding efficiency values ​​were distributed between 10 and 10. 5 Up to 10 7 Between per mole per second. By performing rank transformation, these continuous variables were converted into discrete ranks, and the calculated Spearman correlation coefficient reached 0.85, indicating that the two parameters have a significant positive correlation.

[0068] In the carbon chain length distribution analysis, the side chains were divided into intervals based on the number of carbon atoms, ranging from 4 to 8. The adaptability parameters within each interval were normalized and distributed between 0 and 1. Parameter fitting was performed using 5-fold cross-validation. Initial weight values ​​showed that the interval with a carbon chain length of 6 had the highest weight of 0.45, while the intervals with carbon chain lengths of 4 and 8 had weights of 0.25 and 0.30, respectively. The weight optimization process employed an iterative method, with a convergence threshold of 0.001. In actual calculations, convergence typically required 10 to 15 iterations, with the residual decreasing progressively with each iteration. The optimized weight distribution curve exhibited a distinct bell-shaped characteristic, peaking at a carbon chain length of 6, consistent with the experimentally observed trend in catalytic activity.

[0069] When using partial least squares regression to build the prediction model, three principal components were selected for modeling. The first principal component explained 65% of the variance, mainly reflecting the substrate binding process; the second principal component accounted for 25%, corresponding to conformational changes; and the third principal component accounted for 10%, describing the product release process. The optimal number of principal components determined through cross-validation was three; further increasing the number of principal components did not significantly improve model performance. In the model validation phase, the calculated root mean square error was 0.15, and the significance level of the F-test reached 0.01. The residual distribution exhibited a normal characteristic, and 95% of the predicted values ​​had a relative error of less than 20% compared to the measured values. This prediction accuracy is already of practical value for evaluating enzyme catalytic kinetic parameters. The catalytic efficiency evaluation index system comprehensively considered three key parameters: binding rate constant, substrate conversion efficiency, and product formation rate. The binding rate constant reflects the molecular recognition process, typically on the order of 10^6 moles per second; the substrate conversion efficiency describes the efficiency of chemical bond breaking and formation, usually between 50 and 200 per second; and the product formation rate reflects the integrity of the entire catalytic cycle, generally in the range of 20 to 100 per second. These three parameters are combined using a weighted ratio of 0.4:0.3:0.3 to construct the final evaluation model. Case studies show that this model can accurately predict the catalytic efficiency of substrates with different side chain lengths, providing a quantitative evaluation tool for enzyme engineering.

[0070] S108. Predict the binding rate constants of various gallic acid ester structures, and rank the side chain length distribution and conformational matching degree according to the predicted binding rate constants to obtain the substrate selection scheme.

[0071] Conformational data of gallic ester molecules were collected based on the parameter weights of the catalytic efficiency model. The standardized binding rate between the conformation and the enzyme active site was calculated using molecular dynamics. For the standardized binding rate, principal component analysis was used to extract feature vectors, and the feature vectors were grouped using the K-means clustering algorithm to obtain conformational family distribution data. A dual-index evaluation matrix was established based on the conformational family distribution data, and the positive and negative ideal solutions were calculated using the maximum and minimum value standardization method. For the positive and negative ideal solutions, the relative proximity coefficient was calculated using the Euclidean distance formula. If the proximity coefficient is greater than a preset threshold, the corresponding scheme is included in the candidate scheme sequence. The candidate scheme sequence is then sorted in descending order to obtain the substrate selection scheme.

[0072] Specifically, based on the parameter weights in the catalytic efficiency evaluation model, periodic conformational sampling is performed on gallic ester molecules with different carbon chain lengths, with a sampling time interval of 1 picosecond. The binding rate constant between each conformation and the enzyme active site is calculated using molecular dynamics, and the values ​​are then normalized using minimum-maximum normalization to obtain standardized rate data. For the standardized rate data, principal component analysis is used to extract molecular conformational features, with the cumulative variance contribution rate of the first three feature vectors set to be no less than 90%. K-means clustering is then used to group the feature space samples to obtain conformational families. Based on the conformational family distribution, two evaluation indicators—side chain length distribution and conformational matching degree—are set, and a dual-indicator evaluation matrix is ​​established. Positive and negative ideal solutions are calculated using the maximum-minimum normalization method.

[0073] Based on the positive and negative ideal solutions, the distance from each scheme to the ideal solution is calculated using the Euclidean distance formula. The side chain length weight is set to 0.6, and the conformational matching weight is set to 0.4 using the analytic hierarchy process (AHP). For these weights, the relative proximity coefficient of each scheme is calculated using the distance ratio to the ideal solution. A proximity threshold of 0.75 is set, and candidate schemes meeting the threshold requirement are selected. For these candidate schemes, a multi-dimensional scoring method is used to comprehensively score the structural parameters. The substrate selection priority sequence is obtained by sorting the scoring results in descending order.

[0074] In the optimization and screening of various gallic ester structures, conformational sampling and kinetic analysis are crucial steps. Periodic conformational sampling was performed on gallic ester molecules with carbon chain lengths ranging from 4 to 8, recording a conformation every picosecond, accumulating to 10,000 conformational samples. The binding rate constants obtained through molecular dynamics calculations ranged from 10^5 to 10^7 moles per second, and the normalized rate values ​​after normalization ranged from 0 to 1. Principal component analysis showed that the variance contribution rates of the first three eigenvectors were 60%, 20%, and 15%, respectively, totaling 95%, exceeding the set 90% threshold. These three eigenvectors correspond to the stretching, torsional, and bending motions of the molecules, respectively. K-means clustering divided the conformational samples into three main conformational families, with the largest family accounting for 45% of the total samples, exhibiting strong conformational stability. In constructing the evaluation metrics, the side chain length distribution was described using a Gaussian function, with a mean at 6 carbon atoms and a standard deviation of 1 carbon atom. Conformational matching was calculated based on van der Waals surface area complementarity, with matching values ​​ranging from 0.6 to 0.9. A positive ideal solution corresponds to the case where the length distribution is closest to the center of the Gaussian distribution and has the highest matching degree, while a negative ideal solution corresponds to the case where the distribution deviation is the largest and the matching degree is the lowest. The weight values ​​determined by the analytic hierarchy process (AHP) reflect the dominant role of side chain length in the binding process. A weight ratio of 0.6:0.4 indicates that the contribution of side chain length is slightly greater than that of conformational matching in substrate screening. This weight allocation is consistent with experimentally observed structure-activity relationships, i.e., side chain length has a more significant impact on enzyme activity. In the relative proximity calculation, a threshold of 0.75 reflects a high screening standard. Actual calculations show that approximately 25% of the structures among all candidate molecules meet this threshold requirement. Most of these structures have side chain lengths of 5 to 7 carbon atoms and exhibit good conformational matching characteristics. The multi-dimensional scoring employed five evaluation dimensions: binding rate, conformational stability, spatial matching, energy complementarity, and dynamic adaptability. Each dimension's score ranged from 0 to 100, and a weighted average was used to obtain the overall score. Among the selected priority sequences, the top three structures all possessed side chains with six carbon atoms and exhibited overall scores exceeding 90.

[0075] These structures exhibit excellent characteristics in terms of binding dynamics and spatial matching, with predicted binding rate constants reaching 10. 7 With a conformational match of over 0.85 per mole per second, they indicate good structural complementarity with the enzyme's active site.

[0076] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method for evaluating the catalytic efficiency of tanninases based on enzyme activity detection, characterized in that, The method includes: The interaction trajectories between gallic acid esters with different side chain lengths and the active site of tanninases were simulated, and the side chain length distribution and substrate displacement distance were calculated to obtain conformational matching parameters. The dynamic characteristic parameters of hydrogen bonds are determined by analyzing the influence of side chain length distribution on the number of hydrogen bonds formed and the frequency of hydrogen bond breaking. If the frequency of hydrogen bond breaking in the dynamic characteristic parameters of hydrogen bonds exceeds the preset threshold, the distribution of hydrophobic interaction intensity is determined by analyzing the fluctuation of hydrophobic interaction. The binding efficiency is predicted based on conformational matching parameters, hydrogen bond dynamic characteristic parameters, and hydrophobic interaction intensity distribution, and the predicted binding efficiency value is obtained. The dynamic adaptability parameter is determined based on the side chain swing amplitude and hydrophobic interaction fluctuation characteristics of the enzyme active site. The Spearman correlation coefficient between the dynamic adaptability parameter and the predicted binding efficiency is calculated to evaluate the catalytic efficiency and obtain the evaluation result.

2. The method according to claim 1, characterized in that, The simulation of the interaction trajectories between gallic esters with different side chain lengths and the active site of tanninases, the calculation of side chain length distribution and substrate displacement distance, and the obtaining of conformational matching parameters include: Three-dimensional spatial modeling of the side chain of gallic acid ester molecules and the active center of tanninase was performed using molecular force field, and equilibrium conformation data were obtained by temperature-pressure coupling algorithm. Based on the equilibrium conformation data, the total potential energy of intermolecular interactions is calculated using the group contribution method, and the spatial distribution probability density of the gallate molecule side chain around the active site of tanninase is obtained from the total potential energy of intermolecular interactions. For the spatial distribution probability density, Monte Carlo sampling is performed on the side chain conformation to obtain the conformation ensemble, and the conformation matching degree function value is obtained by calculating the ratio of the cavity volume of the enzyme active center to the steric hindrance parameter from the conformation ensemble. Based on the conformation matching degree function value, a Gaussian distribution function is established to fit the displacement distance distribution curve, and the conformation free energy change value is calculated through the parameters of the distribution curve. Based on the conformational matching degree function value and the free energy change value, the conformational matching degree parameter is obtained.

3. The method according to claim 1, characterized in that, The determination of dynamic characteristic parameters of hydrogen bonds by analyzing the influence of side chain length distribution on the number of hydrogen bonds formed and the frequency of hydrogen bond breaking includes: Based on conformational matching parameters, the intermolecular hydrogen bond formation process is extracted from trajectory data. The spatial distribution characteristics of hydrogen bonds are obtained by performing double Gaussian fitting on the distance and bond angle between hydrogen bond donor and acceptor atoms. The number of hydrogen bonds is counted based on the spatial distribution characteristics of the hydrogen bonds, and the number of hydrogen bond breaks per unit time is calculated by setting a threshold for the change in hydrogen bond length to obtain the breakage frequency data. A curve showing the relationship between side chain length and hydrogen bond breakage frequency was established for the breakage frequency data, and the cumulative breakage probability distribution was obtained by curve integration. The Boltzmann weighted calculation method is used to calculate the free energy of the cumulative fracture probability distribution, and the free energy barrier data is obtained by setting the perturbation step size and controlling the equilibrium time. The free energy barrier data is optimized, the minimum free energy path is iteratively calculated, and the hydrogen bond strength, breakage time and bond length distribution data are linearly normalized in the interval of 0 to 1 to obtain the dynamic characteristic parameters of hydrogen bonds.

4. The method according to claim 1, characterized in that, If the hydrogen bond breaking frequency in the dynamic characteristic parameters of hydrogen bonds exceeds a preset threshold, the distribution of hydrophobic interaction intensity is determined by analyzing the fluctuations in hydrophobic interactions, including: If the hydrogen bond breaking frequency in the hydrogen bond dynamic characteristic parameters exceeds the preset threshold, then the conformational coordinate data are extracted based on the binding process of gallate ester to enzyme active site in molecular dynamics trajectory, and the binding free energy is calculated by temperature integral method. For the conformational coordinate data, the molecular surface area is calculated using the spherical surface area integral method, and the hydrophobic surface distribution map is obtained by scanning the solvent-accessible surface area. Based on the hydrophobic surface distribution map, the interaction energy between hydrophobic residues on the molecular surface is quantified, and the energy fluctuation amplitude is obtained by performing a discrete Fourier transform on the curve of hydrophobic interaction over time. For the energy fluctuation amplitude, the K-means clustering method is used to divide the hydrophobic interaction region on the molecular surface, and the hydrophobic interaction intensity distribution is obtained by integrating the interaction energy of hydrophobic residues in each region using the Monte Carlo method.

5. The method according to claim 1, characterized in that, The method of predicting the binding efficiency based on conformational matching parameters, hydrogen bond dynamic characteristic parameters, and hydrophobic interaction intensity distribution to obtain the predicted binding efficiency value includes: A feature matrix was constructed using conformational matching parameters, hydrogen bond dynamic characteristic parameters, and hydrophobic interaction intensity distribution. Dimensionality-reduced feature data was obtained by principal component analysis. Random sampling is performed based on the dimensionality reduction feature data, and the minimum Gini coefficient is used as the node splitting criterion to construct a random forest training set. For the random forest training set, the voting weights are calculated, and the feature importance sequence and regression prediction value are obtained by weighted averaging. Based on the feature importance sequence and regression prediction values, the tree depth is set to a range of five to fifteen. The optimal parameter combination is obtained through grid search, and the random forest is retrained to obtain the combined efficiency prediction value.

6. The method according to claim 1, characterized in that, The process involves determining the dynamic adaptability parameter based on the side chain wiggle amplitude and hydrophobic interaction fluctuation characteristics of the enzyme active site, calculating the Spearman correlation coefficient between the dynamic adaptability parameter and the predicted binding efficiency, and evaluating the catalytic efficiency, including: By optimizing the binding pocket volume and residue side chain angle of the enzyme active center, the conformational parameters of the optimized enzyme active center are determined. If the pocket deformation rate in the conformational parameters matches the preset adaptation time scale, the substrate rotational freedom and binding rate constant of the optimized conformation are verified to determine the improvement in binding efficiency. Based on the increase in binding efficiency, the side chain swing amplitude and hydrophobic interaction fluctuation characteristics of the enzyme active center were extracted, and the distribution of dynamic free energy change was calculated to obtain the dynamic adaptability parameters. The Spearman correlation coefficient between the dynamic adaptability parameter and the predicted binding efficiency is calculated, the weighted distribution of the side chain length distribution on the binding rate constant is generated, and the catalytic efficiency is evaluated to obtain the evaluation results of the catalytic efficiency.

7. The method according to claim 6, characterized in that, The optimization of the binding pocket volume and residue side chain angles of the enzyme active site determines the conformational parameters of the optimized enzyme active site. If the pocket deformation rate in the conformational parameters matches the preset adaptation time scale, the substrate rotational freedom and binding rate constant of the optimized conformation are verified to determine the improvement in binding efficiency, including: The dihedral angles of the amino acid residue side chains in the binding pocket were randomly perturbed within the torsion angle range using the Monte Carlo method, and the initial configuration set was obtained by configuration sampling of the binding pocket volume through internal coordinate optimization. Based on the initial configuration set, dynamic simulations were performed on each structure under temperature conditions. Carbon atom coordinate matrices were extracted from the trajectory data, and the principal component vectors of pocket deformation motion were obtained by diagonalizing the covariance matrix. Time series analysis was performed on the principal component vectors to obtain pocket deformation rate data by statistically analyzing the time evolution of each vibration mode; Configurations were screened based on pocket deformation rate data. For the selected configurations, the spherical coordinate grid method was used to scan the rotation angle of the substrate molecules in the optimized configurations. The binding rate constant was obtained by Boltzmann weighted averaging, and the improvement in binding efficiency was obtained by comparing the values ​​before and after optimization.

8. The method according to claim 6, characterized in that, Based on the increase in binding efficiency, the side chain twitching amplitude and hydrophobic interaction fluctuation characteristics of the enzyme active site are extracted, and the distribution of dynamic free energy change is calculated to obtain dynamic adaptability parameters, including: A three-dimensional spatial coordinate matrix is ​​obtained by periodic sampling of the conformation of the enzyme active site, and the main motion direction is obtained by calculating the side chain atom covariance based on the three-dimensional spatial coordinate matrix. Statistical analysis methods were used to calculate the displacement amplitude of side chain atoms in the main motion direction, and significant motion sites and key hydrophobic interaction residue pairs were identified through the displacement amplitude. Based on the key hydrophobic interaction residue pairs, a fast Fourier transform of the van der Waals interaction energy is performed, and the hydrophobic interaction characteristic spectrum is obtained through the fast Fourier transform. The conformational transformation is simulated using the free energy perturbation method based on the hydrophobic interaction characteristic spectrum. The dynamic free energy distribution curve is calculated based on the conformational transformation simulation, and the conformational transformation rate is obtained through the dynamic free energy distribution curve. The dynamic adaptability parameter is obtained by normalizing the rate data.

9. The method according to claim 6, characterized in that, The calculation of the Spearman correlation coefficient between the dynamic adaptability parameter and the predicted binding efficiency, the generation of a weighted distribution of side chain length distribution on the binding rate constant, and the evaluation of catalytic efficiency include: Data pairing sequences are obtained based on dynamic adaptation capability parameters and predicted combination efficiency values, and the correlation matrix is ​​calculated using Spearman rank correlation coefficient. The initial weight values ​​of the parameters are obtained by least squares fitting of the correlation matrix, and the initial weight values ​​are updated to the convergence threshold by iterative optimization method to obtain the converged weight values. Based on the convergence weight value, a piecewise integral operation is performed on the binding rate constant, and the integral result is smoothed using a Gaussian kernel function to obtain the weight distribution curve. A partial least squares regression model was established for the weighted distribution curve. The residual distribution of the partial least squares regression model was tested, and the significance level of the model was determined by the F test. A catalytic efficiency evaluation index system was established, which includes three parameters: binding rate constant, substrate conversion efficiency, and product formation rate. The catalytic efficiency evaluation model was obtained by weighted combination.

10. The method according to claim 1, characterized in that, The method further includes: generating a substrate selection scheme based on the predicted binding rate constants of various gallic acid ester structures, specifically including: Conformational data of gallate molecules were collected based on the parameter weights of the catalytic efficiency model, and the normalized binding rate between the conformation and the enzyme active site was obtained by molecular dynamics calculation. For the standardized binding rate, principal component analysis is used to extract feature vectors, and K-means clustering algorithm is used to group the feature vectors to obtain conformational family distribution data; A dual-index evaluation matrix is ​​established based on the conformational family distribution data, and the positive ideal solution and negative ideal solution are calculated using the maximum and minimum value standardization method. For the positive and negative ideal solutions, the relative proximity coefficient is calculated using the Euclidean distance formula. If the proximity coefficient is greater than a preset threshold, the corresponding solution is included in the candidate solution sequence. The candidate solution sequence is then sorted in descending order to obtain the substrate selection solution.