Method and system for regulating the ratio of amino acids to peptides after proteolysis of soy protein

By constructing a multidimensional dynamic covariance matrix and a multi-level enzyme interaction network model, the problem of implicit interaction interference in the multi-enzyme synergistic hydrolysis of soybean protein was solved, enabling precise control of the amino acid to peptide ratio and improving product consistency and predictability.

CN120913658BActive Publication Date: 2026-01-20QINGYUAN HOPE BIOTECHNOLOGY CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511447252.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-01-20
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively identify and resolve latent interactions during the multi-enzyme synergistic hydrolysis of soybean protein, resulting in inaccurate control of the amino acid to peptide ratio and significant fluctuations in product functional properties.

Method used

By collecting coupling parameters throughout the enzymatic hydrolysis process to generate time-series data, a multidimensional dynamic covariance matrix is ​​constructed to identify significant enzyme interactions and build a multi-level enzyme interaction network model, enabling dynamic and precise control of the amino acid to peptide ratio.

Benefits of technology

It improves the consistency and predictability of product functions, reduces batch-to-batch fluctuations in the content of target functional peptides, and enables real-time dynamic control of the enzymatic hydrolysis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913658B_ABST
    Figure CN120913658B_ABST
Patent Text Reader

Abstract

The application relates to the field of biotechnology, and discloses a method and system for regulating the ratio of amino acids and peptides after soybean proteolysis, which comprises collecting enzyme-coupling parameters in the whole proteolysis process to generate time series data, obtaining a multi-dimensional dynamic covariance matrix based on the data, identifying significant enzyme interaction through the matrix, and generating a list; a multi-level enzyme interaction network model is constructed according to the list, and finally, dynamic and accurate regulation of the ratio of amino acids and peptides is realized based on the model; the application can effectively solve the identification problem of hidden interaction interference in a multi-enzyme system, improve the consistency and predictability of product functions, and reduce the batch-to-batch fluctuation of the content of target functional peptides.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology, and more particularly, to a method and system for regulating the ratio of amino acids to peptides after soy protein enzymolysis. BACKGROUND

[0002] In the field of soy protein enzymolysis, as the demand for functional peptides grows, how to accurately control the ratio of amino acids to peptides during enzymolysis has become a key technical problem. Existing enzymolysis techniques can achieve some degree of product control, but when faced with multiple enzyme synergies, there are still many limitations. For example, Chinese patent application CN118389630A proposes a method for improving the content of branched-chain amino acid peptides, which significantly improves the relative content of branched-chain amino acids and the ratio of oligopeptides in the final product, and enhances the biological activity potential of the product. This method mainly focuses on refining the products after enzymolysis through physical and chemical means to improve the content and biological activity of specific peptide segments. Chinese patent application CN119626345A relates to a multifunctional self-assembling peptide recognition method and system, which analyzes peptide sequences through deep learning technology to achieve efficient recognition and classification of self-assembling peptides and various biological functional peptides, and can predict the multiple biological activity functions of peptides, providing a new technical means for functional evaluation of peptides.

[0003] However, existing technologies do not address the identification of hidden interaction interference in multiple enzyme systems. In the process of soy protein multi-enzyme synergistic enzymolysis, there are complex interactions between different enzymes, which go beyond simple additive effects. In particular, in the directed production of specific functional peptides, the phenomenon of "hidden interaction interference" often occurs: the cutting behavior of certain enzymes on specific substrates can be regulated by intermediate products produced by other enzymes in the system. This regulation may be promoting or inhibiting, and often cannot be captured by conventional monitoring methods. For example, a specific peptide segment produced by an endopeptidase A may form a temporary complex with an exopeptidase B in the system, changing its substrate specificity, making the site that is not easily cut become easily hydrolyzed. This hidden interaction is usually not discovered through single enzyme activity determination or conventional component analysis, but it has a significant impact on the ratio of amino acids to different length peptides in the final product, becoming a blind spot for precise control, resulting in fluctuations in product batches even with the same formulation and process parameters. SUMMARY

[0004] In order to overcome the above-mentioned defects of the prior art, the present application provides a method and system for regulating the ratio of amino acids and peptides after soybean protein enzymolysis, which generates time series data by collecting coupling parameters during the whole enzymolysis process, constructs a multi-dimensional dynamic covariance matrix, identifies significant enzyme interactions and generates a list, and then constructs a multi-level enzyme interaction network model, finally realizes dynamic and accurate regulation of the ratio of amino acids and peptides. The present application can effectively solve the problem of identifying hidden interaction interference in a multi-enzyme system, significantly improve the consistency and predictability of product function, and reduce batch-to-batch fluctuations in the content of target functional peptides.

[0005] To achieve the above object, the present application provides the following technical scheme:

[0006] The method for regulating the ratio of amino acids and peptides after soybean protein enzymolysis comprises:

[0007] Collecting enzyme coupling parameters of the target enzymolysis system during the whole enzymolysis process to generate enzyme coupling time series data;

[0008] Obtaining a multi-dimensional dynamic covariance matrix according to the enzyme coupling time series data;

[0009] Identifying significant enzyme interactions based on the multi-dimensional dynamic covariance matrix to generate a list of significant enzyme interactions;

[0010] Constructing a multi-level enzyme interaction network model based on the list of significant enzyme interactions;

[0011] Based on the multi-level enzyme interaction network model, the ratio of amino acids and peptides during the soybean protein enzymolysis process is dynamically regulated.

[0012] Further, the method for obtaining a multi-dimensional dynamic covariance matrix according to enzyme coupling time series data comprises: setting a sliding time window with a time length of T4, taking T5 as the moving interval, slicing the enzyme coupling time series data to obtain N time slices; calculating the covariance between each pair of data in each time slice to obtain a multi-dimensional dynamic covariance matrix.

[0013] Further, the enzyme coupling parameters include activity data of m endopeptidases EN i (i=1,2,...,m) and n exopeptidases EX j (j=1,2,...,n) in the target enzymolysis system, concentration data of p' types of amino acids AA k (k=1,2,...,p') and q types of peptides PEP g (g=1,2,...,q) of different lengths, and microwave fluctuation data of pH value;

[0014] The method for calculating the pairwise covariance between data within each time slice to obtain the multidimensional dynamic covariance matrix includes: for each time slice t (t=1,2,...,N), calculating the endopeptidase EN i (i=1,2,...,m) Activity, exopeptidase EX j (j=1,2,...,n) Activity, amino acid AA k (k=1,2,...,p') concentration, peptide PEP g The covariance between concentration and pH value (g=1,2,...,q) forms the covariance matrix C for time slice t. t ; the covariance matrix C t Each element in the matrix is ​​standardized to a correlation coefficient, resulting in the correlation matrix R for time slice t. t ; the covariance matrix C of N time slices t Construct a dynamic covariance tensor C by combining the time sequences, and combine the correlation matrices R of N time slices. t By combining them in chronological order, a dynamic correlation tensor R is constructed; based on the dynamic covariance tensor C and the dynamic correlation tensor R, a multidimensional dynamic covariance matrix is ​​constructed.

[0015] Furthermore, the method for identifying significant enzyme interactions and generating a list of significant enzyme interactions includes:

[0016] Eigenvalue decomposition is performed on the multidimensional dynamic covariance matrix to obtain the dimension-reduced feature space.

[0017] In the reduced feature space, calculate the significance level of the covariance between endopeptidase activity, exopeptidase activity, amino acid concentration, peptide concentration and pH value.

[0018] Set a significance threshold α, and automatically label covariances with significance levels exceeding the significance threshold as significant enzyme interactions;

[0019] A list of significant enzyme interactions is generated based on the labeled significant enzyme interactions.

[0020] Furthermore, the method for constructing a multi-level enzyme interaction network model based on a list of significant enzyme interactions includes:

[0021] A homology interaction network was constructed between endopeptidases and between exopeptidases as a first-level model.

[0022] Based on a list of significant enzyme interactions, a direct interaction network between endopeptidases and exopeptidases is constructed as a second-level model.

[0023] Based on a list of significant enzyme interactions, a comprehensive indirect regulatory network is constructed as a third-level model.

[0024] Fusion of the first level model, the second level model and the third level model, and construction of a multi-level enzyme interaction network model.

[0025] Further, the method for constructing the homology interaction network between endopeptidases and endopeptidases and between exopeptidases and exopeptidases comprises:

[0026] Obtain sequence information and three-dimensional structure data of all endopeptidases and exopeptidases in the target enzymatic system;

[0027] Based on the sequence information of endopeptidases and exopeptidases, identify enzyme pairs with potential homology interaction;

[0028] Based on the three-dimensional structure data of endopeptidases and exopeptidases, confirm enzyme pairs with homology interaction from enzyme pairs with potential homology interaction, and construct a homology interaction network.

[0029] Further, the method for identifying enzyme pairs with potential homology interaction is: based on the sequence information of endopeptidases and exopeptidases, calculate the sequence similarity between all endopeptidases and the sequence similarity between all exopeptidases, and mark enzyme pairs with sequence similarity exceeding the sequence similarity threshold as enzyme pairs with potential homology interaction.

[0030] Further, the method for confirming enzyme pairs with homology interaction from enzyme pairs with potential homology interaction comprises: based on the three-dimensional structure data of endopeptidases and exopeptidases, calculate the structure similarity of enzyme pairs with homology interaction, and confirm enzyme pairs with structure similarity exceeding the structure similarity threshold as enzyme pairs with homology interaction.

[0031] Further, the method for constructing a comprehensive indirect regulation network comprises:

[0032] Based on the significant enzyme interaction list, construct a bipartite graph B(V E, V S, E), wherein V E is a set of enzyme nodes containing all endopeptidases and exopeptidases; V S is a set of substrate nodes containing all amino acids and peptide segments, and E is a set of enzyme-substrate interaction edges.

[0033] Based on the bipartite graph B, deduce the indirect regulation relationship between enzymes;

[0034] According to the indirect regulation relationship between enzymes, construct an indirect regulation network between endopeptidases and exopeptidases, an indirect regulation network between endopeptidases and endopeptidases, and an indirect regulation network between exopeptidases and exopeptidases, and combine the three networks to form a comprehensive indirect regulation network.

[0035] A soybean protein enzymolysis amino acid and peptide ratio control system for implementing the soybean protein enzymolysis amino acid and peptide ratio control method described above, the system comprising:

[0036] data acquisition module: for collecting the enzyme hydrolysis coupling parameters of the target enzyme hydrolysis system in the whole enzyme hydrolysis process, and generating enzyme hydrolysis coupling time series data;

[0037] slice calculation module: for obtaining a multi-dimensional dynamic covariance matrix according to the enzyme hydrolysis coupling time series data;

[0038] significant interaction identification module: for identifying significant enzyme interactions based on the multi-dimensional dynamic covariance matrix, and generating a significant enzyme interaction list;

[0039] interaction network construction module: for constructing a multi-level enzyme interaction network model based on the significant enzyme interaction list;

[0040] regulation module: for dynamically regulating the ratio of amino acids and peptides in the soybean protein enzyme hydrolysis process based on the multi-level enzyme interaction network model.

[0041] Compared with the prior art, the beneficial effects of the present application are:

[0042] The present application can comprehensively and dynamically reflect various changes in the enzyme hydrolysis process by collecting the coupling parameters of the whole enzyme hydrolysis process and generating time series data. The multi-dimensional dynamic covariance matrix constructed using these data provides a basis for in-depth analysis of the complex interactions between enzymes. The significant enzyme interaction list identified based on the matrix accurately filters out the key factors that have an important influence on the enzyme hydrolysis process. The multi-level enzyme interaction network model further constructed integrates the complex relationships of homology interaction, direct interaction and indirect regulation, and comprehensively analyzes the regulation mechanism in the enzyme hydrolysis system. Finally, the dynamic and accurate regulation based on this model makes the control of the ratio of amino acids and peptides in the soybean protein enzyme hydrolysis process more accurate and meticulous. The present application can capture the implicit interaction interference in the multi-enzyme system in real time, accurately reflect the complex interactions between enzymes and enzymes, and between enzymes and substrates in the enzyme hydrolysis system, improve the explanation and prediction accuracy of the model for the enzyme hydrolysis process, and provide comprehensive model support for the accurate regulation of the enzyme hydrolysis system. The system can adjust the regulation strategy in real time according to the dynamic changes of the enzyme hydrolysis process, thereby effectively improving the controllability of the ratio of amino acids and peptides in the product and the consistency between batches. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, brief descriptions will be given below for the drawings needed to be used in the embodiments or prior art descriptions. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0044] Figure 1A method flow chart of the amino acid and peptide ratio regulation method after soybean protein enzymolysis in the application;

[0045] Figure 2 A principle flow chart of constructing the first level model in the application;

[0046] Figure 3 A method flow chart of constructing the comprehensive indirect regulation network in the application;

[0047] Figure 4 A fusion method schematic diagram of the multi-level enzyme interaction network model in the application;

[0048] Figure 5 A functional module diagram of the amino acid and peptide ratio regulation system after soybean protein enzymolysis in the application. DETAILED DESCRIPTION

[0049] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.

[0050] Embodiment 1

[0051] Please refer to Figure 1 The embodiment provides an amino acid and peptide ratio regulation method after soybean protein enzymolysis, including:

[0052] Step S10, collecting the enzymolysis coupling parameters of the target enzymolysis system in the whole enzymolysis process to generate an enzymolysis coupling time sequence data;

[0053] The enzymolysis coupling parameters include the activity data of m endopeptidases EN i (i=1, 2,..., m) and n exopeptidases EX j (j=1, 2,..., n) in the target enzymolysis system, i and j are index variables of the number of endopeptidase and exopeptidase species respectively, p' kinds of amino acids AA k (k=1, 2,..., p') and q kinds of different length peptide segments PEP gconcentration data of amino acids and peptides (g = 1, 2, …, q) and the fluctuation data of pH value, k and g are index variables of the number of amino acid types and the number of different lengths of peptide segments, respectively. The specific collection method can be realized by setting multiple high-precision sensors in the enzyme reaction system. These sensors can monitor the above parameters in real time and transmit the data to the data processing system. For example, for the monitoring of enzyme activity, fluorescence resonance energy transfer (FRET) technology or enzyme-coupled colorimetric method can be used, which can sensitively reflect the activity change of the enzyme in the reaction process; for the detection of amino acid and peptide concentration, high performance liquid chromatography (HPLC) or mass spectrometry (MS) can be used, which can accurately quantitatively analyze various components in complex mixtures; and for the monitoring of pH value, high-precision and high-frequency pH electrode can be used to capture small pH fluctuations and transient changes in buffer system caused by the change of buffer system, such as the sudden drop of pH caused by the production of organic acid in enzyme hydrolysis. The data collection is carried out at a sampling frequency of 1 minute / minute for enzyme activity, 10 seconds / second for pH value, and 5 minutes / minute for amino acid and peptide concentration. The collected data is sorted and summarized according to the time stamp, and the enzyme hydrolysis coupled time series data is obtained, which records the dynamic change trajectory of each key parameter in the whole enzyme hydrolysis process, and provides detailed basic data support for the subsequent in-depth analysis and regulation of enzyme hydrolysis process.

[0054] In the prior art, traditional enzyme hydrolysis monitoring only collects key enzyme activity and main product concentration, and the sampling interval is very long, such as 30 minutes, which cannot capture the sudden change of enzyme activity caused by pH transient fluctuation. Step S10 solves the problem of loss of enzyme hydrolysis dynamic characteristics caused by data fragmentation in the prior art by high-frequency collection of all parameters, and provides a time-continuous multi-dimensional data basis for subsequent interactive analysis. For example, when the activity of endopeptidase EN2 and the concentration of peptide PEP5 show a positive correlation with a 10-minute lag in the time series, it can be inferred that EN2 has an indirect effect on the generation of PEP5, and the traditional low-frequency sampling may misjudge the correlation as irrelevant.

[0055] Step S20, obtaining a multi-dimensional dynamic covariance matrix according to the enzyme hydrolysis coupled time series data;

[0056] Further, step S20 includes:

[0057] Step S21, setting a sliding time window with a time length of T4, and slicing the enzyme hydrolysis coupled time series data with T5 as the moving interval to obtain N time slices;

[0058] Step S22, standardizing the data in each time slice;

[0059] The activity data of m endopeptidases and n exopeptidases, the concentration data of p' kinds of amino acids and q kinds of different length peptides, and the microwave fluctuation data of pH value are included in each time slice. The setting of the sliding time window is based on the dynamic characteristics of the enzyme hydrolysis process, and the determination of the time length T4 and the moving interval T5 needs to consider many factors such as the speed of enzyme hydrolysis reaction, sampling frequency and data processing complexity. For example, T4 can be set to 5-10 min and T5 can be set to 1-2 min, which can not only ensure the effective capture of key dynamic information of the enzyme hydrolysis process, but also avoid the processing difficulty caused by too large amount of data. In addition, in order to ensure the mathematical feasibility of covariance estimation, the number of samples corresponding to T4 should not be less than twice the total number of variables (m+n+p'+q+1). When the sampling frequency is fixed, T4 will automatically adjust to meet the minimum sample number condition.

[0060] In specific implementation, the enzyme hydrolysis coupled time series data is taken as input, and from the starting point of the time series, the first time window with a time length of T4 is intercepted first to obtain all parameter data in the window and form the first time slice. Then, the second time slice is intercepted by moving T5 from the starting point of the first time window, and the above operation is repeated until the entire time series is traversed, and finally N time slices are obtained. The normalization processing is to scale the data in proportion to make it fall within a specific interval, usually [-1, 1] or [0, 1], which aims to eliminate the influence of dimension difference between different indicators. The data slicing processing by sliding time window can make the data have certain overlap and continuity, because T5 < T4, so it can ensure that part of the common data information is retained between adjacent time slices, thereby better reflecting the continuity and correlation of parameter changes in the enzyme hydrolysis process and avoiding information loss caused by excessive dispersion between data segments. Through the overlapping window design, high-density sampling of the enzyme hydrolysis dynamic process is realized, and the complete evolution trajectory of the event occurring in the target enzyme hydrolysis system is continuously tracked, thereby providing a time-sequentially continuous data set for subsequent covariance analysis.

[0061] In the prior art, static data at fixed time points such as 0, 30 and 60 min after the start of enzyme hydrolysis is usually used for analysis, which cannot capture the rapid interaction within 5-10 min, such as the strong correlation between the increase of EN2 activity and the increase of AA5 concentration within 7-9 min. This step solves the time resolution problem of interactive signals in the dynamic process by using overlapping sliding window, provides time-sequentially accurate interactive data for subsequent dynamic network construction, and can realize time positioning of enzyme-substrate interaction, such as determining that the strong correlation between EN1 and PEP4 occurs within the time slice of 15-20 min of enzyme hydrolysis; and can capture short-duration transient interaction, such as the strong negative correlation between EN4 activity and AA2 concentration within 3 min when the pH drops suddenly.

[0062] Step S23, calculate the covariance between each pair of data in each time slice to obtain a multi-dimensional dynamic covariance matrix.

[0063] Further, step S23 includes:

[0064] Step S231, for each time slice t (t = 1, 2,..., N), calculate the covariance between the endopeptidase EN i (i = 1, 2,..., m) activity, exopeptidase EX j (j = 1, 2,..., n) activity, amino acid AA k (k = 1, 2,..., p') concentration, peptide segment PEP g (g = 1, 2,..., q) concentration and pH value, forming a covariance matrix C t of time slice t.

[0065] Step S232, standardize each element in the covariance matrix C t to a correlation coefficient to obtain a correlation matrix R t of time slice t, wherein the sign of the correlation coefficient represents the mode of action, and the absolute value represents the strength of the action.

[0066] Step S233, combine the covariance matrices C t of the N time slices in time sequence to construct a dynamic covariance tensor C, and combine the correlation matrices R t of the N time slices in time sequence to construct a dynamic correlation tensor R.

[0067] Step S234, based on the dynamic covariance tensor C and the dynamic correlation tensor R, construct a multi-dimensional dynamic covariance matrix.

[0068] The interaction between enzyme activity, substrate concentration and pH during enzymatic hydrolysis has time-varying characteristics. The enzyme is an endopeptidase and an exopeptidase, and the substrate includes amino acids and peptide segments. Traditional static covariance analysis cannot reflect this dynamic correlation. Through time slice level covariance calculation and tensor construction, the time evolution characteristics of the interaction relationship can be converted into quantifiable matrix data, providing a basis for subsequent significant interaction identification. For each time slice t (t = 1, 2,..., N), calculate the covariance between the endopeptidase EN i (i = 1, 2,..., m) activity, exopeptidase EX j (j = 1, 2,..., n) activity, amino acid AA k (k = 1, 2,..., p') concentration, peptide segment PEP g (g = 1, 2,..., q) concentration and pH value, forming a covariance matrix C tThe covariance is calculated using a standard formula to measure the degree of co-variation between variables. The variables refer to endopeptidase activity, exopeptidase activity, amino acid concentration, peptide segment concentration, and pH value. For example, if the covariance between EN1 activity and PEP3 concentration is positive, it indicates that when EN1 activity increases, PEP3 concentration tends to increase, which may imply that EN1 catalyzes PEP3; if the covariance is negative, it may indicate that EN1 inhibits the generation of PEP3. In the prior art, global covariance analysis is often used, such as calculating a single covariance matrix for the entire enzymatic hydrolysis process, which cannot capture the differences in the interaction between the early stage, such as 0-30 min, and the middle and late stages, such as 30-60 min. For example, the covariance between EN3 and PEP5 is 0.5 in the early stage and -0.3 in the middle and late stages. Through time-slice level covariance calculation, the problem of "insufficient time resolution in dynamic interaction process" is solved, which can capture the change in the interaction mode in the enzymatic hydrolysis process, such as the change in the covariance between EX1 and PEP2 from 0.3 to -0.6 when the pH decreases from weak alkaline to neutral, reflecting the change in the mode of pH regulation of enzyme activity.

[0069] The method for standardizing each element in the covariance matrix C t to a correlation coefficient is to divide the covariance value of any two variables by the product of the standard deviations of the two variables. The value range is [-1, 1], which eliminates the influence of the variance of the variable itself and only retains the degree of linear correlation, which is used to identify weak signal interactions ignored by traditional methods and establish a unified interaction strength evaluation standard. The sign indicates the mode of action, representing positive and negative correlations respectively, and the absolute value indicates the strength of the action. The covariance matrix C t of N time slices is combined in time sequence to form a tensor C with dimensions (m+n+p'+q+1) x (m+n+p'+q+1) x N, which reflects the time evolution of the covariance value; similarly, the correlation matrix R t of N time slices is combined to form a tensor R with the same dimensions, which reflects the time variation of the correlation coefficient. The tensor structure can simultaneously retain the time dimension and variable interaction information, for example, the t-th slice in the tensor C corresponds to the covariance matrix of time slice t, which can directly reflect the time evolution of enzyme interaction strength; the sign change of the elements in the tensor R can track the dynamic conversion of the interaction mode, which is positive / negative correlation. By integrating time series through tensor dimension, the enzyme hydrolysis interaction is upgraded from a "static network" to a "dynamic evolution system". The multi-dimensional dynamic covariance matrix integrates the complete covariance information of endopeptidases, exopeptidases, amino acids, peptide segments, and pH in the enzyme hydrolysis process, forming a spatiotemporal coupled interaction network, which provides a structured input for the significant enzyme interaction identification of step S30. Steps S10 and S20 combine "high-frequency acquisition + dynamic slicing" to push the interaction analysis of the enzyme hydrolysis system from the "macro trend" to the "instantaneous dynamic" level, laying the foundation for real-time regulation.

[0070] Step S30, based on the multi-dimensional dynamic covariance matrix, identifying significant enzyme interaction, generating a significant enzyme interaction list;

[0071] Further, step S30 includes:

[0072] Step S31, eigenvalue decomposition is performed on the multi-dimensional dynamic covariance matrix to obtain a reduced dimension feature space;

[0073] Step S32, in the reduced dimension feature space, the significant level of covariance between endopeptidase activity, exopeptidase activity, amino acid concentration, peptide segment concentration and pH value is calculated;

[0074] Step S33, set a significant threshold a, and automatically mark the covariance with a significant level exceeding the significant threshold as a significant enzyme interaction;

[0075] Step S34, according to the marked significant enzyme interaction, generating a significant enzyme interaction list; the significant enzyme interaction list contains interaction entities, action intensity, action mode and time slice identification.

[0076] The covariance matrix between parameters in the enzyme system contains a large amount of redundant information, which needs to be extracted by dimension reduction, and combined with statistical significance test to exclude false correlation caused by random fluctuations, so as to accurately locate the enzyme-enzyme and enzyme-substrate interaction relationship which plays a leading role in the regulation of amino acid and peptide ratio. The purpose of eigenvalue decomposition of the multi-dimensional dynamic covariance matrix is to reduce dimension and filter noise. The covariance matrix with dimension DxD is decomposed into the product of eigenvector matrix and eigenvalue matrix through linear algebra operation, the first K' eigenvectors which can explain more than 85% of the data variance contribution are extracted to form the K' dimensional feature space after dimension reduction, wherein D is the total number of parameters such as enzymes, substrates and pH. For example, when D=50, if the cumulative contribution of the first 5 eigenvalues is 90%, the 50-dimensional data is projected into a 5-dimensional space. Direct analysis of high-dimensional covariance matrix is easily affected by "dimension disaster", and cannot identify key interactions. Through dimension reduction, noise interference can be filtered, and dominant covariance patterns can be focused, solving the problem of difficulty in extracting key patterns in high-dimensional data in the prior art; at the same time, the complexity of subsequent significance calculation is reduced, and the calculation efficiency is improved.

[0077] In the K-dimensional feature space after dimension reduction, the significance level of the covariance between parameters is calculated by statistical hypothesis testing method. Specifically, the original parameters such as endopeptidase EN i activity and peptide segment PEP gCovariance. The null hypothesis "covariance is 0" is established for the covariance between these projected variables, and the p-value of the observed covariance is calculated by t-test or similar statistical methods. The smaller the p-value, the less likely the covariance is generated by random fluctuations. The lack of this statistical verification in the prior art often misjudges the covariance with large absolute value as a significant interaction, i.e. misjudges random fluctuations as real interactions, which will mistakenly include significant lists, and the present application reduces such misjudgments by p-value screening. The setting of the significance threshold a usually uses the standard value 0.05 widely accepted in statistical analysis, which means that at a confidence level of 95%, it is judged whether there is a non-random covariance between variables. For multiple testing situations, further control measures such as Bonferroni correction can be introduced to improve the accuracy of significance judgment. When the p-value is less than or equal to a, it is considered that there is a real biological correlation between the variables; otherwise, it is considered as random fluctuations. Therefore, the covariance with a p-value less than or equal to a is marked as a significant enzyme interaction. In the reduced feature space, indirect interaction paths that cannot be found by traditional methods can be identified. For example, EN7 and PEP 12 have no direct significant covariance, but both are highly correlated with the "AA8 generation" feature vector in the feature space, implying the existence of an EN7→AA8→PEP 12 two-step indirect path. This discovery improves the recognition rate of indirect paths when constructing the indirect regulation network in step S40, revealing the complex regulation mechanism of "substrate mediation" in the enzyme system.

[0078] The significant enzyme interaction list contains four elements: interaction entity, action strength, action mode, and time slice identifier. The interaction entity specifies the enzyme type and the action object, the action strength records the specific value of the covariance, such as covariance = 0.58; the action mode marks positive / negative correlation, such as inhibition marked as negative correlation; the time slice identifier associates the time phase of the interaction occurrence, such as time slice t = 15. Table 1 is an example of the significant enzyme interaction list.

[0079] Table 1 Example of Significant Enzyme Interaction List

[0080] Enzyme type Action object Covariance value Action mode Time slice Endopeptidase EN1 Peptide segment PEP3 0.42 Positive correlation t=5 Exopeptidase EX2 amino acid AA5 -0.38 Negative correlation t=10 Endopeptidase EN3 Peptide segment PEP5 0.65 Positive correlation t=15 Exopeptidase EX4 amino acid AA7 -0.52 Negative correlation t=20 Endopeptidase EN5 Peptide segment PEP8 0.39 Positive correlation t=25

[0081] Step S30 extracts the main covariance patterns through eigenvalue decomposition and combines the significance level calculation and significance threshold setting, which can accurately screen out significant enzyme interactions that have a key impact on the ratio of amino acids and peptides from a large number of complex enzyme interactions. This solves the problem in the prior art that it is difficult to accurately identify key interactions in a complex multi-enzyme system. The generated significant enzyme interaction list records the type, object, strength and pattern of enzyme interaction and other key information, providing a quantitative and structured data basis for subsequent multi-level enzyme interaction network model construction and dynamic enzyme dosage regulation strategy formulation, avoiding model bias and regulation errors caused by ambiguous or incomplete data. Eigenvalue decomposition realizes effective dimensionality reduction of the multi-dimensional dynamic covariance matrix, removes noise and redundant information in the data, and highlights the key covariance patterns. This not only improves the calculation efficiency, but also enables the subsequent analysis to focus more on factors that have an important impact on the enzyme hydrolysis process, avoiding interference from a large amount of useless information. The entire process of step S30 can adapt to the dynamic changes in the enzyme hydrolysis process, because the multi-dimensional dynamic covariance matrix itself is constructed based on time series data, reflecting the evolution characteristics of enzyme interaction over time. Therefore, by repeatedly executing step S30 at different time slices, significant enzyme interactions in the enzyme hydrolysis process can be dynamically tracked and identified, making real-time regulation possible.

[0082] Step S40, based on the significant enzyme interaction list, constructs a multi-level enzyme interaction network model;

[0083] The prior art often only focuses on the direct catalytic action of enzymes, ignoring the functional synergy brought by homology and long-range regulation mediated by substrates, resulting in the inability to comprehensively analyze the complex regulation mechanism in the enzyme hydrolysis process. Step S40 realizes multi-dimensional correlation analysis of evolutionary relationships, physical interactions and metabolic regulation in the enzyme hydrolysis system by integrating the three layers of homology, direct action and indirect regulation, breaking through the limitations of single-dimensional modeling, and can comprehensively capture the complex interaction relationships such as functional synergy of homologous enzyme groups, direct interaction between enzymes, and indirect regulation mediated by substrates in the enzyme hydrolysis process, providing systematic model support for accurately analyzing the regulation mechanism in the enzyme hydrolysis process, and effectively solving the problem in the prior art that the regulation mechanism is not fully analyzed due to single modeling dimension.

[0084] Further, step S40 includes:

[0085] Step S41, construct the homology interaction network between endopeptidases and endopeptidases, and between exopeptidases and exopeptidases as the first level model;

[0086] Please refer to Figure 2 Further, step S41 includes:

[0087] Step S411, obtain sequence information and three-dimensional structure data of all endopeptidases and exopeptidases in the target enzymatic system;

[0088] Step S412, based on the sequence information of the endopeptidases and the exopeptidases, identify enzyme pairs with potential homology interaction;

[0089] The method for identifying enzyme pairs with potential homology interaction is: based on the sequence information of the endopeptidases and the exopeptidases, calculating the sequence similarity between all endopeptidases and the sequence similarity between all exopeptidases, for enzyme pairs with sequence similarity exceeding a sequence similarity threshold, marking as enzyme pairs with potential homology interaction, otherwise marking as enzyme pairs without potential homology interaction.

[0090] Step S413, based on the three-dimensional structure data of the endopeptidases and the exopeptidases, confirming enzyme pairs with homology interaction from enzyme pairs with potential homology interaction, and constructing a homology interaction network; the nodes of the homology interaction network represent specific enzyme molecules, and the edges represent homology interaction, and the weight of the edge is determined by the weighted combination of sequence similarity and structure similarity.

[0091] The method for confirming enzyme pairs with homology interaction from enzyme pairs with potential homology interaction includes: based on the three-dimensional structure data of the endopeptidases and the exopeptidases, calculating the structure similarity of enzyme pairs with homology interaction, enzyme pairs with structure similarity exceeding a structure similarity threshold are confirmed as enzyme pairs with homology interaction, otherwise marking as enzyme pairs without homology interaction.

[0092] Homologous enzymes usually have similar domains and catalytic mechanisms, and through double screening of sequence and structure, potential functionally associated enzyme pairs can be identified, providing an evolutionary relationship-based structure basis for subsequent network models. Traditional homology analysis only relies on sequence alignment, such as BLAST, ignoring three-dimensional structure differences, leading to misjudgment of non-functional homologous pairs, such as enzymes with similar sequences but significant differences in active centers. The present application solves the problem of insufficient accuracy of homology recognition through the double mechanism of "sequence screening + structure verification".

[0093] The sequence information and three-dimensional structure data of all endopeptidases and exopeptidases in the target enzymatic system can be extracted from authoritative databases such as RCSB Protein Database (PDB), UniProt and BRENDA, etc., for obtaining functional annotations of enzymes, such as catalytic type and substrate specificity. Basic local alignment search tool BLAST and multiple sequence alignment tool ClustalOmega are used to calculate the sequence similarity between enzymes. Taking BLAST as an example, the expected value E-value threshold is set to 1x10 -5, the matching length needs to cover more than 70% of the enzyme sequence, and the sequence similarity score in percentage form is calculated. When the sequence similarity between the endopeptidase EN i and EN j exceeds a set sequence similarity threshold, such as 60%, it is marked as an enzyme pair with potential homologous interaction, referred to as a potential homologous pair, excluding false homologous pairs with large structural differences but similar sequences.

[0094] The structure similarity of the potential homologous pair is calculated using the structure superimposition tool TM-align and the three-dimensional structure alignment tool DALI. Taking TM-align as an example, the three-dimensional structure coordinates of the enzyme pair are input, and the TM-score value is calculated, which ranges from 0 to 1. When the value exceeds a preset structure similarity threshold, such as 0.8, it is confirmed to have homologous interaction. In the final constructed homologous interaction network, the nodes are enzyme molecules, and the edge weights are obtained by weighting the sequence similarity and the structure similarity.

[0095] Step S42, based on the significant enzyme interaction list, a direct interaction network between endopeptidases and exopeptidases is constructed as a second-level model;

[0096] Endopeptidases are enzymes that can cleave internal peptide bonds in polypeptide chains, and exopeptidases are enzymes that hydrolyze amino acids one by one from the ends of polypeptide chains. The direct interaction network is used to represent the direct physical contact or functional association between these two types of enzymes. The method for constructing the direct interaction network between endopeptidases and exopeptidases includes extracting enzyme pairs with significant covariance between endopeptidases and exopeptidases from the generated significant enzyme interaction list, taking the action strength and mode of these enzyme pairs as basic data, and using molecular docking technology such as AutoDockVina to perform structure simulation on the enzyme pairs. The three-dimensional structure models of endopeptidases and exopeptidases are superimposed, and the binding energy is calculated. If the binding energy is lower than a set binding energy threshold, such as -5 kcal / mol, it is confirmed that there is a direct physical interaction. By constructing the direct interaction network, the direct synergistic or antagonistic action between endopeptidases and exopeptidases can be identified.

[0097] Step S43, based on the significant enzyme interaction list, a comprehensive indirect regulation network is constructed as a third-level model;

[0098] Please refer to Figure 3 , further, step S43 includes:

[0099] Step S431, based on the significant enzyme interaction list, a bipartite graph B is constructed;

[0100] In addition to direct physical effects, there are a large number of indirect regulations mediated by substrates or intermediates in the enzyme system, but the existing technology lacks systematic modeling of such indirect interactions, which leads to the inability to explain complex phenomena such as "substrate competition" and "product feedback inhibition" in the enzyme degradation process. The present application solves the problem of difficult quantification of indirect regulation relationship through the three-layer architecture of "significance interaction screening-bipartite graph modeling-path analysis".

[0101] Based on the list of significant enzyme interactions, the method for constructing bipartite graph B is: extracting all enzyme-amino acid and enzyme-peptide segment interactions from the list of significant enzyme interactions, defining the extracted enzyme-amino acid and enzyme-peptide segment as enzyme-substrate pairs, and recording the action strength and action mode of the enzyme-substrate pairs; according to the extracted enzyme-substrate pairs, the action strength and the action mode of the enzyme-substrate pairs, a bipartite graph B (VE, VS, E) is constructed, wherein VE is a set of enzyme nodes, including all endopeptidases and exopeptidases; VS is a set of substrate nodes, including all amino acids and peptide segments, and E is a set of enzyme-substrate interaction edges, the weight of the edge is determined by the action strength, and the type of the edge is determined by the action mode. Through the form of bipartite graph, the interaction relationship between enzymes and substrates is directly displayed, making the complex enzyme interaction mechanism more clear and easy to understand. This visualization method not only facilitates researchers to understand the regulation mechanism in the enzyme degradation process, but also provides structured data support for subsequent network analysis, enabling the interaction relationship to be upgraded from "numerical list" to "correlation map", discovering some enzyme-substrate interaction relationships that are easily overlooked in traditional analysis, improving the interaction identification efficiency, and thus providing a new perspective for optimizing the enzyme degradation process. By structuring the dispersed interaction pairs through bipartite graph, the problem of "fragmentation of enzyme-substrate interaction data" in the prior art is solved, making it possible to deduce substrate-mediated indirect regulation.

[0102] Step S432, based on the bipartite graph B, deducing the enzyme-enzyme indirect regulation relationship;

[0103] The enzyme-enzyme indirect regulation relationship includes indirect antagonistic relationship and indirect synergistic relationship. The enzyme pairs sharing the same substrate node in the bipartite graph are traversed, if the endopeptidase and the exopeptidase act on the same substrate such as amino acid or peptide segment, and the action mode is opposite, for example, one promotes and one inhibits, it is considered that there is an indirect antagonistic relationship between the two enzymes mediated by the substrate. If the action mode is the same, for example, both promote or both inhibit, it is considered that there is an indirect synergistic relationship mediated by the substrate. Further, through the path analysis method, a multi-step cascade indirect regulation path is identified. For example, if the endopeptidase EN i influences the production of the amino acid AA k , and the amino acid AA k influences the activity of the exopeptidase EX j , there is an indirect regulation path from the endopeptidase EN i to the exopeptidase EX jThe path strength is the product of all edge weights on the path, and the path type is either promotion or inhibition, determined by the parity of the number of inhibitory relationships on the path, e.g., if there is one inhibition on the path, the overall effect is inhibitory. Based on the enzyme-substrate-enzyme signal transmission logic, all possible indirect paths are identified by path enumeration or traversal algorithms in graph theory. For example, EN3 promotes the generation of peptide segment PEP6, which inhibits the activity of exopeptidase EX4, forming an EN3-PEP6-EX4 secondary indirect inhibitory path with a path strength of 0.3 x (-0.2) = -0.06.

[0104] In step S433, endopeptidase-exopeptidase indirect regulation networks, endopeptidase-endopeptidase indirect regulation networks, and exopeptidase-exopeptidase indirect regulation networks are constructed according to enzyme-enzyme indirect regulation relationships, and the three networks are combined to form a comprehensive indirect regulation network.

[0105] While combining the three networks to form a comprehensive indirect regulation network, the main types of indirect regulation, such as indirect synergy or indirect antagonism, can be inferred and labeled in the network according to the results of path analysis, and the time-varying contribution of each indirect regulation path to the entire enzymatic process is calculated to identify the dominant indirect regulation mechanism at different stages of enzymatic hydrolysis. Among them, the construction of endopeptidase-exopeptidase indirect regulation networks is based on bipartite graphs to extract all indirect paths of endopeptidases to exopeptidases mediated by substrates, with path strength being the product of edge weights and type being determined by the parity of the number of inhibitory relationships; endopeptidase-endopeptidase indirect regulation networks are used to identify synergistic or antagonistic paths between endopeptidases through common substrates; exopeptidase-exopeptidase indirect regulation networks analyze the regulatory relationships between exopeptidases through substrate competition. When combining the three types of networks, all nodes and edges are retained, the path type is marked by different colors and the regulatory mechanism is labeled, and the time-varying contribution is calculated by weighting the significant covariance of each indirect path at different time slices with time slice weights and then dividing by the total time slice weight.

[0106] The prior art indirect regulation analysis often uses global correlation analysis, which cannot distinguish path types and time dynamics, leading to failure to identify the synergistic effect of enzyme families on the same substrate or neglecting the transient regulation of enzymes through substrate competition. The present application systematically analyzes the indirect regulation across enzyme classes through classification modeling and dynamic integration, and clearly determines the contribution of different types of paths in the enzyme digestion process, such as the endopeptidase-exopeptidase path that dominates the balance between peptide segment generation and degradation, the endopeptidase-endopeptidase path that dominates the initial cleavage efficiency of proteins, and the exopeptidase-exopeptidase path that affects the accumulation rate of amino acids. At the same time, time-varying contribution analysis is used to dynamically locate the key regulation stage. For example, if a path has a higher contribution at a specific time period of enzyme digestion, it suggests that this stage is a critical window for intervention. The "time slice weight" can be set according to the characteristics of the enzyme digestion process or expert experience, for example, a higher weight can be given to the key time slice with a higher product generation rate, or a uniform weight can be used. The indirect regulation layer cooperates with the homology network and the direct action network. If the indirect regulation layer is missing, the multi-level network will lack the dimension of indirect regulation, leading to biased key degree index calculation and the inability to implement intervention based on indirect paths. The significant interaction screening of S30 cooperates with the indirect regulation layer to ensure the reliability of the indirect path, and cooperates with the path derivation in the bipartite graph to systematize the fragmented path information and improve the prediction accuracy of indirect regulation. Through the closed-loop process of "significance screening-path derivation-classification merging", the present application provides intervention targets for non-directly acting enzyme pairs for precise regulation, breaking through the traditional regulation limitation of only focusing on direct catalytic enzymes.

[0107] Step S44, fusing the constructed first-level model, second-level model and third-level model, constructing a multi-level enzyme interaction network model.

[0108] Specifically, please refer to Figure 4As shown, the multi-level enzyme interaction network model is obtained by fusing the first-level model, the second-level model and the third-level model, specifically, based on the heterogeneous graph fusion method in graph theory, the first-level homology interaction network, the second-level direct interaction network and the third-level comprehensive indirect regulation network are topologically integrated. Specifically, when fusing, the adjacency matrices of the three networks are expanded along the dimensions in the form of tensor splicing to form a three-dimensional interaction tensor containing three relationship types of homology, direct action and indirect regulation, while retaining the unique node attributes of each layer. By setting a weight mapping function, the edge weights of different levels are normalized to a unified dimension, for example, the sequence similarity of the homology network, the covariance value (real number range) of the direct action network and the path strength (product accumulation value) of the indirect regulation network are standardized by Z-score, and then linearly combined according to the weight coefficients β1=0.3, β2=0.5, β3=0.2 (determined by historical enzyme data training) to generate a comprehensive interaction strength value as the final edge weight in the fused network, wherein β1 corresponds to the weight coefficient of the first-level homology interaction network, which is used to quantify the contribution of sequence similarity and structural similarity between enzyme molecules to the comprehensive interaction strength; β2 corresponds to the weight coefficient of the second-level direct interaction network, which is used to represent the dominant role of the covariance value (direct action strength) calculated based on the significant enzyme interaction list in the comprehensive interaction; β3 corresponds to the weight coefficient of the third-level comprehensive indirect regulation network, which is used to measure the influence of the path strength of the indirect regulation mediated by the substrate on the overall interaction relationship, such as the weighted value of the product of the path edge weight and the time contribution. The multi-level network fusion retains all the indirect path information from the enzyme to the substrate in the indirect regulation network, including the path edge weight, the action mode and the time-varying contribution, providing a structured data basis for calculating the indirect action strength in the subsequent S50.

[0109] The weight coefficients β1, β2, β3 are determined by the following training process based on historical enzyme data: first, a historical data set containing at least 200 groups of industrial enzyme hydrolysis batches is constructed, each group of data containing the original parameters of each level model and the corresponding enzyme hydrolysis product component detection results, such as sequence similarity values of homology network, covariance matrix of direct action network, path strength values of indirect regulation network, and actual proportions of amino acids and peptides. Using the leave-one-out cross-validation strategy, the data set is divided into a training set (199 groups) and a validation set (1 group), and the objective function is optimized by the gradient descent algorithm, and the objective function is the sum of squared Euclidean distances between predicted product components and actual detected product components. During the training process, L1 regularization constraint is applied to β1, β2, β3 to avoid overfitting. When the mean square error on the validation set converges to ≤0.03, the training is stopped, and the weight combination of β1, β2, β3 is finally determined.

[0110] After the multi-level enzyme interaction network model is obtained, the node importance is analyzed by using network topology algorithm, for example, the centrality of enzyme node is calculated by using PageRank algorithm to reflect its global influence in the network; the authority and hub are calculated by using HITS algorithm to measure the importance of the node as an information source and a transfer station; the key role of the node in connecting different network modules is evaluated by using bridging degree. For the substrate node, the key metabolic hub is identified by using the betweenness centrality, such as the amino acid or peptide segment connecting multiple enzymes. Based on the node analysis result, the Louvain algorithm is used for community detection to divide the network into functional modules, such as the endopeptidase cleavage module and the exopeptidase degradation module, and the connection relationship between the modules through the hub substrate is analyzed, such as PEP7 as a key node connecting the cleavage module and the degradation module. Finally, the multi-level enzyme interaction network model is visualized by using Cytoscape and other tools, and the network levels are distinguished by different colors, for example, the homology edge is blue, the direct action edge is red, and the indirect regulation edge is green, the node size corresponds to the centrality, and the edge thickness and color reflect the interaction intensity and type, forming a three-dimensional interactive network model.

[0111] The multi-level enzyme interaction network model solves the problem that the single-level network in the prior art cannot comprehensively reflect the complexity of the enzymatic system. When the traditional single-level network is used, the complex interaction relationship such as functional redundancy of homologous enzyme groups and indirect feedback inhibition in the enzymatic process cannot be captured, resulting in a large prediction error of the model. Through the fusion of multi-level networks, the multi-dimensional correlation of “evolutionary relationship-physical action-metabolic regulation” is realized in the field of enzyme hydrolysis, the interaction relationship in different levels of networks is integrated to form a comprehensive regulation chain, so that the model can accurately reflect the complex interaction between enzymes and enzymes, and between enzymes and substrates in the enzymatic system, and the explanation and prediction accuracy of the model for the enzymatic process are improved, and more comprehensive model support is provided for the precise regulation of the enzymatic system.

[0112] Step S50, based on the multi-level enzyme interaction network model, dynamically regulating the ratio of amino acids and peptides in the soybean protein enzymatic process;

[0113] The prior art often only designs a regulation strategy according to the direct catalytic action of the enzyme, ignores the synergistic effect of homologous enzymes, substrate-mediated indirect regulation and time sequence dynamic change, resulting in a large deviation of the ratio of peptides and amino acids in the actual production. Step S50 constructs a dynamic enzyme dosage regulation strategy based on the multi-level enzyme interaction network model, realizes the precise regulation of the ratio of amino acids and peptides in the soybean protein enzymatic process, and solves the problem of product ratio fluctuation caused by the difficulty in identifying the implicit interaction interference of the multi-enzyme system and the lack of systematicness of the regulation strategy in the prior art.

[0114] Further, step S50 comprises:

[0115] Step S51, based on the multi-level enzyme interaction network model, calculating the key degree index of each enzyme in the target enzymatic system;

[0116] The criticality index represents the overall influence of an enzyme on its substrate. The calculation process involves extracting the direct and indirect effects of each enzyme on various substrates from a multi-level enzyme interaction network model. The direct effect intensity originates from the edge weights of the second-level network; for example, the direct catalytic intensity of EN1 on PEP5 is 0.6. The indirect effect intensity is obtained through indirect path analysis of the third-level network. The calculation of the indirect effect intensity requires traversing all indirect paths and weighting each path according to the formula: "product of path edge weights × reciprocal of path length × time-varying contribution." For example, if a path contains 3 edges with weights of 0.5, 0.4, and 0.3, has 3 edges, and a time-varying contribution of 0.8, then the intensity of this path is 0.5 × 0.4 × 0.3 × (1 / 3) × 0.8 = 0.016. The indirect effect intensity is obtained by summing all paths.

[0117] By combining the strength of direct and indirect effects with the enzyme's network centrality, and optimizing the importance ratio of each factor using historical data, a keyness index is calculated through weighted summation. This keyness index comprehensively reflects the enzyme's direct catalytic ability, indirect regulatory effects, and pivotal role in the network, and is closer to the actual regulatory needs of enzymatic hydrolysis systems than traditional assessment methods that only consider direct effects.

[0118] Determining the importance ratios of each factor through historical data optimization specifically involves collecting historical experimental data covering different enzymatic hydrolysis process conditions. This data includes the intensity of direct and indirect effects, enzyme network centrality, and the corresponding proportions of actual product components. The collected data undergoes preprocessing, including outlier removal and standardization, to eliminate the influence of different indicator units. An optimization algorithm from machine learning is used, with the objective function of minimizing the difference between the predicted and actual product components, to iteratively optimize the importance ratios of the three factors: intensity of direct and indirect effects, and network centrality. During the optimization process, methods such as cross-validation are used to continuously adjust the weights of each factor until the model's prediction accuracy for product components meets the requirements, thus determining the final importance ratios of each factor.

[0119] Step S52: Set the target amino acid to peptide ratio formula TP, and construct a real-time enzyme dosage control model based on the criticality index.

[0120] Existing technologies often employ linear regression models, neglecting dynamic interactions between enzymes. For example, the synergistic inhibitory effect of combining an endopeptidase and an exopeptidase on the formation of specific peptide segments can lead to significant deviations between the product composition and the target formulation in actual production. This invention constructs a real-time enzyme dosage control model based on a criticality index, establishing a precise mapping relationship from enzyme dosage to product composition. This solves the problem that existing control models cannot handle large prediction errors caused by nonlinear interactions and latent interference.

[0121] In a specific implementation, the target proportion formula TP is defined as the expected content proportion of each amino acid and peptide segment in the target product, for example, a certain functional peptide product requires a specific length of peptide segment to reach the expected level, and the amino acid content is controlled within a corresponding range. A deep learning model such as a long short-term memory network (LSTM) or a temporal convolution network (TCN) is used to construct and train the enzyme dosage real-time regulation model. The model input is the enzyme dosage vector and the key index matrix, and the output is the predicted product component vector. The key index matrix is a two-dimensional matrix containing the key index of all enzymes for all substrates. Each row corresponds to an enzyme, and each column corresponds to a substrate. The matrix elements represent the comprehensive influence ability of the enzyme on the substrate. Taking LSTM as an example, it captures the time sequence dependence of the enzyme hydrolysis process through the gating mechanism, such as the lagging effect of increasing the dosage of a certain enzyme on the generation of target peptides in the middle of the enzyme hydrolysis. The network structure includes an input layer, a hidden layer, and an output layer. The parameters are adjusted through historical data training, so that the model learns the complex nonlinear relationship between enzyme dosage and component changes. To handle implicit interaction interference, an "interaction perception layer" is introduced into the model. This layer calculates the interaction weight of different enzyme pairs based on the attention mechanism, automatically identifies strong interaction enzyme pairs, and dynamically adjusts the prediction results based on the abnormal influence of the combination on the components in the historical data when such enzyme pairs are detected simultaneously. For example, the use of a combination of two enzymes leads to a decrease in the yield of a specific peptide segment, and the model compensates for this nonlinear effect in the prediction.

[0122] The objective function is established as the squared Euclidean distance between the predicted components and the target proportion formula. The constrained optimization problem is solved by the sequential quadratic programming (SQP) algorithm, and the constraint conditions include the total enzyme dosage limit and the upper limit of individual enzyme dosage. The enzyme dosage real-time regulation model uses the interaction perception layer to quantify implicit interference and achieve dynamic compensation of enzyme combination effects. Some indirect synergistic pathways of certain enzyme pairs have a dominant influence on target components through substrate mediation, but traditional methods ignore them due to not considering indirect effects. The enzyme dosage real-time regulation model integrates multi-level network information to accurately identify such key regulation pathways, and then adjusts the enzyme dosage in the regulation strategy to achieve precise control of product components.

[0123] Step S53, based on the enzyme dosage real-time regulation model, the dynamic and precise regulation of the proportion of amino acids and peptides in the soybean protein enzyme hydrolysis process is realized.

[0124] Specifically, in the actual production process, real-time collection of key parameters in the enzymatic process is performed, including pH value, temperature, activity change of key enzymes, and concentration change of representative amino acids and peptide segments. The real-time collected data is input into the enzyme dosage real-time regulation model to predict the final product component distribution under the current enzymatic trend. If the prediction result deviates from the target proportion formula TP, the regulation strategy is automatically triggered, and the enzyme dosage supplement scheme that needs to be adjusted is calculated. For example, when the content of a certain peptide segment is insufficient, the model calculates the dosage of the corresponding endopeptidase that needs to be increased. By real-time collection of the enzyme hydrolysis process parameters and input into the enzyme dosage real-time regulation model, the enzyme hydrolysis dynamics can be accurately captured, and the dominant role of the key synergistic path in the middle stage of enzyme hydrolysis, which is ignored by traditional methods, can be identified through the analysis of the indirect regulation path implied in the multi-level network model. Then, the enzyme dosage and process conditions are dynamically adjusted through intelligent intervention strategies. Through the "real-time monitoring-model prediction-intelligent intervention" closed-loop mechanism, the batch consistency is improved, and the product quality fluctuation problem caused by manual regulation lag and fixed parameters in the prior art is effectively solved.

[0125] Embodiment 2

[0126] This embodiment is based on Embodiment 1 and provides a soybean protein enzymolysis amino acid and peptide ratio regulation system, as shown in Figure 5 , which includes:

[0127] A data acquisition module is used to collect the enzyme hydrolysis coupling parameters of the target enzyme hydrolysis system in the entire enzyme hydrolysis process, and generate enzyme hydrolysis coupling time series data;

[0128] A slice calculation module is used to obtain a multi-dimensional dynamic covariance matrix according to the enzyme hydrolysis coupling time series data;

[0129] A significant interaction identification module is used to identify significant enzyme interactions based on the multi-dimensional dynamic covariance matrix, and generate a significant enzyme interaction list;

[0130] An interaction network construction module is used to construct a multi-level enzyme interaction network model based on the significant enzyme interaction list;

[0131] A regulation module is used to dynamically regulate the amino acid and peptide ratio in the soybean protein enzymolysis process based on the multi-level enzyme interaction network model.

[0132] In the significant interaction identification module, the method for obtaining the multi-dimensional dynamic covariance matrix according to the time series data of the enzyme decoupling includes: performing eigenvalue decomposition on the multi-dimensional dynamic covariance matrix to obtain a reduced dimension feature space; calculating the significance level of the covariance between the endopeptidase activity, the exopeptidase activity, the amino acid concentration, the peptide segment concentration and the pH value in the reduced dimension feature space; setting a significance threshold α, and automatically marking the covariance with a significance level exceeding the significance threshold as a significant enzyme interaction; generating a significant enzyme interaction list according to the marked significant enzyme interactions; and the significant enzyme interaction list includes interaction entities, interaction strengths, interaction modes and time slice identifiers.

[0133] In the interaction network construction module, the method for constructing a multi-level enzyme interaction network model based on the significant enzyme interaction list includes:

[0134] constructing a homology interaction network between endopeptidases and endopeptidases, and between exopeptidases and exopeptidases as a first-level model; constructing a direct interaction network between endopeptidases and exopeptidases based on the significant enzyme interaction list as a second-level model; constructing a comprehensive indirect regulation network based on the significant enzyme interaction list as a third-level model; and fusing the constructed first-level model, second-level model and third-level model to construct a multi-level enzyme interaction network model.

[0135] The methods and systems of the present application can be implemented in many ways. For example, the methods and systems of the present application can be implemented through software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is merely for illustration, and the steps of the methods of the present application are not limited to the above specifically described order, unless otherwise specifically stated.

[0136] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of corresponding technical solutions in the prior art are not described in detail to avoid excessive repetition.

[0137] The specific embodiments described above further illustrate the objects, technical solutions and beneficial effects of the present application. It should be understood that the above description is merely a specific embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A method for regulating the ratio of amino acids to peptides after soybean protein hydrolysis, characterized in that, The method includes: Collect the enzyme decoupling parameters of the target enzyme hydrolysis system throughout the entire enzymatic hydrolysis process and generate enzyme decoupling time series data; Based on the enzyme decoupling time series data, a multidimensional dynamic covariance matrix was obtained; Based on the multidimensional dynamic covariance matrix, significant enzyme interactions are identified and a list of significant enzyme interactions is generated. A multi-level enzyme interaction network model is constructed based on a list of significant enzyme interactions. The method for constructing this model includes: constructing homology interaction networks between endopeptidases and between exopeptidases, as a first-level model; constructing a direct interaction network between endopeptidases and exopeptidases based on the list of significant enzyme interactions, as a second-level model; constructing a comprehensive indirect regulatory network based on the list of significant enzyme interactions, as a third-level model; and integrating the constructed first-level, second-level, and third-level models to construct the multi-level enzyme interaction network model. The method for constructing a comprehensive indirect regulatory network includes: constructing a bipartite graph B(VE,VS,E) based on a list of significant enzyme interactions, where VE is the set of enzyme nodes containing all endopeptidases and exopeptidases; VS is the set of substrate nodes containing all amino acids and peptides; and E is the set of enzyme-substrate interaction edges; deriving enzyme-enzyme indirect regulatory relationships based on the bipartite graph; constructing an endopeptidase-exopeptidase indirect regulatory network, an endopeptidase-endopeptidase indirect regulatory network, and an exopeptidase-exopeptidase indirect regulatory network based on these enzyme-enzyme indirect regulatory relationships; and merging these three networks to form a comprehensive indirect regulatory network. Based on a multi-level enzyme interaction network model, the ratio of amino acids to peptides during soybean protein hydrolysis is dynamically regulated.

2. The method for regulating the ratio of amino acids to peptides after soybean protein hydrolysis according to claim 1, characterized in that, The method for obtaining a multidimensional dynamic covariance matrix based on enzyme decoupling time series data includes: setting a sliding time window with a duration of T4, using T5 as the moving interval, slicing the enzyme decoupling time series data to obtain N time slices; calculating the covariance between each pair of data within each time slice to obtain the multidimensional dynamic covariance matrix.

3. The method for regulating the ratio of amino acids to peptides after soybean protein hydrolysis according to claim 2, characterized in that, The enzyme decoupling parameters include m types of endopeptidases (EN) in the target enzymatic hydrolysis system. i and n kinds of exopeptidases EX j Activity data, p'-type amino acids AA k and q different lengths of PEP g Concentration data and pH fluctuation data; The method for calculating the pairwise covariance between data within each time slice to obtain the multidimensional dynamic covariance matrix includes: for each time slice t, calculating the endopeptidase EN i Activity, exopeptidase EX j Activity, amino acid AA k Concentration, Peptide PEP g The covariance between concentration and pH value forms the covariance matrix C for time slice t. t ; the covariance matrix C t Each element in the matrix is ​​standardized to a correlation coefficient, resulting in the correlation matrix R for time slice t. t ; the covariance matrix C of N time slices t Construct a dynamic covariance tensor C by combining the time sequences, and combine the correlation matrices R of N time slices. t By combining them in chronological order, a dynamic correlation tensor R is constructed; based on the dynamic covariance tensor C and the dynamic correlation tensor R, a multidimensional dynamic covariance matrix is ​​constructed.

4. The method for regulating the ratio of amino acids to peptides after soybean protein hydrolysis according to claim 3, characterized in that, The method for identifying significant enzyme interactions and generating a list of significant enzyme interactions includes: Eigenvalue decomposition is performed on the multidimensional dynamic covariance matrix to obtain the dimension-reduced feature space. In the reduced feature space, calculate the significance level of the covariance between endopeptidase activity, exopeptidase activity, amino acid concentration, peptide concentration and pH value. Set a significance threshold α, and automatically label covariances with significance levels exceeding the significance threshold as significant enzyme interactions; A list of significant enzyme interactions is generated based on the labeled significant enzyme interactions.

5. The method for regulating the ratio of amino acids to peptides after soybean protein hydrolysis according to claim 4, characterized in that, The method for constructing homology interaction networks between endopeptidases and between exopeptidases includes: Obtain the sequence information and three-dimensional structure data of all endopeptidases and exopeptidases in the target enzymatic hydrolysis system; Based on the sequence information of endopeptidases and exopeptidases, enzyme pairs with potential homology interactions are identified. Based on the three-dimensional structural data of endopeptidases and exopeptidases, enzyme pairs with homology interactions were identified from enzyme pairs with potential homology interactions, and a homology interaction network was constructed.

6. The method for regulating the ratio of amino acids to peptides after soybean protein hydrolysis according to claim 5, characterized in that, The method for identifying enzyme pairs with potential homology interactions is as follows: based on the sequence information of endopeptidases and exopeptidases, the sequence similarity between all pairs of endopeptidases and the sequence similarity between all pairs of exopeptidases are calculated. For enzyme pairs with sequence similarity exceeding the sequence similarity threshold, they are marked as enzyme pairs with potential homology interactions.

7. The method for regulating the ratio of amino acids to peptides after soybean protein hydrolysis according to claim 6, characterized in that, The method for identifying enzyme pairs with homology interactions from enzyme pairs with potential homology interactions includes: calculating the structural similarity of enzyme pairs with homology interactions based on the three-dimensional structural data of endopeptidases and exopeptidases; enzyme pairs with structural similarity exceeding a structural similarity threshold are identified as enzyme pairs with homology interactions.

8. A system for regulating the ratio of amino acids to peptides after soybean protein hydrolysis, used to implement the method for regulating the ratio of amino acids to peptides after soybean protein hydrolysis as described in any one of claims 1-7, characterized in that, The system includes: The data acquisition module is used to collect the enzyme decoupling parameters of the target enzyme hydrolysis system throughout the entire enzymatic hydrolysis process and generate enzyme decoupling time series data; The slice calculation module is used to obtain a multidimensional dynamic covariance matrix based on enzyme decoupling time series data; The salient interaction identification module, based on a multidimensional dynamic covariance matrix, identifies salient enzyme interactions and generates a list of salient enzyme interactions. The interaction network construction module constructs a multi-level enzyme interaction network model based on a list of salient enzyme interactions. The regulation module, based on a multi-level enzyme interaction network model, dynamically regulates the ratio of amino acids to peptides during soybean protein hydrolysis.

Citation Information

Patent Citations

  • Refining method for increasing content of branched chain amino acid peptide

    CN118389630A

  • Multifunctional self-assembly peptide recognition method and system

    CN119626345A

  • Antibacterial peptide prediction method based on integrated deep learning

    CN120452548A

  • Methods for peptide mass spectrometry fragmentation prediction

    US20210041454A1