Cottonseed protein enzymolysis control method and system
By combining a deep learning model and a Bayesian optimization algorithm with a fuzzy controller in a closed-loop control system, the dynamic mapping problem between enzymatic parameters and functional activity indicators during cottonseed proteolysis was solved, achieving efficient and stable production of cottonseed proteolysis products and reducing production costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies cannot achieve dynamic mapping between enzymatic hydrolysis parameters and functional activity indicators during cottonseed protein hydrolysis, resulting in large fluctuations in the yield of target active peptides, unstable functional activity, and increased production costs.
A deep learning model is used to collect molecular weight distribution and spatial structure data of enzymatic hydrolysis products in real time. The optimal enzymatic hydrolysis parameters are searched through a Bayesian optimization algorithm, and the enzymatic hydrolysis parameters are adjusted in real time using a fuzzy controller. The feasibility is verified by combining the production cost calculation function, and a closed-loop control system is constructed.
This study improved the functional activity and stability of cottonseed protein hydrolysates, reduced production costs, and enhanced the controllability and economy of the hydrolysis process.
Smart Images

Figure CN121780779A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated control technology for enzymatic hydrolysis processes, and more specifically, to a method and system for controlling the enzymatic hydrolysis of cottonseed. Background Technology
[0002] In the industrial production of cottonseed enzymatic hydrolysates, the precise regulation of functional active ingredients such as antimicrobial peptides and ACE-inhibiting peptides faces significant technical bottlenecks. Existing research mainly relies on static single-factor optimization methods, such as conducting single-factor experiments by fixing parameters like enzymatic hydrolysis time and temperature. This approach can only obtain active peptides under specific conditions but has not established a dynamic correlation model between enzymatic hydrolysis parameters and product functional characteristics. This technical deficiency leads to large fluctuations in the yield of target active peptides, failing to meet the requirements of functional stability for industrial production.
[0003] From a process control perspective, minute changes in key parameters such as temperature and pH during enzymatic hydrolysis can lead to alterations in the spatial conformation of peptides, such as the transformation of α-helical structures into random coils, resulting in loss of functional activity. However, existing processes lack real-time monitoring methods and parameter compensation mechanisms for the spatial structure of the product, making it difficult to correct deviations in a timely manner during the reaction.
[0004] In addition, traditional enzymatic hydrolysis processes generally prioritize the degree of hydrolysis, which is inherently in conflict with the functional activity requirements. Overly pursuing the degree of hydrolysis may damage the core structure of active peptides, leading to a decrease in functional activity.
[0005] The root cause of the above problems lies in the fact that existing technologies have failed to construct a multi-objective optimization system oriented towards functional activity. For example, they cannot achieve dynamic mapping between enzymatic hydrolysis parameters and functional activity indicators, resulting in unstable functional activity of enzymatic hydrolysis products and increased production costs.
[0006] In view of this, the present invention proposes a method and system for controlling the enzymatic hydrolysis of cottonseed to solve the above problems. Summary of the Invention
[0007] To overcome the aforementioned deficiencies of the prior art and achieve the above objectives, the present invention provides the following technical solution: a cottonseed protease hydrolysis control method, comprising:
[0008] The molecular weight distribution and spatial structure data of the enzymatic hydrolysis products collected in real time during the cottonseed protein hydrolysis process are input into a pre-trained deep learning model, and the target functional activity index is output.
[0009] Based on the enzymatic hydrolysis parameters in the cottonseed protein hydrolysis process, as well as the preset ideal target functional activity index and the obtained target functional activity index, an optimization objective function is constructed; the optimal enzymatic hydrolysis parameters are obtained by searching the optimization objective function.
[0010] Cottonseed protein was enzymatically hydrolyzed based on optimal hydrolysis parameters. The spatial structure data of the product was monitored in real time during the hydrolysis process, and the optimal hydrolysis parameters were adjusted to obtain the final hydrolysis parameters.
[0011] Furthermore, methods for constructing the objective function include:
[0012] The target functional activity indicators include antibacterial activity indicators and antioxidant activity indicators; the enzymatic hydrolysis parameters include enzyme dosage, substrate concentration, reaction temperature, pH value and reaction time in the cottonseed protein hydrolysis process;
[0013] The enzymatic hydrolysis parameters were used as independent variables.
[0014] The difference between the target functional activity index and the preset ideal target functional activity index is used as the dependent variable;
[0015] The optimized objective function is obtained.
[0016] Furthermore, methods for obtaining optimal enzymatic hydrolysis parameters by searching for optimization algorithms on the objective function include:
[0017] S101: The initial set of sample points for enzymatic hydrolysis parameters is randomly generated using a sampling method;
[0018] S102: Input the initial enzymatic hydrolysis parameter sample point set into the optimization objective function for calculation to obtain the initial objective function value set;
[0019] S103: Construct a Gaussian process surrogate model in the Bayesian optimization algorithm based on the initial enzymatic hydrolysis parameter sample point set and the initial objective function value set;
[0020] S104: The desired improved acquisition function in the Bayesian optimization algorithm is used to calculate on the Gaussian process surrogate model to obtain candidate parameter points;
[0021] S105: Input the candidate parameter points into the optimization objective function for calculation to obtain the new objective function value;
[0022] S106: Add the candidate parameter points and the new objective function value to the initial parameter sample point set and the initial objective function value set to form the updated sample point set and the updated objective function value set;
[0023] S107: Reconstruct the Gaussian process proxy model based on the updated sample point set and the updated objective function value set;
[0024] S108: Repeat steps S104 to S107 until the set maximum number of iterations is reached, output the optimal parameter point corresponding to the historical best objective function value, and map the optimal parameter point to the optimal enzymatic hydrolysis parameter.
[0025] Furthermore, methods for determining the historical optimal objective function value include:
[0026] If the current iteration is the first iteration, set the new objective function value as the historical best objective function value, and record the candidate parameter point as the historical best parameter point;
[0027] If it is not the first iteration, then a judgment is made based on the mathematical properties of the objective function:
[0028] When the objective function is a maximization function, if the new objective function value is greater than the historical best objective function value, then the historical best objective function value and the historical best parameter point are updated.
[0029] When the objective function is to minimize the objective function, if the new objective function value is less than the historical best objective function value, then the historical best objective function value and the historical best parameter point are updated.
[0030] If no update is performed, the original historical best objective function value and historical best parameter point are retained.
[0031] Furthermore, methods for obtaining the final enzymatic hydrolysis parameters include:
[0032] During the enzymatic hydrolysis of cottonseed, the spatial structure data of the hydrolysis products were collected in real time using a circular dichroism chromatograph to obtain real-time spatial structure data.
[0033] Calculate the root mean square error between real-time spatial structure data and preset standard spatial structure data to obtain the structural deviation value;
[0034] Input the structural deviation value into the preset fuzzy controller, and output the adjustment amount of the enzymatic hydrolysis parameters;
[0035] The final enzymatic hydrolysis parameters are obtained by algebraically adding the adjusted enzymatic hydrolysis parameters to the optimal enzymatic hydrolysis parameters.
[0036] Furthermore, the setup method for the fuzzy controller includes:
[0037] The input to the fuzzy controller is the structural deviation value, and the domain of the structural deviation value is from 0 to 1, which is divided into three fuzzy sets: small, medium, and large.
[0038] The output of the fuzzy controller is set as the adjustment amount of the enzymatic hydrolysis parameters; the domain of the adjustment amount of the enzymatic hydrolysis parameters is from -R to R, and is divided into five fuzzy sets: negative large, negative small, zero, positive small, and positive large; among them, the determination of the value of R needs to comprehensively consider the physical boundary of the enzymatic hydrolysis parameters and the quantitative classification requirements of the fuzzy sets;
[0039] Set the fuzzy rule base as follows:
[0040] If the structural deviation value is small, the adjustment amount of the enzymatic hydrolysis parameter is zero;
[0041] If the structural deviation value is medium, then the adjustment amount of the enzymatic hydrolysis parameter is positively small;
[0042] If the structural deviation value is large, then the adjustment amount of the enzymatic hydrolysis parameter should be positive.
[0043] Furthermore, production costs related to the independent variables are taken as the dependent variable; the objective function is obtained by weighted summation of the dependent variables.
[0044] Furthermore, after obtaining the optimal enzymatic hydrolysis parameters, the feasibility of the adjusted enzymatic hydrolysis parameters is verified based on a preset production cost threshold. If feasible, the final enzymatic hydrolysis parameters are obtained; if not feasible, the adjusted enzymatic hydrolysis parameters are corrected based on a preset cost-parameter adjustment mapping table to obtain the corrected enzymatic hydrolysis parameters. The feasibility of the corrected enzymatic hydrolysis parameters is verified. If still not feasible, auxiliary parameters are introduced and added to the independent variables. The auxiliary parameters include surface activity cost and reaction medium modification cost.
[0045] Furthermore, methods for verifying the feasibility of modifying enzymatic hydrolysis parameters include:
[0046] Extract the enzyme dosage, temperature, and pH value from the adjusted enzymatic hydrolysis parameters to form an enzymatic hydrolysis parameter vector;
[0047] The enzymatic hydrolysis parameter vector is input into a preset production cost calculation function to calculate the estimated production cost; the production cost calculation function is obtained by multiple linear regression fitting based on historical enzymatic hydrolysis production data.
[0048] The cost difference is obtained by performing an algebraic subtraction between the estimated production cost and the preset production cost threshold.
[0049] If the cost difference is less than or equal to zero, it is feasible; if the cost difference is greater than zero, it is not feasible.
[0050] Compared with the prior art, the technical effects and advantages of the cottonseed protease hydrolysis control method and system of the present invention are as follows:
[0051] This invention utilizes a deep learning model to process real-time collected molecular weight distribution and spatial structure data to achieve dynamic prediction of target functional activity indicators; it employs a Bayesian optimization algorithm to search for the global optimal solution of the objective function to obtain the optimal enzymatic hydrolysis parameters; it monitors the spatial structure data of the enzymatic hydrolysis products in real time during the enzymatic hydrolysis process and dynamically adjusts the enzymatic hydrolysis parameters through a fuzzy controller; simultaneously, it constructs a production cost calculation function based on historical enzymatic hydrolysis generation data to verify the feasibility of the adjusted enzymatic hydrolysis parameters, and introduces auxiliary parameters during the verification process to correct the adjusted enzymatic hydrolysis parameters, thus obtaining the final enzymatic hydrolysis parameters.
[0052] This invention overcomes the limitations of traditional static single-factor optimization by establishing a dynamic correlation between enzymatic hydrolysis parameters, functional activity, and production cost, achieving closed-loop control from parameter optimization and process regulation to cost verification. Compared with existing technologies, this invention effectively solves the problems of large fluctuations in the yield and unstable functional activity of target active peptides in cottonseed enzymatic hydrolysates, significantly improving the controllability and economy of the enzymatic hydrolysis process, and providing an efficient and reliable technical solution for the precise industrial control of cottonseed enzymatic hydrolysates. Attached Figure Description
[0053] Figure 1 This is a schematic diagram of the cottonseed protease hydrolysis control system according to an embodiment of the present invention;
[0054] Figure 2 This is a flowchart of the cottonseed protease hydrolysis control method according to an embodiment of the present invention;
[0055] Figure 3 This is a flowchart of a method for searching and optimizing the global optimal solution of the objective function according to an embodiment of the present invention;
[0056] Figure 4 This is a flowchart illustrating the method for verifying the feasibility of modifying enzymatic hydrolysis parameters according to an embodiment of the present invention. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be described in detail, clearly, and completely below with reference to the accompanying drawings. It should be particularly noted that the specific embodiments described below are only for better illustrating and explaining the technical solutions of the present invention, and are intended to enable those skilled in the art to better understand and implement the present invention, and should not be construed as limiting the scope of protection of the present invention. Without departing from the spirit and substance of the present invention, those skilled in the art can modify, adjust, or make equivalent substitutions based on the content disclosed in the present invention, and these should all be considered within the scope of protection of the present invention.
[0058] Example 1
[0059] Please see Figure 1 As shown, this embodiment discloses a cottonseed protease hydrolysis control system, including a prediction module, an optimization module, a regulation module, and a verification module. Each module is connected via wired and / or wireless connections to achieve data transmission.
[0060] The prediction module is used to input the molecular weight distribution data and spatial structure data of the enzymatic hydrolysis products collected in real time during the cottonseed protein hydrolysis process into a pre-trained deep learning model, and output the target functional activity index.
[0061] For example, this embodiment provides a method for real-time acquisition of molecular weight distribution data and spatial structure data of enzymatic hydrolysis products in a cottonseed protease hydrolysis control system, as detailed below:
[0062] Molecular weight distribution data were acquired using matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MAMS). In this embodiment, the sampling frequency was once per minute, the acquisition range was 500-5000 Da, and the resolution was set to 0.1 Da. This setting can accurately capture the characteristic molecular weight range of active peptides in the enzymatic hydrolysis products, such as the 1000-3000 Da range commonly found in antimicrobial peptides. Spatial structure data were acquired using circular dichroism spectroscopy (CDS). In this embodiment, the scanning wavelength range was 190-260 nm, and the scanning speed was 50 nm / min. This parameter setting can effectively identify the secondary structural features of peptides, such as α-helices and β-sheets. These structures are closely related to functional activity; for example, the higher the α-helix content, the stronger the antimicrobial activity. The measured values of functional activity indicators were determined through in vitro experiments. For example, antimicrobial activity was measured using the inhibition zone method, and antioxidant activity was determined using DPPH free radical scavenging rate, ensuring the accuracy and comparability of the data.
[0063] Training methods for deep learning models include:
[0064] Collect historical enzymatic hydrolysis products' molecular weight distribution data, spatial structure data, and corresponding measured values of functional activity indicators to form the original training dataset;
[0065] The original training dataset is standardized to obtain a preprocessed training dataset.
[0066] The preprocessed training dataset is divided into a training subset and a validation subset to obtain the dataset partitioning result.
[0067] Construct a neural network architecture containing convolutional layers and long short-term memory layers to obtain the initial model structure;
[0068] The initial model structure is iteratively trained using the training subset from the dataset partitioning results to obtain an intermediate model.
[0069] The intermediate model is evaluated using the validation subset from the dataset partitioning results to obtain the validation loss value.
[0070] When the validation loss value fails to decrease for dt consecutive iterations, training is stopped and the current intermediate model is output as the completed deep learning model.
[0071] For example, this embodiment provides a deep learning training method in a cottonseed protease hydrolysis control system, as detailed below:
[0072] We collected molecular weight distribution data, spatial structure data, and corresponding measured values of functional activity indicators of historical enzymatic hydrolysis products to form the original training dataset.
[0073] Methods for standardizing the original training dataset to obtain a preprocessed training dataset include:
[0074] Calculate the mean and standard deviation of each channel in the molecular weight distribution data of the training dataset; calculate the statistical parameters of all feature dimensions in the spatial structure data; subtract the mean of each channel from the original data of the molecular weight distribution data and divide by the standard deviation to obtain the standardized parameters; for the spatial structure data, process them according to the feature type of their statistical parameters: perform standardization transformation on continuous features and perform min-max normalization transformation on proportional features; after the transformation, generate a standardized feature matrix, at which point the numerical range of all feature dimensions is unified to the interval [-3,3]; persist the standardized parameters and the standardized feature matrix to a JSON configuration file to form a reusable standardization processor, and output a preprocessed training dataset containing the standardized feature matrix and corresponding functional activity index values.
[0075] The preprocessed training dataset is divided into a training subset and a validation subset to obtain the dataset partitioning result.
[0076] Methods for constructing a neural network architecture containing convolutional layers and long short-term memory layers to obtain the initial model structure include:
[0077] Define the dimensional features of the preprocessed training dataset received by the input layer: molecular weight distribution data is input in the form of a 128-dimensional vector, and spatial structure data is input in the form of a 12-dimensional vector;
[0078] A one-dimensional convolutional layer was added to process the molecular weight distribution data. 64 convolutional kernels of size 3 were set, and the ReLU activation function was used to extract local features.
[0079] Connect the max pooling layer for downsampling, and set the pooling size to 2;
[0080] The convolutional output is concatenated with the spatial structure data vector to form a fused feature vector;
[0081] The sequence features are processed by accessing a long short-term memory layer. The first LSTM layer is set with 128 units and returns the complete sequence, while the second LSTM layer is set with 64 units and only returns the final state.
[0082] A fully connected layer is added for feature integration. The first fully connected layer is configured with 32 neurons and activated using ReLU. The output layer is configured with a single neuron to predict functional activity indicators.
[0083] After layer stacking is completed, an untrained weight-initialized neural network architecture is generated. The neural network architecture contains 7 trainable layers and 120,000 parameters, forming an initial model structure that can be directly used for training.
[0084] Methods for iteratively training the initial model structure using a training subset from the dataset partitioning results to obtain an intermediate model include:
[0085] The training subset from the dataset partitioning result is divided into multiple data batches according to a preset batch size, and these batches are sequentially input into the initial model structure for forward propagation computation. After each forward propagation, the mean squared error function is used to calculate the loss value between the predicted and measured values of the model's output functional activity index. Based on this loss value, the backpropagation algorithm is executed, and the Adam optimizer is used to calculate the gradient of the model parameters and update the weights. The above batch processing process is repeated until the complete training subset is traversed, completing one training round. After executing multiple training rounds, the model parameters are continuously optimized, generating an intermediate model with convergent weight parameters. The intermediate model retains all the layer configurations of the initial model structure, but its convolutional kernel weight matrix, long short-term memory layer gating parameters, and fully connected layer weight vectors have all been preliminarily optimized and can be directly used for subsequent validation subset performance evaluation.
[0086] Methods for evaluating the performance of intermediate models and obtaining validation loss values using a validation subset from the dataset partitioning results include:
[0087] The validation subset from the dataset partitioning results is input into the intermediate model for forward propagation computation, while keeping the model parameters frozen. The intermediate model sequentially extracts molecular weight distribution features through convolutional layers, processes sequence features through long short-term memory layers, and outputs predicted values of functional activity indices through fully connected layers. The same mean squared error function as in the training phase is used to calculate the squared difference between the model's predicted values and the measured values of functional activity indices on each sample of the validation subset. The arithmetic mean of the squared differences obtained from all validation subsets is calculated to generate a scalar numerical validation loss value that characterizes the model's generalization ability. The range of the loss value is typically between 0.5 and 0.8. For example, when the average deviation between the predicted and measured values is 0.7, the validation loss value is 0.49. Its magnitude directly reflects the prediction accuracy of the intermediate model on untrained data, providing a quantitative basis for training termination decisions. During the training termination determination process, the initial value of the iteration counter is set to zero. After each training epoch, the comparison result between the current validation loss value and the previous optimal validation loss value is recorded: if the current value is better (in the actual scenario, "optimal" means smaller when the optimization objective is to minimize the loss, and "optimal" means larger when the optimization objective is to maximize the loss), then the optimal validation loss value is updated and the counter is reset; otherwise, the counter is incremented. When the counter accumulates to a preset threshold of 3 times, training is immediately stopped and the current intermediate model is output as the deep learning model that has been trained. The preset threshold is determined through pre-experimentation. For example, under 10 different initial conditions, the average number of epochs in the stable stage of monitoring the validation loss is 2.8 ± 0.4, so it is rounded down to 3 times to ensure timely termination after the model performance stabilizes and avoid the risk of overfitting.
[0088] The optimization module is used to search for the optimal enzymatic hydrolysis parameters based on the preset ideal target functional activity index and the obtained target functional activity index through the optimization algorithm.
[0089] Methods for constructing the objective function include:
[0090] The target functional activity indicators include antibacterial activity indicators and antioxidant activity indicators; the enzymatic hydrolysis parameters include enzyme dosage, substrate concentration, reaction temperature, pH value and reaction time in the cottonseed protein hydrolysis process; the enzymatic hydrolysis parameters are used as independent variables; the difference between the target functional activity indicators and the preset ideal target functional activity indicators is used as the dependent variable; thus, the optimized objective function is obtained.
[0091] Furthermore, production costs related to the independent variables are taken as the dependent variable; the objective function is obtained by weighted summation of the dependent variables.
[0092] For example, this embodiment provides a method for constructing the optimization objective function obtained by searching for the optimal enzymatic hydrolysis parameters using an optimization algorithm, as detailed below:
[0093] The enzyme dosage, substrate concentration, reaction temperature, pH value, and reaction time in the cottonseed proteolysis process were used as independent variables, while the production cost related to the independent variables and the difference between the obtained target functional activity index and the preset ideal target functional activity index were used as dependent variables.
[0094] Enzyme dosage E, substrate concentration S, reaction temperature T, pH value pH, and reaction time t are combined into a five-dimensional independent variable vector x = [E, S, T, pH, t]. x is input into a preset production cost calculation function, which outputs the production cost value C(x). For example, when x = [1.5%, 5.0%, 50℃, 7.0, 4h], C(x) = 28.5 yuan. Simultaneously, x is input into a trained deep learning model, which outputs a predicted value for the target functional activity index.
[0095] Methods for obtaining the objective function by weighted summation of dependent variables include:
[0096] Extract the dependent variable, denoting the actual production cost corresponding to the current enzymatic hydrolysis parameters as C, and the absolute difference between the measured value of the target functional activity index and the ideal target functional activity index as ΔA; call the preset weighting coefficients, where the production cost weight α = 0.7 and the functional activity difference weight β = 0.3. These weighting coefficients are preset by analyzing the correlation coefficient between cost fluctuations and the activity achievement rate in historical production data: for every 1 yuan increase in cost, the probability of a decrease in the activity achievement rate is 0.23; for every 0.1 increase in the activity difference, the customer complaint rate increases by 0.37; perform a weighted summation calculation to generate the scalar form of the optimized objective function value f:
[0097] f = α × C + β × |ΔA|;
[0098] The value of this function is in the range of [12.8, 58.6] in typical scenarios. The smaller the value, the better the overall performance of the enzymatic hydrolysis parameters. For example, when C = 35.2 yuan and |ΔA| = 0.15, f = 0.7 × 35.2 + 0.3 × 0.15 = 24.685, which is better than the benchmark value of 28.4.
[0099] Methods that use Bayesian optimization algorithms to iteratively search for the global optimum of the objective function include:
[0100] S101: The initial set of sample points for enzymatic hydrolysis parameters is randomly generated using a sampling method;
[0101] S102: Input the initial enzymatic hydrolysis parameter sample point set into the optimization objective function for calculation to obtain the initial objective function value set;
[0102] S103: Construct a Gaussian process surrogate model in the Bayesian optimization algorithm based on the initial enzymatic hydrolysis parameter sample point set and the initial objective function value set;
[0103] S104: The desired improved acquisition function in the Bayesian optimization algorithm is used to calculate on the Gaussian process surrogate model to obtain candidate parameter points;
[0104] S105: Input the candidate parameter points into the optimization objective function for calculation to obtain the new objective function value;
[0105] S106: Add the candidate parameter points and the new objective function value to the initial parameter sample point set and the initial objective function value set to form the updated sample point set and the updated objective function value set;
[0106] S107: Reconstruct the Gaussian process proxy model based on the updated sample point set and the updated objective function value set;
[0107] S108: Repeat steps S104 to S107 until the set maximum number of iterations is reached. Output the optimal parameter point corresponding to the historical optimal objective function value as the global optimal solution, and map the optimal parameter point to the optimal enzymatic hydrolysis parameter.
[0108] Methods for obtaining candidate parameter points include:
[0109] X grid points are uniformly generated within the feasible region of the enzymatic hydrolysis parameters covered by the Gaussian process surrogate model; the expected improvement value of each grid point is calculated by calling the Gaussian process surrogate model; the grid point with the largest expected improvement value is selected as the candidate parameter point.
[0110] Please see Figure 3 As shown, this embodiment provides a method for obtaining the global optimal solution of the objective function by iteratively searching using a Bayesian optimization algorithm from the optimal enzymatic hydrolysis parameters obtained through optimization algorithm search, as detailed below:
[0111] Methods for randomly generating an initial sample set of enzymatic hydrolysis parameters using sampling methods include:
[0112] The feasible regions for five independent variables—enzyme dosage, substrate concentration, reaction temperature, pH, and reaction time—were defined as follows: enzyme dosage [0.5%, 3.0%], substrate concentration [4.0%, 8.0%], reaction temperature [40℃, 60℃], pH [6.0, 8.5], and reaction time [2h, 6h]. A Latin hypercube sampling method was used to generate an initial set of 20 sample points in the five-dimensional parameter space, ensuring that each parameter dimension was uniformly divided and that the projections of the sample points did not overlap. In practice, a random permutation sequence of the [0,1] interval was independently generated for each parameter dimension, and after mapping to the actual feasible region, parameter combination points were formed, for example:
[0113] Sample point 1: Enzyme dosage 1.2% / substrate concentration 5.3% / temperature 48℃ / pH 7.2 / time 3.5h;
[0114] Sample point 2: Enzyme dosage 2.1% / substrate concentration 6.7% / temperature 55℃ / pH 6.8 / time 4.8h; ...
[0116] Sample point 20: Enzyme dosage 0.8% / Substrate concentration 7.2% / Temperature 42℃ / pH 8.0 / Time 2.7h;
[0117] The output contains an initial set of enzymatic hydrolysis parameter sample points with 20 sets of parameter combinations. This set satisfies the space filling characteristic, that is, the uniformity deviation of the projection distribution in each dimension is less than 5%, which provides basic data support for subsequent Gaussian process modeling.
[0118] The initial set of enzymatic hydrolysis parameter sample points is input into the optimization objective function for calculation, and the initial objective function value set is obtained.
[0119] Methods for constructing a Gaussian process surrogate model in Bayesian optimization algorithms based on an initial set of enzymatic hydrolysis parameter sample points and an initial set of objective function values include:
[0120] The initial set of sample points for enzymatic hydrolysis parameters is used as the input matrix X, and the initial set of objective function values is used as the output vector y.
[0121] The similarity measure k(x) in the parameter space is defined using the squared exponential kernel function. i ,x j ):
[0122]
[0123] In the formula, x i and x j These are the enzymatic hydrolysis parameter vectors for the i-th and j-th samples, respectively; dLet x be the length-scale hyperparameter of the d-th dimension of the enzymatic hydrolysis parameters. The initial values for each dimension—enzyme dosage, substrate concentration, temperature, pH, and time—are set to 1.0. i,d Let x be the parameter value of the i-th sample in the d-th dimension; j,d Let be the parameter value of the j-th sample in the d-th dimension; Let be the initial value of the signal variance, set to 0.85 for the variance of the initial objective function value set; δ represents the noise variance, fixed at 0.01; exp(·) is the natural exponential function; i represents the i-th sample in the initial enzymatic hydrolysis parameter sample set; j represents the j-th sample in the initial enzymatic hydrolysis parameter sample set; d represents the d-th component of the enzymatic hydrolysis parameter vector; n is the number of samples in the initial enzymatic hydrolysis parameter sample set; δ ij is the Kroneckerdelta function, which is 1 when i = j and 0 otherwise; 5 is the total number of dimensions of the enzymatic hydrolysis parameters.
[0124] The hyperparameters are optimized by maximizing the marginal likelihood function. Specifically, the covariance matrix under the current hyperparameters is calculated; the negative log-marginal likelihood is solved; and the hyperparameters are iteratively updated using the L-BFGS-B algorithm until convergence. Example optimization results: length scale l = [0.8, 1.2, 1.5, 0.7, 2.0], signal variance...
[0125] By combining the optimized kernel function with the training data, a Gaussian process surrogate model is formed that can predict the probability distribution of the objective function value at any enzymatic hydrolysis parameter point.
[0126] Methods for obtaining candidate parameter points by using the expected improvement acquisition function in the Bayesian optimization algorithm on the Gaussian process surrogate model include:
[0127] Identify the historical optimal objective function value f from the current sample point set. min =18.7, corresponding to the enzymatic hydrolysis parameters [1.8%, 5.5%, 50℃, 7.0, 4.0h];
[0128] Generate 1000 uniform grid points within the feasible region of the enzymatic hydrolysis parameters covered by the Gaussian process surrogate model; perform the following operations for each grid point x:
[0129] Invoking a Gaussian process surrogate model to predict the posterior distribution of the objective function value:
[0130] Calculate the mean function μ(x), such as the output of 19.3 for grid point [2.0%, 6.0%, 52℃, 7.2, 3.8h]; calculate the variance function σ(x), such as the output of 0.25 for the same grid point; calculate the improvement amount I(x), such as I(x) = max(18.7 - 19.3, 0) = 0; calculate the expected improvement value EI(x):
[0131]
[0132] In the formula, Φ(·) is the standard normal cumulative distribution function; The standard normal probability density function is used; the EI value corresponding to each grid point is recorded.
[0133] After calculating the EI value for all grid points, select the grid point with the largest expected improvement value as the candidate parameter point. For example, when the expected improvement value of the point [1.5%, 5.2%, 49℃, 7.1, 4.2h] is 0.85, output the point. The obtained candidate parameter points must meet the constraints, such as enzyme dosage + substrate concentration being less than or equal to 9.0%, to ensure parameter feasibility.
[0134] The candidate parameter points are input into the objective function for calculation to obtain the new objective function value.
[0135] The candidate parameter points and the new objective function values are added to the initial parameter sample point set and the initial objective function value set to form the updated sample point set and the updated objective function value set.
[0136] Methods for reconstructing the Gaussian process surrogate model based on updating the sample point set and updating the objective function value set include:
[0137] The reconstruction process uses the same quadratic exponential kernel function as the initial construction, and optimizes the hyperparameters by maximizing the marginal likelihood function, while keeping the noise variance constant, as detailed below:
[0138] L-BFGS-B algorithm is used to iteratively update hyperparameters based on the updated sample point set. The similarity between any two sample points in the parameter space of the updated sample point set is calculated using the same kernel function, yielding the covariance matrix. The hyperparameters are then iteratively updated using the results of the previous model optimization, with the initial hyperparameters taken as the starting point. Calculate the covariance matrix under the current hyperparameters; solve for the negative logarithmic marginal likelihood; the BFGS-B algorithm updates the hyperparameters based on gradient and Hessian approximation until convergence, with the convergence condition set to a function value change of less than 10. -6 Or the gradient norm is less than 10 -5 .
[0139] For example, after optimization, the length-scale hyperparameter l = [0.82, 1.18, 1.53, 0.69, 1.96], the enzyme dosage dimension is slightly increased to 0.82, the substrate concentration dimension is slightly decreased to 1.18, the temperature dimension is slightly increased to 1.53, the pH dimension is slightly increased to 0.69, and the time dimension is slightly decreased to 1.96; signal variance Slightly reduced, reflecting the change in the variance of the objective function value after the update; noise variance is preserved. The negative logarithmic marginal likelihood convergence value decreased from the initial value of 15.3 to 14.8.
[0140] Once the set maximum number of iterations is reached, the optimal parameter point corresponding to the historical best objective function value is output as the global optimal solution; the method for setting the maximum number of iterations includes:
[0141] Since each iteration requires a time-consuming evaluation of the actual objective function, the basic framework needs to be determined based on the available computation time. For example, in the scenario of optimizing cottonseed protein hydrolysis parameters, a single objective function evaluation takes about 15 minutes, and the initial evaluation of 20 sample points takes 5 hours. If the total computation time is allowed to be 20 hours, the maximum number of iterations can be set to 60, and the total number of evaluation points can be 80, ensuring that all calculations are completed within 20 hours.
[0142] Secondly, consider the dimensional complexity of the parameter space. Enzymatic hydrolysis parameters include five dimensions: enzyme dosage, substrate concentration, reaction temperature, pH value, and reaction time. Based on empirical guidelines in Bayesian optimization, the recommended number of iterations is 20 to 50 times the number of parameter dimensions. For a five-dimensional parameter space, 150 iterations can be set as a theoretical reference value, resulting in a total of 170 evaluation points, effectively covering over 99% of the parameter space. However, in practice, this setting must be coordinated with computational resource constraints. If 150 iterations exceed the available time, the time limit should be prioritized.
[0143] A dynamic convergence monitoring mechanism is established as a safety guarantee. During the iteration process, the historical optimal objective function value is tracked in real time. Optimization is terminated early when the change in the objective function value over 10 consecutive iterations is less than a change threshold. For example, the change threshold can be set to 0.3, which is equivalent to 0.5% of the typical objective function value range [12.8, 58.6]. The maximum number of iterations is then used as a hard upper limit, and the actual number of iterations may be reduced due to early convergence.
[0144] During algorithm initialization, both the maximum number of iterations and the convergence threshold must be declared. After each dataset update, the current optimal solution is recorded. Finally, the historical optimal parameter point that satisfies any termination condition is output as the global optimal solution. This approach ensures optimization is completed with limited resources while allowing for timely exit based on convergence, avoiding unnecessary computation.
[0145] Methods for determining the historical optimal objective function value include:
[0146] If the current iteration is the first iteration, set the new objective function value as the historical best objective function value and record the candidate parameter point as the historical best parameter point; if it is not the first iteration, perform a judgment based on the mathematical properties of the optimization objective function: when the optimization objective function is a maximization function, if the new objective function value is greater than the historical best objective function value, then update the historical best objective function value and the historical best parameter point; when the optimization objective function is a minimization function, if the new objective function value is less than the historical best objective function value, then update the historical best objective function value and the historical best parameter point; if no update is performed, then retain the original historical best objective function value and the historical best parameter point.
[0147] Methods for mapping optimal parameter points to optimal enzymatic hydrolysis parameters include:
[0148] Specifically, this process is achieved through a preset parameter mapping function, which accurately converts the optimized dimensionless parameter values into enzymatic hydrolysis parameters with actual physical meaning, such as enzyme dosage, reaction temperature, pH value, reaction time, and substrate concentration.
[0149] For enzyme dosage, assuming the optimized dimensionless enzyme dosage parameter value is MYL, it can be converted into the actual enzyme dosage using the following mapping function:
[0150]
[0151] In the formula, E represents the actual amount of enzyme used; E min and E max These are the lower and upper limits of enzyme dosage, respectively; MYL min and MYL max These are the lower and upper limits of the dimensionless parameter value. For example, in this embodiment, E min It is 0.5%, E max It is 3.0%, MYL min MYL is 0 max The value is 1. This mapping relationship ensures that the enzyme dosage is within a reasonable range and corresponds to the parameter space searched by the optimization algorithm.
[0152] For the reaction temperature, assuming the dimensionless temperature parameter value is FYWD, the conversion formula for the actual reaction temperature is:
[0153]
[0154] In the formula, T is the actual reaction temperature; the maximum actual reaction temperature Tmax is the maximum actual reaction temperature. max The minimum actual reaction temperature is 60℃. min 40℃; FYWD max and FYWD minThese represent the maximum and minimum reaction temperatures, respectively. This linear mapping relationship, based on the temperature characteristics of the enzymatic hydrolysis reaction, ensures that a suitable temperature range can be effectively searched during the optimization process, while also taking into account operability and safety in actual production.
[0155] For pH value, assuming the dimensionless pH parameter is PHZ, the actual pH value is calculated as follows:
[0156]
[0157] In the formula, pH is the actual pH value; the minimum actual pH value is pH. min The maximum actual pH value is 6.0. max It is 8.5 PHZ max and PHZ min These represent the maximum and minimum pH values, respectively. This mapping allows for the accurate conversion of the optimized dimensionless values into the pH values required during actual enzymatic hydrolysis, thus meeting the enzyme's activity requirements and reaction conditions.
[0158] Regarding the reaction time, assuming the dimensionless reaction time parameter value is FYSJ, the actual reaction time is calculated as follows:
[0159]
[0160] In the formula, t is the actual reaction time; the minimum actual reaction time t min The maximum actual reaction time is t = 2 hours. max For 6 hours; FYSJ max and FYSJ min These represent the maximum and minimum response times, respectively. This mapping relationship allows the optimization algorithm to search for the optimal solution within a suitable timeframe, while also taking into account production efficiency and cost factors.
[0161] For substrate concentration, assuming the dimensionless substrate concentration parameter is DWND, the actual substrate concentration is calculated as follows:
[0162]
[0163] In the formula, S is the actual substrate concentration; the minimum actual substrate concentration Smin is the minimum actual substrate concentration. min The maximum actual substrate concentration was 4.0%, and the maximum actual substrate concentration was S. max It is 8.0%; DWND max and DWND min These represent the maximum and minimum substrate concentrations, respectively. In this way, the dimensionless values obtained from the optimization can be converted into actual substrate concentrations, ensuring that the enzymatic hydrolysis reaction proceeds at appropriate substrate concentrations, thereby improving the reaction rate and product quality.
[0164] The design of this parameter mapping function is of great significance. It not only transforms the abstract optimal solution obtained from the Bayesian optimization algorithm into practically operable enzymatic hydrolysis parameters, but also ensures that the parameter values are within a reasonable range, meeting the actual needs and production conditions of the enzymatic hydrolysis reaction. Compared with existing technologies, the parameter mapping method of this invention is more flexible and precise, capable of being adjusted according to different enzymatic hydrolysis systems and target requirements, thereby improving the controllability and optimization effect of the entire enzymatic hydrolysis process. In this way, the quality and production efficiency of enzymatic hydrolysis products can be significantly improved, while reducing production costs, demonstrating significant creativity and practicality.
[0165] Compared with existing technologies, the inventiveness of this embodiment lies in the following: First, it uses a Bayesian optimization algorithm for parameter search, which can search for the global optimal solution more efficiently in a high-dimensional nonlinear space than traditional gradient descent and genetic algorithms, especially for complex biochemical reaction systems such as enzymatic hydrolysis. Second, the constructed optimization objective function considers both production cost and functional activity, achieving multi-objective collaborative optimization, while existing technologies usually only focus on a single objective.
[0166] The control module performs enzymatic hydrolysis of cottonseed protein based on optimal enzymatic hydrolysis parameters. During the enzymatic hydrolysis process, it monitors the spatial structure data of the product in real time and adjusts the optimal enzymatic hydrolysis parameters to obtain the final enzymatic hydrolysis parameters.
[0167] Methods for obtaining the final enzymatic hydrolysis parameters include:
[0168] During the enzymatic hydrolysis of cottonseed, the spatial structure data of the hydrolysis products were collected in real time using a circular dichroism chromatograph to obtain real-time spatial structure data.
[0169] Calculate the root mean square error between real-time spatial structure data and preset standard spatial structure data to obtain the structural deviation value;
[0170] Input the structural deviation value into the preset fuzzy controller, and output the adjustment amount of the enzymatic hydrolysis parameters;
[0171] The final enzymatic hydrolysis parameters are obtained by algebraically adding the adjusted enzymatic hydrolysis parameters to the optimal enzymatic hydrolysis parameters.
[0172] For example, this embodiment provides a method for obtaining the final enzymatic hydrolysis parameters in a cottonseed protease hydrolysis control system, as detailed below:
[0173] During the enzymatic hydrolysis of cottonseed, the spatial structure data of the hydrolysis products were collected in real time using a circular dichroism chromatograph to obtain real-time spatial structure data.
[0174] Methods for calculating the root mean square error between real-time spatial structure data and preset standard spatial structure data to obtain the structural deviation value include:
[0175] Measured values of 12 feature dimensions were extracted from real-time spatial structure data, including α-helix ratio, β-fold ratio, β-turn ratio, random coil ratio, solvent-accessible surface area, gyroscope radius, number of disulfide bonds, hydrophobic core density, structural flexibility index, number of hydrogen bonds, electrostatic potential surface integral, and structural domain contact map density.
[0176] Synchronously acquire reference values for the corresponding dimensions from the preset standard spatial structure data;
[0177] Calculate the squared deviation d between the real-time measurement and the reference value of the k-th spatial structure feature dimension. k :
[0178]
[0179] In the formula, k is the index of the feature dimension of the spatial structure data, representing the dimension of the kth spatial structure data; is the measured value of the k-th spatial structure feature dimension; This is the reference value for the k-th spatial structure feature dimension; for example, when the measured value of the α-spiral ratio is 0.35 and the reference value is 0.40, d1 = 0.0025.
[0180] Sum the squared differences of all dimensions to obtain the sum of squared differences;
[0181] Based on the sum of squared differences, the arithmetic mean square root of the squared differences of all dimensions is calculated to generate a structural deviation value that characterizes the degree of deviation of the overall structure.
[0182] In a typical scenario, when the enzymatic hydrolysis reaction proceeds normally, the structural deviation value remains in the range of 0.05-0.15. For example, by comparing the real-time data [0.35, 0.25, ..., 0.15] with the standard value [0.40, 0.22, ..., 0.18], the squared difference of all dimensions is calculated to be 0.016, and the square root of the arithmetic mean is approximately 0.036.
[0183] The setup methods for a fuzzy controller include:
[0184] The input to the fuzzy controller is the structural deviation value, and the universe of discourse of the structural deviation value is [0,1], which is divided into three fuzzy sets: small, medium, and large.
[0185] The output of the fuzzy controller is set as the adjustment amount of the enzymatic hydrolysis parameter; the universe of discourse of the adjustment amount of the enzymatic hydrolysis parameter is [-R,R], which is divided into five fuzzy sets: negative large, negative small, zero, positive small, and positive large;
[0186] Set the fuzzy rule base as follows:
[0187] If the structural deviation value is small, the adjustment amount of the enzymatic hydrolysis parameter is zero;
[0188] If the structural deviation value is medium, then the adjustment amount of the enzymatic hydrolysis parameter is positively small;
[0189] If the structural deviation value is large, then the adjustment amount of the enzymatic hydrolysis parameter should be positive.
[0190] Methods for inputting structural deviation values into a preset fuzzy controller and outputting enzymatic hydrolysis parameter adjustment amounts include:
[0191] Based on the position of the structural deviation value in the [0,1] domain, calculate its membership degree to the small, medium, and large fuzzy sets; for example, when the structural deviation value is 0.4, the proportion of membership to the small fuzzy set is 0.3, the proportion of membership to the medium fuzzy set is 0.7, and the proportion of membership to the large fuzzy set is 0.0.
[0192] Reasoning is performed using preset fuzzy rules, which include the following rules:
[0193] Rule 1: If the structural deviation value belongs to a small fuzzy set, then the activated output fuzzy set is the zero fuzzy set;
[0194] Rule 2: If the structural deviation value belongs to the middle fuzzy set, then the activated output fuzzy set is the positive small fuzzy set;
[0195] Rule 3: If the structural deviation value belongs to a large fuzzy set, then the activated output fuzzy set is a positive large fuzzy set;
[0196] Based on the proportion of structural deviation values belonging to each fuzzy set, the activated output fuzzy set is truncated; for example, the zero fuzzy set is truncated at a membership degree of 0.3, and the positive small fuzzy set is truncated at a membership degree of 0.7; the truncated output fuzzy set consists of multiple independent membership functions, each corresponding to a rule, and these membership functions may overlap or conflict.
[0197] The truncated fuzzy sets are aggregated into an aggregation function μ using the maximum value principle. agg (y):
[0198] μ agg (y)=max(μ1(y),μ2(y),…,μ r (y));
[0199] In the formula, y is a variable in the universe of discourse [-R,R]; μ r (y) is the membership function after the r-th rule is truncated, and max(·) is the maximum value.
[0200] In the universe of discourse, R represents the maximum absolute range of the enzymatic hydrolysis parameter adjustment, i.e., the upper and lower limits that the fuzzy controller output can adjust. Within the universe of discourse [-R, R], -R represents the maximum downward adjustment range in one step, +R represents the maximum upward adjustment range in one step, and 0 represents no adjustment. The value of R should be consistent with the unit and dimension of the enzymatic hydrolysis parameter being adjusted. For example, when adjusting temperature, it can be expressed as the maximum allowable temperature change value; when adjusting pH, it can be expressed as the maximum allowable acidity / alkalinity change value. It should also be determined in conjunction with historical process experience, the stability of the reaction system, and the sensitivity of the system to parameter changes. If the system is highly sensitive to parameter changes, R should be appropriately reduced to avoid process oscillation; if the system has a large inertia, R can be appropriately increased to accelerate the convergence speed. Usually, process experts first give an empirical upper limit based on a safe and feasible adjustment range, then fine-tune it through experiments or process simulation, and finally use this value as a fixed configuration parameter in the control system to ensure that the adjustment can effectively promote the convergence of the index without destroying the stability of the system.
[0201] The calculated output is the enzymatic parameter adjustment of the aggregation function over the [-R,R] domain.
[0202] output=∫y×μ agg (y)dy / ∫μ agg (y)dy;
[0203] For example, when R = 0.5, the adjustment amount is +0.25, which means that the enzymatic hydrolysis parameters need to be slightly increased; the adjustment amount of the enzymatic hydrolysis parameters is algebraically added to the current enzymatic hydrolysis parameters to obtain the adjusted enzymatic hydrolysis parameters.
[0204] Existing technologies typically rely on fixed parameters for enzymatic hydrolysis, which cannot handle fluctuations and uncertainties during the reaction process, resulting in poor product structural consistency. This invention, however, through real-time monitoring and dynamic compensation, can control the α-helix and β-sheet content of the product within a small fluctuation range, significantly improving the stability of the product structure.
[0205] After obtaining the optimal enzymatic hydrolysis parameters, the feasibility of the adjusted parameters is verified based on a preset production cost threshold. If feasible, the final enzymatic hydrolysis parameters are obtained; if not, the adjusted parameters are corrected based on a preset cost-parameter adjustment mapping table to obtain corrected enzymatic hydrolysis parameters. The corrected parameters are then verified for feasibility. If still infeasible, further adjustments may lead to a significant decrease in enzyme activity and increased costs, as some enzymatic hydrolysis parameters have process boundary limitations. In such cases, auxiliary parameters need to be introduced in addition to the traditional enzymatic hydrolysis parameters. These auxiliary parameters are added as independent variables, allowing the output optimal enzymatic hydrolysis parameters to be corrected during the feasibility verification process if they meet the requirements or are infeasible, ultimately meeting the feasibility verification requirements. These auxiliary parameters include surface activity cost and reaction medium modification cost.
[0206] Please see Figure 4 As shown, the methods for verifying the feasibility of modifying enzymatic hydrolysis parameters include:
[0207] Extract the enzyme dosage, temperature, and pH value from the adjusted enzymatic hydrolysis parameters to form an enzymatic hydrolysis parameter vector;
[0208] The enzymatic hydrolysis parameter vector is input into a preset production cost calculation function to calculate the estimated production cost; the production cost calculation function is obtained by multiple linear regression fitting based on historical enzymatic hydrolysis production data.
[0209] The cost difference is obtained by performing an algebraic subtraction between the estimated production cost and the preset production cost threshold.
[0210] If the cost difference is less than or equal to zero, it is feasible; if the cost difference is greater than zero, it is not feasible.
[0211] For example, this embodiment provides a method for obtaining the final enzymatic hydrolysis parameters in a cottonseed protease hydrolysis control system, as detailed below:
[0212] Extract the enzyme dosage, temperature, and pH value from the adjusted enzymatic hydrolysis parameters to form an enzymatic hydrolysis parameter vector.
[0213] Input the enzymatic hydrolysis parameter vector into the preset production cost calculation function to calculate the estimated production cost.
[0214] Methods for constructing production cost calculation functions include:
[0215] Collect historical enzymatic hydrolysis production data, extract enzyme usage, temperature and pH value, and corresponding actual production costs to form a production dataset;
[0216] A multiple linear regression model was used to fit the production dataset to obtain the production cost calculation function.
[0217] For example, this embodiment provides a method for constructing a production cost calculation function during the process of obtaining the final enzymatic hydrolysis parameters, as follows:
[0218] Historical enzymatic hydrolysis production data were collected, and enzyme usage, temperature, pH value, and corresponding actual production costs were extracted to form a production dataset.
[0219] Methods for obtaining production cost calculation functions by performing multiple linear regression fitting on production datasets include:
[0220] Based on the historical enzymatic hydrolysis production dataset that has been collected, which includes three independent variables—enzyme dosage, temperature, and pH value—and their corresponding actual production costs as dependent variables, the least squares method is used to perform multiple linear regression fitting.
[0221] Let the enzyme dosage be v1, the temperature be v2, the pH value be v3, and the production cost be u. Establish a linear model: u = β0 + β1v1 + β2v2 + β3v3.
[0222] By solving the normal equation system, the optimal estimates of regression coefficients β0, β1, β2 and β3 are determined, so as to minimize the sum of squared residuals between the predicted production cost and the actual production cost for all data points.
[0223] The obtained production cost calculation function expression is:
[0224] Predicted production cost = b0 + b1 × enzyme dosage + b2 × temperature + b3 × pH;
[0225] Where b0, b1, b2, and b3 are the regression coefficients determined by the fitting. The production cost calculation function can directly input new combinations of enzyme dosage, temperature, and pH value, and output the corresponding predicted production cost.
[0226] The estimated production cost is subtracted from the preset production cost threshold using an algebraic subtraction method to obtain the cost difference. The method for setting the production cost threshold includes:
[0227] Extract all actual production cost records from the collected historical enzymatic hydrolysis production data. After sorting these records in ascending order, for example, take the actual production cost value at the 85th percentile of the sorted positions as the threshold benchmark. For example, when the historical data contains 100 records, take the 85th highest actual cost value.
[0228] Based on this, taking into account the current fluctuation range of raw material prices and equipment depreciation coefficient, the threshold benchmark is positively adjusted: when the average purchase price of the main raw materials in the past three months increases by P% compared with the historical data collection period, the threshold benchmark is multiplied by (1+P% / 2); at the same time, based on the service life N (years) of the production equipment, it is multiplied by the depreciation coefficient (1-0.01×min(N,5)).
[0229] The revised cost value is adjusted to the thousandths place and set as the production cost threshold. For example, if the historical 85th percentile cost is 125,800 yuan, the current raw material price has increased by 6%, and the equipment has been used for 3 years, the calculation process is: 125,800 × (1 + 6% / 2) × (1 - 0.01 × 3) = 125,800 × 1.03 × 0.97 ≈ 125,600 yuan. This 126,000 yuan is taken as the official threshold. This setting ensures that the threshold is always higher than the 85th percentile of the historical actual cost and dynamically adapts to cost-influencing factors. When the estimated cost exceeds this value, the solution is deemed unacceptable.
[0230] If the cost difference is less than or equal to zero, the adjusted enzymatic hydrolysis parameters are output as candidate enzymatic hydrolysis parameters; if the cost difference is greater than zero, the pre-defined cost-parameter adjustment mapping table is consulted based on the cost difference to obtain the enzymatic hydrolysis parameter adjustment scheme. The setting method for the cost-parameter adjustment mapping table includes:
[0231] All records where the actual production cost was greater than the corresponding period threshold were screened from the historical enzymatic hydrolysis production dataset. The cost difference was calculated for each record. At the same time, the enzymatic hydrolysis parameter adjustment scheme finally adopted in the case was extracted, including the adjustment amount of enzyme dosage, temperature adjustment and pH value adjustment. These schemes came from the parameter correction values implemented in the historical operation log to reduce costs.
[0232] Group the cost difference in 10% intervals, such as 0%–10%, 10%–20%, etc.
[0233] For each group, the arithmetic mean of the enzyme dosage adjustment, temperature adjustment, and pH adjustment is calculated to form a mapping table entry. For example, historical data analysis shows that in cases where the cost difference is 0% to 10%, the enzyme dosage is reduced by an average of 5%, the temperature is reduced by an average of 2°C, and the pH value is increased by an average of 0.3; when the cost difference is 10% to 20%, the enzyme dosage is reduced by an average of 8%, the temperature is reduced by an average of 3°C, and the pH value is increased by an average of 0.5.
[0234] The cost-parameter adjustment mapping table takes the cost difference range as the input key and outputs the corresponding average enzyme dosage adjustment ratio, absolute temperature adjustment value, and absolute pH adjustment value as the adjustment scheme, ensuring that the scheme is directly applicable to parameter replacement operations.
[0235] Extract the enzyme dosage adjustment ratio from the enzymatic hydrolysis parameter adjustment scheme obtained from the query, for example, -5% means a 5% reduction, temperature adjustment value, for example, -2℃ means a 2℃ reduction, and pH adjustment value, for example, +0.3 means an increase of 0.3 units.
[0236] Perform a proportional calculation on the enzyme dosage in the adjusted enzymatic hydrolysis parameters:
[0237] Corrected enzyme dosage = original enzyme dosage × (1 + adjustment ratio / 100);
[0238] For example, if the original enzyme dosage is 100kg, adjusting the ratio by -5% would correct it to 95kg.
[0239] Perform algebraic addition on the temperature:
[0240] Corrected temperature = Original temperature + Temperature adjustment value;
[0241] For example, if the original temperature is 45℃ and adjusted by -2℃, it will be corrected to 43℃.
[0242] The same algebraic addition is performed on the pH value:
[0243] Corrected pH value = Original pH value + pH adjustment value;
[0244] For example, if the original pH is 5.0, adjusting it by +0.3 will correct it to 5.3;
[0245] The corrected enzyme dosage, temperature, and pH value are combined to form a new vector of enzymatic hydrolysis parameters, which are called the corrected enzymatic hydrolysis parameters. This process ensures that each parameter is replaced with an accurate value according to the mapping table scheme. For example, when the adjustment scheme is (-8%, -3℃, +0.5), the parameter (120kg, 48℃, 4.8) will be replaced with (120×0.92=110.4kg, 48-3=45℃, 4.8+0.5=5.3) to form the corrected enzymatic hydrolysis parameters.
[0246] The feasibility of modifying the enzymatic hydrolysis parameters is verified. If it is still not feasible, auxiliary parameters are introduced and added to the independent variables to reconstruct the optimization objective function. Methods for searching for new optimal enzymatic hydrolysis parameters include:
[0247] In addition to the original five types of enzymatic hydrolysis parameters, surface activity cost SC and reaction medium modification cost MC are added to form a seven-dimensional independent variable vector xz=[E,S,T,pH,t,SC,MC];
[0248] Then the production cost value C(xz) = C(x) + SC + MC, the functional activity difference |ΔA| and the weight coefficient remain unchanged, and the optimization objective function is reconstructed.
[0249] Similarly, the Bayesian optimization algorithm is used to iteratively search for new optimal enzymatic hydrolysis parameters.
[0250] Existing technologies typically focus only on optimizing functional activity while neglecting production cost constraints, rendering the optimized results infeasible in actual production. This invention, by constructing a precise production cost calculation function, achieves synergistic optimization of functional activity and economic efficiency, demonstrating significant inventiveness.
[0251] The 12-dimensional spatial structure data acquired in real time by the circular dichroism chromatograph can accurately capture subtle changes in the spatial structure of peptides. For example, when fluctuations in the enzymatic digestion temperature cause α-helices to transform into random coils, the decrease in the proportion of α-helices in the real-time data will be immediately identified. A structural deviation value is generated by calculating the root mean square error between the 12-dimensional data and the standard structure. This deviation value is processed by a fuzzy controller to output a parameter adjustment amount. For example, when the deviation value is in the middle fuzzy set, the fuzzy controller automatically generates a temperature adjustment amount of +0.25, which is used to correct the current enzymatic digestion temperature in real time through algebraic addition.
[0252] Compared to the problem of large fluctuations in the yield of target active peptides in traditional fixed-parameter processes, this invention, through real-time closed-loop control of 12-dimensional structural features, controls the fluctuation range of α-helix and β-sheet content within a smaller range, significantly reducing the fluctuation amplitude of antimicrobial peptide yield and improving the stability of the IC50 value of ACE inhibitory peptides. The core reason for this is that the fuzzy controller, through a rule base of "zero adjustment for small deviations, small positive adjustment for medium deviations, and large positive adjustment for large deviations," can nonlinearly match the relationship between structural deviations and parameter compensation amounts, avoiding the adjustment lag problem of traditional PID control in nonlinear biochemical reaction scenarios. The full-dimensional monitoring of 12-dimensional structural features, rather than monitoring a single index, ensures the overall stability of the peptide spatial conformation. For example, simultaneously monitoring the hydrophobic core density and the number of hydrogen bonds avoids the defect of focusing only on secondary structure while ignoring changes in tertiary structure.
[0253] Example 2
[0254] Please see Figure 2 As shown, this embodiment provides a method for controlling cottonseed protease hydrolysis, including:
[0255] The molecular weight distribution and spatial structure data of the enzymatic hydrolysis products collected in real time during the cottonseed protein hydrolysis process are input into a pre-trained deep learning model, and the target functional activity index is output.
[0256] Based on the enzymatic hydrolysis parameters in the cottonseed protein hydrolysis process, as well as the preset ideal target functional activity index and the obtained target functional activity index, an optimization objective function is constructed; the optimal enzymatic hydrolysis parameters are obtained by searching the optimization objective function.
[0257] Cottonseed protein was enzymatically hydrolyzed based on optimal hydrolysis parameters. The spatial structure data of the product was monitored in real time during the hydrolysis process, and the optimal hydrolysis parameters were adjusted to obtain the final hydrolysis parameters.
[0258] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0259] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for controlling the enzymatic hydrolysis of cottonseed, characterized in that, include: The molecular weight distribution and spatial structure data of the enzymatic hydrolysis products collected in real time during the cottonseed protein hydrolysis process are input into a pre-trained deep learning model, and the target functional activity index is output. Based on the enzymatic hydrolysis parameters in the cottonseed protein hydrolysis process, as well as the preset ideal target functional activity index and the obtained target functional activity index, an optimization objective function is constructed; the optimal enzymatic hydrolysis parameters are obtained by searching the optimization objective function. Cottonseed protein was enzymatically hydrolyzed based on optimal hydrolysis parameters. The spatial structure data of the product was monitored in real time during the hydrolysis process, and the optimal hydrolysis parameters were adjusted to obtain the final hydrolysis parameters.
2. The method for controlling cottonseed protease hydrolysis according to claim 1, characterized in that, Methods for constructing the objective function include: The target functional activity indicators include antibacterial activity indicators and antioxidant activity indicators; the enzymatic hydrolysis parameters include enzyme dosage, substrate concentration, reaction temperature, pH value and reaction time in the cottonseed protein hydrolysis process; The enzymatic hydrolysis parameters were used as independent variables. The difference between the target functional activity index and the preset ideal target functional activity index is used as the dependent variable; The optimized objective function is obtained.
3. The method for controlling cottonseed protease hydrolysis according to claim 1, characterized in that, Methods for obtaining optimal enzymatic hydrolysis parameters by searching for optimization algorithms on the objective function include: S101: The initial set of sample points for enzymatic hydrolysis parameters is randomly generated using a sampling method; S102: Input the initial enzymatic hydrolysis parameter sample point set into the optimization objective function for calculation to obtain the initial objective function value set; S103: Construct a Gaussian process surrogate model in the Bayesian optimization algorithm based on the initial enzymatic hydrolysis parameter sample point set and the initial objective function value set; S104: The desired improved acquisition function in the Bayesian optimization algorithm is used to calculate on the Gaussian process surrogate model to obtain candidate parameter points; S105: Input the candidate parameter points into the optimization objective function for calculation to obtain the new objective function value; S106: Add the candidate parameter points and the new objective function value to the initial parameter sample point set and the initial objective function value set to form the updated sample point set and the updated objective function value set; S107: Reconstruct the Gaussian process proxy model based on the updated sample point set and the updated objective function value set; S108: Repeat steps S104 to S107 until the set maximum number of iterations is reached, output the optimal parameter point corresponding to the historical best objective function value, and map the optimal parameter point to the optimal enzymatic hydrolysis parameter.
4. The method for controlling cottonseed protease hydrolysis according to claim 3, characterized in that, Methods for determining the historical optimal objective function value include: If the current iteration is the first iteration, set the new objective function value as the historical best objective function value, and record the candidate parameter point as the historical best parameter point; If it is not the first iteration, then a judgment is made based on the mathematical properties of the objective function: When the objective function is a maximization function, if the new objective function value is greater than the historical best objective function value, then the historical best objective function value and the historical best parameter point are updated. When the objective function is to minimize the objective function, if the new objective function value is less than the historical best objective function value, then the historical best objective function value and the historical best parameter point are updated. If no update is performed, the original historical best objective function value and historical best parameter point are retained.
5. The method for controlling cottonseed protease hydrolysis according to claim 1, characterized in that, Methods for obtaining the final enzymatic hydrolysis parameters include: During the enzymatic hydrolysis of cottonseed, the spatial structure data of the hydrolysis products were collected in real time using a circular dichroism chromatograph to obtain real-time spatial structure data. Calculate the root mean square error between real-time spatial structure data and preset standard spatial structure data to obtain the structural deviation value; Input the structural deviation value into the preset fuzzy controller, and output the adjustment amount of the enzymatic hydrolysis parameters; The final enzymatic hydrolysis parameters are obtained by algebraically adding the adjusted enzymatic hydrolysis parameters to the optimal enzymatic hydrolysis parameters.
6. The method for controlling cottonseed protease hydrolysis according to claim 5, characterized in that, The setup methods for a fuzzy controller include: The input to the fuzzy controller is the structural deviation value, and the domain of the structural deviation value is from 0 to 1, which is divided into three fuzzy sets: small, medium, and large. The output of the fuzzy controller is set as the adjustment amount of the enzymatic hydrolysis parameters; the domain of the adjustment amount of the enzymatic hydrolysis parameters is from -R to R, and is divided into five fuzzy sets: negative large, negative small, zero, positive small, and positive large; among them, the determination of the value of R needs to comprehensively consider the physical boundary of the enzymatic hydrolysis parameters and the quantitative classification requirements of the fuzzy sets; Set the fuzzy rule base as follows: If the structural deviation value is small, the adjustment amount of the enzymatic hydrolysis parameter is zero; If the structural deviation value is medium, then the adjustment amount of the enzymatic hydrolysis parameter is positively small; If the structural deviation value is large, then the adjustment amount of the enzymatic hydrolysis parameter should be positive.
7. The method for controlling cottonseed protease hydrolysis according to claim 2, characterized in that: The production cost associated with the independent variable is taken as the dependent variable; the objective function is obtained by weighted summation of the dependent variables.
8. The method for controlling cottonseed protease hydrolysis according to claim 7, characterized in that: After obtaining the optimal enzymatic hydrolysis parameters, the feasibility of the adjusted enzymatic hydrolysis parameters is verified based on the preset production cost threshold. If it is feasible, the final enzymatic hydrolysis parameters are obtained; if it is not feasible, the adjusted enzymatic hydrolysis parameters are corrected based on the preset cost-parameter adjustment mapping table to obtain the corrected enzymatic hydrolysis parameters. The feasibility of modifying the enzymatic hydrolysis parameters is verified. If it is still not feasible, auxiliary parameters are introduced and added to the independent variables. The auxiliary parameters include the surface activity cost and the reaction medium modification cost.
9. The method for controlling cottonseed protease hydrolysis according to claim 8, characterized in that, Methods for verifying the feasibility of modifying enzymatic hydrolysis parameters include: Extract the enzyme dosage, temperature, and pH value from the adjusted enzymatic hydrolysis parameters to form an enzymatic hydrolysis parameter vector; The enzymatic hydrolysis parameter vector is input into a preset production cost calculation function to calculate the estimated production cost; the production cost calculation function is obtained by multiple linear regression fitting based on historical enzymatic hydrolysis production data. The cost difference is obtained by performing an algebraic subtraction between the estimated production cost and the preset production cost threshold. If the cost difference is less than or equal to zero, it is feasible; if the cost difference is greater than zero, it is not feasible.
10. A cottonseed protease hydrolysis control system, used to implement the cottonseed protease hydrolysis control method according to any one of claims 1-9, characterized in that, include: The prediction module is used to input the molecular weight distribution data and spatial structure data of the enzymatic hydrolysis products collected in real time during the cottonseed protein hydrolysis process into a pre-trained deep learning model and output the target functional activity index. The optimization module constructs an optimization objective function based on the enzymatic hydrolysis parameters during the cottonseed protein hydrolysis process, as well as the preset ideal target functional activity index and the obtained target functional activity index; and performs an optimization algorithm search on the optimization objective function to obtain the optimal enzymatic hydrolysis parameters. The control module performs enzymatic hydrolysis of cottonseed protein based on optimal enzymatic hydrolysis parameters. During the enzymatic hydrolysis process, it monitors the spatial structure data of the product in real time and adjusts the optimal enzymatic hydrolysis parameters to obtain the final enzymatic hydrolysis parameters.