An enzyme modification method based on data-driven optimization

By combining molecular docking and point mutation techniques with low-dimensional mutual information coding and a batch Bayesian optimization algorithm based on Gaussian process regression models, the problems of high-dimensional coding and local optima in enzyme modification are solved, achieving high efficiency and low cost in enzyme molecular modification and improving the reliability and efficiency of enzyme modification.

CN116959566BActive Publication Date: 2026-01-23EAST CHINA UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310852648.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-12
Publication Date
2026-01-23
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

Existing enzyme modification methods suffer from poor high-dimensional encoding and representation capabilities, reliance on the quality of the training set for selection of validation targets, and a tendency to get trapped in local optima, resulting in low efficiency and high cost in enzyme modification.

Method used

Molecular docking and point mutagenesis techniques were used to identify mutation hotspot residues. Combined with low-dimensional mutual information encoding and Gaussian process regression models, batch Bayesian optimization algorithms were used to screen and validate targets, constructing a decision space for enzyme modification and achieving iterative screening of globally optimal mutation sequences.

Benefits of technology

This technology enables efficient and low-cost enzyme molecule modification, reducing time and economic investment, improving the reliability and efficiency of enzyme modification, and avoiding resource waste under unreasonable termination conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116959566B_ABST
    Figure CN116959566B_ABST
Patent Text Reader

Abstract

The application discloses an enzyme modification method based on data-driven optimization, comprising the following steps: determining a plurality of mutation hot spot residues by using molecular docking and point mutation technology to generate a mutation space; taking the difference set of the mutation space and a single-point saturation mutation set as a decision space of enzyme modification, and obtaining initial training data according to the single-point saturation mutation sequence and modification characteristics thereof; determining the hyperparameters of low-dimensional mutual information coding and coding the decision space and the initial training data; constructing a proxy model according to the current training data, and then selecting an object for this round of experimental verification from the decision space by using a batch Bayesian optimization algorithm based on a maximum variance change amount to obtain modification characteristics of the verification object; repeating the steps of updating the current training data, constructing the proxy model and obtaining the modification characteristics of the verification object until a condition is met, and obtaining an enzyme modification result. The method can effectively reduce the time cost and economic investment in the enzyme modification process, and can improve the reliability of the enzyme modification method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a data-driven optimization-based enzyme modification method, belonging to the field of machine learning-assisted protein design. Background Technology

[0002] Thanks to the high efficiency, specificity, and environmental friendliness of enzymatic reactions, enzymes have wide applications not only in traditional fields such as chemical engineering, food processing, and environmental science, but also play an irreplaceable role in emerging technologies and products such as gene editing, stem cell technology, and targeted drugs. Numerous studies have found that natural proteases often fail to meet the demands of practical applications in terms of stability, tolerance, and selectivity. Therefore, optimizing and modifying enzyme molecules is not only a key focus of protein science research but also an urgent need for industrial production.

[0003] In the field of protein engineering, commonly used protein modification methods include directed evolution, semi-rational design, and rational design. Directed evolution guides proteins to accumulate beneficial mutations through multiple rounds of repeated mutation, expression, and screening. However, the random introduction of mutations in directed evolution results in a large number of mutants, which is highly detrimental to artificial screening. Semi-rational design selects several sites as modification targets based on prior knowledge such as crystal structure and catalytic mechanisms, thereby improving modification efficiency. However, the success of semi-rational design is closely related to the richness of prior knowledge, leading to significant limitations in its application. Rational design attempts to obtain enzymes with desired properties by precisely controlling the structural space of proteins. However, current limitations stem from the high precision of obtaining the spatial structure of enzyme molecules and the rational understanding of structure-function relationships and catalytic mechanisms, resulting in very limited successful cases of enzyme modification through rational design.

[0004] The high dimensionality and strong coupling of amino acid sequences, along with limited sample size, pose significant challenges to machine learning-based enzyme modification research. Machine learning-based enzyme modification involves three key steps: protein feature extraction, predictive model construction, and validation target selection. Current research has proposed various encoding methods for extracting protein features by incorporating amino sequence arrangement information and physicochemical information of amino acid residues. However, these methods suffer from high dimensionality, poor characterization ability, and unreasonable distribution of quantification values. Furthermore, existing validation target selection methods are highly dependent on the quality of the training set and are prone to getting trapped in local optima, thus failing to guarantee the effectiveness of enzyme modification. Therefore, to achieve efficient and low-cost enzyme molecule modification, in addition to designing encoding methods with stronger characterization capabilities, it is also necessary to introduce more advanced machine learning techniques. Summary of the Invention

[0005] The purpose of this invention is to provide a data-driven optimization method for enzyme modification, thereby enabling efficient enzyme molecule modification at an acceptable cost.

[0006] To achieve the above objectives, the present invention provides a data-driven optimization-based enzyme modification method, comprising:

[0007] Step S1: Several mutation hotspot residues are identified using molecular docking and point mutagenesis techniques, and the mutation space for enzyme modification is generated based on the amino acid sequence and the given maximum number of mutation sites.

[0008] Step S2: Perform enzyme modification characteristic experiments on each unit point saturated mutant sequence based on mutation hotspot residues to obtain the quantitative value of the modification characteristics, and use the unit point saturated mutant sequence and its corresponding quantitative value as the initial training data for the enzyme modification experiment; in addition, use the difference between the mutation space and the set formed by the unit point saturated mutant sequence as the decision space for the enzyme modification experiment.

[0009] Step S3: Based on the initial training data, determine the hyperparameters of the low-dimensional mutual information encoding method using cross-validation. Then, according to the determined hyperparameters, apply the low-dimensional mutual information encoding method to the decision space. Encode the initial training data to obtain the encoded decision space. and current training data;

[0010] Step S4: Based on the current training data, a Gaussian process regression model is constructed using a Gaussian process as a surrogate model. Based on the surrogate model, a batch Bayesian optimization algorithm based on maximizing the change in variance is used to select the verification object for this round of experiments from the encoded decision space. Then, the verification object is decoded into an amino acid sequence and the modification characteristics of the amino acid sequence are experimentally verified to obtain the quantitative value of the modification characteristics of the verification object.

[0011] Step S5: Add all the verification objects of this round of experiments and their quantified values ​​of modification characteristics to the current training data to update the current training data. Update the encoded decision space by deleting the verification objects of this round of experiments, and then go to step S4. Until the termination condition of the batch Bayesian optimization algorithm is met or the maximum number of iterations is reached, take the amino acid sequence corresponding to the optimal result of the quantified value of the modification characteristics of the verification object at this time as the final result of enzyme modification.

[0012] Preferably, step S1 specifically includes:

[0013] Step S11: Based on the protein structure, the substrate binding pocket is determined by molecular docking of the enzyme and the substrate. Then, all mutation hotspot residues related to catalytic activity are identified by site-directed mutagenesis of amino acid residues around the substrate binding pocket and determination of the catalytic activity of the corresponding mutants.

[0014] Step S12: Based on the amino acid sequence and the given maximum number of mutation sites, a mutation space for enzyme modification is generated through mutation hotspot residues. The mutation space is the mutation sequence s of the enzyme. i The collection of enzyme mutant sequences s i It is obtained by mutating at least one hotspot residue in the enzyme sequence, and the mutated sequence s i The number of mutation sites is at most the maximum number of mutation sites mentioned above.

[0015] Preferably, step S2 specifically includes:

[0016] Step S21: Perform enzyme modification characteristic experiments on each site-saturated mutant sequence based on mutation hotspot residues to obtain the quantitative value of the modification characteristics, and use the sequence value of the site-saturated mutant sequence and its corresponding quantitative value as the initial training data for the enzyme modification experiment.

[0017] Step S22: The difference between the mutation space and the set of unit-point saturated mutation sequences is used as the decision space for enzyme modification.

[0018] Preferably, in step S3, the low-dimensional mutual information encoding method includes:

[0019] Step S31: Replace each amino acid sequence in the mutation space S one by one with the amino acid T-scale topological descriptor to obtain the descriptor matrix M for each amino acid sequence.

[0020] Step S32: Calculate the autocovariance matrix C for each amino acid sequence using the following formula:

[0021]

[0022] Where b represents the distance between amino acid residues, and its value is {1,2,…,l}, C b,j With M i,j Let i and j represent the j-th element of the b-th row of the autocovariance matrix C and the j-th element of the i-th row of the descriptor matrix M, respectively. i represents the row position in the descriptor matrix M, j is the group of the descriptor matrix M, and m represents the number of rows in the descriptor matrix M, which is the length of the amino acid sequence.

[0023] Step S33: For each amino acid sequence in the mutation space, concatenate the columns of its autocovariance matrix C, and use principal component analysis to reduce the dimensionality of the concatenation result according to the dimensionality reduction d. The dimensionality reduction result is the encoding result of the mutation space S, which includes the encoded decision space. The encoding results of the sequence values ​​of the initial training data;

[0024] Step S34: Find the encoding result of the sequence value of the initial training data in the encoding result of the mutation space S and use it as the sample input x. i The sample output y is obtained by quantizing the modified characteristics of the initial training data. i Each input sample and each output sample is used as a training sample, thus forming the current training data.

[0025] Preferably, the sample output y is obtained based on the quantized value of the modified characteristics of the initial training data. i Specifically, this includes: optimizing the distribution of the quantized values ​​of the modification characteristics of the initial training data to obtain the optimized distribution result of the quantized values ​​of the modification characteristics of the initial training data, and using the optimized distribution result of the quantized values ​​of the modification characteristics of the initial training data as the sample output y. i .

[0026] Preferably, the distribution of the quantized values ​​of the modified characteristics of the initial training data is optimized using the following formula:

[0027] y = alan(v),

[0028] In the formula, y represents the distribution optimization result of the quantized value, a∈[1,+∞) is the scaling factor of the sample output, and v represents the quantized value of the modified characteristics.

[0029] Preferably, in step S4, a Gaussian process regression model is constructed based on the current training data as a surrogate model, and a batch Bayesian optimization algorithm based on maximizing the change in variance is used to select the validation objects for this round of experiments from the encoded decision space according to the surrogate model, specifically including:

[0030] Step S41: Construct and train a Gaussian process regression model based on the current training data, and obtain the kernel parameter θ of the covariance function and the noise variance from the trained Gaussian process regression model.

[0031] Step S42, based on the type of problem to be solved, from the encoded decision space The sample set worth evaluating is divided from the sample set.

[0032] Step S43, if the sample set X is worth evaluating * If the sample size m is less than or equal to the batch size q, then the sample set X worth evaluating is... * The evaluation sample set obtained in this iteration The evaluation sample set obtained in this iteration All verification objects are decoded into amino acid sequences and the modification characteristics of the amino acid sequences are experimentally verified. After obtaining the quantification value of the modification characteristics, step S5 is executed.

[0033] Preferably, in step S42, when solving the problem for the sample corresponding to the maximum value of the quantized value of the modified characteristic, the following formula is used from the decision space. The sample set worth evaluating is divided from the data:

[0034]

[0035] Where x represents the encoded decision space. The sample in y max z represents the maximum value of the output of the sample in the current training data. α For the lower α quantile of the standard normal distribution, μ(x) and δ n (x) represent the predicted mean and predicted standard deviation of the Gaussian process regression model for sample x, respectively;

[0036] When solving a problem to find the sample corresponding to the minimum quantized value of a modified characteristic, the following formula is used from the decision space. The sample set worth evaluating is divided from the data:

[0037]

[0038] In the formula, y min This represents the minimum value of the sample output in the current training data.

[0039] Preferably, step S4 further includes:

[0040] Step S44, when the sample set X is worth evaluating * The sample size m exceeds 10,000, and samples are randomly selected from the sample set X that is worth evaluating. * Select 10,000 samples from the sample set to form a new sample set X worthy of evaluation. * ;

[0041] Step S45, when the sample set X is worth evaluating * When the sample size m is less than or equal to 10000 and greater than the batch size q, the sample set X worth evaluating is selected. * We sequentially select q samples that maximize the following formula as the evaluation sample set obtained in this iteration, and remove them from the evaluation sample set X. * Delete from the middle, then return to step S43:

[0042]

[0043] In the formula, It is the i-th evaluation sample selected. The sample set X is worth evaluating. * The j-th sample in the dataset, where n represents the number of samples in the current training data, and n+i represents the index of the newly selected ith evaluation sample, Xn It is the set of sample inputs to the current training data, k i,j express and The covariance between them, k n+i,j This represents the union of the set of input samples in the current training data and the set of the first i samples in the evaluation sample set. and The covariance vector between them express With the selected (i+1)th evaluation sample The covariance vector between them, A n+1 For intermediate parameters, K n+i express The covariance matrix, It represents the noise variance, and I denotes the identity matrix. It is based on Calculate the output variance.

[0044] Preferably, in step S5, the termination condition of the batch Bayesian optimization algorithm is:

[0045] When the decision space is large and the objective function is complex, satisfying the following formula three times consecutively is used as the termination condition for the batch Bayesian optimization algorithm:

[0046] |X * |≤0.5%N

[0047] Where N represents the total number of samples in the mutation space, |X * | represents the sample set X that is worth evaluating. * The number of samples;

[0048] In other scenarios, |X * |=0 is used as the termination condition for the batch Bayesian optimization algorithm.

[0049] The data-driven optimization-based enzyme modification method of this invention models enzyme modification as a combinatorial optimization problem of a black-box function, and then introduces batch Bayesian optimization technology from machine learning to solve this problem. Benefiting from the representational power of low-dimensional mutual information encoding, a surrogate model with a well-fitting sequence-function mapping relationship can be constructed. Thanks to the global optimization capability of batch Bayesian optimization, the proposed enzyme modification method can continuously approach the globally optimal mutation sequence in the decision space through an iterative screening process, thereby achieving the modification of specific functions of enzyme molecules. Furthermore, thanks to the termination condition of the batch Bayesian optimization algorithm based on the maximum variance change, the provided enzyme modification method can use a probabilistic method based on current data to determine whether the globally optimal mutation sequence in the decision space has been found, avoiding the waste of experimental resources caused by unreasonable termination conditions. Therefore, the data-driven optimization-based enzyme modification method provided by this invention not only achieves reliable enzyme molecule modification but also reduces the time cost and economic investment in enzyme molecule modification. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the main process of an enzyme modification method based on data-driven optimization provided in an embodiment of the present invention;

[0051] Figure 2 This is a graph showing the change of the encoding result of the maximum modified characteristic quantification value obtained in each iteration of the data-driven optimization enzyme modification in this invention with the number of iterations.

[0052] Figure 3 This is a graph showing how the size of the sample set worth evaluating in this invention changes with the number of iterations. Detailed Implementation

[0053] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solution of the present invention more clearly, and should not be used to limit the scope of protection of the present invention.

[0054] Example 1: A Data-Driven Optimization-Based Enzyme Modification Method

[0055] like Figure 1As shown, this invention discloses a data-driven optimization-based enzyme modification method. By introducing a batch Bayesian optimization algorithm to solve the combinatorial optimization problem of enzyme modification, it reduces the time and economic costs in the modification process and improves the reliability of the enzyme molecule modification method. Benefiting from the representational power of low-dimensional mutual information encoding, a surrogate model with a good sequence-function mapping relationship can be constructed. Benefiting from the global optimization capability of batch Bayesian optimization, the proposed enzyme modification method can continuously approach the globally optimal mutation sequence in the decision space through an iterative screening process, thereby achieving the modification of specific functions of the enzyme molecule. Benefiting from the termination condition of the batch Bayesian optimization algorithm based on the maximum variance change, the provided enzyme modification method can use a probabilistic method based on the current data to determine whether the globally optimal mutation sequence in the decision space has been found, avoiding the waste of experimental resources caused by unreasonable termination conditions.

[0056] The data-driven optimization-based enzyme modification method of the present invention specifically includes the following steps:

[0057] Step S1 involves using molecular docking and point mutagenesis techniques to identify several mutation hotspot residues and generating a mutation space for enzyme modification based on the amino acid sequence and the given maximum number of mutation sites.

[0058] The specific steps of step S1 are as follows:

[0059] Step S11: Based on the protein structure, the substrate binding pocket is determined by molecular docking of the enzyme and the substrate. Then, by site-directed mutagenesis of the amino acid residues around the substrate binding pocket and determination of the catalytic activity of the corresponding mutants, all mutation hotspot residues (i.e. beneficial mutation sites) related to catalytic activity are identified.

[0060] It should be noted that the sequences of the substrate and the unmutated enzyme are given in advance.

[0061] S12, based on the amino acid sequence and a given maximum number of mutation sites, generates the mutation space for enzyme modification through mutation hotspot residues. Where s i Let N represent the mutated sequence of the i-th enzyme in the mutation space, and let N represent the total number of mutated sequences of enzymes in the mutation space.

[0062] Wherein, the mutation space is the mutant sequence s of the enzyme. i The collection of enzyme mutant sequences s i It is obtained by mutating at least one hotspot residue in the enzyme sequence, and the mutated sequence s i The number of mutation sites is at most the maximum number of mutation sites mentioned above. Mutant sequence s i It is a character sequence representing an amino acid sequence.

[0063] Step S2 involves performing enzyme modification characteristic experiments on each site-specific saturated mutant sequence based on mutation hotspot residues to obtain a quantitative value of the modification characteristics. The sequence value of the site-specific saturated mutant sequence and its corresponding quantitative value are used as the initial training data for the enzyme modification experiment. Furthermore, the difference between the mutation space and the site-specific saturated mutant set (i.e., the set composed of site-specific saturated mutant sequences) is used as the decision space for the enzyme modification experiment.

[0064] The specific steps of step S2 are as follows:

[0065] Step S21: Perform enzyme modification characteristic experiments on each site-specific saturated mutant sequence based on mutation hotspot residues to obtain a quantitative value of the modification characteristics, and use the sequence value (denoted as s′) and its corresponding quantitative value (denoted as v) of the site-specific saturated mutant sequence as the initial training data for the enzyme modification experiment. Where s i ' represents the sequence value of the saturated mutation sequence at the i-th unit point in the mutation space, v i The quantized value representing the modification characteristics of the saturated mutation sequence at the i-th unit point in the mutation space;

[0066] A single-point saturation mutant sequence is obtained by mutating one hotspot residue in the enzyme's sequence. The sequence value of a single-point saturation mutant sequence is displayed as a character sequence representing the amino acid sequence, for example:

[0067] 'MAPTLSEQTRQLVRASVPALQKHSVAISATMYRLLFERYPETRSLFELPERQIHKLASALLAYARSIDNPSALQAAIRRMVLSHARAGVQAVHYPLVWECLRDAIKEVLGPDATETLLQAWKEAYDFLAHLLSTKEAQVYAVLAE'.

[0068] The purpose of conducting enzyme modification characteristic experiments is to obtain specific quantitative values ​​of the modification characteristics.

[0069] Step S22: Combine the mutation space S with the unit point saturated mutation set (i.e., the set consisting of unit point saturated mutation sequences). The difference between the sets of and is used as the decision space for enzyme modification experiments.

[0070] Decision space of enzyme modification experiments This serves as the search space for the Bayesian optimization algorithm in subsequent steps.

[0071] Among them, the single-point saturation mutation sequence is a special type of mutation sequence s. iIn other words, the unit-point saturated mutation set is contained within the mutation space S, and each sample point s′ in the unit-point saturated mutation set is also a sample point s in the mutation space S. i .

[0072] Step S3: Based on the initial training data, determine the hyperparameters of the low-dimensional mutual information encoding method (i.e., the encoding hyperparameter l and the reduced dimension d) using cross-validation. Then, according to the determined hyperparameters, apply the low-dimensional mutual information encoding method to the decision space. Encode the initial training data to obtain the encoded decision space. And the current training data.

[0073] Cross-validation is a common method for evaluating the generalization performance of machine learning models. It involves dividing the data into k folds, using each fold as test data for generalization performance, and using the remaining data as training data to calculate k generalization performance values. The average of these k values ​​is then used as the cross-validation result, and the optimal encoding hyperparameter *l* and dimensionality reduction *d* are selected based on this result. In this invention, the hyperparameters of the low-dimensional mutual information encoding method are determined using cross-validation, including: dividing the initial training data into k folds, using each fold as test data for generalization performance, and using the remaining data as training data to calculate the generalization performance; then using the average of the k generalization performance calculation results as the cross-validation result; and selecting the optimal encoding hyperparameter *l* and dimensionality reduction *d* based on this result. *l* is the hyperparameter required for constructing protein features; its larger value indicates a wider range of amino acid interactions considered in the constructed features, which will be used in step S32 below.

[0074] In step S3, the specific steps of the low-dimensional mutual information encoding method are as follows:

[0075] Step S31: Replace each amino acid sequence in the mutation space S one by one with the T-scale topological descriptor of the amino acid to obtain the descriptor matrix M for each amino acid sequence.

[0076] Step S32: Calculate the autocovariance matrix C for each amino acid sequence using the following formula:

[0077]

[0078] Where b represents the distance between amino acid residues, with values ​​{1,2,…,l}, l is the encoding hyperparameter, and C b,j With M i,jLet i and j represent the j-th element of the b-th row of the autocovariance matrix C and the j-th element of the i-th row of the descriptor matrix M, respectively. i represents the row position in the descriptor matrix M, j is the group of the descriptor matrix M, and m represents the number of rows in the descriptor matrix M, which is the length of the amino acid sequence.

[0079] Step S33: For each amino acid sequence in the mutation space, concatenate the columns of its autocovariance matrix C, and use principal component analysis to reduce the dimensionality of the concatenation result according to the dimensionality reduction factor d. The dimensionality reduction result is the encoding result of the mutation space S, which includes the encoded decision space. The encoding results of the sequence values ​​of the initial training data;

[0080] Here, the dimension reduction d is the dimension of the dimensionality reduction result. Each element in the decision space and the sequence value in each sample of the initial training data are elements in the mutation space. Therefore, the encoding results of the encoded decision space and the sequence values ​​of the initial training data can be found in the encoding results of the mutation space. The encoding result is the feature vector.

[0081] The concatenation of columns in the autocovariance matrix C means that the last element of the first column of the autocovariance matrix C is followed by the first element of the second column, and the last element of the second column is followed by the first element of the third column.

[0082] Step S34: Find the encoding result of the sequence value of the initial training data in the encoding result of the mutation space S and use it as the sample input x. i The sample output y is obtained by quantizing the modified characteristics of the initial training data. i Each input sample and each output sample is used as a training sample, thus forming the current training data.

[0083] The sample output y is obtained by quantizing the modified characteristics of the initial training data. i Specifically, this includes: optimizing the distribution of the quantized values ​​of the modification characteristics of the initial training data to obtain the optimized distribution result of the quantized values ​​of the modification characteristics of the initial training data, and using the optimized distribution result of the quantized values ​​of the modification characteristics of the initial training data as the sample output y. i .

[0084] The distribution of the quantized values ​​of the modified characteristics of the initial training data is optimized using the following formula:

[0085] y = aln(v)

[0086] In the formula, y represents the distribution optimization result of the quantized value, a∈[1,+∞) is the scaling factor of the sample output, the magnitude of which does not affect the construction of the surrogate model, but only changes the path of optimization solution, and v represents the quantized value of the modified characteristic. Among them, the quantized value of the modified characteristic is a numerical scalar, such as 1.2.

[0087] Training samples are derived from sample input x i 'With sample output y i Composition, i.e. (x i ',y i Each sample point in the current training data is obtained by transforming sample points in the initial training data, and therefore corresponds to a saturation mutation sequence of a single unit point.

[0088] Step S4: Based on the current training data, a Gaussian process regression model is constructed as a surrogate model using a Gaussian process. Then, based on the surrogate model, a batch Bayesian optimization algorithm based on maximizing the change in variance is used to select the verification object for this round of experiments from the encoded decision space. Subsequently, the verification object is decoded into an amino acid sequence and the modification characteristics of the amino acid sequence are experimentally verified to obtain the quantitative value of the modification characteristics of the verification object.

[0089] Therefore, the constructed Gaussian process regression model is used to reflect the mapping relationship between the encoding result and the quantized value of the modification characteristics. The verification object selected by the batch Bayesian optimization algorithm is the encoding result of a certain mutation sequence in the decision space. However, the experimental verification requires a specific mutation sequence. Therefore, the encoding result needs to be decoded into its mutation sequence. The decoding process transforms the encoding result into a mutation sequence according to the correspondence between the encoding before and after the mutation space.

[0090] Experimental verification of modified properties needs to be carried out in a chemical laboratory. Since enzyme modification is a real experimental process, computer algorithms are only used as a means to guide the experiment. Therefore, corresponding experimental verification needs to be carried out after obtaining the results based on the batch Bayesian optimization algorithm.

[0091] In step S4, a Gaussian process regression model is constructed based on the current training data as a surrogate model. Then, based on the surrogate model, a batch Bayesian optimization algorithm maximizing the change in variance is used to select validation objects for this round of experiments from the encoded decision space. Specifically, this includes:

[0092] Step S41: Construct and train a Gaussian process regression model based on the current training data, and obtain the kernel parameter θ of the covariance function and the noise variance from the trained Gaussian process regression model.

[0093] The construction and training of the Gaussian process regression model refers to sampling the optimal Gaussian process regression model based on the current training data and maximizing the likelihood function of the Gaussian process regression model. The kernel parameter θ of the covariance function is related to the noise variance. It can be obtained directly from a trained Gaussian process regression model.

[0094] Step S42, based on the type of problem to be solved, from the encoded decision space Divide the sample set into those worthy of evaluation. m is a sample set X worth evaluating. * The sample size, i, is the sample set X worth evaluating. * The sample ordinal number;

[0095] In this embodiment, the following formula is used from the encoded decision space Divide the sample set into those worthy of evaluation.

[0096]

[0097] Where x represents the encoded decision space. The sample in y max z represents the maximum value of the output of the sample in the current training data. α For the lower α quantile of the standard normal distribution, μ(x) and δ n (x) represent the predicted mean and predicted standard deviation of the Gaussian process regression model for sample x, respectively;

[0098] In this embodiment, α is set to 99.6%, and its corresponding z α =2.652070.

[0099] In other embodiments, in step S42, the sample set worthy of evaluation is divided using the following formula:

[0100] When solving a problem to find the sample corresponding to the maximum quantized value of a modified characteristic, the following formula is used from the decision space. The sample set worth evaluating is divided from the data:

[0101]

[0102] Where x represents the encoded decision space. The sample in y max z represents the maximum value of the output of the sample in the current training data. α For the lower α quantile of the standard normal distribution, μ(x) and δ n (x) represent the predicted mean and predicted standard deviation of the Gaussian process regression model for sample x, respectively;

[0103] When solving a problem to find the sample corresponding to the minimum quantized value of a modified characteristic, the following formula is used from the decision space. The sample set worth evaluating is divided from the data:

[0104]

[0105] In the formula, y min This represents the minimum value of the sample output in the current training data.

[0106] Step S43, if the sample set X is worth evaluating * If the sample size m is less than or equal to the batch size q, then the sample set X worth evaluating is... * The evaluation sample set obtained in this iteration (i.e., the set of objects to be verified in this round of experiments), and the evaluation sample set obtained in this round of iterations. All verification objects are decoded into amino acid sequences and the modification characteristics of the amino acid sequences are experimentally verified. After obtaining the quantification value of the modification characteristics, step S5 is executed so that the next iteration can be skipped (the next iteration is to start executing step S41 again).

[0107] The batch size q is determined based on external factors such as experimental conditions and optimization costs, and it is a positive integer.

[0108] Step S44, if the sample set X is worth evaluating * The sample size m exceeds 10,000, and samples are randomly selected from the sample set X that is worth evaluating. * Select 10,000 samples from the sample set to form a new sample set X worthy of evaluation. * Then, proceed to step S45;

[0109] Step S45, when the sample set X is worth evaluating * When the sample size m is less than or equal to 10000 and greater than the batch size q, the sample set X worth evaluating is selected. * We sequentially select q samples that maximize the following formula as the evaluation sample set obtained in this iteration. (That is, the verification object of this round of experiments), and it is removed from the sample set X that is worth evaluating. * Delete from the middle, then return to step S43:

[0110]

[0111] In the formula, It is the i-th evaluation sample selected. The sample set X is worth evaluating. *The j-th sample in the dataset, where n represents the number of samples in the current training data, and n+i represents the index of the newly selected ith evaluation sample, X n It is the set of sample inputs to the current training data, k i,j express and The covariance between them, k n+i,j This represents the union of the set of input samples in the current training data and the set of the first i samples in the evaluation sample set. With a sample set X worthy of evaluation * The j-th sample The covariance vector between them express With the selected (i+1)th evaluation sample The covariance vector between them, A n+1 For intermediate parameters, K n+i express The covariance matrix, It represents the noise variance, and I denotes the identity matrix. It is based on Calculate the output variance.

[0112] Step S5: Add all the verification objects of this round of experiments and their quantified values ​​of modification characteristics to the current training data to update the current training data. Update the encoded decision space by deleting the verification objects of this round of experiments. Then go to step S4 to jump to the next iteration. Stop iterating until the termination condition of the batch Bayesian optimization algorithm is met or the maximum number of iterations is reached. At this time, take the amino acid sequence corresponding to the optimal result of the quantified value of the modification characteristics of the verification object at this time as the final result of enzyme modification.

[0113] The termination condition for the batch Bayesian optimization algorithm is:

[0114] When the decision space is large and the objective function is complex, satisfying the following formula three times consecutively is used as the termination condition for the batch Bayesian optimization algorithm:

[0115] |X * ≤0.5%N

[0116] Where N represents the total number of samples in the mutation space, |X * | represents the sample set X that is worth evaluating. * The number of samples;

[0117] Generally, a decision space with more than tens of thousands of samples is considered relatively large. The standard for a complex objective function cannot be explicitly given; it is mainly judged based on the termination condition during the transformation process. After several iterations, |X *When the magnitude of change is very small, the objective function can be considered to be complex.

[0118] In other scenarios, |X * |=0 is used as the termination condition for the batch Bayesian optimization algorithm.

[0119] Experimental results:

[0120] This example, based on Example 1, uses the modification of human GB1 binding protein as an example. Four sites—V39, D40, G41, and V54—were selected as beneficial mutation sites, with a maximum of 4 mutation sites. Therefore, the mutation space for this modification experiment contains 160,000 mutation sequences, and there are 77 single-point saturation mutations at the beneficial mutation sites. For this dataset, this example uses grid search and cross-validation to determine the encoding hyperparameter l=29 and the dimensionality reduction d=8 in the low-dimensional mutual information encoding. Furthermore, since the quantization value of GB1 binding protein does not increase significantly during the modification process, this example sets 'a' in the low-dimensional mutual information encoding to 10. The 77 single-point saturation mutation sequences at the beneficial mutation sites are used as the initial training samples for the modification experiment. Each iteration uses a batch Bayesian optimization algorithm based on the maximum variance change to select 20 validation objects from the decision space for experimental verification of the modification characteristics. The maximum number of iterations is set to 40.

[0121] Figure 2 The horizontal axis represents the number of iterations, and the vertical axis represents the maximum quantization value obtained in each round of experiments. For example... Figure 2 As shown, when the iteration reached the 22nd iteration, the provided enzyme modification method found the globally optimal mutation sequence in the mutation space, with a quantized value of 21.7. At this point, the modification was experimentally validated on 517 samples. Due to the large decision space and complex objective function, this embodiment failed to reach the termination condition for enzyme modification in 40 iterations. At this point, researchers preferred to obtain a satisfactory solution within an acceptable cost. Figure 3 As can be seen, the number of samples in the evaluation set gradually decreases with each iteration, eventually leveling off. The enzyme modification method proposed in this invention allows for the determination of whether further experiments are worthwhile based on the number of samples in the evaluation set, thereby maximizing the modification benefits with limited costs.

[0122] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A data-driven optimization-based enzyme modification method, characterized in that, include: Step S1: Several mutation hotspot residues are identified using molecular docking and point mutagenesis techniques, and the mutation space for enzyme modification is generated based on the amino acid sequence and the given maximum number of mutation sites. Step S2: Perform enzyme modification characteristic experiments on each unit point saturated mutant sequence based on mutation hotspot residues to obtain the quantitative value of the modification characteristics, and use the unit point saturated mutant sequence and its corresponding quantitative value as the initial training data for the enzyme modification experiment; in addition, use the difference between the mutation space and the set formed by the unit point saturated mutant sequence as the decision space for the enzyme modification experiment. Step S3: Based on the initial training data, determine the hyperparameters of the low-dimensional mutual information encoding method using cross-validation. Then, according to the determined hyperparameters, apply the low-dimensional mutual information encoding method to the decision space. Encode the initial training data to obtain the encoded decision space. and current training data; Step S4: Based on the current training data, a Gaussian process regression model is constructed using a Gaussian process as a surrogate model. Based on the surrogate model, a batch Bayesian optimization algorithm based on maximizing the change in variance is used to select the verification object for this round of experiments from the encoded decision space. Then, the verification object is decoded into an amino acid sequence and the modification characteristics of the amino acid sequence are experimentally verified to obtain the quantitative value of the modification characteristics of the verification object. Step S5: Add all the verification objects of this round of experiments and their quantified values ​​of modification characteristics to the current training data to update the current training data. Update the encoded decision space by deleting the verification objects of this round of experiments, and then go to step S4. Until the termination condition of the batch Bayesian optimization algorithm is met or the maximum number of iterations is reached, take the amino acid sequence corresponding to the optimal result of the quantified value of the modification characteristics of the verification object at this time as the final result of enzyme modification.

2. The enzyme modification method based on data-driven optimization according to claim 1, characterized in that, Step S1 specifically includes: Step S11: Based on the protein structure, the substrate binding pocket is determined by molecular docking of the enzyme and the substrate. Then, all mutation hotspot residues related to catalytic activity are identified by site-directed mutagenesis of amino acid residues around the substrate binding pocket and determination of the catalytic activity of the corresponding mutants. Step S12: Based on the amino acid sequence and the given maximum number of mutation sites, a mutation space for enzyme modification is generated through mutation hotspot residues. The mutation space is the mutation sequence s of the enzyme. i The collection of enzyme mutant sequences s i It is obtained by mutating at least one hotspot residue in the enzyme sequence, and the mutated sequence s i The number of mutation sites is at most the maximum number of mutation sites mentioned above.

3. The enzyme modification method based on data-driven optimization according to claim 1, characterized in that, Step S2 specifically includes: Step S21: Perform enzyme modification characteristic experiments on each site-saturated mutant sequence based on mutation hotspot residues to obtain the quantitative value of the modification characteristics, and use the sequence value of the site-saturated mutant sequence and its corresponding quantitative value as the initial training data for the enzyme modification experiment. Step S22: The difference between the mutation space and the set of unit-point saturated mutation sequences is used as the decision space for enzyme modification.

4. The enzyme modification method based on data-driven optimization according to claim 1, characterized in that, In step S3, the low-dimensional mutual information encoding method includes: Step S31: Replace each amino acid sequence in the mutation space S one by one with the amino acid T-scale topological descriptor to obtain the descriptor matrix M for each amino acid sequence. Step S32: Calculate the autocovariance matrix C for each amino acid sequence using the following formula: Where b represents the distance between amino acid residues, and its value is {1,2,…,l}, C b,j With M i,j Let i and j represent the j-th element of the b-th row of the autocovariance matrix C and the j-th element of the i-th row of the descriptor matrix M, respectively. i represents the row position in the descriptor matrix M, j is the group of the descriptor matrix M, and m represents the number of rows in the descriptor matrix M, which is the length of the amino acid sequence. Step S33: For each amino acid sequence in the mutation space, concatenate the columns of its autocovariance matrix C, and use principal component analysis to reduce the dimensionality of the concatenation result according to the dimensionality reduction d. The dimensionality reduction result is the encoding result of the mutation space S, which includes the encoded decision space. The encoding results of the sequence values ​​of the initial training data; Step S34: Find the encoding result of the sequence value of the initial training data in the encoding result of the mutation space S and use it as the sample input x. i The sample output y is obtained by quantizing the modified characteristics of the initial training data. i Each input sample and each output sample is used as a training sample, thus forming the current training data.

5. The enzyme modification method based on data-driven optimization according to claim 4, characterized in that, The sample output y is obtained by quantizing the modified characteristics of the initial training data. i Specifically, this includes: optimizing the distribution of the quantized values ​​of the modification characteristics of the initial training data to obtain the optimized distribution result of the quantized values ​​of the modification characteristics of the initial training data, and using the optimized distribution result of the quantized values ​​of the modification characteristics of the initial training data as the sample output y. i .

6. The enzyme modification method based on data-driven optimization according to claim 5, characterized in that, The distribution of the quantized values ​​of the modified characteristics of the initial training data is optimized using the following formula: y = aln(v) In the formula, y represents the distribution optimization result of the quantized value, a∈[1,+∞) is the scaling factor of the sample output, and v represents the quantized value of the modified characteristics.

7. The enzyme modification method based on data-driven optimization according to claim 1, characterized in that, In step S4, a Gaussian process regression model is constructed based on the current training data as a surrogate model. Then, based on the surrogate model, a batch Bayesian optimization algorithm maximizing the change in variance is used to select validation objects for this round of experiments from the encoded decision space. Specifically, this includes: Step S41: Construct and train a Gaussian process regression model based on the current training data, and obtain the kernel parameter θ of the covariance function and the noise variance from the trained Gaussian process regression model. Step S42, based on the type of problem to be solved, from the encoded decision space The sample set worth evaluating is divided from the sample set. Step S43, if the sample set X is worth evaluating * If the sample size m is less than or equal to the batch size q, then the sample set X worth evaluating is... * The evaluation sample set obtained in this iteration The evaluation sample set obtained in this iteration All verification objects are decoded into amino acid sequences and the modification characteristics of the amino acid sequences are experimentally verified. After obtaining the quantification value of the modification characteristics, step S5 is executed.

8. The enzyme modification method based on data-driven optimization according to claim 7, characterized in that, In step S42, when solving the problem for the sample corresponding to the maximum value of the quantized value of the modified characteristic, the following formula is used from the decision space. The sample set worth evaluating is divided from the data: Where x represents the encoded decision space. The sample in y max z represents the maximum value of the output of the sample in the current training data. α For the lower α quantile of the standard normal distribution, μ(x) and δ n (x) represent the predicted mean and predicted standard deviation of the Gaussian process regression model for sample x, respectively; When solving a problem to find the sample corresponding to the minimum quantized value of a modified characteristic, the following formula is used from the decision space. The sample set worth evaluating is divided from the data: In the formula, y min This represents the minimum value of the sample output in the current training data.

9. The enzyme modification method based on data-driven optimization according to claim 7, characterized in that, Step S4 further includes: Step S44, when the sample set X is worth evaluating * The sample size m exceeds 10,000, and samples are randomly selected from the sample set X that is worth evaluating. * Select 10,000 samples from the sample set to form a new sample set X worthy of evaluation. * ; Step S45, when the sample set X is worth evaluating * When the sample size m is less than or equal to 10000 and greater than the batch size q, the sample set X worth evaluating is selected. * We sequentially select q samples that maximize the following formula as the evaluation sample set obtained in this iteration, and remove them from the evaluation sample set X. * Delete from the middle, then return to step S43: In the formula, It is the i-th evaluation sample selected. The sample set X is worth evaluating. * The j-th sample in the dataset, where n represents the number of samples in the current training data, and n+i represents the index of the newly selected ith evaluation sample, X n It is the set of sample inputs to the current training data, k i,j express and The covariance between them, k n+i,j This represents the union of the set of input samples in the current training data and the set of the first i samples in the evaluation sample set. and The covariance vector between them express With the selected (i+1)th evaluation sample The covariance vector between them, A n+1 For intermediate parameters, K n+i express The covariance matrix, It represents the noise variance, and I denotes the identity matrix. It is based on Calculate the output variance.

10. The enzyme modification method based on data-driven optimization according to claim 7, characterized in that, In step S5, the termination condition of the batch Bayesian optimization algorithm is: When the decision space is large and the objective function is complex, satisfying the following formula three times consecutively is used as the termination condition for the batch Bayesian optimization algorithm: |X * |≤0.5%N Where N represents the total number of samples in the mutation space, |X * | represents the sample set X that is worth evaluating. * The number of samples; In other scenarios, |X * |=0 is used as the termination condition for the batch Bayesian optimization algorithm.

Citation Information

Patent Citations

  • Affinity modification system and method for antibody / macromolecular drug

    CN115171774A