Protein mutation optimization method based on protein language model and sorting loss

By constructing a saturated mutation dataset and introducing the extreme gradient boosting algorithm with the Bradley-Terry ranking loss function, the problems of sample scarcity and insufficient screening accuracy in protein mutation screening are solved, and efficient and accurate mutant screening is achieved, which is suitable for functional optimization of various protein families.

CN120748489AActive Publication Date: 2025-10-03SHEN ZHEN SHI AI DI BEI KE SHENG WU YI YAO YOU XIAN GONG SI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511180061.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-10-03
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing machine learning algorithms face the problems of scarce sample data and insufficient screening accuracy in protein mutation screening, resulting in low screening efficiency and extended R&D cycles. In particular, it is difficult to accurately capture the relative order between mutants in complex protein optimization tasks.

Method used

By constructing a saturated mutation dataset, combining the protein language model and the Bradley-Terry ranking loss function, the extreme gradient boosting algorithm is used to optimize the screening process, and a multi-round screening mode is adopted to quickly lock in highly active mutants and optimize the model's ability to sort relative order.

Benefits of technology

It significantly improves the accuracy and efficiency of protein mutation screening, reduces the number of screening rounds, and reduces experimental costs and time expenditures. It is suitable for different types of protein families and function optimization tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748489A_ABST
    Figure CN120748489A_ABST
Patent Text Reader

Abstract

The invention discloses a protein mutation optimization method based on a protein language model and sorting loss. The method comprises the following steps: constructing a saturation mutation data set; initializing parameters; generating a high-dimensional embedding expression of the protein sequence by using the protein embedding model; in the first round of screening, saturated mutation is selected for screening according to a mutation selection strategy; in each later round of screening, a Bradley-Terry sorting loss function is introduced into a limit gradient lifting algorithm, and mutation with the highest predicted activity value is selected for screening; checking the number of screening rounds to determine whether to finish or not, if the maximum number of screening rounds is not reached, returning to carry out the next round of screening, otherwise, finishing the screening process; obtaining a high-activity mutation set; the method overcomes the problem of insufficient samples caused by high experimental cost and time-consuming sample generation in the traditional method, and supports multi-objective optimization, such as improvement of activity, stability and expression level at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a protein mutation optimization method, in particular to a protein mutation optimization method based on a protein language model and ranking loss. Background Art

[0002] In protein engineering, protein mutation screening is a crucial step in optimizing protein function. Protein language models, pre-trained on massive amounts of protein sequences, encode sequences into high-dimensional embedding vectors that embody functional properties, supporting mutation screening. Existing machine learning algorithms leverage these embeddings to train functional prediction models, establishing a mapping between sequence embeddings and protein function. These algorithms have demonstrated strong potential for applications such as predicting the affinity of mutants in antibody optimization and identifying mutations with enhanced catalytic efficiency in enzyme engineering.

[0003] However, these algorithms also face several limitations in practical applications. First, the scarcity of sample data is a major challenge. Current machine learning algorithms rely heavily on large amounts of experimental data for training. However, obtaining this data requires costly manual experimentation, and the sample generation process is time-consuming, making it difficult to obtain sufficient sample data. This lack of sample data directly impacts the efficiency of screening algorithms, making it difficult to quickly identify highly active mutants. Especially in complex protein optimization tasks, the lack of sample data often requires multiple rounds of iteration to obtain valid information, thus extending the R&D cycle.

[0004] Secondly, in terms of screening accuracy and efficiency, existing machine learning algorithms generally use mean squared error and cross-entropy loss functions. These loss functions primarily focus on the absolute error between the predicted and true values, ignoring the importance of the relative order between mutants. When the functional differences between mutants are subtle or the functional landscape is complex, these loss functions can lead to underestimation or omission of highly active mutants. This means that the algorithm may not accurately capture the relative order between mutants, leading to incorrect evaluation of some mutants with high potential activity, thereby reducing screening accuracy. Summary of the Invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a protein mutation optimization method based on protein language model and ranking loss to optimize the sample scarcity problem, improve screening efficiency and reduce the number of screening rounds.

[0006] Technical solution: The protein mutation optimization method of the present invention comprises the following steps: S1. Construct a saturation mutation dataset; S2, initialization parameters, including setting the experimental dataset, screening rounds, the number of mutations evaluated in each round, and the mutation selection strategy; S3. Generate high-dimensional embedding representation of protein sequences in saturation mutation dataset using protein language model; S4, in the first round of screening, saturation mutations were selected for screening according to the initial strategy; S5. In each subsequent round of screening, the Bradley-Terry ranking loss function is introduced into the extreme gradient boosting algorithm to select the mutation with the highest predicted activity value for screening; S6. Check the number of screening rounds to determine whether it is finished. If the maximum number of screening rounds has not been reached, return to the next round of screening; otherwise, end the screening process; S7. Obtain a set of highly active mutations.

[0007] Preferably, the saturation mutation dataset described in S1 processes some sites of the protein sequence through continuous and discrete saturation mutation methods to obtain corresponding saturation mutation samples, and measures their activity value labels through experiments. All saturation mutation samples constitute a saturation mutation dataset.

[0008] Preferably, the mutation selection strategy in S2 includes an initial strategy and a selection strategy; the initial strategy serves as the mutation selection strategy type for the first round of screening, and the selection strategy serves as the mutation selection strategy type for subsequent rounds of screening.

[0009] Preferably, the initial strategy is manual selection, and the selection strategy is a top-n strategy.

[0010] Preferably, the high-dimensional embedding representation in S3 is implemented by the following formula: ; in, is the encoding function of the protein language model, Indicates that each element contains a mutation sequence, Indicates the corresponding embedding of the mutation sequence.

[0011] Preferably, S4 includes: randomly selecting M saturation mutations from the saturation mutation dataset using the initial strategy, directly using the activity value label measured in advance as the true activity value for the i-th selected saturation mutation, updating the experimental dataset, and using the updated experimental dataset for the next round of screening.

[0012] Preferably, the goal of the Bradley-Terry ranking loss function described in S5 is to minimize the ranking error, by increasing the difference between the predicted values ​​of the mutants so that the probability of correct ranking is close to 1, and then guiding the extreme gradient boosting algorithm to construct a decision tree and calculate the weight of the leaf node through the first-order derivative and the second-order derivative, determine the node splitting method, and optimize the relative order.

[0013] Preferably, let , the first-order derivative and second-order derivative formulas are as follows: ; ; in, 、 represents the first-order derivative, reflecting the correctness of the pairwise sorting, 、 represents the second-order derivative, and Follow changes, reflecting the dynamic curvature of the sorting task, Indicates mutant The predicted value of Indicates mutant The predicted value of .

[0014] Preferably, the extreme gradient boosting algorithm uses the first-order derivative and the second-order derivative to calculate the gain Gain to evaluate the contribution of the node split to the reduction of the negative log-likelihood.

[0015] Through multiple rounds of screening iterations in step S6, the final set of highly active mutations is obtained, that is, the top M mutations with the largest activity values ​​selected in the final round.

[0016] Beneficial effects: Compared with the existing technology, the present invention has the following significant advantages: 1. High-dimensional sequence embedding is generated by the protein language model to capture the functional characteristics of the sequence, providing rich feature representation for functional prediction. Combined with the multi-round screening mode, high-activity mutations can be quickly locked with a small amount of experimental data, thereby alleviating the problem of sample scarcity; 2. By introducing the Bradley-Terry ranking loss function to optimize the protein language model, the optimized model can more accurately rank mutations, thereby improving the screening accuracy of high-activity mutations, and the efficiency of the screening task is also significantly accelerated; 3. By introducing protein saturated mutation samples, the model can learn more diverse data, thereby improving its performance in actual screening tasks and reducing the number of multiple rounds of screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 Schematic diagram for comparing the performance of the present invention. DETAILED DESCRIPTION

[0018] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0019] like Figure 1 As shown, the protein mutation optimization method based on protein language model and ranking loss includes the following steps: S1. Construct a saturation mutation dataset. The specific steps are as follows: Given a wild-type protein sequence and mutation datasets , mutation dataset , where each mutation yes For example, if there is a protein sequence with a length of 100, and each point can mutate to 20 amino acids, then its mutation dataset have , using continuous saturation mutation and discrete saturation mutation methods to process some points, obtain the corresponding saturation mutation samples, and measure their activity value labels through experiments. The dataset composed of these saturation mutation samples is called saturation mutation dataset , expressed as .

[0020] Continuous saturation mutagenesis involves selecting a continuous range of sites (for example, sites 301-400 of AsCas12f and sites 161-200 of β-lactamase) and processing them to generate the corresponding saturation mutation samples. Given the strong reliance of current multi-round iterative methods on random sampling, discrete mutation samples were used to select randomly distributed mutation sites (for example, experimental verification was conducted on AsCas12f and β-lactamase datasets, randomly selecting 10% of the saturation mutation sites in the amino acid sequence).

[0021] The construction of a saturation mutation dataset not only enhances the model's ability to explore the protein functional landscape but also provides a more reliable data foundation for subsequent multiple iterations. By systematically integrating saturation mutation samples, the present invention can more efficiently identify highly active mutations in the early stages of screening, thereby reducing experimental costs and time. Based on the current multi-round iterative model framework, the optimization effect of saturation mutation samples was verified on two protein datasets, AsCas12f and β-lactamase. The AsCas12f dataset contains 7,941 mutation samples, 18% (1,436 of which are highly active); the β-lactamase dataset contains 4,978 mutation samples, 7.9% (393 of which are highly active). For example, in the β-lactamase dataset, comparing the extreme gradient boosting algorithm with ranking loss and protein saturation mutation samples with the extreme gradient boosting algorithm with ranking loss, the application of protein saturation mutation samples reduced the number of screening rounds from 4.4 to 3.8, fully demonstrating its important value in optimizing data diversity. More importantly, this strategy has wide applicability and can seamlessly adapt to different types of protein families and functional optimization tasks, providing a new technical path for the intelligent development of protein engineering.

[0022] High activity is typically quantified using the activity value of the wild-type protein as a standard. If the activity value of the mutant is greater than that of the wild-type protein, it is defined as high activity. The protein datasets used in this invention all use the activity value of the wild-type protein as the standard for high activity.

[0023] S2. Initialize the parameters, including setting the screening rounds, the number of mutations to be evaluated in each round, the experimental dataset, and the mutation selection strategy. The specific steps are as follows: Setting up the experimental dataset , set the screening round R=10, the number of mutations evaluated in each round M=16, where the experimental dataset D is the dataset constructed for each experiment, and each element in D contains the mutation sequence , the true activity value , the mutation sequence corresponds to the embedded .

[0024] Here, R refers to the number of iterations set in the experiment. The choice of M is usually determined by the actual scenario needs. For example, given a protein, the number of mutations that need to be screened is the final number. Here, M = 16 is a value obtained after statistically analyzing a large number of experimental results and considering the experimental costs in real-world scenarios. The reference range is {8, 16, 20}.

[0025] The protein language model can be freely selected according to the specific situation. In this embodiment, the protein language model selected is the ESM-2 model (15 billion parameters). As the first round of mutation selection strategy type, select strategy S as the subsequent round of mutation selection strategy type; initial strategy The random selection method is used, that is, the first round of screening is Randomly select M saturation mutations for subsequent screening, and use the top n strategy, that is, select the top n mutations with the highest activity values ​​in the prediction results, and set n=M=16. For manual selection, the experimental dataset D can be set to a non-empty set and the selected mutations can be added.

[0026] S3. Generate a high-dimensional embedding representation of the protein sequence of the saturation mutation dataset using the protein language model. The specific steps are as follows: For the saturation mutation set , calculate the embedding set , where each mutation The embedding of is generated by the following formula: ; in, is the encoding function of the protein language model, Indicates that each element contains a mutation sequence, Indicates the corresponding embedding of the mutation sequence. The experimental protein language model uses the ESM-2 model (15 billion parameters) to generate a 5120-dimensional protein sequence embedding, and finally obtains the embedding matrix , providing feature input for subsequent model training.

[0027] In the first round of screening (round 1 screening), the initial strategy is used ,from Randomly select M saturation mutations in the saturation mutation, and directly use the activity value label measured in advance as the true activity value for the i-th saturation mutation selected , update the experimental dataset D and add , use the updated experimental dataset D for the next round of screening.

[0028] S5. In each subsequent round of screening (1-R rounds of screening), the Bradley-Terry ranking loss function is introduced into the extreme gradient boosting algorithm (XGBoost) to select the mutation with the highest predicted activity value for screening. The specific steps are as follows: The Bradley-Terry ranking loss is a commonly used ranking loss function that effectively captures ranking relationships in pairwise comparisons and is particularly suitable for multi-candidate ranking tasks. The default loss function used in the Random Forest algorithm and the Extreme Gradient Boosting algorithm is the Mean Squared Error (MSE) loss function. This is a pointwise loss widely used in regression tasks. It measures the model's prediction error by calculating the average of the squared differences between the predicted and true values. The following uses the Extreme Gradient Boosting algorithm as an example to describe the workflow before and after the Bradley-Terry ranking loss function is introduced into the algorithm.

[0029] The mean square error loss function MSE is defined as: ; Where n is the number of samples, is the true target value of sample i, is the predicted value of sample i.

[0030] The extreme gradient boosting algorithm minimizes the loss function through gradient optimization. It needs to provide the first-order gradient and the second-order gradient to indicate the steepest descent direction and the approximate curvature to accelerate convergence. , the first-order and second-order gradients are as follows: First-order gradient: ; The gradient reflects the difference between the predicted value and the true value. If the gradient is positive ( ), indicating that the predicted value is too high and the model should be reduced , if the gradient is negative ( ), indicating that the predicted value is too low and the model should be enlarged .

[0031] Second-order gradient (Hessian): ; The Hessian is a constant that indicates that the MSE is a quadratic function with a fixed curvature. The constant Hessian simplifies the optimization process because it does not depend on the predicted or true values.

[0032] After introducing the Bradley-Terry ranking loss function: Given a wild-type protein sequence and mutation datasets. The mutation set is , where each mutation yes A single point mutation of . The embedding set calculated for the mutation set , predicted using the extreme gradient boosting algorithm .

[0033] Based on the Bradley-Terry ranking loss function, assuming that the mutant Mutants The predicted value of Before, the probability of correct sorting is: ; Defined as the negative log-likelihood, for all positive sequence pairs ( Should be ranked forward): ; in, and Indicates mutant and mutants The label value of .

[0034] The goal of the Bradley-Terry ranking loss function is to minimize the ranking error by increasing , correct sorting probability Close to 1. Different from the absolute error of regression tasks, Bradley-Terry ranking loss focuses on relative order and is suitable for scenarios that require priority ranking.

[0035] The first-order derivative (gradient) and second-order derivative (Hessian) are used to guide the extreme gradient boosting algorithm to build a decision tree and determine the node splitting method. , The first and second derivative formulas are as follows: ; ; in, 、 represents the first-order derivative, reflecting the correctness of the pairwise sorting. If , that is, the sorting error, Negative, indicating an increase ; Positive, prompt decrease ; 、 represents the second-order derivative, and Follow Changes reflect the dynamic curvature of the sorting task, and smaller values ​​reflect the dynamic curvature of the sorting task. Indicates mutant The predicted value of Indicates mutant The predicted value of .

[0036] The extreme gradient boosting algorithm uses these derivatives to calculate the gain Gain and evaluate the node split pair The reduced contribution is given by the following formula: .

[0037] Splitting with a large Gain value makes the ordering of samples in the left and right child nodes more consistent, for example, concentrating high-priority samples on one side. The tree structure is shaped by derivatives to optimize the relative order.

[0038] After the tree is built, The derivative of is used to calculate the leaf node weight and update the prediction score of the mutation , the weight of leaf node j is: ; The weights reflect the need to adjust the order of samples within the leaf nodes, and the dynamic nature of the Hessian adapts the calculation to the pairwise relationship. The predicted value of sample i is updated as: ; in, Is the learning rate, which controls the update amplitude. make Increase, improve ; The score difference is optimized by weighting to ensure that high-priority samples receive higher scores.

[0039] The improved algorithm f is trained on the experimental dataset D, that is, the extreme gradient boosting algorithm with the Bradley-Terry ranking loss function is used to train the mutation dataset so that f can predict the activity values ​​of all mutations. : ;in, .

[0040] for Use the selection strategy S to select the top M mutations with the largest activity values. As in the first round of screening, for the selected i-th mutation, if the mutation belongs to the saturation mutation dataset, the pre-determined activity value label is directly used as the true activity value If the mutation belongs to the non-saturated mutation dataset, its true activity value is determined experimentally. , and finally update the experimental dataset D, adding . Use only f to predict the activity values ​​of all saturation mutations .

[0041] like Figure 2 As shown, the average rounds required for the four methods to achieve 100% high activity at different dataset sites are shown: extreme gradient boosting algorithm + ranking loss + protein saturation mutation samples, random forest algorithm + protein saturation mutation samples, extreme gradient boosting algorithm + ranking loss, and random forest algorithm; including AsCas12f dataset 351-400 sites, β-lactamase dataset 247-287 sites, AsCas12f dataset random 10% sites, and β-lactamase dataset random 10% sites.

[0042] As can be seen, the introduction of the Bradley-Terry ranking loss significantly improves the model's sensitivity to subtle activity differences, enabling more accurate identification of highly active mutations, particularly in tasks with complex activity landscapes. For example, in the β-lactamase dataset, a comparison was made between the extreme gradient boosting algorithm combined with ranking loss and protein-saturated mutation samples and the random forest algorithm combined with protein-saturated mutation samples (combining mean squared error (MSE) and random forest algorithm with protein-saturated mutation samples). The results showed that the improved extreme gradient boosting algorithm reduced the number of screening rounds from 7.3 to 4.1, fully demonstrating its advantages in ranking accuracy and efficiency. Furthermore, this method achieves efficient ranking by relying on embeddings generated by a protein language model, rather than complex protein structural information. It is highly versatile and scalable, offering new technical possibilities for the field of protein engineering and laying a solid foundation for future multi-objective optimization tasks.

[0043] S6. Check the number of screening rounds to determine whether it is finished. If the maximum number of screening rounds has not been reached, return to the next round of screening; otherwise, end the screening process.

[0044] Set the maximum number of screening rounds R, and determine whether the number of screening rounds is reached, that is, determine the relationship between r and R: when r is less than R, r is incremented by 1 and the process returns to the previous step; otherwise, the screening process ends.

[0045] S7. Through multiple rounds of screening iterations in the above steps, the final set of highly active mutations is obtained, i.e., the top M mutations with the largest activity values ​​selected in the final round.

Claims

1. A protein mutation optimization method based on protein language model and ranking loss, characterized in that: The following steps are involved: S1. Construct a saturation mutation dataset; S2, initialization parameters, including setting the experimental dataset, screening rounds, the number of mutations evaluated in each round, and the mutation selection strategy; S3. Generate high-dimensional embedding representation of protein sequences using protein language models; S4, in the first round of screening, saturation mutations were selected for screening according to the initial strategy; S5. In each subsequent round of screening, the Bradley-Terry ranking loss function is introduced into the extreme gradient boosting algorithm to select the mutation with the highest predicted activity value for screening; S6. Check the number of screening rounds to determine whether it is finished. If the maximum number of screening rounds has not been reached, return to the next round of screening; otherwise, end the screening process; S7. Obtain a set of highly active mutations.

2. The protein mutation optimization method according to claim 1, characterized in that The saturation mutation dataset described in S1 processes some sites of the protein sequence through continuous and discrete saturation mutation methods to obtain corresponding saturation mutation samples, and measures their activity value labels through experiments. All saturation mutation samples constitute a saturation mutation dataset.

3. The protein mutation optimization method according to claim 1, characterized in that: The mutation selection strategy in S2 includes an initial strategy and a selection strategy; the initial strategy serves as the mutation selection strategy type for the first round of screening, and the selection strategy serves as the mutation selection strategy type for subsequent rounds of screening.

4. The protein mutation optimization method according to claim 3, characterized in that: The initial strategy is manual selection, and the selection strategy is the top-n strategy, which selects the top n mutations with the highest activity values ​​in the prediction results.

5. The protein mutation optimization method according to claim 1, characterized in that: The high-dimensional embedding representation described in S3 is achieved through the following formula: ; in, is the encoding function of the protein language model, Indicates that each element contains a mutation sequence, Indicates the corresponding embedding of the mutation sequence.

6. The protein mutation optimization method according to claim 1, characterized in that: The S4 includes: randomly selecting M saturation mutations from the saturation mutation dataset using the initial strategy, directly using the activity value label measured in advance as the true activity value for the i-th selected saturation mutation, updating the experimental dataset, and using the updated experimental dataset for the next round of screening.

7. The protein mutation optimization method according to claim 1, characterized in that: The goal of the Bradley-Terry ranking loss function described in S5 is to minimize the ranking error by increasing the difference between the predicted values ​​of the mutants so that the probability of correct ranking is close to 1. The first-order and second-order derivatives are then used to guide the extreme gradient boosting algorithm to construct a decision tree and calculate the weights of leaf nodes, determine the node splitting method, and optimize the relative order.

8. The protein mutation optimization method according to claim 7, characterized in that: make , the first-order derivative and second-order derivative formulas are as follows: ; ; in, 、 represents the first-order derivative, reflecting the correctness of the pairwise sorting, 、 represents the second-order derivative, and Follow changes, reflecting the dynamic curvature of the sorting task, Indicates mutant The predicted value of Indicates mutant The predicted value of .

9. The protein mutation optimization method according to claim 8, characterized in that: The extreme gradient boosting algorithm uses the first-order and second-order derivatives to calculate the gain Gain, which evaluates the contribution of node splitting to the reduction of negative log-likelihood.

10. The protein mutation optimization method according to claim 1, characterized in that: Through multiple rounds of screening iterations in step S6, the final set of highly active mutations is obtained, that is, the top M mutations with the largest activity values ​​selected in the final round.

Citation Information

Patent Citations

  • Mutant protein screening method and device and related equipment

    CN119296640A

  • Method and device for predicting protein mutation effect

    CN119580837A

  • Large language model driven data augmentation for protein machine learning

    US20250218545A1

  • Using protein large language models to improve protein activity

    WO2025080941A1