Signal peptide-protein combination secretion efficiency prediction method and system based on language model

By using a language model-based approach, signal peptide-protein sequence features are extracted and feature vectors are extracted using a protein language model. Combining one-dimensional convolution and feature concatenation, the problem of predicting the secretion efficiency of signal peptide-protein combinations is solved, thereby improving prediction accuracy and protein expression levels.

CN116631513BActive Publication Date: 2026-01-06SHENZHEN TAILI BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310611713.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-26
Publication Date
2026-01-06
Estimated Expiration
2043-05-26

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict the secretion efficiency of signal peptide and target protein combinations, resulting in low protein yield, poor stability, and uneven extraction of signal peptide and protein sequence features, which affects prediction performance.

Method used

By using a language model-based approach, the feature sequences of the signal peptide-protein sequence are extracted, and amino acid and sequence feature vectors are extracted using a protein language model. Combined with one-dimensional convolution and feature concatenation, the data are input into a prediction model for classification or regression prediction, thus balancing the information of the signal peptide and protein sequence.

Benefits of technology

It improves the prediction accuracy of signal peptide-protein combination secretion efficiency, avoids information masking and structural feature loss caused by inconsistent sequence lengths, and achieves more efficient protein expression level prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631513B_ABST
    Figure CN116631513B_ABST
Patent Text Reader

Abstract

The application discloses a signal peptide-protein combination secretion efficiency prediction method and system based on a language model. The method comprises the following steps: (1) dividing a signal peptide-protein sequence to be predicted into translation units, and taking the first M amino acid sequences at the N terminal of each translation unit as signal peptide characteristic sequences; (2) inputting the signal peptide characteristic sequences into a pre-trained protein language model to obtain amino acid residue characteristic vectors and / or protein sequence characteristic vectors; (3) splicing the amino acid residue characteristic vectors and / or the protein sequence characteristic vectors to obtain a secretion characteristic vector of the signal peptide-protein sequence; and (4) inputting the secretion characteristic vector of the signal peptide-protein sequence into a prediction model to predict the secretion efficiency grade of the signal peptide-protein sequence. The application improves the prediction accuracy of the prediction model for the secretion efficiency of the signal peptide-protein combination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics, and more specifically, relates to a method and system for predicting the secretion efficiency of signal peptide-protein combinations based on a language model. Background Technology

[0002] Many antibody drugs or therapeutic proteins face problems such as low yield, poor stability, and low activity during research and production. Increasing protein yield is a pressing issue for the industry. Currently, in the large-scale preparation of antibody drugs, animal cell expression systems are often used to produce secretory proteins. Signal peptides, as an amino acid sequence linked to the N-terminus of secretory proteins, can control the protein secretion pathway and thus affect the expression yield of antibody proteins. Therefore, the efficient expression of recombinant proteins is closely related to signal peptides. Many prokaryotic and eukaryotic signal peptides are functionally interchangeable even between different species, and natural signal peptides are not necessarily the most effective. Integrating signal peptides from different species with target proteins can also mediate increased antibody secretion in CHO cells. Based on these findings, we can screen different signal peptides from different species and fuse them into the proteins to be expressed, thereby increasing the final expression level of secretory proteins.

[0003] Lars et al. fused therapeutic antibodies and FC fusion proteins with 16 different signal peptides and analyzed secretion efficiency through transient and stable transduction. Compared with control signal peptides, signal peptides from multiple species, even the natural immunoglobulin G signal peptide, could not achieve higher efficiency. The natural signal peptides using human albumin and human penicillin yielded the best results. Ryan et al. clustered databases of immunoglobulin heavy and light chain signal peptides based on sequence similarity, ultimately classifying heavy chain signal peptides into 8 categories and light chain signal peptides into 2 categories. They then fused these into the five best-selling therapeutic antibodies to analyze the effect of the signal peptides on expression levels. The optimized signal peptides, compared to the original signal peptides, could increase Rituxan production by two times. The high cost of interleukin-21 production limits its application. Hee Jun Cho et al. increased production by 10-fold by optimizing the IL-21 codon. To further improve production, they searched the literature for five signal peptides and evaluated their effects on IL-21 expression levels. The human penicillin signal peptide was found to increase production by another three-fold. Wei-Li Ling et al. swapped different backbones between Trastuzumab and Pertuzumab and systematically compared the effects of myeloma and natural signal peptides on yield in 168 antibody variants. In most cases, the myeloma signal peptide yielded higher yields than the natural signal peptide. Furthermore, they established a logistic regression model based on the amino acid sequence of the signal peptide and target protein to predict yield, marking the first attempt to predict yield using the amino acid sequence properties of the signal peptide and protein. Stefano et al. manually extracted 156 features of the signal peptide and performed high-throughput analysis on the expression levels of 11,643 signal peptides fused with AmyQ (an α-amylase from Bacillus amyloliquefaciens) in Bacillus subtilis proteins. The manually extracted signal peptide features were then used to predict protein yield using a random forest model.

[0004] Despite the widespread importance of signal peptides in industry and the years of research they have received, finding the most suitable signal peptide for a target protein often involves repeated experiments. Previous studies have investigated the relationship between signal peptide binding yield and target protein production using amino acid count characteristics, and have also studied the relationship between signal peptide characteristics and production yield by immobilizing a single protein in a prokaryotic system. However, these studies often involve manual extraction of characteristics, and the binding of the signal peptide to the target protein is often too singular. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a method, electronic device, or non-transitory computer-readable storage medium for predicting the secretion efficiency of signal peptide-protein combinations based on a language model. The purpose is to balance the characteristics of the signal peptide and protein sequences by extracting the sequence length of the signal peptide-protein sequence as the signal peptide characteristic sequence, and to enrich the structural features of the extracted signal peptide characteristic sequence through a protein language model, thereby solving the technical problem of accurately predicting the secretion efficiency of signal peptide-protein combinations in existing technologies.

[0006] To achieve the above objectives, according to one aspect of the present invention, a method for predicting the secretion efficiency of signal peptide-protein combinations based on a language model is provided, comprising the following steps:

[0007] (1) Divide the signal peptide-protein sequence to be predicted into translation units, and extract the first M amino acid sequences at the N-terminus of each translation unit as the signal peptide characteristic sequence.

[0008] (2) Input the signal peptide feature sequence of each translation unit obtained in step (1) into the pre-trained protein language model to obtain the amino acid residue feature vector and / or protein sequence feature vector of the translation unit.

[0009] (3) The amino acid residue feature vector and / or protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted obtained in step (2) are spliced ​​together to obtain the secretion feature vector of the signal peptide-protein sequence.

[0010] (4) Input the secretion feature vector of the signal peptide-protein sequence obtained in step (3) into the prediction model to predict the secretion efficiency level of the signal peptide-protein sequence.

[0011] Preferably, in the method for predicting the secretion efficiency of the signal peptide-protein combination, the length of the signal peptide characteristic sequence, M, in step (1) is between 80 and 200, preferably between 100 and 150.

[0012] Preferably, in the signal peptide-protein combination secretion efficiency prediction method, the translation unit in step (1) refers to the smallest independent unit that translates an mRNA sequence into an amino acid sequence, and generally a protein subunit is the translation unit.

[0013] Preferably, the signal peptide-protein combination secretion efficiency prediction method uses a protein language model that includes, but is not limited to, ESM-1, ESM-2, AminoBERT, and a natural language deep learning model trained using amino acid sequences, wherein the framework of the natural language deep learning model is BERT and its derivative frameworks, or GPT and its derivative frameworks.

[0014] The protein language model is based on a language model and uses amino acid sequences as training data to perform self-supervised pre-training of the model, resulting in amino acid feature vectors and protein sequence feature vectors.

[0015] Preferably, in the method for predicting the secretion efficiency of the signal peptide-protein combination, the amino acid residue feature vector and / or protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted are spliced ​​together and normalized to obtain the secretion feature vector of the signal peptide-protein sequence.

[0016] Preferably, in the method for predicting the secretion efficiency of the signal peptide-protein combination, the amino acid residue feature vector and / or protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted are concatenated using one of the following methods:

[0017] First, the amino acid residue feature vectors and / or protein sequence feature vectors of each translation unit are directly concatenated into a secretion feature vector, wherein the dimension of the secretion feature vector is the sum of the dimensions of the amino acid residue feature vectors and / or protein sequence feature vectors of each translation unit.

[0018] Secondly, the amino acid residue feature vectors and / or protein sequence feature vectors of each translation unit are convolved using a one-dimensional convolution model and then concatenated to form a secretion feature vector.

[0019] Preferably, in the signal peptide-protein combination secretion efficiency prediction method, the one-dimensional convolutional model includes a one-dimensional convolutional layer and a pooling layer; the amino acid residue feature vector and / or protein sequence feature vector are convolved in the sequence length direction through the one-dimensional convolutional layer; the output of the convolutional layer is passed through a pooling layer to reduce the information dimension and prevent overfitting; the output of the pooling layer is converted into a one-dimensional vector, and multiple vectors that have been converted into one-dimensional vectors are directly concatenated into a secretion feature vector.

[0020] Preferably, the prediction model of the signal peptide-protein combination secretion efficiency prediction method is a classification model or a regression model; preferably, a classification model, such as a support vector machine, random forest, or multilayer perceptron model.

[0021] Preferably, the prediction model of the signal peptide-protein combination secretion efficiency prediction method is trained according to the following method:

[0022] (4-1) Obtain training data on the secretion efficiency of signal peptide-protein combinations of the same type as the signal peptide-protein combination to be predicted, wherein the signal peptide-protein combination and the signal peptide-protein combination to be predicted have the same number of translation units.

[0023] (4-2) The training data on secretion efficiency of signal peptide-protein combination obtained in step (4-1) is classified according to secretion efficiency or the secretion efficiency value is directly predicted after normalization of the data to obtain a training sample set; preferably, it is classified according to secretion efficiency.

[0024] (4-3) The prediction model is trained using the training sample set obtained in step (4-2) with the goal of minimizing the value of the loss function until it converges;

[0025] When the prediction model is a classification model, the cross-entropy between the predicted secretion efficiency and the actual secretion efficiency is preferably used as the loss function Loss, denoted as:

[0026]

[0027] Where N is the number of training samples, K is the number of classification categories, and if the i-th sample belongs to category c, then y ic =1, otherwise y ic =0, p ic To predict the probability that sample i belongs to category c;

[0028] When the prediction model is a regression model, the mean squared error or mean absolute error between the predicted secretion efficiency and the actual secretion efficiency is preferably used as the loss function Loss, denoted as:

[0029]

[0030] Where N is the number of samples, y i Let i be the secretion efficiency of the i-th sample. Let m be the predicted secretion efficiency of the i-th sample, and m be the norm. If m = 1, the loss function is the mean absolute error; if m = 2, the loss function is the mean squared error.

[0031] According to another aspect of the present invention, a system is provided as an electronic device or a non-transitory computer-readable storage medium, wherein the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the signal peptide-protein combination secretion efficiency prediction method provided by the present invention.

[0032] The non-transitory computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the signal peptide-protein combination secretion efficiency prediction method provided by the present invention.

[0033] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0034] This invention truncates the signal peptide-protein sequence to an appropriate length, balancing signal peptide and protein sequence information. Using a protein language model, it extracts the amino acid physicochemical characteristics and sequence context features of the signal peptide feature sequence, inputting these into a prediction model. The neural network then predicts the expression level of the signal peptide paired with the target protein, thus replacing experiments to find the most suitable signal peptide for the target protein. This invention avoids the problem of excessively long protein sequences obscuring key features of the signal peptide sequence. Furthermore, by using a protein language model, it enriches the structural features of the truncated protein sequence, avoiding the problem of missing structural features due to incomplete sequences, and improving the accuracy of the prediction model in predicting the secretion efficiency of the signal peptide-protein combination. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the process for predicting the secretion efficiency of signal peptide-protein combinations based on language models provided by the present invention.

[0036] Figure 2 This is a schematic diagram of the prediction model structure used in the language model-based signal peptide-protein combination secretion efficiency prediction method provided in Embodiments 1 to 3 of the present invention;

[0037] Figure 3 This is a schematic diagram of the prediction model structure used in the language model-based signal peptide-protein combination secretion efficiency prediction method provided in Embodiment 4 of the present invention. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0039] One of the difficulties in predicting the relationship between signal peptide and target protein binding yield lies in the challenge of effectively extracting the yield characteristics of the signal peptide and target protein. The sequences of the signal peptide and target protein are relatively independent, while the length of the signal peptide is typically much shorter than that of the target protein. Artificially extracting sequence features fails to effectively characterize the complex protein assembly and secretion processes mediated by the signal peptide, resulting in poor prediction performance. Conversely, considering only sequence features leads to a significant weighting of the target protein compared to the signal peptide, with sequence length causing a weight imbalance. Furthermore, the lack of consideration for the association between the signal peptide and target protein further contributes to the poor prediction results.

[0040] The present invention provides a method for predicting the secretion efficiency of signal peptide-protein combinations based on language models, such as... Figure 1 As shown, it includes the following steps:

[0041] (1) Divide the signal peptide-protein sequence to be predicted into translation units. For each translation unit, extract the first M amino acid sequences from the N-terminus as the signal peptide characteristic sequence; M is 80 to 200, preferably 100 to 150; a translation unit refers to the smallest independent unit that translates an mRNA sequence into an amino acid sequence. Generally, protein subunits are translation units. For example, in recombinant proteins such as antibodies, the light chain and heavy chain are translated separately, and each is a translation unit.

[0042] By truncating the sequence length of the translation unit, the input weights of the signal peptide and the protein sequence in the input signal are balanced and consistent, so as not to cause information masking because the length of the protein sequence is much longer than the length of the signal peptide, and at the same time, the sequence information of the correlation between the signal peptide and the protein sequence can be extracted.

[0043] (2) Input the signal peptide feature sequence of each translation unit obtained in step (1) into the pre-trained protein language model to obtain the amino acid residue feature vector and / or protein sequence feature vector of the translation unit; the pre-trained protein language model, such as publicly published language models such as ESM-1, ESM-2, AminoBERT, etc., can also be trained by using deep learning model frameworks in natural language such as BERT and its derivative frameworks, GPT and its derivative frameworks, etc.

[0044] The protein language model is based on a language model and uses amino acid sequences as training data to perform self-supervised pre-training of the model. The resulting amino acid feature vector and protein sequence feature vector not only include the physicochemical characteristics of amino acids, but also obtain the contextual and structural information of the sequence, thereby enabling more effective extraction of the interaction relationship between the signal peptide and the target protein amino acid sequence.

[0045] Pre-trained protein language models, trained on approximately 280 million natural protein sequences from protein sequence libraries such as UniParc, can learn the physicochemical properties of amino acids, evolutionary information about amino acid sequences, and structural information without requiring manual annotation. Pre-trained language models address the problem of neural networks being unable to effectively extract features due to insufficient data, thus necessitating the collection and extraction of new features.

[0046] (3) The amino acid residue feature vector and / or protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted obtained in step (2) are spliced ​​together, preferably after normalization, to obtain the secretion feature vector of the signal peptide-protein sequence.

[0047] The amino acid residue feature vectors and / or protein sequence feature vectors of each translation unit of the signal peptide-protein sequence to be predicted can be concatenated using one of the following methods:

[0048] Firstly, the amino acid residue feature vectors and / or protein sequence feature vectors of each translation unit are directly concatenated into a secretion feature vector, wherein the dimension of the secretion feature vector is the sum of the dimensions of the amino acid residue feature vectors and / or protein sequence feature vectors of each translation unit.

[0049] Secondly, the amino acid residue feature vectors and / or protein sequence feature vectors of each translation unit are convolved using a one-dimensional convolution model and then concatenated to form a secretion feature vector. The one-dimensional convolution model includes a one-dimensional convolutional layer and a pooling layer. The amino acid residue feature vectors and / or protein sequence feature vectors are convolved in the sequence length direction through the one-dimensional convolutional layer. The output of the convolutional layer is passed through a pooling layer to reduce the information dimensionality and prevent overfitting. The output of the pooling layer is converted into a one-dimensional vector, and multiple vectors that have been converted into one-dimensional vectors are directly concatenated to form a secretion feature vector.

[0050] (4) Input the secretion feature vector of the signal peptide-protein sequence obtained in step (3) into the prediction model to predict the secretion efficiency level of the signal peptide-protein sequence. The prediction model is a classification model or a regression model; a classification model is preferred, such as a support vector machine, random forest, or multilayer perceptron model.

[0051] When the prediction model is a classification model, the specific steps are as follows:

[0052] The secretion feature vector is used as input to predict the secretion efficiency level through a prediction model.

[0053] When the prediction model is a multilayer perceptron model, the final output layer of the multilayer perceptron is the secretion efficiency level to be predicted.

[0054] The prediction model is trained using the following method:

[0055] (4-1) Obtain training data on the secretion efficiency of signal peptide-protein combinations of the same type as the signal peptide-protein combination to be predicted. The signal peptide-protein combination and the signal peptide-protein combination to be predicted have the same number of translation units. For example, an antibody has two translation units, a light chain and a heavy chain. The signal peptide-protein combination used in the training data should have two translation units.

[0056] Training data can be collected from literature or experimental data generated in the laboratory.

[0057] (4-2) The training data on the secretion efficiency of the signal peptide-protein combination obtained in step (4-1) is either classified according to secretion efficiency or the secretion efficiency value is directly predicted after data normalization to obtain a training sample set; preferably, it is classified according to secretion efficiency, for example, into two categories: expressed and not expressed, or into three categories: low expression, medium expression, and high expression, etc. Since the secretion efficiency of specific signal peptide-protein combinations varies greatly, classification prediction can achieve a high accuracy rate without affecting the judgment utility.

[0058] (4-3) The prediction model is trained using the training sample set obtained in step (4-2) with the goal of minimizing the value of the loss function until it converges;

[0059] When the prediction model is a classification model, the cross-entropy between the predicted secretion efficiency and the actual secretion efficiency is preferably used as the loss function Loss, denoted as:

[0060]

[0061] Where N is the number of training samples, K is the number of classification categories, and if the i-th sample belongs to category c, then y ic =1, otherwise y ic =0, p ic To predict the probability that sample i belongs to category c.

[0062] When the prediction model is a regression model, the mean squared error or mean absolute error between the predicted secretion efficiency and the actual secretion efficiency is preferably used as the loss function Loss, denoted as:

[0063]

[0064] Where N is the number of samples, y i Let i be the secretion efficiency of the i-th sample. Let m be the predicted secretion efficiency of the i-th sample, and m be the norm. If m = 1, the loss function is the mean absolute error; if m = 2, the loss function is the mean squared error.

[0065] The following is an example:

[0066] The original training set data came from the literature "Optimization of heavy chain and light chainsignal peptides for high level expression of therapeutic antibodies in CHO cells[J]. PloS one,2015,10(2):e0116878.", with an expression level greater than 100 tagged as "expressed" or 1, and an expression level less than 1 tagged as "not expressed" or 0, for a total of 87 entries (Table 1). The test set data came from our company's test results, with an expression level greater than 1 tagged as "expressed" or 1, and an expression level less than 1 tagged as "not expressed" or 0, for a total of 30 entries (Table 2).

[0067] Table 1: Training set data on signal peptide-target protein secretion efficiency

[0068]

[0069]

[0070] Table 2. Data from the signal peptide-target protein secretion efficiency test set.

[0071]

[0072]

[0073] The evaluation criterion for the model is the accuracy of its predictions on the test set.

[0074]

[0075] Where All represents the total number of all samples, TP represents the number of samples that are actually expressed but are also predicted to be expressed (correct classification), TN represents the number of samples that are actually not expressed but are also predicted to be not expressed (correct classification), and Accuracy is the accuracy rate.

[0076] Example 1

[0077] The language model-based method for predicting the secretion efficiency of signal peptide-protein combinations provided in this embodiment, targeting antibody secretion efficiency levels, includes the following steps:

[0078] (1) Divide the signal peptide-protein sequence to be predicted into translation units, and for each translation unit, extract the first M amino acids from the N-terminus as the signal peptide characteristic sequence; specifically:

[0079] The antibody light chain and heavy chain linked by the signal peptide are treated as two separate sequences, each serving as a translation unit.

[0080] M is set to 149. Sequence language models will automatically pad for lengths less than 149.

[0081] (2) Input the signal peptide feature sequence of each translation unit obtained in step (1) into a pre-trained protein language model to obtain the amino acid residue feature vector and / or protein sequence feature vector of the translation unit; specifically:

[0082] Input the signal peptide feature sequence of each translation unit obtained in step (1) into the pre-trained esm2_t33_650M_UR50D protein language model; obtain the protein sequence feature V, the dimension of V is 1280.

[0083] (3) The amino acid residue feature vector or protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted obtained in step (2) are concatenated to obtain the secretion feature vector of the signal peptide-protein sequence; specifically:

[0084] The protein sequence features of the signal peptide linking the antibody light and heavy chains are spliced ​​together, resulting in a feature V' with dimensions of 2560. The mean μ of each feature dimension in the training set is calculated. train With standard deviation σ train The features of the training and test sets are standardized, resulting in a new feature vector S. train =(V' train -μ train ) / σ train ,,S test =(V' test -μ train ) / σ train As a secretion feature vector.

[0085] (4) Input the secretion feature vector of the signal peptide-protein sequence obtained in step (3) into the prediction model to predict the secretion efficiency level of the signal peptide-protein sequence. Specifically:

[0086] This embodiment uses the Support Vector Machine (SVM) classification model for classification.

[0087] The training process for the prediction model is as follows:

[0088] To find the optimal parameters, we use a grid search method to find the optimal parameters for the SVM. The parameter grid of the SVM is param_grid={"kernel":('linear','poly','rbf','sigmoid'),"C":[1e-1,1e0,1e1,1e2],"gamma":np.logspace(-2,5,2)}, and the scoring function is roc_auc.

[0089] The optimal parameters were finally found in the training set to be {'C':1.0,'gamma':0.01,'kernel':'sigmoid'}. Using this optimal parameter model to predict the test set data, the prediction accuracy was 66.7%.

[0090] Example 2

[0091] The language model-based method for predicting the secretion efficiency of signal peptide-protein combinations provided in this embodiment, targeting antibody secretion efficiency levels, includes the following steps:

[0092] (1) Divide the signal peptide-protein sequence to be predicted into translation units, and for each translation unit, extract the first M amino acids from the N-terminus as the signal peptide characteristic sequence; specifically:

[0093] The antibody light chain and heavy chain linked by the signal peptide are treated as two separate sequences, each serving as a translation unit.

[0094] M is set to 149. Sequence language models will automatically pad for lengths less than 149.

[0095] (2) Input the signal peptide feature sequence of each translation unit obtained in step (1) into a pre-trained protein language model to obtain the amino acid residue feature vector and / or protein sequence feature vector of the translation unit; specifically:

[0096] Input the signal peptide feature sequence of each translation unit obtained in step (1) into the pre-trained esm2_t33_650M_UR50D protein language model; obtain the protein sequence feature V, the dimension of V is 1280.

[0097] (3) The amino acid residue feature vector or protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted obtained in step (2) are concatenated to obtain the secretion feature vector of the signal peptide-protein sequence; specifically:

[0098] The protein sequence features of the signal peptide linking the antibody light and heavy chains are spliced ​​together, resulting in a feature V' with dimensions of 2560. The mean μ of each feature dimension in the training set is calculated. train With standard deviation σ train The features of the training and test sets are standardized, resulting in a new feature vector S. train =(V' train -μ train ) / σ train ,,S test =(V' test -μ train ) / σ trainAs a secretion feature vector.

[0099] (4) Input the secretion feature vector of the signal peptide-protein sequence obtained in step (3) into the prediction model to predict the secretion efficiency level of the signal peptide-protein sequence. Specifically:

[0100] This embodiment uses a random forest classification model.

[0101] The training process for the prediction model is as follows:

[0102] To find the optimal parameters, we use a grid search method to find the optimal parameters for the random forest. The parameter grid of the random forest is param_grid={'n_estimators':[10,100],"max_depth":[5,None],'max_features':['sqrt','log2',None],"criterion":["gini","entropy",'log_loss'],}, and the scoring function is roc_auc.

[0103] Finally, the optimal parameters were found in the training set as {'criterion':'entropy','max_depth':5,'max_features':None,'n_estimators':100}. The model with the optimal parameters was used to predict the test set data, and its prediction accuracy was 56.7%.

[0104] Example 3

[0105] The method for predicting the secretion efficiency of signal peptide-protein combinations based on language models provided in this embodiment is as follows: Figure 2 As shown, the steps for determining antibody secretion efficiency levels include:

[0106] (1) Divide the signal peptide-protein sequence to be predicted into translation units, and for each translation unit, extract the first M amino acids from the N-terminus as the signal peptide characteristic sequence; specifically:

[0107] The antibody light chain and heavy chain linked by the signal peptide are treated as two separate sequences, each serving as a translation unit.

[0108] M is set to 149. Sequence language models will automatically pad for lengths less than 149.

[0109] (2) Input the signal peptide feature sequence of each translation unit obtained in step (1) into a pre-trained protein language model to obtain the amino acid residue feature vector and / or protein sequence feature vector of the translation unit; specifically:

[0110] Input the signal peptide feature sequence of each translation unit obtained in step (1) into the pre-trained esm2_t33_650M_UR50D protein language model; obtain the protein sequence feature V, the dimension of V is 1280.

[0111] (3) The amino acid residue feature vector or protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted obtained in step (2) are concatenated to obtain the secretion feature vector of the signal peptide-protein sequence; specifically:

[0112] The protein sequence features of the signal peptide linking the antibody light and heavy chains are spliced ​​together, resulting in a feature V' with dimensions of 2560. The mean μ of each feature dimension in the training set is calculated. train With standard deviation σ train The features of the training and test sets are standardized, resulting in a new feature vector S. train =(V' train -μ train ) / σ train ,,S test =(V' test -μ train ) / σ train As a secretion feature vector.

[0113] (4) Input the secretion feature vector of the signal peptide-protein sequence obtained in step (3) into the prediction model to predict the secretion efficiency level of the signal peptide-protein sequence. Specifically:

[0114] This embodiment uses a multilayer perceptron (MLP) for classification.

[0115] The training process for the prediction model is as follows:

[0116] To find the optimal parameters, we use a grid search method to find the optimal parameters for the multilayer perceptron. The parameter grid of the multilayer perceptron is param_grid={'hidden_layers':[0,1,2],"neurons":[8,16,32],'optimizers':['SGD'],"learning_rate":[1e-3,1e-2,1e-1,1],"epoches":

[10] ,,"batch_sizes":

[16] }, and the scoring function is accuracy.

[0117] The optimal parameters were finally found in the training set as {'hidden_layers':1,"neurons":8,'optimizers':'SGD',"learning_rate":1e-3,"epoches":10,,"batch_sizes":16}. Using the model with the optimal parameters to predict the test set data, the prediction accuracy was 76.7%.

[0118] Example 4

[0119] The method for predicting the secretion efficiency of signal peptide-protein combinations based on language models provided in this embodiment is as follows: Figure 3 As shown, the steps for determining antibody secretion efficiency levels include:

[0120] (1) Divide the signal peptide-protein sequence to be predicted into translation units, and for each translation unit, extract the first M amino acids from the N-terminus as the signal peptide characteristic sequence; specifically:

[0121] The antibody light chain and heavy chain linked by the signal peptide are treated as two separate sequences, each serving as a translation unit.

[0122] M is set to 149. Sequence language models will automatically pad for lengths less than 149.

[0123] (2) Input the signal peptide feature sequence of each translation unit obtained in step (1) into a pre-trained protein language model to obtain the amino acid residue feature vector and / or protein sequence feature vector of the translation unit; specifically:

[0124] Input the signal peptide feature sequence of each translation unit obtained in step (1) into the pre-trained esm2_t33_650M_UR50D protein language model; obtain the protein sequence feature V, the dimension of V is 1280.

[0125] (3) The amino acid residue feature vector or protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted obtained in step (2) are concatenated to obtain the secretion feature vector of the signal peptide-protein sequence; specifically:

[0126] The amino acid residue feature vectors and / or protein sequence feature vectors of each translation unit are convolved using a one-dimensional convolution model and then concatenated to form a secretion feature vector. The one-dimensional convolution model includes a one-dimensional convolutional layer and a pooling layer. The amino acid residue feature vectors and / or protein sequence feature vectors are convolved along the sequence length direction using the one-dimensional convolutional layer. The output of the convolutional layer is passed through a pooling layer to reduce the information dimensionality and prevent overfitting. The output of the pooling layer is converted into a one-dimensional vector, and the two converted one-dimensional vectors are directly concatenated to form the secretion feature vector. Figure 3 As shown, the concatenated feature V' has a dimension of 2560. The mean μ of each feature dimension in the training set is calculated. train With standard deviation σ train The features of the training and test sets are standardized, resulting in a new feature vector S. train =(V' train -μ train ) / σ train ,,S test =(V' test -μ train ) / σ train As a secretion feature vector.

[0127] (4) Input the secretion feature vector of the signal peptide-protein sequence obtained in step (3) into the prediction model to predict the secretion efficiency level of the signal peptide-protein sequence. Specifically:

[0128] This embodiment uses a multilayer perceptron (MLP) for classification.

[0129] The training process for the prediction model is as follows:

[0130] To find the optimal parameters, we use a grid search method to find the optimal parameters for the one-dimensional convolutional model. The MLP hyperparameters of the one-dimensional convolutional model are fixed as those in Example 3. Other hyperparameter grids related to one-dimensional convolution are: param_grid = {'kernel_sizes':[3,5],"out_channels":[8,16,32,64]} and the scoring function is accuracy.

[0131] The optimal parameters were ultimately found to be {'kernel_sizes':5,"out_channels":32} in the training set. Using this optimal parameter model to predict the test set data, the prediction accuracy was 76.7%.

[0132] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A signal peptide-protein combination secretion efficiency prediction method based on a language model, characterized by, The method comprises the following steps: (1) dividing the signal peptide-protein sequence to be predicted into translation units, and taking the first M amino acid sequences at the N-terminal of each translation unit as a signal peptide feature sequence; (2) inputting the signal peptide feature sequence of each translation unit obtained in step (1) into a pre-trained protein language model to obtain an amino acid residue feature vector and / or a protein sequence feature vector of the translation unit; (3) performing feature splicing on the amino acid residue feature vector and / or the protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted obtained in step (2) to obtain a secretion feature vector of the signal peptide-protein sequence; (4) inputting the secretion feature vector of the signal peptide-protein sequence obtained in step (3) into a prediction model to predict the secretion efficiency level of the signal peptide-protein sequence.

2. The signal peptide-protein combination secretion efficiency prediction method according to claim 1, wherein The length of the signal peptide feature sequence in step (1) is M, which is between 80 and 200, preferably 100 to 150.

3. The signal peptide-protein combination secretion efficiency prediction method according to claim 1, wherein The translation unit in step (1) refers to the smallest independent unit of mRNA sequence translation into amino acid sequence, and generally refers to a protein subunit.

4. The signal peptide-protein combination secretion efficiency prediction method according to claim 1, wherein The pre-trained protein language model includes but is not limited to ESM-1, ESM-2, AminoBERT, and a natural language deep learning model trained using an amino acid sequence, and the framework of the natural language deep learning model is BERT and its derivative framework, or GPT and its derivative framework. The protein language model is based on a language model, uses an amino acid sequence as training data, performs self-supervised pre-training of the model, and obtains an amino acid feature vector and a protein sequence feature vector.

5. The signal peptide-protein combination secretion efficiency prediction method according to claim 1, wherein The amino acid residue feature vector and / or the protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted is subjected to feature splicing and normalization processing to obtain a secretion feature vector of the signal peptide-protein sequence.

6. The signal peptide-protein combination secretion efficiency prediction method according to claim 1, wherein The amino acid residue feature vector and / or the protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted is subjected to feature splicing and normalization processing to obtain a secretion feature vector of the signal peptide-protein sequence. One of the following methods is used for feature splicing of the amino acid residue feature vector and / or the protein sequence feature vector of each translation unit of the signal peptide-protein sequence to be predicted: One is to directly splice the amino acid residue feature vector and / or the protein sequence feature vector of each translation unit into a secretion feature vector, and the dimension of the secretion feature vector is the sum of the dimensions of the amino acid residue feature vector and / or the protein sequence feature vector of each translation unit; 7. The signal peptide-protein combination secretion efficiency prediction method according to claim 6, wherein The other is to convolve the amino acid residue feature vector and / or the protein sequence feature vector of each translation unit using a one-dimensional convolution model and then splice it into a secretion feature vector. The one-dimensional convolution model includes a one-dimensional convolution layer and a pooling layer; the amino acid residue feature vector and / or the protein sequence feature vector is convolved in the sequence length direction through the one-dimensional convolution layer; the convolution layer output is reduced in information dimension through a pooling layer to prevent overfitting; the pooling layer output is converted into a one-dimensional vector, and multiple one-dimensional vectors are directly spliced into a secretion feature vector.

8. The signal peptide-protein combination secretion efficiency prediction method according to claim 1, wherein, The prediction model is a classification model or a regression model; preferably a classification model, such as a support vector machine, a random forest, or a multi-layer perceptron model.

9. The signal peptide-protein combination secretion efficiency prediction method according to claim 1, wherein, The prediction model is trained according to the following method: (4-1) obtaining signal peptide-protein combination secretion efficiency training data of the same type as the signal peptide-protein combination to be predicted, the signal peptide-protein combination to be predicted having the same number of translation units as the signal peptide-protein combination to be predicted; (4-2) classifying the signal peptide-protein combination secretion efficiency training data obtained in step (4-1) according to secretion efficiency or directly predicting the secretion efficiency value after normalizing the data, to obtain a training sample set; preferably classifying according to secretion efficiency; (4-3) training the prediction model using the training sample set obtained in step (4-2) until convergence, with the goal of minimizing the value of the loss function; When the prediction model is a classification model, the cross-entropy of the model-predicted secretion efficiency and the true secretion efficiency is preferably used as the loss function Loss, denoted as: where N is the number of training samples, K is the number of classification categories, y ic = 1 if the i-th sample belongs to the c-th category, otherwise y ic = 0, p ic is the probability that the i-th sample belongs to the c-th category. When the prediction model is a regression model, the mean square error or the mean absolute error of the model-predicted secretion efficiency and the true secretion efficiency is preferably used as the loss function Loss, denoted as: where N is the number of samples, y i is the secretion efficiency of the i-th sample, is the predicted secretion efficiency of the i-th sample, and m is the norm, e.g. m = 1 for the mean absolute error or m = 2 for the mean squared error.

10. A system, being an electronic device or a non-transitory computer-readable storage medium, characterized in that, The electronic device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the steps of the signal peptide-protein combination secretion efficiency prediction method according to any one of claims 1 to 9; The non-transitory computer-readable storage medium has a computer program stored thereon, and the computer program is executable by a processor to implement the steps of the signal peptide-protein combination secretion efficiency prediction method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • System and method for increasing synthesized protein stability

    CN113727994A

  • Method and system for predicting amino acid sequence in antibody protein CDR region

    CN113838523A