A method for predicting the acid-base stability of a protein, an electronic device, and a storage medium
By constructing a protein acid-base stability prediction model based on bacterial tolerance pH, the problem of limitations in protein acid-base stability research data is solved, and efficient screening of protein stability is achieved.
Patent Information
- Application Number
- CN202311265944.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-09-27
AI Technical Summary
The existing technology lacks a protein acid-base stability database that is directly used, resulting in limited research data on protein acid-base stability, and it is impossible to quickly and accurately screen out stable proteins that meet production needs.
By collecting the pH value and exocrine protein sequences of bacteria and their survival tolerance, training classification or regression models, building a protein acid-base stability prediction model, using the bacteria tolerant pH value as the tolerance pH value of exocrine proteins, obtaining protein sequence and their tolerance pH value data, and making stability predictions.
Accurate classification of protein acid and base stability or prediction of the upper and lower limits of tolerant pH values is achieved, improving the accuracy and efficiency of protein stability screening.
Smart Images

Figure CN118072829B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and more specifically, relates to a method for predicting the acid-base stability of proteins, an electronic device, and a storage medium. Background Art
[0002] In the field of protein design, the focus is often on the structure and functional activity of proteins. However, in actual production applications, proteins are also required to have a certain stability, including acid-base stability. How to quickly and accurately screen out stable proteins that meet production needs from a large number of designed candidate active proteins has gradually become an important link in the field of protein design.
[0003] For the analysis of protein acid-base stability, there is currently no publicly available database that can be directly used, resulting in the inability to produce research results that meet application requirements due to data limitations in the study of protein acid-base stability. Summary of the Invention
[0004] In view of the above-mentioned defects or improvement requirements of the prior art, the present invention provides a method for predicting the acid-base stability of proteins, an electronic device, and a storage medium, aiming to use the pH value tolerated by bacteria as the pH value tolerated by the excreted proteins of the bacteria, so as to collect protein pH value tolerance information from the existing bacterial survival database for prediction algorithms, thereby solving the technical problem of data limitations caused by the lack of a protein acid-base stability database.
[0005] To achieve the above object, according to one aspect of the present invention, a method for predicting the acid-base stability of proteins is provided, including the following steps:
[0006] (1) Collect bacteria, the tolerated pH values they survive, and the excreted protein sequences of the corresponding bacteria;
[0007] (2) Use the pH value tolerated by the bacteria obtained in step (1) as the pH value tolerated by the excreted proteins of the bacteria to obtain protein sequence and its tolerated pH value data;
[0008] (3) Use the protein sequence and its tolerated pH value data obtained in step (2) as training samples to train a classification model or a regression model to obtain a trained protein acid-base stability prediction model;
[0009] (4) Input the protein sequence whose acid-base stability is to be predicted into the protein acid-base stability prediction model obtained in step (3) to obtain the protein acid-base stability prediction result.
[0010] Preferably, in the method for predicting the acid-base stability of proteins, in step (1), the following method is used to determine whether a protein is an excreted protein:
[0011] Determine whether the full sequence of the protein contains a signal peptide. If it contains a signal peptide, determine that the protein sequence is an exosomal protein sequence; otherwise, determine it as a non-exosomal protein sequence.
[0012] Preferably, in the method for predicting the acid-base stability of the protein, in step (1), the sequence after removing the signal peptide is used as the exosomal protein sequence.
[0013] Preferably, in the method for predicting the acid-base stability of the protein, a signal peptide prediction tool is used to predict the signal peptide site, and the sequence downstream of the signal peptide site is used as the sequence after removing the signal peptide.
[0014] Preferably, in the method for predicting the acid-base stability of the protein, in step (3), when the protein acid-base stability prediction model is a classification model, the proteins are classified according to the tolerated pH value range to obtain classification labels; the protein sequence and its classification labels are used as training samples.
[0015] When the protein acid-base stability prediction model is a regression model, the upper limit or the lower limit of the tolerated pH value of the protein is used as the dependent variable, and the protein sequence and its upper limit or lower limit of the tolerated pH value are used as training samples.
[0016] Preferably, in the method for predicting the acid-base stability of the protein, in step (3), the protein acid-base stability prediction model preferably adopts a neural network model, specifically a multi-layer perceptron, a support vector machine, or a logistic regression model.
[0017] Preferably, in the method for predicting the acid-base stability of the protein, in step (4), the protein sequence is encoded and then input into the protein acid-base stability prediction model.
[0018] Preferably, in the method for predicting the acid-base stability of the protein, the protein sequence is encoded using a protein language model.
[0019] According to another aspect of the present invention, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for predicting the acid-base stability of the protein provided by the present invention are implemented.
[0020] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for predicting the acid-base stability of the protein provided by the present invention are implemented.
[0021] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0022] Since extracellular proteins can maintain a certain activity under the pH conditions for bacterial survival, the pH conditions for bacterial survival and the pH values tolerated by bacteria can be used as the pH values tolerated by their extracellular proteins, which are used as indicators for measuring the acid-base stability of bacterial extracellular proteins. In the present invention, by using the pH values tolerated by bacteria as the pH values tolerated by the excreted proteins of the bacteria, protein sequences and their tolerated pH value data are obtained, and accurate samples of the acid-base stability of proteins are obtained, which are used for predicting the acid-base stability of both excreted proteins and intracellular proteins, and can accurately classify the acid-base stability of proteins, or predict the upper and lower limits of the pH values tolerated by proteins. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a schematic flow chart of the method for predicting the acid-base stability of proteins provided by the present invention;
[0024] Figure 2 is a schematic diagram of the implementation process of the method for predicting the acid-base stability of proteins provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0026] The method for predicting the acid-base stability of proteins provided by the present invention, as Figure 1 shown, includes the following steps:
[0027] (1) Collect bacteria, the pH values tolerated by them for survival, and the excreted protein sequences of the corresponding bacteria;
[0028] Judge whether a protein is an excreted protein according to the following method:
[0029] Judge whether the full sequence of the protein contains a signal peptide. If it contains a signal peptide, the protein sequence is judged as an excreted protein sequence; otherwise, it is judged as a non-excreted protein sequence. Specifically, signal peptide prediction tools such as SignalP 6.0 can be used to judge whether the protein sequence contains a signal peptide.
[0030] Preferably, the sequence without the signal peptide is used as the excreted protein sequence; the signal peptide on the finally folded protein is cleaved, so it does not affect the acid-base stability of the protein. It needs to be removed according to the prediction result of the signal peptide site, and the extracellular protein sequence with the signal peptide is screened out and its signal peptide is excised, so as to form a data set for analyzing the acid-base stability of the protein for bacterial survival tolerance to pH.
[0031] Preferably, signal peptide prediction tools such as SignalP 6.0 are used to predict the signal peptide site, and the sequence downstream of the signal peptide site is used as the sequence without the signal peptide.
[0032] (2) The pH value tolerated by the bacteria obtained in step (1) is used as the pH value tolerated by the excreted protein of the bacteria, and data on the protein sequence and its tolerated pH value are obtained;
[0033] Since extracellular proteins can maintain a certain activity under the pH conditions for bacterial survival, the pH conditions for bacterial survival and the pH value tolerated by the bacteria can be used as the pH value tolerated by their extracellular proteins, which is used as an index for measuring the acid-base stability of bacterial extracellular proteins.
[0034] (3) Using the protein sequence and its tolerated pH value data obtained in step (2) as training samples, a classification model or a regression model is trained to obtain a trained protein acid-base stability prediction model;
[0035] When the protein acid-base stability prediction model is a classification model, the protein is classified according to the tolerated pH value range to obtain classification labels; the protein sequence and its classification labels are used as training samples; the model preferably uses a neural network model, specifically a multi-layer perceptron;
[0036] When the protein acid-base stability prediction model is a regression model, the upper limit or the lower limit of the tolerated pH value of the protein is used as the dependent variable, and the protein sequence and its upper limit or lower limit of the tolerated pH value are used as training samples;
[0037] The protein acid-base stability prediction model is a multi-layer perceptron, a support vector machine, or a logistic regression model.
[0038] (4) The protein sequence whose acid-base stability is to be predicted is input into the protein acid-base stability prediction model obtained in step (3) to obtain the protein acid-base stability prediction result.
[0039] Preferably, the protein sequence is encoded and then input into the protein acid-base stability prediction model; specifically, the protein sequence is encoded using a protein language model.
[0040] The following are examples:
[0041] The protein acid-base stability prediction method provided in this embodiment is as follows Figure 2 shown, and includes the following steps:
[0042] (1) Collect bacteria, the tolerated pH values for their survival, and the excreted protein sequences of the corresponding bacteria;
[0043] Obtain the acid-base stability analysis data set, and the specific process is as follows:
[0044] (1-1) Obtain the data of bacteria and their survival pH conditions from the BacDive database. Import the Python library of BacDive, and search for bacteria data according to the BacDive_id. The data fields include BacDive_id (the unique identifier of the bacteria in the BacDive database), NCBI_tax_id (the unique identifier of the species in the NCBI database), culture pH (including the culture pH, minimum growth pH, maximum growth pH, and optimal growth pH of the bacteria), and select the minimum growth pH as the limit pH value for the acid tolerance of the bacteria, with the field name min_pH_growth.
[0045] (1-2) Download all protein data of bacteria from the NCBI's RefSeq database (The NCBI Reference Sequence, proposed by NCBI, aiming to provide non-redundant and manually selected reference sequences for all common organisms). The data fields include refseq_id (the unique identifier of the protein in the RefSeq database), protein_seq (protein sequence), NCBI_tax_id, and BacDive_id. Then, connect the bacteria data and protein data according to the BacDive_id and NCBI_tax_id, remove the data rows where min_pH_growth is empty, and group by protein_seq, and select the data row with the smallest min_pH_growth within the group to obtain approximately 1.66 million data rows.
[0046] (1-3) Determine whether a protein is an excreted protein according to the following method:
[0047] Use signalP 6.0 to predict whether all protein sequences contain signal peptides and the sites of signal peptides. Only retain the proteins with signal peptides and excise the signal peptides according to the sites of signal peptides, and finally generate a protein acid-base stability analysis data set based on BacDive, with approximately 190,000 data rows.
[0048] Taking advantage of the relationship between the bacterial growth environment and the stability of its proteins, the data in the bacterial database BacDive was fully utilized, making it the main data source for constructing the protein stability classification algorithm. Among them, there are approximately 190,000 pieces of data on acid-base stability analysis.
[0049] (2) Use the pH tolerance value of the bacteria obtained in step (1) as the pH tolerance value of the excreted proteins of the bacteria to obtain protein sequences and their pH tolerance value data;
[0050] Integrate the BacDive database and the NCBI database to obtain the dataset required for constructing the acid-base stability prediction model. The data includes protein sequences and their pH tolerance values, and the pH tolerance value is derived from the highest pH value for bacterial growth in the BacDive database.
[0051] (3) Use the protein sequences and their pH tolerance value data obtained in step (2) as training samples to train a classification model or a regression model to obtain a trained protein acid-base stability prediction model;
[0052] In this embodiment, the protein acid-base stability prediction model uses a classification model to classify the proteins according to the pH tolerance value range to obtain classification labels; specifically as follows:
[0053] Preprocessing of acid-base stability analysis data. According to the pH tolerance value, the data is divided into two categories with 3 as the critical value. A pH value greater than 3 indicates acid intolerance, and a pH value less than or equal to 3 indicates acid tolerance. Then, the samples of each category are balanced by random sampling to ensure that the data volume of each category is not very different. According to the sequence similarity, the balanced data is divided into a training set and a test set to ensure that the sequences in the test set are not similar to those in the training set. Use the protein sequences and their classification labels as training samples;
[0054] The classification model in this embodiment uses a neural network model, specifically a multi-layer perceptron, and its working process is as follows:
[0055] Protein embedding. Use the protein language model "facebook / esm2_t33_650M_UR50D" (downloaded from the website "https: / / huggingface.co / models") to embed the protein sequences in all the preprocessed data. The embedding of each protein sequence is represented as a 1280-dimensional vector.
[0056] Training of the acid-base stability binary classification model. The model uses a multi-layer perceptron in the neural network, with an input layer (dimension 1280), an output layer (dimension 2), and two hidden layers (dimensions 640 and 320 respectively). The ReLU activation function is used. The protein embedding is input into the model, and the acid-base stability binary classification model is trained with the tolerance pH value category as the label. During training, the cross-entropy loss function and the Adam optimizer (learning rate 0.00005) are used. The model is saved after 100 training iterations.
[0057] Use the protein language model to embed all protein sequences. The protein language model can effectively extract the general features of proteins, thus better predicting the stability of proteins based on these features.
[0058] Adopt machine learning models, such as multi-layer perceptron, support vector machine, logistic regression, etc. The model can learn more complex non-linear relationships from the training data, thereby improving its modeling ability and accuracy.
[0059] (4) Input the protein sequence whose acid-base stability is to be predicted into the protein acid-base stability prediction model obtained in step (3) to obtain the protein acid-base stability prediction result.
[0060] The protein sequence in this embodiment is encoded and then input into the protein acid-base stability prediction model; specifically, the protein sequence is encoded using the protein language model as follows:
[0061] Use the model to predict the tolerance pH value category of new proteins. Use the protein language model "facebook / esm2_t33_650M_UR50D" to embed the new protein sequence, input this embedding into the model, and finally the model will output the tolerance pH value category of this protein.
[0062] The trained protein acid-base stability classification or regression model all has the ability of rapid prediction and can realize the high-throughput screening of protein stability.
[0063] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting the acid-base stability of a protein, characterized in that, Including the following steps: (1) Collect bacteria, the pH values they can tolerate for survival, and the exoprotein sequences of the corresponding bacteria; (2) Use the pH values tolerated by the bacteria obtained in step (1) as the pH values tolerated by the exoproteins of the bacteria to obtain protein sequences and their pH tolerance data; Determine whether a protein is an exoprotein according to the following method: Determine whether the full sequence of the protein contains a signal peptide. If it contains a signal peptide, determine that the protein sequence is an exoprotein sequence; otherwise, determine it as a non-exoprotein sequence; Use the sequence after removing the signal peptide as the exoprotein sequence; (3) Use the protein sequences and their pH tolerance data obtained in step (2) as training samples to train a classification model or a regression model to obtain a trained protein acid-base stability prediction model; (4) Input the protein sequence whose acid-base stability is to be predicted into the protein acid-base stability prediction model obtained in step (3) to obtain the protein acid-base stability prediction result.
2. The protein acid-base stability prediction method according to claim 1, characterized in that Use a signal peptide prediction tool to predict the signal peptide site, and use the sequence downstream of the signal peptide site as the sequence after removing the signal peptide.
3. The protein acid-base stability prediction method according to claim 1, wherein In step (3), when the protein acid-base stability prediction model is a classification model, classify the protein according to the pH tolerance range to obtain a classification label; Use the protein sequence and its classification label as training samples; When the protein acid-base stability prediction model is a regression model, use the upper limit or lower limit of the pH value tolerated by the protein as the dependent variable, and use the protein sequence and its upper limit or lower limit of the pH value tolerated as training samples.
4. The method for predicting the acid-base stability of a protein according to claim 1, wherein In step (3), the protein acid-base stability prediction model preferably uses a neural network model, specifically a multi-layer perceptron, a support vector machine, or a logistic regression model.
5. The protein acid-base stability prediction method according to claim 1, wherein, In step (4), the protein sequence is encoded and then input into the protein acid-base stability prediction model.
6. The protein acid-base stability prediction method according to claim 1, wherein Encode the protein sequence using a protein language model.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the protein acid-base stability prediction method according to any one of claims 1 to 6.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the protein acid-base stability prediction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Protein solubility prediction method based on combined machine learning model
CN114582423A