A method and device for processing a protein free energy change prediction model

By constructing a protein free energy change prediction model that combines protein semantic model and gradient enhancement decision tree model, and using multi-dimensional molecular structural characteristics for prediction, the problem of difficult prediction accuracy in the existing technology is solved, and higher prediction accuracy and model flexibility are achieved.

CN116844633BActive Publication Date: 2025-05-16BEIJING DP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310841986.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-10
Publication Date
2025-05-16
Estimated Expiration
2043-07-10

AI Technical Summary

Technical Problem

The existing protein free energy change prediction methods are based on one-dimensional molecular sequences and cannot effectively reflect the molecular structure characteristics of two-dimensional and three-dimensional space, making it difficult to improve the prediction accuracy.

Method used

A free energy change prediction model combining protein semantic model and gradient enhancement decision tree model is constructed. By pre-training and supervised training of protein semantic model, multi-dimensional molecular structural features are used to predict, and multiple protein semantic model selection is provided to improve the flexibility and robustness of the model.

Benefits of technology

Through the utilization of multi-dimensional molecular structural characteristics, the prediction accuracy of protein free energy change is improved, the generalization and robustness of the model are enhanced, and convenient application mode is provided to meet the needs of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844633B_ABST
    Figure CN116844633B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention relates to a method and device for processing a protein free energy change prediction model, the method comprising: constructing a free energy change prediction model; constructing a first and a second data set; pre-training a protein structure feature recognition module based on the first data set; after the pre-training is completed, the free energy change prediction model is trained based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module; after the model training is completed, a model selection identifier, a model application mode and model application data are received; and model parameters of the free energy change prediction model are configured according to the model selection identifier; if the model application mode is the first mode, single-sequence free energy change prediction and stability level evaluation are performed; if the model application mode is the second mode, multi-sequence free energy change prediction, stability level evaluation and multi-sequence stability ranking are performed. The present invention can improve the prediction accuracy of free energy change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and device for processing a protein free energy change prediction model. Background Art

[0002] Protein free energy change (Delta Delta G, DDG) refers to the free energy change caused by the change of the stability of the protein in the system as its amino acid sequence changes. Specifically, the protein free energy change is the difference in free energy between the mutated protein and the original protein. A positive protein free energy change means that the stability of the mutated protein is weakened compared to the original protein; a negative protein free energy change means that the stability of the mutated protein is enhanced compared to the original protein. Protein free energy change is an important parameter for evaluating the effect of amino acid mutations on protein stability. Accurate prediction of protein free energy change is of great significance to protein evolution research.

[0003] At present, most conventional protein free energy change prediction methods are based on one-dimensional molecular sequences. However, one-dimensional molecular sequences cannot reflect the more complex two-dimensional and three-dimensional spatial molecular structure characteristics, which leads to the fact that the accuracy of conventional prediction methods cannot be further improved.

[0004] In recent years, with the in-depth development of artificial intelligence technology, more and more research has begun to focus on protein free energy prediction methods based on artificial intelligence technology. How to combine artificial intelligence technology to provide a prediction model that improves the accuracy of protein free energy change prediction is the technical problem that the present invention needs to solve. Summary of the invention

[0005] The purpose of the present invention is to provide a method, device, electronic device and computer-readable storage medium for processing a protein free energy change prediction model in view of the defects of the prior art. The present invention first constructs a free energy change prediction model that can predict protein free energy change by using a protein semantic model that can perform multi-dimensional (one-dimensional, two-dimensional, three-dimensional) molecular structure feature recognition on a one-dimensional molecular sequence and a gradient boosting decision tree model; then, the protein semantic model is pre-trained so that it can obtain more and more accurate multi-dimensional molecular structure features to improve the generalization of the model; then, after the model parameters of the protein semantic model are solidified, the gradient boosting decision tree model is trained based on a supervised model training method to improve the prediction accuracy of the model; in addition, the present invention also provides multiple groups of model selection (ESM-1b model, ESM-1v model, ESM-if1 model, ESM2 model and ProtBERT model), and train a set of corresponding protein semantic model parameters and gradient boosting decision tree model parameters for each model selection during model training, while improving the model prediction accuracy, it can also maintain good robustness; in addition, the present invention also provides two types of application modes for users to use: 1) In the first mode, the free energy change of the original-mutated protein sequence pair input by the user is predicted and the stability of the mutated protein is graded according to the prediction results; 2) In the second mode, the free energy change of an original protein sequence input by the user and its corresponding multiple mutant protein sequences are predicted separately and the stability of each mutated protein is graded according to the prediction results and the mutated protein sequences are sorted according to the evaluation results. The free energy change prediction model provided by the present invention can improve the prediction accuracy of free energy change based on multidimensional molecular structure characteristics, the training method of the present invention can further improve the generalization of the model, the model selection mechanism provided by the present invention can further improve the flexibility and robustness of the model, and the two types of application modes provided by the present invention can further improve the convenience of using the model.

[0006] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present invention provides a method for processing a protein free energy change prediction model, the method comprising:

[0007] Constructing a free energy change prediction model; the free energy change prediction model is used to predict protein free energy changes according to the input original and mutant protein sequences; the free energy change prediction model includes a protein structure feature recognition module, a feature splicing module and a free energy change prediction module;

[0008] Constructing a pre-training data set and a prediction training data set as the corresponding first and second data sets; and pre-training the protein structure feature recognition module based on the first data set; after the pre-training is completed, model training is performed on the free energy change prediction model based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module;

[0009] After the model training is completed, the model selection identifier, model application mode and corresponding model application data input by the user are received; and the model parameters of the free energy change prediction model are configured according to the model selection identifier; and the model application mode is identified; if the model application mode is the first mode, a single sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result; if the model application mode is the second mode, a multi-sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result, and a multi-sequence stability ranking is performed according to the assessment result; the model application mode includes a first mode and a second mode; when the model application mode is the first mode, the corresponding model application data includes a first original protein sequence and a first mutant protein sequence; when the model application mode is the second mode, the corresponding model application data includes a second original protein sequence and multiple second mutant protein sequences.

[0010] Preferably, the original protein sequence is the molecular sequence of the original protein, the mutant protein sequence is the molecular sequence of the mutant protein, and the mutant protein is a protein obtained by protein mutation of the original protein; the mutation types of the protein mutation include substitution mutation, deletion mutation and insertion mutation; the mutant protein carries one or more mutant residues, and the mutant residues are residues generated by the protein mutation;

[0011] The input end of the protein structure feature recognition module is connected to the input end of the free energy change prediction model, and the output end is connected to the input end of the feature splicing module; the output end of the feature splicing module is connected to the input end of the free energy change prediction module; the output end of the free energy change prediction module is connected to the output end of the free energy change prediction model;

[0012] The protein structure feature recognition module is a protein semantic model implemented based on the transformer model structure; the protein semantic model selection range of the protein structure feature recognition module includes ESM-1b model, ESM-1v model, ESM-if1 model, ESM2 model and ProtBERT model;

[0013] The free energy change prediction module is implemented based on a gradient boosting decision tree model;

[0014] The protein structure feature recognition module is used to perform protein semantic feature recognition processing on the input original and mutant protein sequences respectively to obtain the corresponding original and mutant sequence feature tensors and send them to the feature splicing module; the feature splicing module is used to perform feature tensor splicing on the original and mutant sequence feature tensors to obtain the corresponding spliced ​​feature tensors and send them to the free energy change prediction module; the free energy change prediction module is used to perform free energy change regression prediction based on the spliced ​​feature tensor to output the corresponding predicted free energy change data;

[0015] The first data set includes a plurality of first model data sets; the first model data sets correspond one-to-one to the models in the selection range of the protein semantic model;

[0016] The second data set includes a plurality of first data records; the first data records include a first training original protein sequence, a first training mutant protein sequence and a first label free energy variation data;

[0017] The identification value range of the model selection identification is composed of multiple first model identifications; the first model identifications correspond one-to-one to the models in the protein semantic model selection range.

[0018] Preferably, the constructing of the pre-training data set and the prediction training data set as the corresponding first and second data sets specifically includes:

[0019] The model training data sets conventionally used by each model in the protein semantic model selection range of the protein structure feature recognition module are used as the corresponding first model data sets; and all the obtained first model data sets are used to form the corresponding first data set; the model training data sets used by each model are composed of multiple protein sequences;

[0020] The original-mutated protein sequence pairs with free energy change information from the KAGGLE database and the VariBench database are extracted as the corresponding first training original protein sequence and the first training mutant protein sequence, and the corresponding free energy change information is used as the corresponding first label free energy change data, and the first training original protein sequence, the first training mutant protein sequence and the first label free energy change data are used to form the corresponding first data record; and all the first data records are used to form the corresponding first original data set; and the first original data set is deduplicated to obtain the corresponding second original data set; and the first data records in the second original data set whose first label free energy change data does not meet the preset free energy change value range are deleted to obtain the corresponding third original data set; and the third original data set is enhanced based on the sliding window method to obtain the corresponding second data set.

[0021] Furthermore, performing data enhancement processing on the third original data set based on a sliding window method to obtain the corresponding second data set specifically includes:

[0022] Step 41, initializing the first enhanced data set to be empty; and taking the first first data record in the third original data set as the corresponding current record;

[0023] Step 42, using the first training original protein sequence, the first training mutant protein sequence and the first label free energy variable data currently recorded as the corresponding first sequence, second sequence and first free energy variable data DDG;

[0024] Step 43, traversing each of the mutated residues in the second sequence; and during the traversal, taking the currently traversed mutated residue as the corresponding first residue Z i ; and based on a preset first number X sliding window sizes L x , in the second sequence, the first residue Z i As the center, the front and rear N=(L x -1) / 2 residues are associated to obtain the corresponding first associated residue sequence S x {Z i-N …Z i-1 ,Z i ,Z i+1 …Z i+N}; and according to each of the first associated residue sequences S x , the first sequence and the second sequence dynamically modulate the first free energy variable DDG data to obtain the corresponding second free energy variable data DDG x; And each of the second free energy variable data DDG x as a new first label free energy variable data; and the first training original protein sequence corresponding to the first and second sequences, the first training mutant protein sequence and each of the new first label free energy variable data form a new first data record; and the obtained X new first data records form a record corresponding to the current mutant residue Z i The first data subset corresponding to the first sequence; and at the end of the traversal, all the first data subsets obtained are added to the first enhanced data set; the residue index i is an integer greater than 1, and the residue index i is the residue ranking index of the corresponding mutation residue in the second sequence; the first number X is an integer greater than 0; each sliding window size L x is an odd number greater than or equal to 3; N is an integer greater than or equal to 1;

[0025] Step 44, identifying whether the current record is the last first data record in the third original data set; if so, going to step 45; if not, taking the next first data record in the third original data set as the new current record and going to step 42;

[0026] Step 45, merging the first enhanced data set with the third original data set to obtain a corresponding first merged data set; performing data record deduplication processing on the first merged data set; and using the first merged data set after deduplication as the corresponding second data set.

[0027] Preferably, the pre-training of the protein structure feature recognition module based on the first data set specifically includes:

[0028] Based on each of the first model data sets in the first data set, the corresponding model in the protein semantic model selection range is trained, and when the current model training is completed, the model parameters of the current model are extracted and stored as the corresponding first protein semantic model parameters; and the pre-training is confirmed to be completed when all models in the protein semantic model selection range have completed training.

[0029] Preferably, the step of performing model training on the free energy change prediction model based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module specifically includes:

[0030] Step 61, performing parameter solidification processing on each model within the protein semantic model selection range of the protein structure feature recognition module;

[0031] Step 62, taking the first model within the selection range of the protein semantic model as the currently selected model of the protein structure feature recognition module;

[0032] Step 63, traverse each of the first data records in the second data set; and during the traversal, use the first data record currently traversed as the corresponding current data record; and use the currently selected model to perform protein semantic feature recognition processing on the first training original protein sequence and the first training mutant protein sequence of the current data record to obtain the corresponding first original sequence feature tensor and the first mutant sequence feature tensor; and use the feature splicing module to perform feature tensor splicing on the first original sequence feature tensor and the first mutant sequence feature tensor to obtain the corresponding first spliced ​​feature tensor; and use the first label free energy variable data of the current data record as the corresponding second label free energy variable data, and use the first spliced ​​feature tensor and the second label free energy variable data to form a corresponding second data record; and at the end of the traversal, all the obtained second data records form a corresponding third data set;

[0033] Step 64, performing gradient boosting decision tree model training on the free energy change prediction module based on the third data set in a preset K-Fold cross validation manner;

[0034] Step 65, when the training of the gradient boosting decision tree model is completed, the model parameters of the free energy change prediction module are extracted and stored as the first gradient boosting decision tree model parameters corresponding to the currently selected model; and whether the currently selected model is the last model in the selection range of the protein semantic model is confirmed; if so, go to step 66; if not, take the next model in the selection range of the protein semantic model as the new currently selected model and return to step 63 for training;

[0035] Step 66, confirming that the model training is completed.

[0036] Preferably, configuring the model parameters of the free energy change prediction model according to the model selection identifier specifically includes:

[0037] The model corresponding to the identification value of the model selection identifier in the protein semantic model selection range is used as the currently selected model of the protein structure feature recognition module; and the first protein semantic model parameters and the first gradient boosting decision tree model parameters corresponding to the currently selected model stored locally are extracted as the corresponding first and second model parameters; and based on the first model parameters, the model parameters of the protein structure feature recognition module of the free energy change prediction model are configured; and based on the second model parameters, the model parameters of the free energy change prediction module of the free energy change prediction model are configured.

[0038] Preferably, the performing single sequence free energy change prediction based on the model application data and the free energy change prediction model and performing stability level evaluation based on the prediction results specifically includes:

[0039] The first original protein sequence and the first mutant protein sequence of the model application data are input into the free energy change prediction model to predict the protein free energy change to obtain the corresponding first predicted free energy change data; and the preset first correspondence table is queried according to the first predicted free energy change data, and the first stability level field of the first correspondence record in which the first free energy value range field in the table satisfies the first predicted free energy change data is extracted as the corresponding first stability level; and the first stability level is output as the result of this stability level assessment; the first correspondence table is a correspondence table for reflecting the correspondence between the free energy value range and the stability level; the first correspondence table includes multiple first correspondence records; the first correspondence record includes the first free energy value range field and the first stability level field; the smaller the free energy value of the first free energy value range field, the higher the stability level of the corresponding first stability level field.

[0040] Preferably, the performing of multi-sequence free energy change prediction according to the model application data and the free energy change prediction model, performing stability grade evaluation according to the prediction results, and performing multi-sequence stability ranking according to the evaluation results specifically includes:

[0041] A group of corresponding protein sequence pairs are formed by each of the second mutant protein sequences and the second original protein sequence of the model application data; and the second original protein sequence and the second mutant protein sequence of each of the protein sequence pairs are input into the free energy change prediction model to perform protein free energy change prediction to obtain corresponding second predicted free energy change data; and according to each of the second predicted free energy change data, the preset first correspondence table is queried, and the first stability level field of the first correspondence record in the table whose first free energy value range field satisfies the current second predicted free energy change data is extracted as the corresponding second stability level; and all the second mutant protein sequences are sorted in descending order of the second stability level to obtain a corresponding mutant protein sequence sorting queue; and the mutant protein sequence sorting queue is output as the stability sorting result of this time.

[0042] A second aspect of the embodiment of the present invention provides a device for implementing the processing method of the protein free energy change prediction model described in the first aspect, the device comprising: a model construction module, a model training module and a model application module;

[0043] The model building module is used to build a free energy change prediction model; the free energy change prediction model is used to predict protein free energy changes based on the input original and mutant protein sequences; the free energy change prediction model includes a protein structure feature recognition module, a feature splicing module and a free energy change prediction module;

[0044] The model training module is used to construct a pre-training data set and a prediction training data set as the corresponding first and second data sets; and pre-train the protein structure feature recognition module based on the first data set; after the pre-training is completed, the free energy change prediction model is trained based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module;

[0045] The model application module is used to receive a model selection identifier, a model application mode and corresponding model application data input by a user after model training is completed; and configure model parameters of the free energy change prediction model according to the model selection identifier; and identify the model application mode; if the model application mode is the first mode, a single sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result; if the model application mode is the second mode, a multi-sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result, and a multi-sequence stability ranking is performed according to the assessment result; the model application mode includes a first mode and a second mode; when the model application mode is the first mode, the corresponding model application data includes a first original protein sequence and a first mutant protein sequence; when the model application mode is the second mode, the corresponding model application data includes a second original protein sequence and multiple second mutant protein sequences.

[0046] A third aspect of an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0047] The processor is used to be coupled to the memory, read and execute instructions in the memory, so as to implement the method steps described in the first aspect above;

[0048] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

[0049] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions. When the computer instructions are executed by a computer, the computer executes the instructions of the method described in the first aspect above.

[0050] The embodiment of the present invention provides a processing method, device, electronic device and computer-readable storage medium for a protein free energy change prediction model; the present invention first constructs a free energy change prediction model that can predict protein free energy change by using a protein semantic model that can perform multi-dimensional (one-dimensional, two-dimensional, three-dimensional) molecular structure feature recognition on a one-dimensional molecular sequence and a gradient boosting decision tree model; then, the protein semantic model is pre-trained so that it can obtain more and more accurate multi-dimensional molecular structure features to improve the generalization of the model; then, after the model parameters of the protein semantic model are solidified, the gradient boosting decision tree model is trained based on a supervised model training method to improve the prediction accuracy of the model; in addition, the present invention also provides multiple groups of model selection (ESM-1b model, ESM-1v ... Model, ESM-if1 model, ESM2 model and ProtBERT model), and train a set of corresponding protein semantic model parameters and gradient boosting decision tree model parameters for each model selection during model training, while improving the model prediction accuracy, it can also maintain good robustness; in addition, the present invention also provides two types of application modes for users to use: 1) In the first mode, the free energy change of the original-mutated protein sequence pair input by the user is predicted and the stability of the mutated protein is graded according to the prediction results; 2) In the second mode, the free energy change of an original protein sequence input by the user and its corresponding multiple mutant protein sequences are predicted separately, and the stability of each mutated protein is graded according to the prediction results, and the mutated protein sequences are sorted according to the evaluation results. The free energy change prediction model provided by the present invention predicts free energy changes based on multidimensional molecular structure characteristics, which improves the prediction accuracy; the present invention combines pre-training and supervised training to improve the generalization of the model; the model selection mechanism of the present invention improves the flexibility and robustness of the model; the two types of application modes of the present invention improve the convenience of using the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A schematic diagram of a method for processing a protein free energy change prediction model provided in Example 1 of the present invention;

[0052] Figure 2 A module structure diagram of a free energy change prediction model provided in Example 1 of the present invention;

[0053] Figure 3 A module structure diagram of a processing device for a protein free energy change prediction model provided in the second embodiment of the present invention;

[0054] Figure 4 A schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0056] Embodiment 1 of the present invention provides a method for processing a protein free energy change prediction model, such as Figure 1 As shown in the schematic diagram of a processing method of a protein free energy change prediction model provided in Example 1 of the present invention, the method mainly comprises the following steps:

[0057] Step 1, construct a free energy change prediction model;

[0058] The free energy change prediction model is used to predict protein free energy changes based on the input original and mutant protein sequences; here, the original protein sequence is the molecular sequence of the original protein, the mutant protein sequence is the molecular sequence of the mutant protein, and the mutant protein is the protein obtained by protein mutation of the original protein; the mutation types of protein mutation include substitution mutation, deletion mutation and insertion mutation; the mutant protein contains one or more mutant residues, and the mutant residues are the residues generated by protein mutation;

[0059] like Figure 2 As shown in the module structure diagram of the free energy change prediction model provided in the first embodiment of the present invention, the free energy change prediction model includes a protein structure feature recognition module, a feature splicing module and a free energy change prediction module; the input end of the protein structure feature recognition module is connected to the input end of the free energy change prediction model, and the output end is connected to the input end of the feature splicing module; the output end of the feature splicing module is connected to the input end of the free energy change prediction module; the output end of the free energy change prediction module is connected to the output end of the free energy change prediction model;

[0060] The protein structure feature recognition module is a protein semantic model implemented based on the transformer model structure; the protein semantic model selection range of the protein structure feature recognition module includes the ESM-1b model, ESM-1v model, ESM-if1 model, ESM2 model and ProtBERT model; the protein structure feature recognition module is used to perform protein semantic feature recognition processing on the input original and mutant protein sequences respectively to obtain the corresponding original and mutant sequence feature tensors and send them to the feature splicing module;

[0061] Here, the above models are all public protein semantic models, and their corresponding training data sets and model pre-training methods based on unsupervised methods are also public. The present embodiment will subsequently pre-train each model based on its public training data sets and training methods; any of the above protein semantic models used in the embodiment of the present invention can identify more multi-dimensional (one-dimensional, two-dimensional, three-dimensional) molecular structure features based on the input one-dimensional molecular sequence; in addition, the protein semantic model selection range of the protein structure feature recognition module in the embodiment of the present invention can also be based on application requirements. Model deletion or introduction of new models, the collection method and training method of the training data set of the newly added model are similar to those of the above models, and are also implemented based on the model training data set and model training method of the model itself;

[0062] The feature splicing module is used to perform feature tensor splicing on the original and mutant sequence feature tensors to obtain the corresponding spliced ​​feature tensors and send them to the free energy change prediction module; here, the feature splicing module of the embodiment of the present invention can be implemented based on a simple linear network, or based on a simple tensor splicing function;

[0063] The free energy change prediction module is implemented based on the Gradient Boosting Decision Tree (GBDT) model; the free energy change prediction module is used to perform free energy change regression prediction based on the spliced ​​feature tensor and output the corresponding predicted free energy change data.

[0064] Step 2, constructing a pre-training data set and a prediction training data set as the corresponding first and second data sets; and pre-training the protein structure feature recognition module based on the first data set; after the pre-training is completed, the free energy change prediction model is trained based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module;

[0065] Specifically comprising: step 21, constructing a pre-training data set and a prediction training data set as corresponding first and second data sets;

[0066] The first data set includes a plurality of first model data sets; the first model data sets correspond one-to-one to the models in the selection range of the protein semantic model; the second data set includes a plurality of first data records; the first data records include a first training original protein sequence, a first training mutant protein sequence and a first label free energy change data;

[0067] Specifically, it includes: step 211, taking the model training data sets conventionally used by each model in the protein semantic model selection range of the protein structure feature recognition module as the corresponding first model data set; and forming the corresponding first data set from all the obtained first model data sets;

[0068] Among them, the model training datasets used by each model consist of multiple protein sequences;

[0069] Here, as mentioned above, the models in the protein semantic model selection range of the protein structure feature recognition module are specifically the ESM-1b model, the ESM-1v model, the ESM-if1 model, the ESM2 model and the ProtBERT model, and the model training data sets of these models, i.e., the first model data sets, are all public;

[0070] Step 212, extracting each original-mutated protein sequence pair with free energy change information from the KAGGLE database and the VariBench database as the corresponding first training original protein sequence and the first training mutant protein sequence, and using the corresponding free energy change information as the corresponding first label free energy change data, and forming a corresponding first data record from the obtained first training original protein sequence, the first training mutant protein sequence and the first label free energy change data; and forming a corresponding first original data set from all the obtained first data records; and performing data record deduplication processing on the first original data set to obtain a corresponding second original data set; and performing data record deletion processing on the first data record whose first label free energy change data does not meet the preset free energy change value range in the second original data set to obtain a corresponding third original data set; and performing data enhancement processing on the third original data set based on a sliding window method to obtain a corresponding second data set;

[0071] Here, the well-known KAGGLE database and VariBench database contain a large number of protein molecule sequences, among which there are a large number of original-mutated protein sequence pairs with free energy change information. The data information of such sequence pairs includes: original protein sequence, mutant protein sequence and corresponding free energy change information. The embodiment of the present invention extracts the original-mutated protein sequence pairs with free energy change information from the KAGGLE database and the VariBench database to form a first original data set;

[0072] After obtaining the first original data set, the embodiment of the present invention first performs deduplication processing on it to obtain the second original data set, that is, the first data records with high similarity (the original protein sequence is highly similar, the mutant protein sequence is highly similar, and the free energy change information is highly similar) in the data set are filtered out;

[0073] After obtaining the second original data set, the embodiment of the present invention filters out the first data records whose free energy change information does not meet the preset requirements to obtain a third original data set; here, the preset requirements of the embodiment of the present invention are the free energy change value range, and the free energy change value range is usually 0<|DDG|≤10;

[0074] It can be seen from the disclosed free energy calculation method that when calculating the free energy of a molecular structure, point-by-point free energy calculation is conventionally adopted and the free energy calculation results of all nodes are processed by weighted full addition. When performing point-by-point calculation, a domain space is constructed with the current point (atom or residue) as the center and the free energy of the current point is calculated based on the associated objects (atoms or residues) in the domain space. In the molecular sequence, this domain space is also called a sliding window sequence centered on the current point, that is, if the size of the domain space of a certain point or the length of the sliding window sequence is adjusted under the premise that the molecular sequence remains unchanged, different free energy calculation results will be generated; and the free energy change is the difference between the free energy corresponding to the mutated protein sequence and the free energy corresponding to the original protein sequence. Then, when the original protein sequence, the mutated protein sequence, and the positions of the mutated residues on the mutated protein sequence are all known, different free energy change outputs can be generated by modulating the sliding window sequence length of the mutated residues; the embodiment of the present invention is based on this principle to achieve training data enhancement, that is: data enhancement processing is performed on the third original data set based on the sliding window method to obtain the corresponding second data set, specifically including:

[0075] Step A1, initializing the first enhanced data set to be empty; and taking the first first data record in the third original data set as the corresponding current record;

[0076] Step A2, taking the currently recorded first training original protein sequence, first training mutant protein sequence and first label free energy variable data as the corresponding first sequence, second sequence and first free energy variable data DDG;

[0077] Step A3, traversing each mutant residue in the second sequence; and during the traversal, taking the currently traversed mutant residue as the corresponding first residue Z i ; and based on a preset first number X sliding window sizes L x , in the second sequence with the first residue Z i As the center, the front and rear N=(L x -1) / 2 residues are associated to obtain the corresponding first associated residue sequence S x {Z i-N …Z i-1 ,Z i ,Z i+1 …Z i+N}; and according to each first associated residue sequence S x , the first sequence and the second sequence dynamically modulate the first free energy variable DDG data to obtain the corresponding second free energy variable data DDG x ; and each second free energy change data DDG xas a new first label free energy change data; and the first training original protein sequence corresponding to the first and second sequences, the first training mutant protein sequence and each new first label free energy change data form a new first data record; and the obtained X new first data records form a sequence corresponding to the current mutant residue Z i The corresponding first data subset; and at the end of the traversal, all the obtained first data subsets are added to the first enhanced data set;

[0078] Wherein, the residue index i is an integer greater than 1, and the residue index i is the residue ranking index of the corresponding mutant residue in the second sequence; the first number X is an integer greater than 0; each sliding window size L x is an odd number greater than or equal to 3; N is an integer greater than or equal to 1;

[0079] Step A4, identifying whether the current record is the last first data record in the third original data set; if so, going to step A5; if not, taking the next first data record in the third original data set as the new current record and going to step A2;

[0080] Step A5, merging the first enhanced data set and the third original data set to obtain a corresponding first merged data set; performing data record deduplication processing on the first merged data set; and using the first merged data set after deduplication as the corresponding second data set;

[0081] Step 22, pre-training a protein structure feature recognition module based on the first data set;

[0082] Specifically, the method includes: training the corresponding model in the protein semantic model selection range based on each first model data set in the first data set, and extracting the model parameters of the current model at the end of the current model training and storing them as the corresponding first protein semantic model parameters; and confirming the end of pre-training when all models in the protein semantic model selection range have completed training;

[0083] Here, as mentioned above, the various models in the protein semantic model selection range of the protein structure feature recognition module are specifically the ESM-1b model, the ESM-1v model, the ESM-if1 model, the ESM2 model and the ProtBERT model, and the model pre-training methods of these models based on the unsupervised method are also public. The present embodiment pre-trains the models based on the model training data sets disclosed by each model, i.e., the first model data set and the training method; it should be noted that in order to ensure that each model can achieve the same or similar model performance level after training, the embodiment of the present invention will train the above-mentioned protein semantic models separately based on the same model training requirements;

[0084] Step 23, after the pre-training is completed, the free energy change prediction model is trained based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module;

[0085] Specifically, it includes: step 231, performing parameter solidification processing on each model within the protein semantic model selection range of the protein structure feature recognition module;

[0086] Step 232, taking the first model in the protein semantic model selection range as the currently selected model of the protein structure feature recognition module;

[0087] Step 233, traverse each first data record in the second data set; and during the traversal, take the currently traversed first data record as the corresponding current data record; and use the currently selected model to perform protein semantic feature recognition processing on the first training original protein sequence and the first training mutant protein sequence of the current data record to obtain the corresponding first original sequence feature tensor and the first mutant sequence feature tensor; and use the feature splicing module to perform feature tensor splicing on the first original sequence feature tensor and the first mutant sequence feature tensor to obtain the corresponding first spliced ​​feature tensor; and use the first label free energy variable data of the current data record as the corresponding second label free energy variable data, and use the first spliced ​​feature tensor and the second label free energy variable data to form a corresponding second data record; and at the end of the traversal, all the obtained second data records form a corresponding third data set;

[0088] Here, the third data set includes a plurality of second data records, and the second data record includes a first concatenated feature tensor and a second label free energy variable data; when the free energy variable prediction module is trained based on the third data set in a subsequent step, the first concatenated feature tensor is used as a model input for free energy variable regression prediction, and the second label free energy variable data is used as label data to calculate the loss and guide the model optimization during the regression prediction process;

[0089] Step 234, performing gradient boosting decision tree model training on the free energy change prediction module based on the third data set in a preset K-Fold cross validation manner;

[0090] Here, when the embodiment of the present invention performs training based on the K-Fold cross-validation method, K is set to 5 by default;

[0091] Step 235, when the training of the gradient boosting decision tree model is completed, the model parameters of the free energy change prediction module are extracted and stored as the first gradient boosting decision tree model parameters corresponding to the currently selected model; and whether the currently selected model is the last model in the protein semantic model selection range is confirmed; if so, go to step 236; if not, the next model in the protein semantic model selection range is used as the new currently selected model and return to step 233 for training;

[0092] Step 236, confirm that the model training is completed.

[0093] In summary, from the above steps 21-23, it can be seen that the embodiment of the present invention will train and save a set of corresponding first protein semantic model parameters and first gradient boosting decision tree model parameters for each model selection during model training.

[0094] Step 3, after the model training is completed, the model selection identifier, model application mode and corresponding model application data input by the user are received; and the model parameters of the free energy change prediction model are configured according to the model selection identifier; and the model application mode is identified; if the model application mode is the first mode, a single sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level evaluation is performed according to the prediction result; if the model application mode is the second mode, a multi-sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level evaluation is performed according to the prediction result, and a multi-sequence stability ranking is performed according to the evaluation result;

[0095] Specifically, it includes: step 31, receiving a model selection identifier, a model application mode and corresponding model application data input by a user;

[0096] Among them, the identification value range of the model selection identification is composed of multiple first model identifications; the first model identifications correspond one-to-one to the models in the protein semantic model selection range; the model application mode includes a first mode and a second mode; when the model application mode is the first mode, the corresponding model application data includes a first original protein sequence and a first mutant protein sequence; when the model application mode is the second mode, the corresponding model application data includes a second original protein sequence and multiple second mutant protein sequences;

[0097] Step 32, configuring model parameters of the free energy change prediction model according to the model selection identifier;

[0098] Specifically, it includes: taking the model corresponding to the identification value of the model selection identification in the protein semantic model selection range as the currently selected model of the protein structure feature recognition module; extracting the first protein semantic model parameters and the first gradient boosting decision tree model parameters corresponding to the currently selected model stored locally as the corresponding first and second model parameters; and configuring the model parameters of the protein structure feature recognition module of the free energy change prediction model based on the first model parameters; and configuring the model parameters of the free energy change prediction module of the free energy change prediction model based on the second model parameters;

[0099] Step 33, and identifying the model application mode;

[0100] Step 34, if the model application mode is the first mode, a single sequence free energy change prediction is performed based on the model application data and the free energy change prediction model, and a stability level assessment is performed based on the prediction result;

[0101] Specifically, it includes: step 341, inputting the first original protein sequence and the first mutant protein sequence of the model application data into the free energy change prediction model to perform protein free energy change prediction to obtain corresponding first predicted free energy change data;

[0102] Step 342, querying a preset first correspondence table according to the first predicted free energy variable data, extracting the first stability level field of the first correspondence record whose first free energy value range field satisfies the first predicted free energy variable data as the corresponding first stability level;

[0103] Among them, the first correspondence table is a correspondence table for reflecting the correspondence between the free energy value range and the stability level; the first correspondence table includes multiple first correspondence records; the first correspondence record includes a first free energy value range field and a first stability level field; the smaller the free energy value of the first free energy value range field, the higher the stability level of the corresponding first stability level field;

[0104] Step 343, and outputting the first stability level as the stability level evaluation result of this time;

[0105] Step 35, if the model application mode is the second mode, then perform multi-sequence free energy change prediction according to the model application data and the free energy change prediction model, perform stability level evaluation according to the prediction results, and perform multi-sequence stability ranking according to the evaluation results;

[0106] Specifically, it includes: step 351, forming a group of corresponding protein sequence pairs from each second mutant protein sequence of the model application data and the second original protein sequence;

[0107] Step 352, inputting the second original protein sequence and the second mutant protein sequence of each protein sequence pair into the free energy change prediction model to perform protein free energy change prediction to obtain corresponding second predicted free energy change data;

[0108] Step 353, querying the preset first correspondence table according to each second predicted free energy variable data, extracting the first stability level field of the first correspondence record whose first free energy value range field satisfies the current second predicted free energy variable data as the corresponding second stability level;

[0109] Step 354, sorting all second mutant protein sequences in descending order of the second stability level to obtain a corresponding mutant protein sequence sorting queue;

[0110] Step 355, and output the mutant protein sequence sorting queue as the stability sorting result.

[0111] Figure 3 The module structure diagram of a processing device for a protein free energy change prediction model provided in the second embodiment of the present invention, the device is a terminal device or server that implements the aforementioned method embodiment, and can also be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiment, for example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 3 As shown, the device includes: a model building module 201, a model training module 202 and a model application module 203.

[0112] The model building module 201 is used to build a free energy change prediction model; the free energy change prediction model is used to predict protein free energy changes based on the input original and mutant protein sequences; the free energy change prediction model includes a protein structure feature recognition module, a feature splicing module and a free energy change prediction module.

[0113] The model training module 202 is used to construct a pre-training data set and a prediction training data set as the corresponding first and second data sets; and pre-train the protein structure feature recognition module based on the first data set; after the pre-training is completed, the free energy change prediction model is trained based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module.

[0114] The model application module 203 is used to receive the model selection identifier, model application mode and corresponding model application data input by the user after the model training is completed; and configure the model parameters of the free energy change prediction model according to the model selection identifier; and identify the model application mode; if the model application mode is the first mode, a single sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result; if the model application mode is the second mode, a multi-sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result, and a multi-sequence stability ranking is performed according to the assessment result; the model application mode includes a first mode and a second mode; when the model application mode is the first mode, the corresponding model application data includes a first original protein sequence and a first mutant protein sequence; when the model application mode is the second mode, the corresponding model application data includes a second original protein sequence and multiple second mutant protein sequences.

[0115] The processing device of a protein free energy change prediction model provided in an embodiment of the present invention can execute the method steps in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here.

[0116] It should be noted that it should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the model building module can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The function of the above-mentioned module is determined. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in the processor element or instructions in the form of software.

[0117] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more digital signal processors (DSP), or one or more field programmable gate arrays (FPGA). For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0118] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the above method embodiments are generated. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above-mentioned computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above-mentioned computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) methods. The above-mentioned computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The above-mentioned available medium can be a magnetic medium (such as a floppy disk, a hard disk, a tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0119] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server for implementing the method of the aforementioned embodiment, or may be a terminal device or a server for implementing the method of the aforementioned embodiment connected to the aforementioned terminal device or server. Figure 4As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver 303. Various instructions may be stored in the memory 302 to complete various processing functions and implement the processing steps described in the aforementioned embodiment method. Preferably, the electronic device involved in the embodiment of the present invention also includes: a power supply 304, a system bus 305 and a communication port 306. The system bus 305 is used to realize the communication connection between components. The above-mentioned communication port 306 is used for connecting and communicating between the electronic device and other peripherals.

[0120] exist Figure 4 The system bus 305 mentioned in the figure can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to realize the communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include random access memory (RAM), and may also include non-volatile memory (Non-Volatile Memory), such as at least one disk storage.

[0121] The above-mentioned processor can be a general-purpose processor, including a central processing unit CPU, a network processor (Network Processor, NP), a graphics processing unit (Graphics Processing Unit, GPU), etc.; it can also be a digital signal processor DSP, an application-specific integrated circuit ASIC, a field programmable gate array FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0122] It should be noted that an embodiment of the present invention further provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the method and processing process provided in the above embodiments.

[0123] An embodiment of the present invention further provides a chip for executing instructions, wherein the chip is used to execute the processing steps described in the aforementioned method embodiment.

[0124] The embodiment of the present invention provides a processing method, device, electronic device and computer-readable storage medium for a protein free energy change prediction model; the present invention first constructs a free energy change prediction model that can predict protein free energy change by using a protein semantic model that can perform multi-dimensional (one-dimensional, two-dimensional, three-dimensional) molecular structure feature recognition on a one-dimensional molecular sequence and a gradient boosting decision tree model; then, the protein semantic model is pre-trained so that it can obtain more and more accurate multi-dimensional molecular structure features to improve the generalization of the model; then, after the model parameters of the protein semantic model are solidified, the gradient boosting decision tree model is trained based on a supervised model training method to improve the prediction accuracy of the model; in addition, the present invention also provides multiple groups of model selection (ESM-1b model, ESM-1v ... Model, ESM-if1 model, ESM2 model and ProtBERT model), and train a set of corresponding protein semantic model parameters and gradient boosting decision tree model parameters for each model selection during model training, while improving the model prediction accuracy, it can also maintain good robustness; in addition, the present invention also provides two types of application modes for users to use: 1) In the first mode, the free energy change of the original-mutated protein sequence pair input by the user is predicted and the stability of the mutated protein is graded according to the prediction results; 2) In the second mode, the free energy change of an original protein sequence input by the user and its corresponding multiple mutant protein sequences are predicted separately, and the stability of each mutated protein is graded according to the prediction results, and the mutated protein sequences are sorted according to the evaluation results. The free energy change prediction model provided by the present invention predicts free energy changes based on multidimensional molecular structure characteristics, which improves the prediction accuracy; the present invention combines pre-training and supervised training to improve the generalization of the model; the model selection mechanism of the present invention improves the flexibility and robustness of the model; the two types of application modes of the present invention improve the convenience of using the model.

[0125] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0126] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0127] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for processing a protein free energy change prediction model, characterized in that: The method comprises: Constructing a free energy change prediction model; the free energy change prediction model is used to predict protein free energy changes according to the input original and mutant protein sequences; the free energy change prediction model includes a protein structure feature recognition module, a feature splicing module and a free energy change prediction module; Constructing a pre-training data set and a prediction training data set as the corresponding first and second data sets; and pre-training the protein structure feature recognition module based on the first data set; after the pre-training is completed, model training is performed on the free energy change prediction model based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module; After the model training is completed, the model selection identifier, model application mode and corresponding model application data input by the user are received; and the model parameters of the free energy change prediction model are configured according to the model selection identifier; and the model application mode is identified; if the model application mode is the first mode, a single sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result; if the model application mode is the second mode, a multi-sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result, and a multi-sequence stability ranking is performed according to the assessment result; the model application mode includes a first mode and a second mode; when the model application mode is the first mode, the corresponding model application data includes a first original protein sequence and a first mutant protein sequence; when the model application mode is the second mode, the corresponding model application data includes a second original protein sequence and a plurality of second mutant protein sequences; The single sequence free energy change prediction is specifically as follows: inputting the first original protein sequence and the first mutant protein sequence of the model application data into the free energy change prediction model to perform protein free energy change prediction to obtain corresponding first predicted free energy change data; The multi-sequence free energy change prediction is specifically as follows: each of the second mutant protein sequences and the second original protein sequence of the model application data forms a group of corresponding protein sequence pairs; and the second original protein sequence and the second mutant protein sequence of each of the protein sequence pairs are input into the free energy change prediction model to perform protein free energy change prediction to obtain corresponding second predicted free energy change data.

2. The method for processing the protein free energy change prediction model according to claim 1, characterized in that: The original protein sequence is the molecular sequence of the original protein, the mutant protein sequence is the molecular sequence of the mutant protein, and the mutant protein is a protein obtained by protein mutation of the original protein; the mutation types of the protein mutation include substitution mutation, deletion mutation and insertion mutation; the mutant protein carries one or more mutant residues, and the mutant residues are residues generated by the protein mutation; The input end of the protein structure feature recognition module is connected to the input end of the free energy change prediction model, and the output end is connected to the input end of the feature splicing module; the output end of the feature splicing module is connected to the input end of the free energy change prediction module; the output end of the free energy change prediction module is connected to the output end of the free energy change prediction model; The protein structure feature recognition module is a protein semantic model implemented based on the transformer model structure; the protein semantic model selection range of the protein structure feature recognition module includes ESM-1b model, ESM-1v model, ESM-if1 model, ESM2 model and ProtBERT model; The free energy change prediction module is implemented based on a gradient boosting decision tree model; The protein structure feature recognition module is used to perform protein semantic feature recognition processing on the input original and mutant protein sequences respectively to obtain the corresponding original and mutant sequence feature tensors and send them to the feature splicing module; the feature splicing module is used to perform feature tensor splicing on the original and mutant sequence feature tensors to obtain the corresponding spliced ​​feature tensors and send them to the free energy change prediction module; the free energy change prediction module is used to perform free energy change regression prediction based on the spliced ​​feature tensor to output the corresponding predicted free energy change data; The first data set includes a plurality of first model data sets; the first model data sets correspond one-to-one to the models in the selection range of the protein semantic model; The second data set includes a plurality of first data records; the first data records include a first training original protein sequence, a first training mutant protein sequence and a first label free energy variation data; The identification value range of the model selection identification is composed of multiple first model identifications; the first model identifications correspond one-to-one to the models in the protein semantic model selection range.

3. The method for processing the protein free energy change prediction model according to claim 2, characterized in that: The constructing of the pre-training data set and the prediction training data set as the corresponding first and second data sets specifically includes: The model training data sets conventionally used by each model in the protein semantic model selection range of the protein structure feature recognition module are used as the corresponding first model data sets; and all the obtained first model data sets are used to form the corresponding first data set; the model training data sets used by each model are composed of multiple protein sequences; The original-mutated protein sequence pairs with free energy change information from the KAGGLE database and the VariBench database are extracted as the corresponding first training original protein sequence and the first training mutant protein sequence, and the corresponding free energy change information is used as the corresponding first label free energy change data, and the first training original protein sequence, the first training mutant protein sequence and the first label free energy change data are used to form the corresponding first data record; and all the first data records are used to form the corresponding first original data set; and the first original data set is deduplicated to obtain the corresponding second original data set; and the first data records in the second original data set whose first label free energy change data does not meet the preset free energy change value range are deleted to obtain the corresponding third original data set; and the third original data set is enhanced based on the sliding window method to obtain the corresponding second data set.

4. The method for processing the protein free energy change prediction model according to claim 3, characterized in that: The performing data enhancement processing on the third original data set based on a sliding window method to obtain the corresponding second data set specifically includes: Step 41, initializing the first enhanced data set to be empty; and taking the first first data record in the third original data set as the corresponding current record; Step 42, using the first training original protein sequence, the first training mutant protein sequence and the first label free energy variable data currently recorded as the corresponding first sequence, second sequence and first free energy variable data DDG; Step 43, traversing each of the mutated residues in the second sequence; and during the traversal, taking the currently traversed mutated residue as the corresponding first residue Z i ; and based on a preset first number X sliding window sizes L x , in the second sequence, the first residue Z i Take the center as the center and N=(L x -1) / 2 residues are associated to obtain the corresponding first associated residue sequence S x {Z i-N …Z i-1 ,Z i ,Z i+1 …Z i+N }; and according to each of the first associated residue sequences S x , the first sequence and the second sequence dynamically modulate the first free energy variable DDG data to obtain the corresponding second free energy variable data DDG x ; And each of the second free energy variable data DDG x as a new first label free energy variable data; and the first training original protein sequence corresponding to the first and second sequences, the first training mutant protein sequence and each of the new first label free energy variable data form a new first data record; and the obtained X new first data records form a record corresponding to the current mutant residue Z i The first data subset corresponding to the first sequence; and at the end of the traversal, all the first data subsets obtained are added to the first enhanced data set; the residue index i is an integer greater than 1, and the residue index i is the residue ranking index of the corresponding mutation residue in the second sequence; the first number X is an integer greater than 0; each sliding window size L x is an odd number greater than or equal to 3; N is an integer greater than or equal to 1; Step 44, identifying whether the current record is the last first data record in the third original data set; if so, going to step 45; if not, taking the next first data record in the third original data set as the new current record and going to step 42; Step 45, merging the first enhanced data set with the third original data set to obtain a corresponding first merged data set; performing data record deduplication processing on the first merged data set; and using the first merged data set after deduplication as the corresponding second data set.

5. The method for processing the protein free energy change prediction model according to claim 2, characterized in that: The pre-training of the protein structure feature recognition module based on the first data set specifically includes: Based on each of the first model data sets in the first data set, the corresponding model in the protein semantic model selection range is trained, and when the current model training is completed, the model parameters of the current model are extracted and stored as the corresponding first protein semantic model parameters; and the pre-training is confirmed to be completed when all models in the protein semantic model selection range have completed training.

6. The method for processing the protein free energy change prediction model according to claim 5, characterized in that: The performing model training on the free energy change prediction model based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module specifically includes: Step 61, performing parameter solidification processing on each model within the protein semantic model selection range of the protein structure feature recognition module; Step 62, taking the first model within the selection range of the protein semantic model as the currently selected model of the protein structure feature recognition module; Step 63, traverse each of the first data records in the second data set; and during the traversal, use the first data record currently traversed as the corresponding current data record; and use the currently selected model to perform protein semantic feature recognition processing on the first training original protein sequence and the first training mutant protein sequence of the current data record to obtain the corresponding first original sequence feature tensor and the first mutant sequence feature tensor; and use the feature splicing module to perform feature tensor splicing on the first original sequence feature tensor and the first mutant sequence feature tensor to obtain the corresponding first spliced ​​feature tensor; and use the first label free energy variable data of the current data record as the corresponding second label free energy variable data, and use the first spliced ​​feature tensor and the second label free energy variable data to form a corresponding second data record; and at the end of the traversal, all the obtained second data records form a corresponding third data set; Step 64, performing gradient boosting decision tree model training on the free energy change prediction module based on the third data set in a preset K-Fold cross validation manner; Step 65, when the training of the gradient boosting decision tree model is completed, the model parameters of the free energy change prediction module are extracted and stored as the first gradient boosting decision tree model parameters corresponding to the currently selected model; and whether the currently selected model is the last model in the selection range of the protein semantic model is confirmed; if so, go to step 66; if not, take the next model in the selection range of the protein semantic model as the new currently selected model and return to step 63 for training; Step 66, confirming that the model training is completed.

7. The method for processing the protein free energy change prediction model according to claim 6, characterized in that: The configuring of model parameters of the free energy change prediction model according to the model selection identifier specifically includes: The model corresponding to the identification value of the model selection identifier in the protein semantic model selection range is used as the currently selected model of the protein structure feature recognition module; and the first protein semantic model parameters and the first gradient boosting decision tree model parameters corresponding to the currently selected model stored locally are extracted as the corresponding first and second model parameters; and based on the first model parameters, the model parameters of the protein structure feature recognition module of the free energy change prediction model are configured; and based on the second model parameters, the model parameters of the free energy change prediction module of the free energy change prediction model are configured.

8. The method for processing the protein free energy change prediction model according to claim 2, characterized in that: The performing single sequence free energy change prediction according to the model application data and the free energy change prediction model and performing stability level evaluation according to the prediction results specifically includes: The first original protein sequence and the first mutant protein sequence of the model application data are input into the free energy change prediction model to perform protein free energy change prediction to obtain the corresponding first predicted free energy change data; and according to the first predicted free energy change data, the preset first correspondence table is queried, and the first stability level field of the first correspondence record in which the first free energy value range field in the table satisfies the first predicted free energy change data is extracted as the corresponding first stability level; and the first stability level is output as the result of this stability level assessment; the first correspondence table is a correspondence table for reflecting the correspondence between the free energy value range and the stability level; the first correspondence table includes multiple first correspondence records; the first correspondence record includes the first free energy value range field and the first stability level field; the smaller the free energy value of the first free energy value range field, the higher the stability level of the corresponding first stability level field.

9. The method for processing the protein free energy change prediction model according to claim 8, characterized in that: The method of predicting free energy changes of multiple sequences according to the model application data and the free energy change prediction model, evaluating the stability level according to the prediction results, and ranking the stability of multiple sequences according to the evaluation results specifically includes: A group of corresponding protein sequence pairs are formed by each of the second mutant protein sequences and the second original protein sequence of the model application data; and the second original protein sequence and the second mutant protein sequence of each of the protein sequence pairs are input into the free energy change prediction model to perform protein free energy change prediction to obtain the corresponding second predicted free energy change data; and according to each of the second predicted free energy change data, the preset first correspondence table is queried, and the first stability level field of the first correspondence record in the table whose first free energy value range field satisfies the current second predicted free energy change data is extracted as the corresponding second stability level; and all the second mutant protein sequences are sorted in descending order of the second stability level to obtain the corresponding mutant protein sequence sorting queue; and the mutant protein sequence sorting queue is output as the stability sorting result of this time.

10. A device for executing the method for processing the protein free energy change prediction model according to any one of claims 1 to 9, characterized in that: The device comprises: a model building module, a model training module and a model application module; The model building module is used to build a free energy change prediction model; the free energy change prediction model is used to predict protein free energy changes based on the input original and mutant protein sequences; the free energy change prediction model includes a protein structure feature recognition module, a feature splicing module and a free energy change prediction module; The model training module is used to construct a pre-training data set and a prediction training data set as the corresponding first and second data sets; and pre-train the protein structure feature recognition module based on the first data set; after the pre-training is completed, the free energy change prediction model is trained based on the second data set on the premise of solidifying the model parameters of the protein structure feature recognition module; The model application module is used to receive a model selection identifier, a model application mode and corresponding model application data input by a user after model training is completed; and configure model parameters of the free energy change prediction model according to the model selection identifier; and identify the model application mode; if the model application mode is the first mode, a single sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result; if the model application mode is the second mode, a multi-sequence free energy change prediction is performed according to the model application data and the free energy change prediction model, and a stability level assessment is performed according to the prediction result, and a multi-sequence stability ranking is performed according to the assessment result; the model application mode includes a first mode and a second mode; when the model application mode is the first mode, the corresponding model application data includes a first original protein sequence and a first mutant protein sequence; when the model application mode is the second mode, the corresponding model application data includes a second original protein sequence and multiple second mutant protein sequences.

11. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is used to couple with the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 9; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for predicting binding free energy of protein and ligand based on progressive neural network

    CN110910951A

  • Method and device for predicting binding free energy of protein and ligand molecules

    CN112466410A