A recognition system for naturally disordered functional regions

By designing an identification system that utilizes protein language model and multi-task learning framework, the problem of insufficient comprehensive and accurate prediction of protein disordered functional regions in the prior art is solved, and accurate identification and prediction of multiple disordered functional regions are achieved.

CN119724369BActive Publication Date: 2025-05-13SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510228071.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-13
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The prior art lacks annotation of ion or other small molecule binding functions when predicting disordered functional regions of proteins, and fails to effectively utilize protein language models in deep learning technology, resulting in insufficient comprehensive and accurate prediction results.

Method used

A natural disordered functional region recognition system was designed, and sequence listing vectors were obtained using protein language models, and disordered regions bound to proteins, nucleic acids, lipids, ions and other small molecules were predicted through a multi-task learning framework, a multi-scale feature processing module and a multi-function classifier, as well as a single feature processing module and a single classifier to predict flexible connection regions.

Benefits of technology

Accurate identification of disordered regions and flexible connecting regions in protein sequences is achieved, covering a variety of disordered functional regions, improving the comprehensiveness and stability of prediction, and significantly better than the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724369B_ABST
    Figure CN119724369B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of protein function prediction, and is an identification system for natural disordered functional regions. The system receives a protein sequence, converts it into a sequence representation vector based on a protein language model, and outputs it to the disordered binding region and flexible connection region prediction models respectively. The disordered binding region prediction model consists of a multi-scale feature processing module and a multifunctional classifier. By predicting the probability that each residue in the sequence is located in the disordered region and binds to five types of molecules such as proteins and nucleic acids, the disordered binding region and its binding molecule type are inferred. The flexible connection region prediction model consists of a single feature processing module and a single classifier, which can predict the probability that each residue in the sequence is located in the flexible connection region. The beneficial effects of the present invention: On the blind test data set provided by CAID, a platform for predicting and evaluating natural disordered proteins and their functions, the prediction accuracy ranks among the top, so this method can achieve accurate prediction of disordered functional regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of protein function prediction, and in particular relates to a recognition system for natural disordered functional regions. Background Art

[0002] With the rapid development of bioinformatics and structural biology, protein function prediction has become an important topic in modern biomedical research. Proteins are core macromolecules in living organisms, and the prediction of their functions is directly related to the study of disease mechanisms, the discovery of drug targets, and the development of disease diagnosis and treatment. Proteins are believed to rely on stable three-dimensional structures when they function, forming a sequence-structure-function research paradigm. However, studies have found that many proteins do not have a fixed structure under physiological conditions and are highly flexible. Their functions rely on disordered states. This type of protein is intrinsically disordered proteins / regions, referred to as IDPs / IDRs. They play a key role in multiple biological processes such as cell signal transduction, gene expression regulation, and protein-protein interactions, and they are closely related to many diseases. Therefore, accurately predicting the intrinsic disordered functional regions of proteins is of great significance for understanding the mechanisms of these diseases and developing new treatment strategies. About 3,000 intrinsically disordered proteins were studied in the Disprot database 9.4 version, and more than 1,000 and 300 annotations on binding functions and flexible connection functions were found, respectively. There are more than 200 million sequences in the Uniprot database, including reviewed and unreviewed data, which have not been annotated with IDPs / IDRs binding function and flexible connection function.

[0003] Researchers can study intrinsically disordered proteins through experimental methods such as nuclear magnetic resonance (NMR) spectroscopy, circular dichroism (CD) spectroscopy, fluorescence spectroscopy, mass spectrometry, etc. to obtain their structure, dynamic information, and interaction functions. The accuracy is relatively high, but it is time-consuming and labor-intensive, the experimental cost is high, and it is random. There are also molecular dynamics simulation methods to study the function and structure of intrinsically disordered proteins, which require a lot of computing resources, high computing costs, limited force fields, huge conformational sampling space, and time scale restrictions.

[0004] Computational methods are used to predict disordered functional regions of proteins, including a variety of tools. ANCHOR2 predicts disordered regions that bind to proteins based on biophysical scoring functions; MoRFCHiBi uses support vector machines and Bayesian rules, combined with physicochemical sequence properties and conservation information, to predict molecular recognition features. In addition, DisoRDPbind and DeepDISOBind use logistic regression and deep convolutional networks, respectively, to predict disordered regions that bind to RNA, DNA, and proteins; while flDPnn uses fully connected layers and random forest methods to process rich protein sequence information to predict disordered regions and flexible connection regions that bind to proteins, RNA, and DNA, respectively. DisoFLAG uses a protein language model and a graph-based interactive model to predict disordered regions and flexible connection regions that bind to proteins, DNA, RNA, ions, and lipids, respectively. Other tools such as DisoLipPred, APOD, and DFLpred focus on the rapid prediction of disordered regions and flexible connection regions that bind to lipids, respectively.

[0005] These computational methods are fast, but they also have some problems. Very few methods consider disordered regions bound to ions or other small molecules, and most methods lack the protein language model that uses deep learning technology. Machine learning-based methods consider residue, local window, and full sequence level information from a feature perspective, while the network framework based on deep learning methods hardly considers information at different levels of the sequence. It only uses a single network framework to process sequence features, and does not capture local and global information of the sequence. There are not many methods that use multi-task learning, and most of them predict a certain type of disordered functional region separately without considering multiple disordered functional regions together.

[0006] The reasons for this include the following: (1) There are few annotations on the binding functions of IDPs / IDRs to ions and other small molecules. The previous data in the Disprot database mainly focused on the disordered regions that bind to proteins and nucleic acids, and lacked research on the binding to ions and other small molecules. (2) The prediction of disordered functional regions was not linked to the protein language model, and the utilization rate was low. (3) The characteristics of different functions of intrinsically disordered proteins were not taken into account, and using only a single model could not capture all the information. Summary of the invention

[0007] In order to solve the problems existing in the prior art, the present invention provides a recognition system for naturally disordered functional regions.

[0008] The technical solution adopted by the present invention to solve the technical problem is as follows: a recognition system for naturally disordered functional regions, comprising:

[0009] A protein language model that receives a protein sequence and outputs a corresponding sequence representation vector;

[0010] A disordered binding region prediction model has a built-in multi-scale feature processing module and a multi-functional classifier for capturing local and global information. The output end of the multi-scale feature processing module is connected to the multi-functional classifier, and the multi-functional classifier has multiple classifiers connected in parallel. The input end of the disordered binding region prediction model is connected to the output end of the protein language model. The prediction model receives a protein sequence representation vector as a sequence feature, and performs feature processing and multiple disordered binding function identification in sequence through the multi-scale feature processing module and the multi-functional classifier, and outputs the probability that each residue in the protein sequence is located in the disordered region and binds to proteins, nucleic acids, lipids, ions and other small molecules respectively;

[0011] The flexible connection region prediction model has a built-in single feature processing module and a single classifier that capture global scale information. The output end of the single feature processing module is connected to the single classifier. The input end of the flexible connection region prediction model is connected to the output end of the protein language model. The prediction model receives the protein sequence representation vector as the sequence feature, performs feature processing and flexible connection function identification, and outputs the probability that each residue in the protein sequence is located in the flexible connection region.

[0012] Preferably, the protein language model is ProtT5-XL-U50; the one-dimensional convolutional network CNN and the bidirectional long short-term memory network BiLSTM in the multi-scale feature processing module are set in parallel; the single feature processing module is a bidirectional long short-term memory network BiLSTM; five classifiers in the multifunctional classifier are set in parallel and executed in sequence; and a single classifier is set to a classifier.

[0013] Preferably, each classifier has a built-in residual multilayer perceptron Residual MLP, and the residual multilayer perceptrons in different classifiers have consistent frameworks and hyperparameters, but different weights.

[0014] Preferably, the residual multilayer perceptron Residual MLP includes multiple linear layers and activation functions ReLU, and uses skip connections to enhance the training effect of the network.

[0015] Preferably, the residual multi-layer perceptron Residual MLP is connected to a normalized layer, a Sigmoid function is built in the normalized layer, and the Sigmoid function normalizes its output.

[0016] Preferably, the CNN and BiLSTM of the multi-scale feature processing module are followed by a linear layer. After receiving the output results of the CNN and BiLSTM, the linear layer performs dimensionality reduction, concatenates the dimensionality reduction results of the two, and outputs shared basic features.

[0017] Preferably, the one-dimensional convolutional network CNN includes: two one-dimensional CNN layers with a kernel size of 3 and a ReLU nonlinear activation function located in the middle; the BiLSTM consists of a forward LSTM and a reverse LSTM, and LSTM is a long short-term memory network.

[0018] Preferably, the protein language model ProtT5-XL-U50 generates a 1024-dimensional feature vector for each residue in the protein sequence, the input and output dimensions of the multi-scale feature processing module and the single feature processing module are both 1024, and the input dimension of the multi-functional classifier and the single classifier is 1024 and the output dimension is 1.

[0019] Preferably, the disordered combined region prediction model and the flexible connected region prediction model are trained using the pytorch framework with a batch size of 1, SGD optimizer and a learning rate of 5e -4 .

[0020] Preferably, the training process of the disordered combined region prediction model and the flexible connected region prediction model is as follows:

[0021] S1. Extract disordered protein sequences containing six functions including protein binding, nucleic acid binding, lipid binding, ion binding, other small molecule binding and flexible connection from the DisProt database, perform data preprocessing, and divide them into training set, validation set and test set;

[0022] S2. Input the sequences in the training set into the protein language model, obtain a 1024-dimensional feature vector for each residue, and obtain a sequence representation vector consisting of an L*1024-dimensional feature matrix for each sequence, where L represents the length of the protein sequence;

[0023] S3. Select a prediction model, input the sequence representation vector into the disordered binding region model or the flexible connection region model, calculate the result of the corresponding prediction model through forward propagation, and use the loss function to evaluate the error, then calculate the gradient through back propagation, and use the stochastic gradient descent method to update the weight. In the framework of pytorch, iterate the above process to optimize the performance. When the loss function on the validation set is the smallest, the model weight corresponding to this iteration is the optimal weight, and the training is completed.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows: the recognition system of natural disordered functional regions utilizes sequence representation vectors obtained from the protein language model, and adopts a multi-task learning framework, a multi-scale feature processing module and a multifunctional classifier to predict disordered regions bound to proteins, nucleic acids, lipids, ions and other small molecules, and a single feature processing module and a single classifier that capture global information to predict flexible connection regions. Good results have been achieved on multiple data sets, and accurate prediction of functional annotations can be achieved with more comprehensive prediction functions. The stability is better than the existing prediction technology DisoFLAG, which is the latest and most effective technology. After testing, considering various data sets and functions, the effect of the present invention has been significantly improved and advanced compared with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A flow chart of the system for identifying naturally disordered functional regions;

[0026] Figure 2 This is the framework diagram of the residual MLP in the classifier;

[0027] Figure 3 A calculation diagram of the AUC index of the recognition system DPF-sys-Protein of the present invention and other corresponding predictors on the CAID2 disordered binding function dataset;

[0028] Figure 4 This is a graph of the calculation of the APS index of the identification system DPF-sys-Protein and other corresponding predictors of the present invention on the CAID2 disordered binding function dataset;

[0029] Figure 5 AUC calculation diagram of the recognition system DPF-sys-Linker of the present invention and other corresponding predictors on the CAID2 flexible link function dataset;

[0030] Figure 6 It is a calculation graph of the APS index of the identification system DPF-sys-Linker of the present invention and other corresponding predictors on the CAID2 flexible connection function dataset;

[0031] Figure 7 A calculation diagram of the AUC index of the recognition system DPF-sys-Protein of the present invention and other corresponding predictors on the CAID3 disordered binding function dataset;

[0032] Figure 8 This is a graph of the calculation of the APS index of the identification system DPF-sys-Protein of the present invention and other corresponding predictors on the CAID3 disordered binding function dataset;

[0033] Fig. 9 It is a calculation diagram of the AUC index of the recognition system DPF-sys-Linker of the present invention and other corresponding predictors on the CAID3 flexible connection function dataset;

[0034] Fig.10 It is a calculation graph of the APS index of the identification system DPF-sys-Linker of the present invention and other corresponding predictors on the CAID3 flexible connection function dataset;

[0035] Fig.11 A calculation diagram of the AUC index of the recognition system DPF-sys-Protein of the present invention and other corresponding predictors in the disordered function test of protein binding on the TE210 test set;

[0036] Fig.12 A calculation diagram of the APS index of the recognition system DPF-sys-Protein of the present invention and other corresponding predictors in the disordered function test of binding to proteins on the TE210 test set;

[0037] Fig.13 A calculation diagram of the AUC index of the recognition system DPF-sys-Protein of the present invention and other corresponding predictors in the disordered function test of binding to proteins on the TE83 test set;

[0038] Fig.14 This is a graph showing the calculation of the APS index of the recognition system DPF-sys-Protein of the present invention and other corresponding predictors in the disordered function test of protein binding on the TE83 test set. DETAILED DESCRIPTION

[0039] In order to facilitate the understanding of the present invention, the present invention is described in more detail below in conjunction with the accompanying drawings and specific embodiments. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in this specification. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present invention more thorough and comprehensive.

[0040] The recognition system of natural disordered functional regions uses the semantic information provided by the protein language model to consider the various specific types of disordered binding functions and flexible connection functions of proteins separately, and combines different deep learning networks to predict disordered regions and flexible connection regions that bind to proteins, nucleic acids, lipids, ions, and other small molecules. Among them, other small molecules refer to small molecules other than the four types of molecules mentioned above, including vitamins, alcohols, organic acids, etc. The specific steps are as follows:

[0041] 1. Data preparation:

[0042] Disprot is a gold standard database of intrinsically disordered proteins and regions (IDPs / IDRs), providing rich high-quality functional annotations for IDPs and IDRs, including disordered binding regions and flexible connection regions, based on the gene ontology GO and the intrinsic disordered protein ontology IDPO. Based on the GO terms for protein, nucleic acid, ion, lipid, and other small molecule binding and the IDPO terms for flexible connection regions, 1073 sequences with at least one of the six functional annotations were obtained from Disprot9.4 using the annotation code provided by CAID2 (https: / / github.com / BioComputingUP / caid2-reference / blob / master / src / references.ipynb), which takes into account the hierarchical structure of the functional annotations and ensures the integrity and accuracy of the data labels.

[0043] Data preprocessing, using CD-hit to cluster sequences at 25% similarity, and the clusters were divided into training, validation, and test sets at a ratio of 3:1:1; sequences with more than 25% similarity to the disordered binding and flexible connection function datasets released by CAID2 were removed from the training set to maintain the reliability of the evaluation. A training set, a validation set, and an independent test set were established, containing 552, 227, and 210 sequences, respectively, denoted as TR552, VA227, and TE210.

[0044] To further evaluate the stability of the prediction model, newly added sequences with disordered binding or flexible connection functions were collected from Disprot9.6, and sequences overlapping with the CAID2 dataset were deleted to obtain an independent test set of 83 sequences, recorded as TE83.

[0045] 2. Feature extraction:

[0046] Protein language models (PLMs) trained on large-scale protein sequence datasets through self-supervised learning based on deep learning and natural language processing (NLP) techniques can capture a variety of protein information about structure, function, and evolution. ProtT5-XL-U50, trained on the BFD dataset and fine-tuned on the UniRef50 dataset, achieved state-of-the-art results on multiple NLP benchmarks by exploiting a transformer architecture with an encoder-decoder structure and a BERT-style denoising objective to corrupt and reconstruct individual tokens at a mask rate of 15%.

[0047] All sequences were input into the ProtT5-XL-U50 protein language model, and a 1024-dimensional feature vector was obtained for each residue. Therefore, for each sequence, an L*1024-dimensional feature matrix was obtained, where L represents the length of the protein sequence.

[0048] 3. Prediction Model

[0049] Establish as Figure 1 The deep learning architecture shown in the figure is used to predict disordered regions and flexible connection regions in proteins that bind to proteins, nucleic acids, lipids, ions and other small molecules. First, the residue representation of each sequence obtained by ProtT5-XL-U50 is input into the multi-scale feature processing module to further obtain the shared basic binding features. Then, for a variety of specific types of disordered binding functions, a disordered binding region prediction model for multi-task prediction is established. The network framework of multi-task learning is used to integrate the one-dimensional convolutional network CNN that captures local information and the global information network BiLSTM that extracts long-range relationships of sequences as a multi-scale feature processing module to learn the underlying features of the binding function. Five ResidualMLP classifiers are used to identify the specificity of different types of binding functions. Finally, five scores between 0 and 1 are obtained for each residue in the sequence through the Sigmoid layer, which represent the probability that the residue is located in the disordered region and binds to proteins, nucleic acids, lipids, ions, and other small molecules. For the flexible connection function, a flexible connection region prediction model for single task prediction is established, a BiLSTM single feature processing module and a Residual MLP are designed, and finally a score between 0 and 1 is obtained for each residue in the sequence, representing the probability that the residue is located in the flexible connection region.

[0050] The specific modules are explained as follows:

[0051] Multi-scale feature processing module: Combine BiLSTM and CNN parallel settings into a shared network to obtain the underlying features of the disordered binding function. BiLSTM and CNN are classic and efficient algorithms that are widely used in the processing of sequence features. Long short-term memory network (LSTM) is a variant of recurrent neural network (RNN) that can learn long-term dependencies in sequences. BiLSTM considers the bidirectional long-range information of the sequence and consists of a forward LSTM and a reverse LSTM. We use two one-dimensional CNN layers with a kernel size of 3 and a ReLU nonlinear activation function in the middle to capture local features. The results of both paths are reduced in dimension through a linear layer and then concatenated. The features obtained from the two different types of networks more comprehensively characterize the disordered binding function, and training with data of different binding types together can obtain more general underlying features.

[0052] Single feature processing module: Since the flexible connection region mainly acts as a segment connecting two domains, its recognition is different from that of the disordered binding region. Recognition should consider information from the domain, i.e., more long-range information, so local features do not significantly help in this process. For the annotation of the flexible connection region data, we train it separately and only use BiLSTM as a feature processor without sharing the underlying features with the disordered binding function.

[0053] Classifier: For the disordered binding function, the shared basic features extracted by the multi-scale feature processing module are input into five specific binding residual multilayer perceptrons Residual MLP, and their outputs are normalized by the Sigmoid function to obtain five specific disordered binding probability scores ranging from 0 to 1. These five branches mine the specific binding patterns of disordered regions that bind to proteins, nucleic acids, lipids, ions and other small molecules from the shared basic features. Similarly, the flexible connection region features generated by its single feature processing module are also passed to a residual multilayer perceptron Residual MLP to further extract the specificity of the flexible connection region and then normalize it to calculate its probability score. Each residual multilayer perceptron Residual MLP has a consistent framework and parameter settings, but different network weights. The specific framework is as follows Figure 2 ,The features updated by the multi-scale feature processing module or the single feature processing module will pass through the linear layer, and then through two modules containing ReLU activation function and linear layer. The output of each module is added back to the input data through skip connection, and finally passes through a linear layer for output.

[0054] The input and output dimensions of the multi-scale feature processing module and the single feature processing module are both 1024. The input dimension of each classifier in the multi-functional classifier and the single classifier is 1024 and the output dimension is 1. The input and output dimensions of the BiLSTM module in the multi-scale feature processing module are 1024, and the hidden layer dimension is 512. The input and output dimensions of the first layer of CNN in the CNN module are 1024 and 512, and the input and output dimensions of the second layer of CNN are 512 and 1024. After passing through the BiLSTM and CNN modules, the input and output dimensions of the linear layer are 1024 and 512, and then the two results are spliced, so the output dimension of the multi-scale feature processing module is 1024. The input and output dimensions of the BiLSTM module of the single feature processing module are 1024, and the hidden layer dimension is 512. The input and output dimensions of the first linear layer in each Residual MLP are 1024 and 512, the input and output dimensions of the last linear layer are 512 and 1, and the input and output dimensions of other layers are 512.

[0055] 4. Model training:

[0056] A100-PCIE-40GB GPU is used for training. When training the disordered connection region prediction model, the five types of disordered connection losses are added together to get the total loss, where each loss is the binary cross entropy loss of the corresponding function; for the flexible connection region prediction model, a separate binary cross entropy loss is used.

[0057] Model loss function design:

[0058] ;

[0059] ;

[0060] ;

[0061] in, represent Binary cross entropy loss function for unordered features of type; Represents the residue A vector of true labels for the unordered features of type; Represents the residue The probability vector of the model predictions for the type of unordered function; represent No. The first component, The true labels of the residues; represent No. The first component, The predicted probability of residues; is the number of residues. Binding total loss represents the total loss function of disordered binding function, which in this application refers to the total loss function obtained by adding the binary cross entropy loss functions of disordered functions of the five binding types. They represent the five binding types, including binding to proteins, nucleic acids, lipids, ions and other small molecules. (ProteinBinding)、 (Nucleic acid Binding) (Lipid Binding) (Ion Binding)and (Other small molecule Binding), corresponding to the disordered functions binding to proteins, nucleic acids, lipids, ions and other small molecules. DFL (Disorder Flexible Linker) represents the flexible linking function. represents the flexible connection function loss function. Binary cross entropy loss function representing the type of flexible connection function.

[0062] Two models were trained using the pytorch framework with a batch size of 1, SGD optimizer and a learning rate of 5e -4 . Data is input into the prediction model to generate output results. The model calculates the predicted value through forward propagation and evaluates the error using the loss function. Then the gradient is calculated through backpropagation, and the weights are updated using the stochastic gradient descent method. The process is repeated for iteration to optimize performance. When the loss function on the validation set is the smallest, we select the network corresponding to that iteration. On the A100-PCIE-40GB GPU, it takes 2.5 hours to train the disordered junction region prediction model and 3.5 hours to train the flexible connection region prediction model. The loss function of the flexible connection region prediction model converges slower than that of the disordered junction region prediction model.

[0063] In addition to collecting data with annotations of protein binding, nucleic acid binding, and lipid binding in the Disprot database, the present invention also collects data with disordered binding functions of ions and other small molecules and incorporates them into training; the disordered binding functions and flexible connection functions of various specific types of proteins are considered separately, and different prediction models are designed based on their respective characteristics. For the prediction of various specific types of disordered binding functions, networks of different scales are integrated to capture information and make predictions; for the flexible connection function, a global scale network is used for processing and prediction.

[0064] 5. Model testing:

[0065] For the convenience of subsequent description, the recognition system of natural disordered functional regions is recorded as DPF-sys (Disorder Protein functions System), among which the disordered region prediction models for protein binding, nucleic acid binding, lipid binding, ion binding and other small molecules are recorded as DPF-sys-Protein, DPF-sys-Nuc, DPF-sys-Lipid, DPF-sys-Ion, DPF-sys-Oth, respectively, and the flexible connection region prediction model is recorded as DPF-sys-Linker. The results are good on test sets such as TE210 and TE83, and can achieve accurate recognition of multiple types of disordered binding regions and flexible connection regions.

[0066] Working principle description: Combined with Figure 1Understanding, a protein sequence is input into the protein language model ProtT5-XL-U50 to obtain the representation vector of each residue in the sequence, which contains the semantic information of the protein sequence. Then the representation vector of a sequence is input into the disordered binding region prediction model and the flexible connection region prediction model respectively. When input into the disordered binding region prediction model, the representation vector will pass through the multi-scale feature processing module of CNN and BiLSTM fusion, and then pass through five residual multi-layer perceptrons Residual MLP and Sigmoid function, and finally output five scores between 0 and 1 for each residue in the sequence, representing the probability of the residue being located in the disordered region and binding to proteins, nucleic acids, lipids, ions, and other small molecules respectively. When input into the flexible connection region prediction model, the representation vector will only pass through the BiLSTM single feature processing module and a residual multi-layer perceptron Residual MLP and Sigmoid function, and finally output a score between 0 and 1 for each residue in the sequence, representing that the residue is located in the flexible connection region. Finally, it can be known which region of the protein sequence is a disordered region and has binding functions or flexible connection functions of different molecular types.

[0067] In specific applications, the identification system, parameter description:

[0068] (1) -t specifies the prediction type. There are three options: binding, linker, and all, which represent multiple types of disordered binding regions, flexible connection regions, or both.

[0069] (2) -i specifies the input FASTA file; -d specifies the processor to be used. You can choose gpu or cpu. If you choose gpu, you need to append a number to indicate the GPU card index, such as gpu0.

[0070] The specific operation mode of the recognition system is as follows:

[0071] 1. Installation steps:

[0072] (1) Install Anaconda, the installation address is https: / / www.anaconda.com / ;

[0073] (2) Download the network weight file of the protein language model ProtT5-XL-U50 (https: / / huggingface.co / Rostlab / prot_t5_xl_uniref50 / resolve / main / pytorch_model.bin?download=true) and copy it to / DPF-sys / utils / prot_t5_xl_uniref50 / file;

[0074] (3) Locate the software installation directory: cd DPF-sys;

[0075] (4) Create a virtual environment and configure it: conda env create -f DPF_sys.yml;

[0076] (5) Activate the environment: conda activate DPF_sys.

[0077] 2. Predicting multiple types of disordered binding regions

[0078] (1) Command: python predict.py -t binding -I example.fasta -d gpu0; pythonpredict.py -t binding -i example.fasta -d gpu0;

[0079] (2) Results: Generate files binding_scores.txt and binding_binary.txt and save them in the / DPF-sys / data_save / example / result / directory. binding_scores.txt contains the prediction scores of different disordered binding functions for each sequence. binding_binary.txt contains the binary classification results of different disordered binding functions for each sequence;

[0080] File contents:

[0081] Line 1: >Sequence ID;

[0082] Line 2: protein sequence (1-character amino acid code);

[0083] Row 3: prediction results of disordered regions bound to proteins;

[0084] Row 4: prediction results of disordered regions binding to nucleic acids;

[0085] Row 5: prediction results of disordered regions bound to lipids;

[0086] Row 6: prediction results of disordered regions bound to ions;

[0087] Row 7: Prediction results of disordered regions binding to other small molecules.

[0088] 3. Predicting flexible connection areas

[0089] (1) Command: python predict.py -t linker -i example.fasta -d gpu0;

[0090] (2) Results: Generate files linker_scores.txt and linker_binary.txt and save them in the / DPF-sys / data_save / example / result / directory. linker_scores.txt contains the prediction score of the flexible connection function of each sequence. linker_binary.txt contains the binary classification result of the flexible connection function of each sequence;

[0091] File contents:

[0092] Line 1: >Sequence ID;

[0093] Line 2: protein sequence (1-character amino acid code);

[0094] Row 3: Prediction results of the flexible connection area.

[0095] 4. Predict disordered multi-binding and flexible connection regions

[0096] (1) Command: python predict.py -t all -i example.fasta -d gpu0;

[0097] (2) Results: Generate the files all_scores.txt and all_binary.txt and save them in the / DPF-sys / data_save / example / result / directory. all_scores.txt contains the prediction scores of the disordered function for each sequence. all_binary.txt contains the binary classification results of the disordered function for each sequence;

[0098] File contents:

[0099] Line 1: >Sequence ID;

[0100] Line 2: protein sequence (1-character amino acid code);

[0101] Row 3: prediction results of disordered regions bound to proteins;

[0102] Row 4: prediction results of disordered regions binding to nucleic acids;

[0103] Row 5: prediction results of disordered regions bound to lipids;

[0104] Row 6: prediction results of disordered regions bound to ions;

[0105] Row 7: Prediction results of disordered regions binding to other small molecules.

[0106] Row 8: Prediction results of flexible connection areas.

[0107] VI. Index Evaluation

[0108] The identification of naturally disordered functional regions systematically predicts six propensity scores that quantify the likelihood of each disorder.

[0109] The AUC and APS indicators are mainly used to evaluate the quality of the prediction effect. AUC is the area enclosed by the Receiver Operating Characteristic curve (ROC), and its value range is [0,1]. Randomly select a positive sample and a negative sample. The probability that the current classification algorithm ranks this positive sample before the negative sample based on the calculated score value is the AUC value. AUC>0.5 indicates that the probability value predicted by the model is consistent with the probability of the true label, and the larger the AUC value, the higher the model prediction accuracy; AUC=0.5 indicates that the model prediction result is highly random and the prediction result has no reference value; AUC<0.5 indicates that the model prediction result is opposite to the true result. AUC measures the performance of the model under all possible classification thresholds, so it is not affected by a single threshold and is relatively stable. On extremely unbalanced data sets, AUC may not accurately reflect the performance of the model. APS is an indicator used to evaluate the performance of a binary classification model, especially when there is a class imbalance. It combines the concepts of precision and recall to provide a more comprehensive model evaluation. APS is obtained by calculating the area under the precision-recall curve. It considers the precision at each recall level and averages them. The APS value is between 0 and 1. The higher the score, the better the model. A high APS value means that the model can maintain a high precision rate under multiple thresholds. It only cares about positive examples, not negative examples. If the number of positive examples is small, both Precision and Recall are easily affected, the accuracy is not that high, and it is not robust. AUC and APS can be combined to look at these two indicators.

[0110] In addition, other metrics are also considered. The maximum harmonic mean (F1-max) between precision and recall under all thresholds, for each unordered feature, we use the threshold determined by F1-max to make binary predictions. We use the MCC metric to measure the quality of binary predictions.

[0111] 7. Model prediction performance test

[0112] To evaluate the performance of our method on disordered functions, we compared it with other predictors that performed well on CAID2. These predictors include the latest DisoFLAG, which uses deep learning to predict disordered regions and flexible connection regions that bind to proteins, DNA, RNA, ions, and lipids, respectively. flDPnn combines deep learning and machine learning to predict disordered regions and flexible connection regions that bind to proteins, RNA, and DNA, respectively. DeepDISOBind uses deep learning, while DisoRDPbind uses machine learning to predict disordered regions that bind to RNA, DNA, and proteins, respectively. AHCHOR2 predicts disordered regions that bind to proteins based on energy estimation methods. MoRFCHiBi (lightweight and web versions) predicts disordered regions that bind to proteins by combining machine learning and statistical methods. DisoLipPred is used to predict disordered regions that bind to lipids, while APOD and DFLpred aim to quickly predict flexible connection regions. Based on deep learning and protein language models, our method predicts disordered regions and flexible connection regions that bind to proteins, nucleic acids, lipids, ions, and other small molecules, respectively, covering the widest range of disordered functions.

[0113] The official website of CAID, a platform for predicting and evaluating natural disordered proteins and their functions, provides the second and third rounds of blind test data sets (CAID2, CAID3), both of which include disordered binding functional region data sets and flexible connection functional region data sets. For the CAID2 and CAID3 disordered binding functional region data sets, the protein-bound disordered region prediction model DPF-sys-Protein of the present invention is used for prediction. This is also because the annotations of the CAID2 and CAID3 disordered binding functional data sets are mostly disordered functional regions bound to proteins. For the CAID2 and CAID3 flexible connection functional region data sets, the flexible connection region prediction model DPF-sys-Linker of the present invention is used for prediction.

[0114] Test results see Figure 3-10 For the disordered binding functional region datasets of CAID2 and CAID3, the AUCs of DPF-sys-Protein in the present invention were 0.861 and 0.892, respectively, ranking second and first, and the second value was only 0.018 different from the first in the test, and the APSs were 0.466 and 0.66, respectively; for the CAID2 and CAID3 flexible connection functional region datasets, the AUCs of DPF-sys-Linker in the present invention were 0.825 and 0.839, respectively, ranking first and eighth, and the APS indicators were 0.198 and 0.427, respectively, both ranking first, which was a significant improvement compared with the second place.

[0115] The recognition system of the present invention is tested with the TE210 and TE83 test sets formed in the data preparation stage by the other predictors mentioned above. Figure 11-14 The following table shows the calculation diagrams of AUC and APS indicators when the recognition system of the present invention and the other corresponding predictors are tested on the TE210 and TE83 test sets respectively, taking the evaluation of disordered functional regions binding to proteins as an example; other test items are not shown one by one, and the calculation principles are similar. Finally, the results of the TE210 test set are shown in Table 1, and the results of the TE83 test set are shown in Table 2.

[0116] According to Tables 1 and 2, compared with the highest values ​​of the corresponding indicators in other methods, on the TE210 test set, the present invention improves the AUC of disordered functional regions bound to proteins, lipids and ions by 2%, 4% and 2%, respectively. The APS of disordered functional regions bound to proteins and lipids also increases by 14% and 77%, respectively. On the TE83 test set, the AUC of disordered functional regions bound to proteins and flexible connection functional regions increases by 4% and 1%, respectively. The APS of disordered functional regions bound to proteins and ions increases by 45% and 45%, respectively. The present invention has a significant improvement in the APS index of disordered functional regions bound to proteins, lipids and ions, and the AUC index of disordered functional regions bound to proteins, lipids, ions and flexible connection functions also increases. The prediction effect for disordered functional regions bound to nucleic acids is good, and the gap with other methods is small. At present, this method is the only one that can predict disordered functional regions bound to small molecules, and the effect is good. The TE83 test set does not contain positive samples of disordered functional regions bound to small molecules, so the AUC and APS indicators cannot be calculated. In summary, this method has good comprehensive prediction effect on natural disordered functional regions.

[0117] In both test sets, the present invention performs best in disordered functional regions binding to proteins. Although the present invention has lower indicators than DisoFLAG in disordered functional regions binding to lipids and ions on the TE83 test set, and lower indicators than DisoFLAG in flexible connection functional regions on the TE210 test set, the performance gap of DisoFLAG in these categories is large on these two data sets, indicating that its stability is not as good as that of the present invention.

[0118] Table 1 TE210 test set results statistics

[0119]

[0120] Table 2 TE83 test set results statistics

[0121]

[0122] In view of various data sets and functions, the effects of the present invention have achieved significant improvement and progress compared with the prior art.

[0123] Analysis of other test results: To further measure the effectiveness of the prediction model, the present invention also conducted selection and ablation experiments on the input features and network modules of the prediction model, and demonstrated the prediction of protein-bound disordered regions and flexible connection regions evaluated on the TE210 dataset as a representative. The rest of the test results are similar.

[0124] Table 3 Ablation experiment of input features of TE210 dataset

[0125]

[0126] Table 4 Network module ablation experiment on TE210 dataset

[0127]

[0128] Table 5 Ablation experiment of TE210 dataset prediction model

[0129]

[0130] From the ablation experiment table 3 of the input features, ESM-1b and ESM2 are two other protein language models. AF2_str refers to the structure predicted by the structure prediction software Alphafold2, extracting structural features such as secondary structure, main chain angle, and predicted local distance difference test value and alignment error value. AF2_str&ProtT5-XL-U50 refers to the feature vector obtained by splicing the two features of AF2_str and ProtT5-XL-U50. Whether it is predicting the disordered region bound to the protein or the flexible connection region, the representation vector provided by the ProtT5-XL-U50 protein language model has the best effect. The protein structure information AF2_str and ProtT5-XL-U50 are not as good as ProtT5-XL-U50 alone, indicating that the information contained in the structural information and the protein language model is redundant and noisy.

[0131] The ablation experiment of the network module refers to replacing the multi-scale feature processing module in the disordered combined region prediction model with a single bidirectional long short-term memory network BiLSTM or convolutional network CNN, and replacing the single feature processing module of the flexible connection region prediction model with a combination of the bidirectional long short-term memory network BiLSTM and the convolutional network CNN or a convolutional network. Separate training means that the multi-task learning framework is not used, and data of different combination types are trained separately. As can be seen from Table 4, the effect of using a multi-scale feature processing module for the disordered combined region prediction model and training data of different combination types together will be better, and using a single feature processing module for the flexible connection region prediction model is better.

[0132] For the ablation experiment of the prediction model, the deep network refers to superimposing the original network framework and using jump connections to allow the input features to pass through more layers of networks, with more parameters and a deeper network framework. The attention network refers to a model that uses the attention mechanism, which has a wide range of applications in different fields. According to Table 5, the effects of other models are not as good as the disordered combined region and flexible connection region prediction models.

[0133] Overall, the two types of prediction models in the identification system of naturally disordered functional regions have achieved the optimal effect and can realize accurate prediction of disordered binding regions and flexible connection regions.

[0134] Those skilled in the art should recognize that the above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. As long as they are within the spirit of the present invention, appropriate changes and modifications to the above embodiments are within the scope of protection claimed by the present invention.

Claims

1. A system for identifying naturally disordered functional regions, characterized in that include: A protein language model that receives a protein sequence and outputs a corresponding sequence representation vector; A disordered binding region prediction model has a built-in multi-scale feature processing module and a multi-functional classifier for capturing local and global information. The output end of the multi-scale feature processing module is connected to the multi-functional classifier, and the multi-functional classifier has multiple classifiers connected in parallel. The input end of the disordered binding region prediction model is connected to the output end of the protein language model. The prediction model receives a protein sequence representation vector as a sequence feature, and performs feature processing and multiple disordered binding function identification in sequence through the multi-scale feature processing module and the multi-functional classifier, and outputs the probability that each residue in the protein sequence is located in the disordered region and binds to proteins, nucleic acids, lipids, ions and other small molecules respectively; The flexible connection region prediction model has a built-in single feature processing module and a single classifier that capture global scale information. The output end of the single feature processing module is connected to the single classifier. The input end of the flexible connection region prediction model is connected to the output end of the protein language model. The prediction model receives the protein sequence representation vector as the sequence feature, performs feature processing and flexible connection function identification, and outputs the probability that each residue in the protein sequence is located in the flexible connection region.

2. The system for identifying natural disordered functional regions according to claim 1, characterized in that: The protein language model is ProtT5-XL-U50; In the multi-scale feature processing module, the one-dimensional convolutional network CNN and the bidirectional long short-term memory network BiLSTM are set in parallel; the single feature processing module is the bidirectional long short-term memory network BiLSTM; In the multifunctional classifier, five classifiers are set in parallel and executed sequentially; in the single classifier, one classifier is set.

3. The system for identifying natural disordered functional regions according to claim 2, characterized in that: Each classifier has a built-in residual multilayer perceptron. The residual multilayer perceptrons in different classifiers have consistent frameworks and hyperparameters, but different weights.

4. The system for identifying natural disordered functional regions according to claim 3, characterized in that: The residual multilayer perceptron Residual MLP contains multiple linear layers and activation function ReLU, and uses skip connections to enhance the training effect of the network.

5. The system for identifying natural disordered functional regions according to claim 3, characterized in that: The residual multi-layer perceptron Residual MLP is connected to a normalized layer, and a Sigmoid function is built in the normalized layer. The Sigmoid function normalizes its output.

6. The system for identifying natural disordered functional regions according to claim 2, characterized in that: The CNN and BiLSTM of the multi-scale feature processing module are both followed by a linear layer. After receiving the output results of CNN and BiLSTM, the linear layer performs dimensionality reduction, concatenates the dimensionality reduction results of the two, and outputs shared basic features.

7. The system for identifying natural disordered functional regions according to claim 2, characterized in that: The one-dimensional convolutional network CNN includes: two one-dimensional CNN layers with a kernel size of 3 and a ReLU nonlinear activation function in the middle; BiLSTM consists of a forward LSTM and a reverse LSTM, and LSTM is a long short-term memory network.

8. The system for identifying natural disordered functional regions according to any one of claims 2 to 7, characterized in that: The protein language model ProtT5-XL-U50 generates a 1024-dimensional feature vector for each residue in the protein sequence. The input and output dimensions of the multi-scale feature processing module and the single feature processing module are both 1024. The input dimension of the multi-functional classifier and the single classifier is 1024 and the output dimension is 1.

9. The system for identifying natural disordered functional regions according to claim 1, characterized in that: The disordered combined region prediction model and the flexible connected region prediction model were trained using the pytorch framework with a batch size of 1, SGD optimizer and a learning rate of 5e. -4 .

10. The system for identifying natural disordered functional regions according to claim 9, characterized in that: The training process of the disordered combined region prediction model and the flexible connection region prediction model: S1. Extract disordered protein sequences containing six functions including protein binding, nucleic acid binding, lipid binding, ion binding, other small molecule binding and flexible connection from the DisProt database, perform data preprocessing, and divide them into training set, validation set and test set; S2. Input the sequences in the training set into the protein language model, obtain a 1024-dimensional feature vector for each residue, and obtain a sequence representation vector consisting of an L*1024-dimensional feature matrix for each sequence, where L represents the length of the protein sequence; S3. Select a prediction model, input the sequence representation vector into the disordered binding region model or the flexible connection region model, calculate the result of the corresponding model through forward propagation, and use the loss function to evaluate the error, then calculate the gradient through back propagation, and use the stochastic gradient descent method to update the weight. In the framework of pytorch, iterate the above process to optimize the performance. When the loss function on the validation set is the smallest, the model weight corresponding to this iteration is the optimal weight, and the training is completed.

Citation Information

Patent Citations

  • Protein disorder region prediction method based on deep learning and language model

    CN119252318A

  • Nucleic acid binding protein recognition method based on protein map and protein language model

    CN119252348A