Protein Disordered Flexible Linker Prediction Method Based on Multi-Task Learning

By constructing a multi-task learning protein disordered flexible linker prediction network model, the problems of low prediction accuracy and high computing resource consumption in the prior art are solved, efficient and accurate protein disordered flexible linker prediction is achieved, and cost is reduced.

CN116864010BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310751702.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-25
Publication Date
2025-07-01
Estimated Expiration
2043-06-25

AI Technical Summary

Technical Problem

In the prior art, the prediction accuracy of protein disordered flexible linkers is low, and the computing resources are consumed largely, which increases costs.

Method used

Using a multi-task learning method, a protein disordered flexible connector prediction network model including a shared layer, a disordered area tower layer and a disordered flexible connector tower layer is constructed. Through iterative training and loss function update, the prediction accuracy is improved, and only the evolutionary information and physical and chemical characteristics of the protein are used to reduce computing resource consumption.

Benefits of technology

It improves the accuracy of protein disordered flexible linker prediction, reduces computing resource consumption and cost, and avoids the problem of pre-training information loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116864010B_ABST
    Figure CN116864010B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting protein disordered flexible linkers based on multi-task learning, which includes the following steps: obtaining a multi-task data set; dividing the data set for predicting protein disordered flexible linkers; unifying the lengths of the training data set for predicting protein disordered regions, the training data set for predicting disordered flexible linkers, and the test data set; constructing a protein feature representation matrix; constructing a network model for predicting protein disordered flexible linkers; iteratively training the network model for predicting protein disordered flexible linkers; and obtaining the prediction results of protein disordered flexible linkers. When constructing the network model for predicting protein disordered flexible linkers, the present invention uses disordered region data to increase the information content, uses multi-task learning to reduce the information loss during the training process, improves the abundance of information in the network model, and effectively improves the accuracy of identifying protein disordered flexible linkers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and particularly relates to a method for predicting protein disordered flexible linkers based on multi-task learning, which can be applied to the discovery of drug action sites. Background Art

[0002] A disordered flexible linker is a specific region in intrinsically disordered proteins that connects two or more structurally defined domains. Disordered flexible linkers are usually composed of amino acid sequences, which are highly flexible and lack secondary structure. They play an important role in regulating protein-protein interactions and facilitating functional pathways, such as protein-protein interactions, signal transduction, and control of gene expression.

[0003] Traditional prediction of protein disordered flexible linkers is to obtain the protein spectrum by nuclear magnetic resonance technology and analyze the protein spectrum to predict protein disordered flexible linkers. This method is costly and has a long prediction cycle.

[0004] Existing prediction of protein disordered flexible linkers is based on computational methods. The technical idea is to construct a training data set for disordered flexible linkers, perform protein feature representation on the training set data, build a prediction network for disordered flexible linkers, train the disordered flexible linker network using the training set data, and then predict protein disordered flexible linkers. For example, Pang et al. published an article titled "TransDFL: identification of disordered flexible linkers in proteins by transfer learning" in "Genomics, Proteomics & Bioinformatics" in 2022, proposing a method for predicting protein disordered flexible linkers based on transfer learning. This method first collects disordered region data and disordered flexible linker data. Protein feature representation is performed through seven commonly used physicochemical properties (steric parameters, polarizability, volume, hydrophobicity, isoelectric point, helix probability, and sheet probability), position-specific scoring matrices, secondary structure features, and solvent accessibility. Then, a prediction network for disordered flexible linkers is built. The disordered flexible linker prediction network is pre-trained using disordered region data and fine-tuned using disordered flexible linker data to obtain a trained disordered flexible linker prediction network, and this network is used to predict protein disordered flexible linkers. This method uses disordered region data to increase the amount of information. However, the pre-training-fine-tuning method will lose the pre-trained information during the fine-tuning process, thereby affecting the prediction accuracy. At the same time, this method uses a large number of protein features, consuming a large amount of computing resources during the training and prediction stages and increasing the cost. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects existing in the above-mentioned prior art, and to provide a method for predicting protein disordered flexible linkers based on multi-task learning, so as to solve the technical problem of low prediction accuracy existing in the prior art.

[0006] To achieve the above purpose, the technical solutions adopted by the present invention include the following steps:

[0007] (1) Obtain a multi-task data set:

[0008] The M amino acid sequences X of disordered region proteins obtained IDRs and their corresponding disordered region labels Y IDRs constitute a training data set for predicting protein disordered regions At the same time, the N amino acid sequences X of disordered flexible linker proteins obtained DFLs and their corresponding disordered flexible linker labels Y DFLs constitute a prediction data set D for protein disordered flexible linkers DFLs ={X DFLs , Y DFLs}, where M≥1000 and N≥200;

[0009] (2) Divide the prediction data set for protein disordered flexible linkers:

[0010] Based on sequence similarity, divide the prediction data set D for protein disordered flexible linkers DFLs to obtain a training data set for predicting disordered flexible linkers of proteins containing N train proteins and a test data set containing N test proteins where N = N train + N test , N train > N test ;

[0011] (3) Unify the lengths of the proteins in the training data set for predicting protein disordered regions, the training data set for predicting disordered flexible linkers, and the test data set:

[0012] Unify the lengths of the proteins in the training data set for predicting protein disordered regions the training data set for predicting disordered flexible linkers and the test data set to obtain a training data set for predicting protein disordered regions with a protein length of G the training data set for predicting disordered flexible linkers and the test data set

[0013] (4) Construct a protein feature representation matrix:

[0014] Construct The corresponding dimensions are M×G×D, N train ×G×D, N test The protein feature representation matrix of the disordered region training dataset with dimensions of ×G×D The protein feature representation matrix of the disordered flexible linker training dataset The protein feature representation matrix of the disordered flexible linker test dataset where D > 30;

[0015] (5) Construct a protein disordered flexible linker prediction network model:

[0016] Construct a protein disordered flexible linker prediction network model O including a shared layer and a disordered region tower layer and a disordered flexible linker tower layer connected to the output end of the shared layer and arranged in parallel. Among them, the shared layer includes an embedding layer and a Transformer Encoder layer stacked in sequence; both the disordered region tower layer and the disordered flexible linker tower layer include a fully connected network and a Sigmod classifier stacked in sequence;

[0017] (6) Iteratively train the protein disordered flexible linker prediction network model:

[0018] (6a) Initialize the iteration number as t, the maximum iteration number as T, the trainable parameters of the shared layer, the disordered region tower layer, and the disordered flexible linker tower layer in the current prediction network model are w1, w2, w3 respectively, and let t = 0;

[0019] (6b) Use the protein feature representation matrix of the disordered region training dataset The protein feature representation matrix of the disordered flexible linker training dataset as the input of the prediction network model O for forward propagation to obtain the disordered region prediction result of the disordered region training set The disordered flexible linker prediction result of the disordered flexible linker training set

[0020] (6c) Update the trainable parameters w1, w2, w3 of the shared layer, the disordered region tower layer, and the disordered flexible linker tower layer in the current prediction network model through the disordered region loss value L IDRs and the disordered flexible linker loss value L DFLs to obtain the prediction network model O of this iteration t ;

[0021] (6d) Judge whether t = T holds. If so, obtain the trained prediction network model O * , otherwise, let t = t + 1, O = O * , and execute step (6b);

[0022] (7) Obtain the prediction results of protein disordered flexible linkers:

[0023] Use the protein feature representation matrix of the disordered flexible linker test data set as the input of the trained prediction network model O * to perform forward propagation, and obtain the disordered flexible linker prediction of the protein sequences in the disordered flexible linker test data set

[0024] Compared with the prior art, the present invention has the following advantages:

[0025] (1) The prediction network model constructed by the present invention includes a shared layer and a disordered region tower layer and a disordered flexible linker tower layer connected to the output end thereof and arranged in parallel. During the training process of the model, the disordered region data participates in the calculation of the loss function in each iteration, expanding the information volume without losing information, and effectively improving the prediction accuracy.

[0026] (2) The present invention abandons the redundant feature collection and integration in the protein features used, and only adopts the evolutionary information and physicochemical characteristics of proteins, solving the problems of large resource consumption and high cost usually brought by complex feature engineering in existing research. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is the implementation flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0028] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the present invention does not belong to the method for diagnosing and treating diseases.

[0029] Refer to Figure 1 , the present invention includes the following steps:

[0030] Step 1) Obtain a multi-task data set:

[0031] Combine the obtained M amino acid sequences X of disordered region proteins IDRs and their corresponding disordered region labels Y IDRs to form a protein disordered region prediction training data set At the same time, combine the obtained N amino acid sequences X of disordered flexible linker proteins DFLs and their corresponding disordered flexible linker labels Y DFLs to form a protein disordered flexible linker prediction data set D DFLs ={X DFLs , Y DFLs ​}, where M ≥ 1000 and N ≥ 200; in this embodiment, M = 3000 and N = 248.

[0032] Step 2) Divide the protein disordered flexible linker prediction data set:

[0033] 2a) Use the BlastClust software package to cluster the protein amino acid sequences X DFLs in the disordered flexible linker data set D DFLs with a sequence similarity greater than a into one class, and obtain the clustering result C = {C1, C2, …, C c , …, C k}, where a ≥ 25%, k ≥ 100, and C c represents the protein amino acid sequences included in the c-th class in the clustering result C; in this embodiment, a = 25% and k = 194;

[0034] 2b) Divide C into two parts by class, and obtain the disordered flexible linker prediction training data set train containing N proteins and the test data set test containing N proteins, where N = N train + N test , N train > N test ; in this embodiment, N train = 166 and N test = 82.

[0035] Step 3) Unify the lengths of the proteins in the protein disorder region prediction training data set, the disordered flexible linker prediction training data set, and the test data set:

[0036] Since the protein lengths are not unified and the gradient descent algorithm cannot update the parameters, the proteins in the protein disorder region prediction training data set the disordered flexible linker prediction training data set and the test data set are unified in length. Specifically: for protein sequences with a length less than G, they are padded with 0s, and for the parts after protein sequences with a length exceeding G, they are truncated to obtain the protein disorder region prediction training data set the disordered flexible linker prediction training data set and the test data set where G > 1000; in this embodiment, G = 1500.

[0037] Step 4) Construct a protein feature representation matrix:

[0038] 4a) Use the PSI-BLAST tool to calculate The evolutionary information of each protein in 1 to obtain an evolutionary information matrix of the intrinsically disordered region training set with dimensions M×G×D train ×G×D 1 to obtain an evolutionary information matrix of the intrinsically disordered region training set with dimensions N test ×G×D 1 where D Evolutionary information matrix of proteins in the intrinsically disordered flexible linker training set Evolutionary information matrix of proteins in the intrinsically disordered flexible linker test set where D 1 ≥20; in this embodiment, D 1 = 20;

[0039] 4b) Use the AAindex tool to calculate the physicochemical information of each protein in 2 to obtain an evolutionary information matrix of the intrinsically disordered region training set with dimensions M×G×D train ×G×D 2 to obtain an evolutionary information matrix of the intrinsically disordered region training set with dimensions N test ×G×D 2 where D Evolutionary information matrix of proteins in the intrinsically disordered flexible linker training set Evolutionary information matrix of proteins in the intrinsically disordered flexible linker test set where D 2 ≥13, and the selected physicochemical information has indexes CIDH920101, EISD860103, NISK860101, QIAN880105, ROBB760101, ROBB760108, ROBB760112, ROBB760113, CORJ870103, CORJ870106, CORJ870107, CORJ870108, MIYS990104 in AAindex; in this embodiment, D 2 = 13;

[0040] 4c) Concatenate and along the third dimension to obtain protein feature representation matrices of the intrinsically disordered region training data set with dimensions M×G×D, N train ×G×D, N test ×G×D Protein feature representation matrix of the intrinsically disordered flexible linker training data set Protein feature representation matrix of the intrinsically disordered flexible linker test data set where D = D 1 + D 2 ; in this embodiment, D = 33.

[0041] Step 5) Construct a protein intrinsically disordered flexible linker prediction network model:

[0042] Construct a protein intrinsically disordered flexible linker prediction network model O that includes a shared layer and an intrinsically disordered region tower layer and an intrinsically disordered flexible linker tower layer connected to the output end of the shared layer and arranged in parallel. Among them, the shared layer includes an embedding layer and a Transformer Encoder layer stacked in sequence; the intrinsically disordered region tower layer includes a fully connected network and a Sigmod classifier stacked in sequence; the intrinsically disordered flexible linker tower layer includes a fully connected network and a Sigmod classifier stacked in sequence; in this embodiment, the dimension of the embedding layer in the shared layer is 64, and the fully connected network in the intrinsically disordered region tower layer and the intrinsically disordered flexible linker tower layer only includes one hidden layer, and the number of hidden layer units is 32.

[0043] Step 6) Iteratively train the protein intrinsically disordered flexible linker prediction network model:

[0044] 6a) Initialize the iteration number as t, the maximum iteration number as T, the trainable parameters of the shared layer, the intrinsically disordered region tower layer, and the intrinsically disordered flexible linker tower layer in the current prediction network model are w1, w2, and w3 respectively, and let t = 0;

[0045] 6b) Use the protein feature representation matrix of the intrinsically disordered region training data set and the protein feature representation matrix of the intrinsically disordered flexible linker training data set as the input of the prediction network model O for forward propagation to obtain the prediction result of the intrinsically disordered region of the intrinsically disordered region training set and the prediction result of the intrinsically disordered flexible linker of the intrinsically disordered flexible linker training set The implementation steps are as follows:

[0046] 6b1) The embedding layer in the shared layer performs embedding representation on the protein feature representation matrix of the intrinsically disordered region training data set and the protein feature representation matrix of the intrinsically disordered flexible linker training data set respectively to obtain the embedding vectors of the intrinsically disordered region training set and the embedding vectors of the intrinsically disordered flexible linker training set The Transformer Encoder layer performs forward propagation on respectively to obtain the shared layer hidden vectors of the intrinsically disordered region training set and the shared layer hidden vectors of the intrinsically disordered flexible linker training set

[0047] 6b2) The fully connected networks in the intrinsically disordered region tower layer and the intrinsically disordered flexible linker tower layer perform operations on the shared layer hidden vectors of the intrinsically disordered region training set and the shared layer hidden vectors of the intrinsically disordered flexible linker training set Perform forward propagation respectively to obtain the hidden vectors of the disordered region tower layers of the disordered region training set The hidden vectors of the tower layers of the disordered flexible linker training set The Sigmod classifier Perform disordered region prediction and disordered flexible linker prediction respectively to obtain the disordered region prediction results of the disordered region training set The disordered flexible linker prediction results of the disordered flexible linker training set

[0048] 6c) Through the disordered region loss value L IDRs and the disordered flexible linker loss value L DFLs Update the trainable parameters w1, w2, and w3 of the shared layer, disordered region tower layer, and disordered flexible linker tower layer in the current prediction network model to obtain the prediction network model O for this iteration t ; The implementation steps are as follows:

[0049] 6c1) Adopt the binary cross-entropy loss function and pass through and as well as and Calculate the disordered region loss value L IDRs and the disordered flexible linker loss value L DFLs :

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056] where respectively represent the prediction results and labels of the m-th protein in the disordered region training dataset; represents the prediction results and labels of the n-th protein in the disordered flexible linker training dataset;

[0057] 6c2) Calculate the overall loss L of the prediction network model O through L IDRs and L DFLs and calculate the partial derivatives η1, η2, and η3 of L with respect to w1, w2, and w3:

[0058] L = L DFLs + L IDRs

[0059]

[0060]

[0061]

[0062] 6c) Adopt the gradient descent method, and update the trainable parameters w1, w2, and w3 of the shared layer, the disordered region tower layer, and the disordered flexible connector tower layer through the partial derivatives η1, η2, and η3 of w1, w2, and w3 to obtain the prediction network model O of this iteration. t 。

[0063] 6d) Judge whether t = T holds. If so, obtain the trained prediction network model O. * , otherwise, let t = t + 1, O = O * , and execute step 6b).

[0064] Step 7) Obtain the prediction result of the protein disordered flexible connector:

[0065] Use the protein feature representation matrix of the disordered flexible connector test data set as the input of the trained prediction network model O * to perform forward propagation, and obtain the disordered flexible connector prediction of the protein sequence in the disordered flexible connector test data set .

[0066] The technical effects of the present invention will be further described below in combination with simulation experiments:

[0067] 1. Experimental conditions and content:

[0068] The simulation experiment is carried out on Python 3.7 combined with the pytroch1.7 framework on an Ubuntu platform with an Intel(R) Xeon(R) Gold 5115 CPU (20 cores), a main frequency of 2.40 GHz, and a memory of 48G.

[0069] A comparative simulation is performed on the disordered flexible connector prediction results of the present invention and the existing disordered flexible connector prediction methods, and the results are shown in Table 1.

[0070] 2. Analysis of experimental results:

[0071] Table 1

[0072] Model Name ROC-AUC Prior Art 0.783 The Present Invention 0.802

[0073] Referring to Table 1, the ROC-AUC of the protein disordered flexible linker of the method of the present invention is 0.802, and the index is higher than that of the prior art method, which proves that the method of the present invention improves the recognition accuracy of the protein disordered flexible linker. ROC-AUC is the area under the curve of the receiver operating characteristic curve ROC, and the larger the value, the better the performance.

Claims

1. A method for predicting protein disordered flexible linkers based on multi-task learning, characterized in that, The steps include the following: (1) Obtain a multi-task data set: The obtained M unordered region protein amino acid sequences X IDRs and their corresponding unordered region labels Y IDRs constitute a protein unordered region prediction training data set At the same time, the obtained N unordered flexible linker protein amino acid sequences X DFLs and their corresponding unordered flexible linker labels Y DFLs constitute a protein unordered flexible linker prediction data set D DFLs ={X DFLs , Y DFLs}, where M≥1000 and N≥200; (2) Divide the protein disordered flexible linker prediction data set: Protein disordered flexible linker prediction dataset D based on sequence similarity DFLs is partitioned to obtain a protein disordered flexible linker prediction training dataset containing N train proteins and a test dataset containing N test proteins where N = N train + N test , N train > N test ; (3) Unify the lengths of the proteins in the protein disorder region prediction training data set, the training data set for disordered flexible linker prediction, and the test data set: Training dataset for predicting protein disordered regions Training dataset for predicting disordered flexible linkers and test datasets Unify the lengths of the proteins in the training dataset for predicting protein disordered regions with a protein length of G, obtaining a training dataset for predicting protein disordered regions Training dataset for predicting disordered flexible linkers and test datasets (4) Construct a protein feature representation matrix: Construct The corresponding dimensions are M×G×D, N train ×G×D, N test Unordered region training dataset protein feature representation matrix of ×G×D Disordered flexible linker training dataset protein feature representation matrix Disordered flexible linker test dataset protein feature representation matrix where D > 30; (5) Construct a protein disordered flexible linker prediction network model: Construct a protein disordered flexible linker prediction network model O including a shared layer and a disorder region tower layer and a disordered flexible linker tower layer connected to the output end of the shared layer and arranged in parallel, where the shared layer includes an embedding layer and a Transformer Encoder layer stacked in sequence; both the disorder region tower layer and the disordered flexible linker tower layer include a fully connected network and a Sigmod classifier stacked in sequence; (6) Iteratively train the protein disordered flexible linker prediction network model: (6a) Initialize the iteration number as t, the maximum iteration number as T, the trainable parameters of the shared layer, the disorder region tower layer, and the disordered flexible linker tower layer in the current prediction network model are w1, w2, and w3 respectively, and let t = 0; (6b) The protein feature representation matrix of the disordered region training data set The protein feature representation matrix of the disordered flexible linker training data set As the input of the prediction network model O for forward propagation, the disordered region prediction result of the disordered region training set is obtained The disordered flexible linker prediction result of the disordered flexible linker training set (6c) Update the trainable parameters w1, w2, and w3 of the shared layer, the disordered region tower layer, and the disordered flexible linker tower layer in the current prediction network model through the disordered region loss value L IDRs and the disordered flexible linker loss value L DFLs to obtain the prediction network model O of this iteration t ; (6d) Determine whether t = T holds. If so, obtain the trained prediction network model O * , otherwise, set t = t + 1, O = O * , and execute step (6b); (7) Obtain the protein disordered flexible linker prediction result: Unordered flexible linker test dataset protein feature representation matrix As the input of the trained prediction network model O * Perform forward propagation to obtain the unordered flexible linker prediction of the protein sequence in the unordered flexible linker test dataset in the unordered flexible linker prediction 2. The method according to claim 1, characterized in that The protein intrinsically disordered flexible linker prediction dataset D described in step (2) DFLs is partitioned, and the implementation steps are as follows: (2a) Using the BlastClust software package, for the disordered flexible linker dataset D DFLs of the protein amino acid sequence X DFLs with a sequence similarity greater than a are grouped into one class, obtaining the clustering result C = {C1, C2, …, C c , …, C k}, where a ≥ 25%, k ≥ 100, and C c represents the protein amino acid sequences contained in the c-th class protein set in the clustering result C; (2b) Divide C into two parts by category, obtaining a prediction training data set of disordered flexible linkers containing N train protein sequences and a test data set containing N protein sequences test protein sequences 3. The method according to claim 1, wherein The unification of the lengths of the proteins in the training dataset for predicting protein disordered regions, the training dataset for predicting disordered flexible linkers, and the test dataset described in step (3) is specifically as follows: For protein sequences with a length less than G, they are padded with 0s, and for the parts of protein sequences with a length exceeding G, they are truncated, resulting in a training dataset for predicting disordered regions with a protein length of G Training dataset for predicting disordered flexible linkers and the test dataset where G > 1000 4. The method according to claim 1, characterized in that, The implementation steps of constructing the protein feature representation matrix described in step (4) are: (4a) Calculate the evolutionary information of each protein using the PSI-BLAST tool to obtain an evolutionary information matrix of the disordered region training set with corresponding dimensions of M×G×D 1 、N train ×G×D 1 、N test ×G×D 1 of the disordered region training set Evolutionary information matrix of the disordered flexible linker training set proteins Evolutionary information matrix of the disordered flexible linker test set proteins where D 1 ≥20; (4b) Calculate the physicochemical information of each protein using the AAindex tool to obtain a physicochemical information matrix of the intrinsically disordered region training set with dimensions M×G×D 2 、N train ×G×D 2 、N test ×G×D 2 and a physicochemical information matrix of the intrinsically disordered flexible linker training set and a physicochemical information matrix of the intrinsically disordered flexible linker test set where D 2 ≥13, and the selected physicochemical information is indexed as CIDH920101, EISD860103, NISK860101, QIAN880105, ROBB760101, ROBB760108, ROBB760112, ROBB760113, CORJ870103, CORJ870106, CORJ870107, CORJ870108, MIYS990104 in AAindex;​ (4c) concatenate with respectively in the third dimension to obtain the protein feature representation matrices of the disordered region training datasets with dimensions M×G×D, N train ×G×D, N test ×G×D protein feature representation matrix of the disordered flexible linker training dataset protein feature representation matrix of the disordered flexible linker test dataset where D = D 1 + D 2 .

5. The method according to claim 1, wherein For the prediction network model O described in step (5), where: the dimension of the embedding layer in the shared layer is 64, and the fully connected network in both the disorder region tower layer and the disordered flexible linker tower layer only includes one hidden layer, and the number of hidden layer units is 32 for both.

6. The method according to claim 1, characterized in that, The protein feature representation matrix of the disordered region training data set described in step (6b) The protein feature representation matrix of the disordered flexible linker training data set Is used as the input of the prediction network model O for forward propagation. The implementation steps are as follows: (6b1) The embedding layer in the shared layer performs embedding representations on the protein feature representation matrix of the disordered region training dataset and the protein feature representation matrix of the disordered flexible linker training dataset respectively, to obtain the embedding vectors of the disordered region training set and the embedding vectors of the disordered flexible linker training set The Transformer Encoder layer performs forward propagation respectively to obtain the shared layer hidden vectors of the disordered region training set and the shared layer hidden vectors of the disordered flexible linker training set (6b2) Shared layer hidden vectors of the fully connected network in the disordered area tower layer and the disordered flexible connection body tower layer for the disordered area training set Shared layer hidden vectors of the disordered flexible connection body training set Perform forward propagation respectively to obtain the disordered area tower layer hidden vectors of the disordered area training set Tower layer hidden vectors of the disordered flexible connection body training set Sigmod classifier for Perform disordered area prediction and disordered flexible connection body prediction respectively to obtain the disordered area prediction results of the disordered area training set Disordered flexible connection body prediction results of the disordered flexible connection body training set 7. The method according to claim 1, characterized in that, The loss value L of passing through the disordered region described in step (6c) IDRs and the loss value L of the disordered flexible connector DFLs Update the trainable parameters w1, w2, and w3 of the shared layer, disordered region tower layer, and disordered flexible connector tower layer in the current prediction network model. The implementation steps are as follows: (6c1) The binary cross-entropy loss function is adopted, and by and as well as and calculate the loss value L of the disordered region IDRs , the loss value L of the disordered flexible linker DFLs : Among them, respectively represent the prediction result and label of the m-th protein in the disordered region training dataset; represents the prediction result and label of the n-th protein in the disordered flexible linker training dataset; (6c2)Through L IDRs and L DFLs calculate the overall loss L of the prediction network model O, and calculate the partial derivatives η1, η2, and η3 of L with respect to w1, w2, and w3: L = L DFLs + L IDRs (6c3) The gradient descent method is adopted to update the trainable parameters w1, w2, and w3 of the shared layer, the disordered area tower layer, and the disordered flexible connection body tower layer through the partial derivatives η1, η2, and η3 of w1, w2, and w3, and the prediction network model O of this iteration is obtained. t .

Citation Information

Patent Citations

  • Protein solubility prediction method based on multi-dimensional sequence embedding

    CN113223620A

  • Intelligent caching method for multi-task optimization

    CN114490447A