A method and system for predicting liquid-liquid phase separation proteins based on multiple features

By extracting multivariate features and using graph attention networks and integrated learners, the problems of high cost, low throughput and low prediction accuracy of liquid-liquid phase separation protein recognition in the prior art are solved, and higher prediction performance and stability are achieved.

CN119811507BActive Publication Date: 2025-06-17YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510292965.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-17
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing liquid-liquid phase protein identification methods have problems such as high cost, low throughput and low prediction accuracy. The second-generation prediction tools are limited to limited sequence characteristics, so the accuracy and stability of prediction still need to be improved.

Method used

The liquid-liquid phase separation protein prediction method is adopted based on multiple characteristics, and the sequence characteristics, spatial structural characteristics and secondary structural characteristics are extracted, and the graph attention network and integrated learner are used to mine and predict deep structural characteristics.

Benefits of technology

It effectively improves the prediction performance of liquid-liquid phase protein separation, enhances the accuracy and stability of prediction, and has higher prediction performance than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811507B_ABST
    Figure CN119811507B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting liquid-liquid phase separation proteins based on multiple features. When making predictions, this method simultaneously extracts sequence features, secondary structure features, and spatial structure features, constructs a residue contact map based on the spatial structure, and uses a graph attention network to extract deep structural features from the spatial structure graph constructed according to the residue contact map and sequence embedding. Based on the characteristic that the protein phase separation behavior is closely related to its structure, the prediction performance is effectively improved. In addition, this solution uses a three-layer stacked graph attention network to extract the structural features of proteins, combines the physicochemical features extracted from the sequence, and predicts liquid-liquid phase separation proteins through a stacked ensemble learning model, which can further improve the model prediction performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer bioinformatics, and particularly relates to a method and system for predicting liquid-liquid phase separation proteins based on multiple features. Background Art

[0002] Liquid-liquid phase separation (LLPS) refers to the formation of phase separation droplets with different components and properties by the interaction of certain biomacromolecules (such as proteins and RNAs) within cells. These droplets are similar to the separation state of oil droplets in water, forming a unique substructure within the cell. Identifying liquid-liquid phase separation proteins is of great significance in the research and treatment of diseases. Existing methods for identifying liquid-liquid phase separation proteins include traditional experimental methods (fluorescence microscopy, atomic force microscopy, nuclear magnetic resonance, etc.), first-generation prediction tools (PLAAC, LARKS, CatGRANULE, PScore, etc.), and second-generation prediction tools (FuzDrop, etc.).

[0003] Traditional experimental methods identify key proteins related to LLPS by directly observing the phase separation behavior. However, although experimental means can provide accurate qualitative and quantitative information, the high cost and low throughput of experimental operations have obvious limitations in large-scale data screening.

[0004] First-generation prediction tools use computational methods, most of which were not initially developed specifically for predicting phase separation tendency. Although they have been proven to be able to predict phase separation proteins later, the features used by these methods are simple and cannot capture complex phase separation driving factors, resulting in low prediction accuracy and being only suitable for preliminary screening.

[0005] Second-generation prediction tools are based on a large amount of experimental data and complex physicochemical features, combined with machine learning algorithms, and can perform large-scale prediction of LLPS proteins. However, these methods are limited to limited sequence features, and the accuracy and stability of prediction still need to be improved. Summary of the Invention

[0006] The object of the present invention is to propose a method and system for predicting liquid-liquid phase separation proteins based on multiple features for the problems existing in the prior art.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A method for predicting liquid-liquid phase separation proteins based on multiple features, the method comprising:

[0009] Preparing a positive and negative sample data set;

[0010] Obtaining the sequence data and structure data of each positive and negative sample;

[0011] Extract sequence features based on sequence data;

[0012] Generate sequence embeddings based on sequence data;

[0013] Construct a residue contact map based on structural data, obtain a spatial structure map based on the residue contact map and sequence embeddings, and extract the spatial structure features of the protein through a three-layer stacked graph attention network based on the spatial structure map;

[0014] Extract secondary structure features based on structural data;

[0015] Concatenate the sequence features, spatial structure features, and secondary structure features to obtain a concatenated feature set;

[0016] Use the concatenated feature set to train an ensemble learning machine composed of a base learner and a meta learner. The base learner takes the concatenated features as input and outputs secondary features, and the meta learner takes the secondary features as input and outputs the prediction result of the protein liquid-liquid phase separation behavior.

[0017] In the above liquid-liquid phase separation protein prediction method based on multiple features, obtain proteins with liquid-liquid phase separation behavior and unmodified as the positive sample dataset from the liquid-liquid phase separation database;

[0018] Use proteins that have never been reported to have liquid-liquid phase separation behavior in multiple organisms as the negative sample dataset.

[0019] In the above liquid-liquid phase separation protein prediction method based on multiple features, obtain the sequence data of all proteins from the UniProt database according to the UniProt ID of the proteins in the dataset, and obtain the structural data of all proteins from the AlphaFold database.

[0020] In the above liquid-liquid phase separation protein prediction method based on multiple features, extracting sequence features based on sequence data specifically includes:

[0021] Use the ESpritz and SEG algorithms to calculate the scores of intrinsically disordered regions and low complexity regions respectively;

[0022] Calculate the granule formation propensity score of the protein through the catGRANULE algorithm;

[0023] Use the PLAAC algorithm to calculate the prion-like propensity score of the protein;

[0024] Calculate the length, charged residue ratio, net charge per residue, Kappa value, Omega value, polyproline II helix propensity, average hydrophobicity, isoelectric point, proportion of residues contributing to chain extension, and proportion of residues promoting disordered regions of each protein sequence through the Localcider toolkit;

[0025] Use PScore to calculate the frequency of π-π interactions;

[0026] Obtain protein solubility through the Protein-sol software package.

[0027] In the above method for predicting liquid-liquid phase separation proteins based on multiple features, extracting secondary structure features based on structural data specifically includes directly obtaining secondary structure features such as α-helix, β-sheet, and β-turn of proteins through the DSSP software package.

[0028] In the above method for predicting liquid-liquid phase separation proteins based on multiple features, constructing a residue contact map based on structural data, obtaining a spatial structure map based on the residue contact map and sequence embedding, and extracting the spatial structure features of proteins through a three-layer stacked graph attention network based on the spatial structure map, specifically including:

[0029] Construct a residue contact map according to the structural data. The contact map of a protein with length L is represented as an L-order square matrix C = {c ij}, i, j = 1, 2, …, L:

[0030] (1)

[0031] If the Euclidean distance between the α-carbon atoms of two residues i, j is less than 8 , then it is defined that these two residues are in contact, represented as 1, otherwise 0;

[0032] Define the spatial structure map of the protein as G = (V, E), where V represents the set of nodes, each node corresponds to a residue of the protein, the initial node feature is the sequence embedding obtained by using the SeqVec model with sequence data as input, and E is an adjacency matrix, derived from the residue contact map;

[0033] That is, the spatial structure map consists of two parts: sequence embedding and adjacency matrix (i.e., residue contact map);

[0034] The three-layer stacked graph attention network takes the spatial structure map with initial node features as input and updates the node features through the following formula to output spatial structure features:

[0035] (2)

[0036] where K represents the number of heads of the multi-head attention mechanism, is a learnable linear transformation matrix, is the set of 1-order neighbors of node i, || represents the feature concatenation operation, is the activation function, is the normalized attention coefficient calculated by the k-th attention mechanism.

[0037] First, train the graph attention network alone to fix the parameters, and then use the trained graph attention network to extract spatial structure features for the training of subsequent base learners and meta-learners.

[0038] In the above liquid-liquid phase separation protein prediction method based on multiple features, the base learners include Random Forest, XGBoost, and LightGBM. Random Forest, XGBoost, and LightGBM respectively perform probability prediction based on the input spatial structure graph and output probability values , probability value , probability value , and splice the output results of the three base learners to obtain the secondary features:

[0039]

[0040] where , respectively represent the prediction probabilities of the i-th sample in the three base learners, and N is the total number of samples;

[0041] The meta-learner includes logistic regression.

[0042] In the above liquid-liquid phase separation protein prediction method based on multiple features, this method further includes using oversampling technology to increase the number of positive samples to balance the number of positive and negative samples.

[0043] A liquid-liquid phase separation protein prediction method based on multiple features, the method includes:

[0044] Obtain the sequence data and structure data of the protein to be predicted;

[0045] Extract sequence features based on the sequence data;

[0046] Generate sequence embeddings based on the sequence data;

[0047] Construct a residue contact map based on the structure data, obtain a spatial structure graph based on the residue contact map and sequence embeddings, and use a trained three-layer stacked graph attention network to extract the spatial structure features of the protein based on the spatial structure graph;

[0048] Extract secondary structure features based on the structure data;

[0049] Splice the sequence features, spatial structure features, and secondary structure features to obtain a spliced feature set;

[0050] Using the spliced feature set as input, the trained ensemble learner is used to output the prediction result of the liquid-liquid phase separation behavior of the protein to be predicted.

[0051] A liquid-liquid phase separation protein prediction system based on multiple features, including a first type of feature extraction module, a second type of feature extraction module, a feature splicing module, and a prediction module;

[0052] The first type of feature extraction module is used to extract sequence features based on sequence data;

[0053] The second type of feature extraction module is used to extract secondary structure features based on structural data, construct a residue contact map based on structural data, obtain a spatial structure map based on the residue contact map and the sequence embedding generated based on sequence data, and extract the spatial structure features of the protein based on the spatial structure map through a graph attention module;

[0054] The feature splicing module is used to splice the sequence features, spatial structure features, and secondary structure features to obtain a spliced feature set;

[0055] The prediction module is used to output the prediction result of the liquid-liquid phase separation behavior of the protein based on the spliced feature set.

[0056] The advantages of the present invention are as follows:

[0057] This solution extracts sequence features, secondary structure features, and spatial structure features simultaneously. In addition, a residue contact map is constructed based on the spatial structure, and a graph attention network is used to extract the spatial structure map constructed based on the residue contact map and sequence embedding for in-depth structural feature mining. Based on the characteristic that the protein phase separation behavior is closely related to its structure, the prediction performance can be effectively improved;

[0058] This solution uses a three-layer stacked graph attention network to extract the structural features of the protein, combines the physicochemical features extracted from the sequence, and predicts the liquid-liquid phase separation protein through a stacked ensemble learning model, which can further improve the model prediction performance. Description of the Drawings

[0059] Figure 1 It is a flow chart of the implementation method of the liquid-liquid phase separation protein prediction method based on multiple features of the present invention;

[0060] Figure 2 It is a model framework diagram of the liquid-liquid phase separation protein prediction method based on multiple features of the present invention;

[0061] Figure 3 It is a partial feature comparison box plot between positive and negative data sets in the liquid-liquid phase separation protein prediction method based on multiple features of the present invention;

[0062] Figure 4 The contribution of features extracted manually

[0063] Figure 5 The prediction flow chart of the liquid-liquid phase separation protein prediction method based on multiple features of the present invention

[0064] Figure 6 The ROC curve graph of the liquid-liquid phase separation protein prediction method based on multiple features of the present invention and other existing technologies

[0065] Figure 7 The system block diagram of the liquid-liquid phase separation protein prediction system based on multiple features of the present invention

[0066] Reference numerals: The first type of feature extraction module 01; the second type of feature extraction module 02; the figure attention module 021; the feature splicing module 03; the prediction module 04 Detailed implementation manners

[0067] This solution provides a liquid-liquid phase separation protein prediction method based on multiple features, as Figure 1 and Figure 2 shown, and the specific steps are as follows

[0068] The first step is to construct a data set: Screen and obtain highly reliable natural proteins with liquid-liquid phase separation behavior that have been experimentally verified or computationally obtained and are unmodified from databases such as PhaSepDB, LLPSDB, PhaSePro, and DrLLPS as the positive data set

[0069] Collect proteins that have never been reported to have liquid-liquid phase separation behavior from 10 representative species including humans, Caenorhabditis elegans, Drosophila melanogaster, Escherichia coli, Arabidopsis thaliana, mice, rats, Saccharomyces cerevisiae, Schizosaccharomyces pombe, and Xenopus laevis as the negative data set

[0070] Remove the proteins whose structures cannot be queried in known databases, and finally obtain the final data set through redundancy removal by the CD-HIT tool

[0071] The second step is to obtain sequence data and structure data: Obtain the FASTA format sequence data of all proteins in the data set from the UniProt (Universal Protein Resource) database according to the UniProt ID (the identifier used to uniquely identify proteins in the UniProt database) of the proteins in the data set, and obtain the PDB structure data files of all proteins from the AlphaFold database

[0072] Step 3, sequence feature extraction: Based on the sequence data, use the ESpritz and SEG algorithms to calculate the scores of intrinsically disordered regions (IDRs) and low complexity regions (LCRs) respectively, to reflect the characteristics of protein sequences in terms of flexibility, dynamics, and repetitive sequences;

[0073] Calculate the granule-formation propensity score of the protein through the catGRANULE algorithm;

[0074] Calculate the prion-like propensity score (LLR) of the protein using the PLAAC algorithm;

[0075] Calculate the length (Length), fraction of charged residues (FCR), net charge per residue (NCPR), Kappa value (measuring the distribution pattern of charged residues in the sequence), Omega value (reflecting the distribution pattern between charged residues and proline residues and other amino acid residues), polyproline II helix propensity (PPIIPropensity), mean hydropathy, isoelectric point, fraction of residues contributing to chain extension (Fraction Expanding), fraction of residues promoting disordered regions (DisorderPromotingFraction) for each protein sequence through the Localcider toolkit;

[0076] Calculate the frequency of π-π interactions (Pi interaction) using PScore;

[0077] Obtain the protein solubility (Percent-Sol) through the Protein-sol software package.

[0078] A partial feature comparison between the positive and negative datasets is as Figure 3 shown, Figure 3 reflecting the distribution differences of the values of each feature between the positive and negative datasets. It can be seen from the figure that these feature distribution differences are obvious and contribute to the classification prediction of positive and negative classes.

[0079] Step 4, structural feature extraction: Based on the structural data file PDB, directly obtain the secondary structure features of the protein such as alpha-helix (Alpha_Helix), beta-sheet (Beta_Sheet), and beta-turn (Beta_Turn) through the DSSP software package.

[0080] In addition, further construct a residue contact map according to the structural data file PDB. The contact map of a protein with length L is represented as an L-order square matrix C = {c ij}, i, j = 1, 2,..., L.

[0081]

[0082] That is, if the Euclidean distance between the alpha-carbon atoms of two residues i and j of a protein is less than 8 , then these two residues are defined as in contact with each other, denoted as 1, otherwise 0.

[0083] Secondly, the spatial structure diagram of the protein is defined as G=(V, E), where V represents the set of nodes, each node corresponding to a residue of the protein, and the adjacency matrix is derived from the residue contact map.

[0084] A structural feature extraction module composed of a three-layer stacked graph attention network (GAT) is adopted to extract protein structure features of a fixed length through a global max-pooling operation. GAT introduces an attention mechanism when aggregating neighbor node features, and can assign different weight coefficients according to the influence degree of neighbor nodes on the target node, so as to more effectively capture the connections between residues that are far apart but have spatial dependence relationships. Specifically as follows:

[0085] The initial input is the node feature , and the feature dimension of each node is . In this scheme, the protein amino acids are modeled as vectors of a fixed dimension d = 1024 through SeqVec embedding. After 3 layers of GAT processing, new node features are generated. The node feature update formula is:

[0086]

[0087] where K represents the number of heads of the multi-head attention mechanism, is a learnable linear transformation matrix, is the set of first-order neighbors of node i, || represents the feature concatenation operation, is the activation function, is the normalized attention coefficient calculated through the k-th attention mechanism, and is defined as follows:

[0088]

[0089] j is defined in the feature update formula , that is, an element in the set of first-order neighbors of node i.

[0090] In this scheme, the number of heads of the multi-head attention mechanism is set to K = 3 in the first two layers of GAT, and set to 1 in the last layer of GAT. The node features of the last layer go through a global max-pooling operation to output a feature vector with a fixed length of 128.

[0091] Step 5: Concatenate the obtained sequence features and structural features, and then use SMOTE oversampling to increase the number of positive samples and balance the number of positive and negative samples. The contribution degrees of the manually extracted features - sequence features and secondary structure features are as Figure 4 shown.

[0092] Step 6: Input the concatenated feature samples into the ensemble learning model for training and validation. This solution constructs a two-level model framework combining a base learner and a meta-learner. First, use Random Forest, XGBoost, and LightGBM as base learners. These three models are all algorithms based on decision trees and have strong non-linear modeling capabilities and robustness. Among them:

[0093] Random Forest realizes the classification and regression of input features through the integration of multiple decision trees, which can effectively reduce the overfitting risk of a single model. Its core idea is to construct multiple decision trees for the input feature X and output the prediction result by means of majority voting or average :

[0094]

[0095] where T is the number of decision trees, is the prediction result of the t-th tree.

[0096] XGBoost is an algorithm based on the gradient boosting framework, which adopts a weighted iterative optimization strategy to approximately iteratively optimize the first and second derivatives of the loss function , and its prediction result is expressed as:

[0097]

[0098] where is the learning rate is the incremental prediction model of the t-th tree.

[0099] LightGBM is an efficient implementation based on histogram decision trees. By grouping the data and accelerating the gradient update, it can quickly process large-scale high-dimensional data. Its optimization process is similar to XGBoost, but it further improves the splitting point selection strategy to improve efficiency. By training the input feature samples X respectively, these three base learners generate the predicted probability values corresponding to each sample , , .

[0100] Then, the output results of the above base learners are used as secondary features to construct the training data for the input meta-learner. Specifically, the predicted probabilities output by the base learners are concatenated to form a new secondary feature matrix Z:

[0101]

[0102] where , represent the predicted probabilities of the i-th sample in the three base learners respectively, and N is the total number of samples.

[0103] Logistic Regression is used as the meta-learner to further model and optimize the secondary features. Logistic Regression is defined as:

[0104]

[0105] where are the parameters of Logistic Regression, which are optimized by minimizing the cross-entropy loss function:

[0106]

[0107] where N is the number of samples, and are the true label and predicted probability of the i-th sample respectively.

[0108] During this process, to ensure the generalization performance of the model, the K-fold cross-validation technique is used to validate the base learners and the meta-learner. By dividing the dataset into K subsets, using one subset as the validation set and the remaining subsets as the training set, the model is trained and validated cyclically, and finally the validation results are integrated to evaluate the stability and performance of the model. Finally, the trained ensemble model is applied to unseen test feature samples. Initial prediction results , , are generated by the base learners, and then the secondary feature matrix Z is optimized and integrated by the meta-learner to obtain the final prediction result .

[0109] As Figure 5 shown, the testing process is as follows. The test feature samples are used as the proteins to be predicted:

[0110] Obtain the sequence data and structure data of the protein to be predicted;

[0111] Extract sequence features based on the sequence data, and the specific method is the same as the training process;

[0112] Generate sequence embeddings based on the sequence data;

[0113] Construct a residue contact map based on structural data, obtain a spatial structure map based on the residue contact map and sequence embedding, and use a trained three-layer stacked graph attention network to extract the spatial structure features of the protein based on the spatial structure map;

[0114] Extract secondary structure features based on structural data;

[0115] The specific methods for extracting spatial structure features and secondary structure features are the same as the training process.

[0116] Concatenate the sequence features, spatial structure features, and secondary structure features to obtain a concatenated feature set;

[0117] Using the concatenated feature set as input, use the above-trained ensemble learner to output the prediction result of the liquid-liquid phase separation behavior of the protein to be predicted 。

[0118] To verify the performance and effectiveness of this solution, this solution method is compared with the following existing methods:

[0119] PLAAC, used to predict the prion-like propensity of proteins, and evaluate its ability to form liquid-liquid phase separation by calculating the composition and distribution of specific amino acids in the protein sequence;

[0120] CatGRANUL, used to predict the granule formation propensity of proteins, and evaluate its ability to form liquid-liquid phase separation by analyzing the physicochemical properties and structural features of the protein sequence;

[0121] PScore, used to predict the liquid-liquid phase separation propensity of proteins, and evaluate its ability to form liquid-liquid phase separation by calculating the frequency of π-π interactions in the protein sequence;

[0122] FuzDrop, this method combines multiple sequence features and structural features, and evaluates the ability of proteins to form liquid-liquid phase separation through a deep neural network model.

[0123] The area under the ROC curve (AUROC) is used as the performance evaluation index of the model, and the ROC curve is drawn. The results are as Figure 6 shown. It can be seen that the ROC curve of this solution method forms a shape close to a right angle. Compared with other methods, it has the best curve shape, the model has the lowest false positive rate and the highest true positive rate, showing better model performance.

[0124] Furthermore, as Figure 7 shown, this solution also provides a liquid-liquid phase separation protein prediction system based on multiple features, including a first type of feature extraction module 01, a second type of feature extraction module 02, a feature concatenation module 03, and a prediction module 04;

[0125] The first type of feature extraction module 01 is used to extract sequence features based on sequence data. The extraction method is the same as that in the above model training process and will not be elaborated here.

[0126] The second type of feature extraction module 02 is used to extract secondary structure features based on structural data, construct a residue contact map based on the structural data, obtain a spatial structure map based on the residue contact map and the sequence embedding generated based on the sequence data, and extract the spatial structure features of the protein through the graph attention module 021. The graph attention module 021 is used to implement a three-layer stacked graph attention network, and extract spatial structure features through this graph attention network.

[0127] The feature splicing module 03 is used to splice the sequence features, spatial structure features and secondary structure features to obtain a spliced feature set.

[0128] The prediction module 04 is used to implement the above-mentioned ensemble learner and output the prediction result of the protein liquid-liquid phase separation behavior based on the spliced feature set.

[0129] This solution uses a three-layer stacked graph attention network to extract the structural features of proteins, combines the physicochemical features extracted from the sequence, and predicts liquid-liquid phase separation proteins through a stacked ensemble learning model, effectively improving the prediction accuracy of the system for whether a protein undergoes phase separation. Compared with various existing protein liquid-liquid phase separation prediction methods, it has higher prediction performance.

[0130] The specific embodiments described in this article are only illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

[0131] It should be noted that there is no absolute sequence relationship among the above steps when they are put into use. It only means that this embodiment is in the sequence given above. Those skilled in the art can change the order of the steps and still will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for predicting liquid-liquid phase separation proteins based on multivariate features, characterized in that: The method includes: Obtain proteins with liquid-liquid phase separation behavior and without modification from the liquid-liquid phase separation database as positive sample data sets; Proteins that have never been reported to have liquid-liquid phase separation behavior in various organisms are used as negative sample datasets; Obtain sequence data and structural data of each positive and negative sample; Extract sequence features based on sequence data: The scores of intrinsic disordered regions and low-complexity regions were calculated using the ESpritz and SEG algorithms, respectively; The particle formation propensity score of the protein was calculated by the catGRANULE algorithm; The prion-like propensity score of the protein was calculated using the PLAAC algorithm; The Localcider toolkit was used to calculate the length, proportion of charged residues, net charge per residue, Kappa value, Omega value, polyproline II helical propensity, average hydrophobicity, isoelectric point, proportion of residues that contribute to chain extension, and proportion of residues that promote disordered regions for each protein sequence; The frequency of π-π interactions was calculated using PScore; Protein solubility was obtained using the Protein-sol software package; Generate sequence embeddings based on sequence data; A residue contact graph is constructed based on the structural data, a spatial structure graph is obtained based on the residue contact graph and sequence embedding, and the spatial structural features of the protein are extracted based on the spatial structure graph through a three-layer stacked graph attention network, specifically including: The residue contact map is constructed based on the structural data. The contact map of a protein with a length of L is represented by an L-order square matrix C = {c ij }, i,j=1,2,…,L: (1) If the Euclidean distance between the alpha carbon atoms of two residues i and j is less than 8 , then the two residues are defined as contacting each other, which is represented by 1, otherwise it is 0; The spatial structure graph of the protein is defined as G = (V, E), where V represents a node set, each node corresponds to a residue of the protein, and the sequence embedding obtained by using the SeqVec model with sequence data as input is the initial node feature —, and E is an adjacency matrix. Derived from residue contact map; The three-layer stacked graph attention network takes the spatial structure graph with initial node features as input, and updates the node features to output spatial structure features through the following formula: (2) Among them, K represents the number of heads of the multi-head attention mechanism, is a learnable linear transformation matrix, is the first-order neighbor set of node i, || represents the feature concatenation operation, is the activation function, is the normalized attention coefficient calculated by the kth attention mechanism; Extraction of secondary structural features based on structural data; Splicing the sequence features, spatial structure features and secondary structure features to obtain a splicing feature set; Using the splicing feature set to train an integrated learner consisting of three base learners and a meta learner, the base learner takes the splicing feature as input and outputs secondary features, and the meta learner takes the secondary features as input and outputs a prediction result of protein liquid-liquid phase separation behavior; The three base learners make probability predictions based on the input spatial structure graph and output probability values ​​respectively. , probability value , probability value , concatenating the output results of the three base learners to obtain the secondary features: in , Respectively represent the predicted probability of the i-th sample in the three base learners, and N is the total number of samples; The meta-learner includes logistic regression.

2. The method for predicting liquid-liquid phase separation proteins based on multivariate features according to claim 1, characterized in that: The sequence data of all proteins were obtained from the UniProt database according to the UniProt ID of the proteins in the dataset, and the structural data of all proteins were obtained from the AlphaFold database.

3. The method for predicting liquid-liquid phase separation proteins based on multivariate features according to claim 1, characterized in that: Extracting secondary structural features based on structural data specifically includes directly obtaining secondary structural features of proteins including α-helix, β-sheet, and β-turn through the DSSP software package.

4. The method for predicting liquid-liquid phase separation proteins based on multivariate features according to claim 1, characterized in that: The three base learners include Random Forest, XGBoost, and LightGBM.

5. The method for predicting liquid-liquid phase separation proteins based on multivariate features according to claim 1, characterized in that: The method also includes using oversampling technology to increase the number of positive samples to balance the number of positive and negative samples.

6. A multi-feature liquid-liquid phase separation protein prediction system based on the multi-feature liquid-liquid phase separation protein prediction method according to any one of claims 1 to 5, comprising a first-type feature extraction module, a second-type feature extraction module, a feature splicing module and a prediction module: The first type of feature extraction module is used to extract sequence features based on sequence data; The second type of feature extraction module is used to extract secondary structure features based on structural data, construct a residue contact map based on the structural data, obtain a spatial structure map based on the residue contact map and sequence data, and extract spatial structural features of the protein based on the spatial structure map through a graph attention module; The feature splicing module is used to splice the sequence features, spatial structure features and secondary structure features to obtain a splicing feature set; The prediction module is used to output the prediction result of protein liquid-liquid phase separation behavior based on the splicing feature set.

Citation Information

Patent Citations

  • LLPS protein identification method, apparatus and device, and storage medium

    CN117688431A

  • Protein phase separation prediction model construction method, protein phase separation prediction model detection method and protein phase separation prediction model detection system

    CN118136117A