A method and system for predicting phase separation driving residues

By constructing a classification model that utilizes protein sequence and functional features, the probability of phase separation can be predicted and driving residues can be identified. This solves the problem of quantifying the contribution of individual amino acids and enables accurate analysis of phase separation studies and disease associations.

CN117012269BActive Publication Date: 2026-01-02SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310763821.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2026-01-02
Estimated Expiration
2043-06-26

AI Technical Summary

Technical Problem

Existing technologies struggle to quantify the contribution of individual amino acids to protein phase separation, making it difficult to precisely manipulate phase separation and identify the impact of pathogenic mutations on phase separation. This limits research on phase separation function and disease association analysis.

Method used

By constructing a classification model, the probability of phase separation is predicted by utilizing the sequence properties and functional characteristics of specific sequences in proteins. Furthermore, the impact of pathogenic mutations on phase separation is quantified by identifying driving residues through successive truncation of protein sequences.

Benefits of technology

It enables precise identification and quantification of phase separation-driven residues, provides an analytical framework for the relationship between phase separation and disease, and improves the accuracy of phase separation function research and the understanding of disease pathogenesis mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117012269B_ABST
    Figure CN117012269B_ABST
Patent Text Reader

Abstract

The application discloses a method and system for predicting phase separation driving residues, which integrates multidimensional sequence information and functional information of proteins, such as word vectors, sequence evolution information, amino acid composition, etc., to construct a machine learning model for predicting phase separation proteins and identifying driving residues. Based on the changes in phase separation prediction probability corresponding to different amino acid sequence changes, the contribution of each amino acid to the phase separation ability of the protein can be quantified. Based on PSPHunter, the phase separation probability of the protein missing each unit is calculated. The lower the score, the greater the impact of the unit on phase separation. The amino acid corresponding to the largest change in phase separation ability after truncation is considered to be a phase separation driving residue, and the region with relatively large influence and forming a valley is considered to be a driving region. The application realizes the prediction of phase separation proteins and the identification of driving residues, and can be further used for quantifying the influence of pathogenic mutations on phase separation and providing an analysis framework for analyzing the relationship between phase separation and diseases.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of phase separation prediction, and more particularly, to a method and system for predicting phase separation driving residues. BACKGROUND

[0002] Liquid-liquid phase separation is one of the important biophysical mechanisms that mediates the formation of membraneless compartments for macromolecules such as proteins and nucleic acids. In the past few years, liquid-liquid phase separation of biological molecules has become a unified biophysical mechanism for understanding important biological processes such as transcriptional control, autophagy, chromatin formation, and chromatin structure organization. Although some sequence features that can promote protein liquid-liquid phase separation have been known, how to quantitatively decode the contribution of individual amino acids to protein phase separation is still largely unknown, which limits the functional study of phase separation.

[0003] Currently, perturbing phase separation is the main strategy to study its function. The most widely used perturbation is the use of a chemical called 1,6-hexanediol, which is the only tool available for globally disrupting phase separation. Since 1,6-hexanediol widely interferes with weak hydrophobic interactions of protein-protein or protein-RNA, manipulating the phase separation of specific biological molecules is not specific. Directly knocking out / knocking down the expression of specific proteins is another option. However, it is often difficult to explain whether the functional output caused by this operation is caused by the perturbation of phase separation or other non-phase separation functions. Since the intrinsic disordered regions (IDRs) of phase separation proteins seem to be highly related to the protein phase separation ability, truncation of IDRs is an important strategy to disrupt protein phase separation. Typically, these IDRs are rich in charged and multivalent interaction centers. Therefore, another strategy to perturb phase separation is to replace the core amino acids with other types of amino acids, such as alanine. However, these perturbations always involve large fragments or a large number of amino acids. It is necessary to minimize the fragment or mine the combination of a small number of amino acids to precisely manipulate phase separation.

[0004] In addition, phase separation is an important missing mechanism for explaining pathogenic mutations. It has been reported that a variety of diseases, including neurodegenerative diseases, skin diseases, syndactyly syndromes, and malignancies, are related to phase separation, and it is not clear how disease mutations affect phase separation. The potential hypothesis is that the combination of small-effect mutations perturbs phase separation, which can lead to "genetic missingness" in complex disease susceptibility. However, without quantitative analysis of the impact of individual amino acids (driving residues) on phase separation, it is difficult to identify core mutations of phase separation among a large number of clinical mutation candidates. Therefore, it is necessary to identify pathogenic mutations in driving residues that can affect phase separation. SUMMARY

[0005] The primary object of the present application is to provide a method for predicting phase separation driving residues, realizing the prediction of phase separation proteins and the identification of driving residues, which can be further used to quantify the influence of pathogenic mutations on phase separation and provide an analysis framework for analyzing the relationship between phase separation and diseases.

[0006] A further object of the present application is to provide a system for predicting phase separation driving residues.

[0007] To solve the above technical problems, the technical solutions of the present application are as follows:

[0008] A method for predicting phase separation driving residues, comprising the following steps:

[0009] S1: Collecting experimentally verified phase separation proteins and background proteins, and establishing a training data set and an independent test data set according to the phase separation proteins and background proteins;

[0010] S2: Constructing a classification model, wherein the input of the classification model is sequence property-based features and functional features calculated according to a specific sequence in a protein, and the output of the classification model is the phase separation probability of the specific sequence;

[0011] S3: Training the classification model using the training data set and testing the classification model using the independent test data set to obtain a final classification model;

[0012] S4: Truncating the full-length protein sequence into fragments of a specific length one by one, and for each fragment, calculating the phase separation probability of the protein sequence missing the fragment using the final classification model to form a curve showing the influence of each amino acid on the phase separation ability of the protein;

[0013] S5: Obtaining the specific positions of the driving residues according to the regions in the curve that deviate greatly from the mean value of the curve.

[0014] Preferably, step S1 specifically comprises the following steps:

[0015] Obtaining experimentally verified phase separation proteins from PhaSepDB, LLPSDB, DrLLPS and PhaSePro databases;

[0016] Extracting human proteins by using the keyword "human" in the UniProt database and the screening criterion "Reviewed";

[0017] Obtaining background proteins by excluding the experimentally verified phase separation proteins from the human proteins;

[0018] Obtaining positive samples by filtering the phase separation proteins to limit the sequence similarity between any two proteins to be lower than a preset threshold;

[0019] randomly select the same number of proteins in the background protein as the positive sample as the negative sample;

[0020] The pairs of positive samples and negative samples are divided into a training data set and an independent test data set.

[0021] Preferably, the sequence attribute features in step S2 include amino acid composition, evolutionary conservation, predicted functional site annotation, and word vector, and the functional features include protein functional annotation information and network properties.

[0022] Preferably, the amino acid composition is specifically:

[0023] Each amino acid is divided into polar amino acid, charged amino acid, hydrophobic amino acid, and GP amino acid, wherein the polar amino acid includes N, Q, S, and T, the charged amino acid includes R, K, D, and E, the hydrophobic amino acid includes L, A, V, I, F, Y, M, H, W, and C, and the GP amino acid includes G and P.

[0024] The percentage of all amino acids in each category of amino acid in the total number of amino acids and the percentage of consecutive two amino acids are calculated respectively to generate a 20-dimensional vector representing the amino acid composition feature, representing the amino acid composition.

[0025] Preferably, the evolutionary conservation is specifically:

[0026] A sequence profile PSSM containing evolutionary information of protein sequences is used. For each protein sequence to be searched, the PSI-BLAST program is used to search the NR database of NCBI with parameters j=3 and e=0.001 to obtain a matrix representing the substitution frequency of all 20 amino acids at a specific position.

[0027] Each row of the matrix is subjected to Z-score conversion, and then for each column, the average score of each amino acid is calculated, and the matrix is compressed into a 20-dimensional vector representing the evolutionary properties of each type of amino acid. Then, the average scores of each category are summarized according to the polar, charged, hydrophobic, and GP amino acids, and a four-dimensional vector is used to represent them.

[0028] Distribution of amino acid conservation scores: Each amino acid of the query sequence obtains a conservation score, which is a relative entropy that can represent its evolutionary characteristics. Five statistical features are calculated, including the maximum, minimum, first quartile, second quartile, and third quartile. These five statistical features are used as a five-dimensional vector to represent the amino acid conservation score.

[0029] Search query: The HMM profiles of each protein are first Z-score normalized by row and then averaged by column, which yields a 20-dimensional vector for each amino acid. This 20-dimensional vector is further condensed into 4 dimensions according to the amino acid categories of polarity, charge, hydrophobicity, and GP.

[0030] Preferably, the predicted functional site annotation is specifically:

[0031] All potential phosphorylation sites, methylation sites, S-nitrosylation sites, and palmitoylation sites of each protein are extracted using the PTM predictor GPS; for the predicted mutation information, we perform saturation mutagenesis on each amino acid using Rhapsody, and according to the output of Rhapsody, we calculate the percentage of deleterious and neutral mutations, with the length of the protein representing the size of the protein, and each protein represented by a 14-dimensional vector.

[0032] Preferably, the word vector is specifically:

[0033] Each protein sequence is treated as a sentence, with its subsequences as words, and a distributed representation is constructed using the word2vec method.

[0034] Preferably, the protein function annotation information is specifically:

[0035] Each protein is annotated with experimentally validated PTM and mutation information, four types of post-translational modification information, including phosphorylation, acetylation, ubiquitination, and methylation, are downloaded from the PhosphoSitePlus database, the PTM frequency of each modification is defined as the number of annotated sites divided by the length of the protein sequence, mutation information is extracted from the HuVarBase database, by searching for the uniport ID of each protein, variants, including pathogenic variants and neutral variants, are obtained; protein expression abundance information is extracted from PAXdb, the age of each queried protein is inferred using phylogenetic analysis, and is extracted from ProteinHistorian, where 'PPODv4_Jaccard_families' is selected as the protein family database and 'Wagner parsimony' is selected as the ancestral reconstruction algorithm, it is checked whether these target proteins are enriched in essential genes, housekeeping genes, the basic gene list of genes is the intersection generated by whole genome single guide RNA screening and haploid gene capture screening, and the housekeeping gene list includes genes expressed in all tissues, each protein is represented by an 11-dimensional vector.

[0036] Preferably, the network attribute is specifically:

[0037] The PPI in the hippie database was used to establish a network, and three PS-related network features were calculated for each protein sequence to be searched, including the shortest distance between the protein sequence to be searched and its nearest PS, the average shortest distance between the protein sequence to be searched and the known PS, and the proportion of PSs near the protein sequence to be searched; based on the complete PPI network, four general attributes were also calculated, including degree, interactivity, clustering coefficient and average neighbor degree.

[0038] A system for predicting phase separation driving residues, comprising:

[0039] A data module that collects experimentally verified phase separation proteins and background proteins, and establishes a training data set and an independent test data set according to the phase separation proteins and background proteins;

[0040] A model construction module that constructs a classification model, the input of which is a sequence attribute-based feature and a functional feature calculated according to a specific sequence in a protein, and the output of which is the phase separation probability of the specific sequence;

[0041] A training and testing module that trains the classification model using the training data set and tests the classification model using the independent test data set to obtain a final classification model;

[0042] A calculation module that truncates a full-length protein sequence into fragments of a specific length one by one, and for each fragment, calculates the phase separation probability of the protein sequence missing the fragment using the final classification model to form a curve showing the influence of each amino acid on the phase separation ability of the protein;

[0043] A driving residue positioning module that obtains the specific position of the driving residue according to the area in the curve that deviates greatly from the mean value of the curve.

[0044] Compared with the prior art, the technical scheme of the present application has the beneficial effects that:

[0045] The present application realizes the prediction of phase separation proteins and the identification of driving residues by integrating the protein sequences and functional features in existing phase separation protein resources. The present application can be further used to quantify the influence of pathogenic mutations on phase separation and provide an analysis framework for analyzing the relationship between phase separation and diseases. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 The figure is a schematic diagram of the method of the present application.

[0047] Figure 2 The figure is a schematic diagram of the system module of the present application.

[0048] Figure 3 A schematic diagram for evaluating phase variation and detecting phase separation driving regions for the present application.

[0049] Figure 4 An experimental verification diagram for identifying phase separation driving regions for the present application.

[0050] Figure 5 An intention for predicting phase separation proteome by integrating multiple aspects of information for the present application.

[0051] Figure 6 A schematic diagram for phase separation proteins tend to be highly related to diseases.

[0052] Figure 7 A schematic diagram for pathogenic mutations. DETAILED DESCRIPTION

[0053] The accompanying drawings are included to provide a further understanding of the present application and are incorporated in and constitute a part of this patent document.

[0054] In order to better illustrate the present embodiment, some components in the drawings may be omitted, enlarged or reduced, and do not represent the actual size of the product.

[0055] It is understandable to some skilled in the art that some well-known structures and their descriptions in the drawings may be omitted.

[0056] The technical solutions of the present application will be further described below in combination with the drawings and embodiments.

[0057] Embodiment 1

[0058] The present embodiment provides a method for predicting phase separation driving residues, as shown in the following steps: Figure 1

[0059] S1: Collecting experimentally verified phase separation proteins and background proteins, and establishing a training data set and an independent test data set according to the phase separation proteins and background proteins;

[0060] S2: Constructing a classification model, the input of the classification model being sequence property-based features and functional features calculated according to a specific sequence in a protein, and the output of the classification model being a phase separation probability of the specific sequence;

[0061] S3: Training the classification model using the training data set and testing the classification model using the independent test data set to obtain a final classification model;

[0062] S4: Truncating a full-length protein sequence into fragments of a specific length one by one, for each fragment, calculating a phase separation probability of a protein sequence missing the fragment using the final classification model, and forming a curve showing the influence of each amino acid on the phase formation ability of the protein; ​

[0063] S5: obtaining the specific position of the driving residues according to the area deviating from the curve mean value in the curve.

[0064] Step S1 specifically comprises the following steps:

[0065] From the PhaSepDB, LLPSDB, DrLLPS and PhaSePro databases, 167 experimentally verified phase separation proteins (PS167) are obtained;

[0066] 20168 human proteins are extracted by using the keyword "human" in the UniProt database and the screening standard "Reviewed";

[0067] 20001 background proteins (BG20001) are obtained by excluding the experimentally verified phase separation proteins from the human proteins;

[0068] 135 positive samples (PS135) are obtained by further filtering the phase separation proteins by limiting the sequence similarity between any two proteins to be lower than 30%;

[0069] The same number of proteins as the number of positive samples are randomly selected from the background proteins as negative samples, and this process is repeated 100 times to avoid bias in generating negative samples.

[0070] The pairs of positive samples and negative samples are divided into training data sets and independent test data sets.

[0071] Taking 135 positive samples as an example, 70% (95 phase separation proteins) and the corresponding 95 negative samples are established as the training data set, and the remaining 30% (40 phase separation proteins) and 40 negative samples are set as the independent test data set.

[0072] The sequence attribute features in step S2 include amino acid composition, evolutionary conservation, predicted functional site annotation and word vector, and the functional features include protein function annotation information and network properties.

[0073] The amino acid composition is specifically:

[0074] According to Quiroz's research on how intrinsically disordered proteins encode phase behavior, each amino acid is divided into polar amino acid, charged amino acid, hydrophobic amino acid and GP amino acid, wherein the polar amino acid includes N, Q, S and T, the charged amino acid includes R, K, D and E, the hydrophobic amino acid includes L, A, V, I, F, Y, M, H, W and C, and the GP amino acid includes G and P; N, Q, S, T, R, K, D, E, L, A, V, I, F, Y, M, H, W, C, G and P are different amino acids.

[0075] The percentage of all amino acids in each category and the percentage of two consecutive amino acids in each category are calculated respectively, and a 20-dimensional vector representing the amino acid composition feature is generated to represent the amino acid composition. Specifically, each two consecutive amino acids is regarded as a unit, and the category information is further assigned to the unit. Then, the percentage of the polar-polar, polar-charge and other two categories is calculated to represent the second amino acid composition feature.

[0076] The evolutionary conservation is specifically:

[0077] For a sequence profile PSSM containing evolutionary information of protein sequences, for each protein sequence to be searched, the PSI-BLAST program is used to search the NR database of NCBI with parameters j=3 and e=0.001, and a matrix is obtained, which represents the substitution frequency of all 20 amino acids at a specific position.

[0078] In order to describe the characteristics of each protein, each row of the matrix is subjected to Z-score conversion, and then for each column, the average score of each amino acid is calculated, and the matrix is compressed into a 20-dimensional vector representing the evolutionary properties of each type of amino acid. Then, the average score of each category is summarized according to the polarity, charge, hydrophobicity and GP amino acids, and a four-dimensional vector is used to represent them.

[0079] Distribution of amino acid conservation score: Each amino acid of the query sequence obtains a conservation score, which is a relative entropy that can represent its evolutionary characteristics. In order to represent the distribution of the conservation score of each protein, five statistical features are calculated, including the maximum, minimum, first quartile, second quartile and third quartile. The five statistical features are used as a five-dimensional vector to represent the amino acid conservation score.

[0080] Search query: The HMM profile of each protein is first normalized by row Z-score, and then averaged by column, which generates a 20-dimensional vector for each amino acid. According to the amino acid categories of polarity, charge, hydrophobicity and GP, the 20-dimensional vector is further condensed into 4-dimensional.

[0081] The predicted functional site annotation is specifically:

[0082] Sequence-derived structural features can be used to reflect the structural preference of a protein, the intrinsic disordered information, RNA-binding proteins, post-translational modifications (PTMs) have been approved to be related to phase separation. In the current study, the secondary structure state of amino acids and the relative solvent accessibility of more than 20% accessible amino acids were assigned by the SPIDER2 program. The disordered amino acids are the percentage of helical amino acids, sheet amino acids, coil amino acids, contactable amino acids, disordered amino acids, RNA-binding amino acids and DNA-binding amino acids. Using a PTM predictor, GPS extracts all potential phosphorylation sites, methylation sites, S-nitrosylation sites and palmitoylation sites for each protein; for the predicted mutation information, we use Rhapsody to perform saturation mutagenesis on each amino acid, according to the output of Rhapsody, calculate the percentage of deleterious and neutral mutations, and use the length of the protein to represent the size of the protein, and each protein is represented by an X-dimensional vector.

[0083] The word vector is specifically:

[0084] Each protein sequence is regarded as a sentence, and its subsequence is a word, and a distributed representation is constructed using the word2vec method.

[0085] Generally, a corpus is needed to train the word embedding model by skip-gram or bag-of-words algorithm. By testing three data sets (or databases), i.e. PS135, PS167 and SwissProt, to select the corpus. Each protein sequence is converted into three sequences, which are composed of non-overlapping 3-grams. The word2vec is realized using the word bag model provided by the Gensim software package (https: / / radimrehurek.com / gensim / ). The maximum distance between the current and predicted words is set to 70, and the dimension of the word vector is set to 60. Considering the additivity of word embedding, we represent a specific sequence by adding the word vectors of all 3-grams.

[0086] The protein function annotation information is specifically:

[0087] Each protein was annotated with experimentally validated PTM and mutation information, four types of post-translational modification information were downloaded from PhosphoSitePlus database, including phosphorylation, acetylation, ubiquitination and methylation, the PTM frequency of each modification was defined as the number of annotated sites divided by the length of protein sequence, mutation information was extracted from HuVarBase database, which is a comprehensive database integrating 1000 Genomes, ClinVar, COSMIC, Humsavar and SwissVar, by searching the uniport ID of each protein, finally 774863 variants were obtained from 18318 proteins. Among these variants, 702,048 were pathogenic variants and 72,815 were neutral variants; protein expression abundance information was extracted from PAXdb (2017, Organism: WHOLE_ORGANISM), the age of each queried protein was inferred using phylogenetic analysis and extracted from ProteinHistorian, where 'PPODv4_Jaccard_families' was selected as the protein family database and 'Wagner parsimony' was selected as the ancestral reconstruction algorithm, whether these target proteins were enriched in essential genes, housekeeping genes was checked, the basic gene list of genes was the intersection generated by whole genome single guide RNA screening and haploid gene capture screening, the housekeeping gene list included 8874 genes expressed in all tissues, and each protein was represented by an x-dimensional vector.

[0088] The network attributes are specifically:

[0089] Since phase separation proteins generally have similar biological functions, these proteins are expected to be densely distributed in the protein-protein interaction (PPI) network, based on this assumption, the PPI in the hippie database was used to establish the network, for each protein sequence to be searched, three network features related to PS were calculated, including the shortest distance between the protein sequence to be searched and its nearest PS, the average shortest distance between the protein sequence to be searched and the known PS, and the proportion of PS near the protein sequence to be searched; based on the complete PPI network, four general attributes were also calculated, including degree, interactivity, clustering coefficient and average neighbor degree.

[0090] Classification model. After extracting the above features, models were developed to predict the phase separation of proteins. In this work, six machine learning algorithms were evaluated, including support vector machine (SVM), naive Bayes classifier (NB), neural network (NN), random forest (RF), LightGBM, and XGBoost. All these algorithms were implemented with the scikit-learn package. In SVM, the parameters c and g of the radial basis function were set to 2 and 0.125, respectively, and the number of trees in RF was set to 500. To integrate the features from multiple aspects, including voting on the binary scores of the first layer, averaging the probability scores of the first layer, directly integrating different features into one model, and two-layer stacking models. By comparing the performance, the two-layer stacking model based on XGBoost was finally selected. To evaluate these models, five-fold cross-validation was performed using the training set, and independent testing was performed using the test set. Widely used metrics were calculated, including recall, precision, F1 score, accuracy (ACC), and Matthews correlation coefficient (MCC).

[0091] Driving region detection. Take a specific protein as an example, the full-length protein sequence is truncated by a specific length fragment by turns (step size is 1), and the length of the fragment in this embodiment is 20. Each 20 consecutive amino acids is regarded as a unit to represent the intermediate amino acids. After evaluating the influence of all truncation possibilities on phase separation, a curve of the influence of each amino acid (excluding the first and last 10 amino acids) on the phase separation ability of the protein can be obtained. The region with the greater deviation from the mean value is considered to be the driving region that most affects the phase separation ability of the protein. Specifically, the deviation of each amino acid from the average phase separation probability is calculated, and the number of amino acids with the greatest deviation is determined based on the length of the sequence. In this embodiment, the number of control driving residues is considered to be between 20-40, so when the sequence length is greater than 2000, 1% of the total number of amino acids is taken, when the sequence length is between 1000-2000, 2% of the total number of amino acids is taken, when the sequence length is between 500-1000, 4% of the total number of amino acids is taken, and when the sequence length is less than 500, 5% of the total number of amino acids is taken. Then, the amino acids with the greatest deviation are connected according to the sequence position, and the specific position of the driving region can be obtained.

[0092] Example 2

[0093] This embodiment provides a system for predicting phase separation driving residues, as shown in Figure 2 The system comprises:

[0094] a data module, which collects experimentally verified phase separation proteins and background proteins, and establishes a training data set and an independent test data set according to the phase separation proteins and background proteins;

[0095] a model construction module, which constructs a classification model, an input of the classification model being sequence attribute-based features and functional features calculated according to a specific sequence in a protein, and an output of the classification model being a phase separation probability of the specific sequence;

[0096] a training and testing module, which trains the classification model by using the training data set, tests the classification model by using the independent test data set, and obtains a final classification model;

[0097] a calculation module, which truncates a full-length protein sequence into fragments of a specific length one by one, calculates, for each fragment, a phase separation probability of a protein sequence in which the fragment is deleted by using the final classification model, and forms a curve of the influence of each amino acid on the phase separation ability of a protein;

[0098] a driving residue positioning module, which obtains specific positions of driving residues according to a region in the curve that deviates from a mean value of the curve to a greater extent.

[0099] Embodiment 3

[0100] Based on Embodiments 1 and 2, the present embodiment provides the following specific embodiments:

[0101] The present embodiment is a machine learning algorithm (PSPHunter) using protein sequence features, which constructs a machine learning model for predicting phase separation proteins and identifying driving residues by integrating multi-dimensional sequence information and functional information of proteins, such as word vectors, sequence evolution information, and amino acid composition. Based on the changes in the phase separation prediction probability corresponding to different amino acid sequences, the contribution of each amino acid to the phase separation ability of a protein can be quantified. Specifically, each 20 consecutive amino acids is regarded as a unit to represent the intermediate amino acids. The phase separation probability of the protein missing each unit is calculated based on PSPHunter. The lower the score, the greater the influence of the unit on phase separation. The amino acid corresponding to the greatest change in phase separation ability after truncation is considered to be a phase separation driving residue, and the region that has a relatively large influence and forms a valley is considered to be a driving region.

[0102] The method of the present embodiment can detect known phase separation-related amino acids.

[0103] To quantify the contribution of each amino acid to the phase separation of proteins, the present embodiment designs a series of descriptors that convert local features into global features, such as word embedding vector features that encode amino acid composition, position-specific evolutionary information scoring matrix, and hidden Markov model atlas, etc. It makes PSPHunter sensitive to sequence changes. Based on the changes in the phase separation prediction probability corresponding to different amino acid sequence changes, we can quantify the contribution of each amino acid to the phase separation ability of the protein. Specifically, we regard every 20 consecutive amino acids as a unit to represent the intermediate amino acids. Based on PSPHunter, the phase separation probability of the protein missing each unit is calculated. The lower the score, the greater the impact of the unit on phase separation. After truncation, the amino acid corresponding to the largest change in phase separation ability is considered to be a phase separation driving residue, and the region that has a relatively large impact and forms a valley is considered to be a driving region Figure 3 a).

[0104] Basu et al. reported that alanine extension or reduction of the polyalanine region of the syndrome-associated protein HOXD13 accounted for the enhancement or reduction of phase separation ability, respectively. Using PSPHunter to evaluate the effect of increasing and reducing alanine on HOXD13, it was found to be consistent with the experimental results Figure 3 b). It was reported that the mutation of charged amino acids in the core pluripotency factor OCT4 to alanine and the mutation of phenylalanine in the DEAD-box helicase 4 (DDX4) to alanine significantly destroyed its phase separation ability, which was consistent with the prediction of PSPHunter Figure 3 c, d). Interestingly, PSPHunter can also reproduce the phase separation results when IDR is fused with mutant proteins, further proving that PSPHunter can accurately evaluate the effect of amino acid changes on the phase separation ability of proteins Figure 3 c, d).

[0105] To verify the reliability of PSPHunter in predicting phase separation driving residues of phase separation proteins, first compare the phase separation driving residues identified by PSPHunter with the key phase separation regions reported in the literature Figure 3 e). It is found that driving residues can cover most of the known phase separation regions Figure 3 f). Driving residues are closer to the consistent phase separation region than random fragments Figure 3 g). For specific regions, such as the proline-rich (PXX) region of the UBQLN2 protein, the glycine-rich region of the DNA-binding protein TDP43, the RGG region of the sm-like protein LSm4, and the linker connecting the first two SH3 domains of the NCK1 protein have been proven to be the key phase separation regions of the corresponding proteins. According to the distribution of the trough (driving region), it can be observed that they are highly consistent with the known phase separation regions Figure 3h, i, j, k, l). Other typical known phase separation protein driving residues are listed in the extended data figure S2. As PSPHunter provides quantified phase separation influence for each amino acid, it is able to determine the most critical amino acids for phase separation. Overall, this indicates that PSPHunter is able to finely identify driving residues and covers the phase separation driving region of most typical phase separation proteins.

[0106] Experimental validation of identified phase separation driving residues.

[0107] In addition, to demonstrate the function of potential driving residues in phase separation. The driving residues and non-driving residues were predicted by PSPHunter and further experimental validation was performed. It is expected that the truncation of driving residues will interfere with the phase separation ability of the protein, while the truncation of non-driving residues will have little effect on the phase separation ability of the protein Figure 4 a). For this purpose, we chose GATA binding protein 3 (GATA3), a typical phase separation protein, for further validation. It has been reported in the literature that the IDR (aa 105-260) of GATA3 exhibits a liquid-like behavior in HEK293 cells. Therefore, GATA3 protein was scanned by PSPHunter to determine the driving residues. According to the difference in phase separation ability, we predicted that the amino acids (aa) at positions 322-327 were driving residues, while the fragment of 88-93 aa was a non-driving residue control Figure 4 b). Notably, the predicted driving residues were not located in the IDR region of GATA3, but were close to its nucleic acid binding region. By truncating only the 6 predicted driving residues, the number of GATA3 puncta was significantly reduced Figure 4 e), while the dynamics of these puncta was also reduced Figure 4 f). We further verified these results in vitro. Without the complex microenvironment in vivo, we believe that in vitro experiments can better reflect the driving characteristics of the deleted amino acids. Photobleaching experiments showed that GATA3 with driving residue truncation was difficult to recover after photobleaching, while GATA3 with control amino acid truncation and wild-type GATA3 recovered quickly, respectively Figure 4 c, d). The same phenomenon was observed for the core pluripotency factor SOX2. We found that driving residues 19-24 aa could significantly affect the dynamics of SOX2 compared to control amino acids 266-268 aa.

[0108] PSPHunter applied to the identification of driving residues of OCT4 and RYBP proteins.

[0109] The PSPHunter algorithm has been successfully applied to identify the driving residues of the core pluripotency factor OCT4 Figure 5 a) and the PcG family protein RYBP (Figure 5 b) the driving residues. Taken together, these results indicate that PSPHunter can accurately identify the phase separation driving residues of proteins.

[0110] PSPHunter can predict phase separation proteome by integrating multi-faceted information

[0111] PSPHunter aims to predict phase separation proteins and driving residues. By linking phase separation proteins to human diseases and pathogenic mutations, PSPHunter can be further used to explore the role of phase separation in modulating diseases Figure 6 a) In this study, phase separation proteins were collected from PhaSepDB, LLPSDB, DrLLPS and PhaSePro, and scaffold proteins that can drive phase separation were manually extracted. This approach defined 167 phase separation proteins (PS167). Further, it was divided into training dataset and independent dataset.

[0112] Hydrophobic, electrostatic, π-π stacking and cation-π stacking, and other multivalent interactions drive macromolecules to phase separation, indicating that the driving force of phase separation may exist in the amino acid composition. It was found that sequence features showed significant differences between phase separation proteins and background proteins. In addition, protein expression, network properties, protein evolution information and other functional features were introduced in the prediction of phase separation proteins. Both sequence and functional feature models produced good classification performance Figure 6 b) By integrating sequence and functional features, the complementarity between sequence and functional information leads to performance improvement Figure 6 b) Compared with the state-of-the-art prediction models, PSPHunter shows advantages on various datasets Figure 6 c).

[0113] To evaluate the performance of PSPHunter in predicting phase separation proteins, PSPHunter was first applied to the whole proteome. A total of 898 phase separation proteins (phase separation proteome) were identified, of which 746 were newly predicted phase separation proteins. Then 99 RNA granules identified by b-isox method were collected, 78% of which can be covered by our phase separation proteome. Compared with random proteins, typical phase separation related proteins such as processing bodies, stress granules, disordered proteins, etc. have significantly higher probability of phase separation Figure 6 d) The distribution of PSPHunter scores of the whole proteome is consistent with that of the four-layer protein dataset ranked based on weighted experimental evidence. Overall, PSPHunter demonstrates advantages from a computational perspective.

[0114] To further validate the reliability of PSPHunter, the predicted proteins were tested for their phase separation ability by experiments. The predicted proteins were divided into higher ranked phase separation proteins and lower ranked phase separation proteins according to their PSPHunter scores. Functional analysis showed that the higher ranked proteins tend to be enriched in processes related to phase separation, such as RNA processing and regulation, while the lower ranked proteins tend to be enriched in G protein-coupled receptor (GPCR)-related processes. Further stratification and random selection of candidate proteins from potential and non-phase separation proteins were performed for experimental validation. These candidate proteins were purified in vitro. Phase diagrams showed that potential phase separation proteins formed green spots, and fluorescence recovery after photobleaching (FRAP) analysis showed the dynamics of these spots after photobleaching. In contrast, non-phase separation proteins had irregular overall shapes and were difficult to recover after photobleaching. In vivo experiments of these proteins showed the same phenomenon Figure 6 e, f). In summary, it demonstrated that the PSPHunter algorithm has the ability to identify reliable phase separation proteins.

[0115] 80% of phase separation proteins are highly related to diseases.

[0116] So far, only a small fraction of known diseases have been proven to be caused by phase separation dysfunction. Therefore, most diseases have not been directly linked to pathogenic mechanisms involving phase separation. So far, phase separation of condensates related to specific diseases has provided important new insights into the biological regulation of condensates and the pathogenic mechanisms of diseases. Therefore, we next asked what types of diseases are most relevant to phase separation.

[0117] Using all potential phase separation proteins as input, disease enrichment analysis was performed. Known phase separation and disease associations were revealed, including those related to neurological diseases (frontotemporal dementia, tauopathies). Interestingly, prostate cancer and leukemia were also significantly enriched. To further determine whether these diseases are highly related to phase separation, all proteins involved in the corresponding diseases were extracted. The results showed that the phase separation probability of these proteins was indeed significantly higher than that of random proteins, indicating that these diseases may be closely related to phase separation.

[0118] To further investigate which types of diseases are related to phase separation, disease-related proteins were collected from different categories, including the integrated resource for cancer information, the comprehensive disease database (DisGeNET), the Mendelian disease database (OMIM), and the rare disease database (Orphanet). According to the predicted phase separation protein group by PSPHunter, we found that nearly 80% of them are disease-related proteins Figure 7 a). In particular, most phase separation proteins are related to cancer, especially breast cancer and female reproductive organ cancer, cancer of the digestive system and urinary system,Figure 7 b) To further confirm the relationship between phase separation and diseases, the correlation between the phase separation probability and the number of diseases involved was investigated globally. Indeed, a positive correlation was observed (Fig. 4B). Figure 7 c) The segmented statistics further suggest that the more the number of diseases involved, the higher the phase separation probability of the protein. For example, most phase separation proteins, such as P53, TAU and MECP2, are involved in a large number of disease types (Fig. 4C). Figure 7 c, d) Taken together, these results confirm a hypothesis that condensate dysregulation can be a potential pathogenic mechanism across a broad spectrum of human diseases.

[0119] The same or similar reference numerals in different drawings denote the same or similar components;

[0120] The terms describing the positional relationship in the drawings are only used for illustrative description and should not be understood as a limitation to the present patent;

[0121] Obviously, the above-mentioned embodiments of the present application are only examples for clearly illustrating the present application, and are not intended to limit the implementation modes of the present application. Based on the above description, other different forms of changes or variations can be made by those skilled in the art. Here, it is not necessary and also impossible to enumerate all the implementation modes. Any modification, equivalent replacement and improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the claims of the present application.

Claims

1. A method of predicting phase separation driving residues, characterized by, The method comprises the following steps: S1: collecting experimentally verified phase separation proteins and background proteins, and establishing a training data set and an independent test data set according to the phase separation proteins and the background proteins; S2: constructing a classification model, wherein the input of the classification model is sequence attribute-based features and functional features calculated according to a specific sequence in a protein, and the output of the classification model is a phase separation probability of the specific sequence; S3: training the classification model by using the training data set and testing the classification model by using the independent test data set to obtain a final classification model; S4: truncating a full-length protein sequence into fragments of a specific length one by one, for each fragment, calculating a phase separation probability of a protein sequence in which the fragment is deleted by using the final classification model, and forming a curve of the influence of each amino acid on the phase separation ability of a protein; S5: obtaining specific positions of driving residues according to a region with a large deviation degree from a mean value in the curve; The sequence attribute-based features in step S2 include amino acid composition, evolutionary conservation, predicted functional site annotation and word vector, and the functional features include protein functional annotation information and network attribute; The evolutionary conservation is specifically as follows: For a sequence profile PSSM containing protein sequence evolution information, for each protein sequence to be searched, a matrix is obtained by searching the NR database of NCBI using the PSI-BLAST program with parameters j=3 and e=0.001, and the matrix represents the replacement frequency of all 20 kinds of amino acids at a specific position; Each row of the matrix is subjected to Z-score conversion, then the average score of each amino acid is calculated for each column, the matrix is compressed into a 20-dimensional vector to represent the evolutionary attribute of each type of amino acid, then the average scores of each category are summarized according to polarity, charge, hydrophobicity and GP amino acids, and a four-dimensional vector is used to represent the average scores. Distribution of amino acid conservation scores: each amino acid of the query sequence obtains a conservation score, and the conservation score is a relative entropy capable of representing the evolutionary characteristics, and five statistical characteristics including maximum, minimum, first quartile, second quartile and third quartile are calculated, and the five statistical characteristics are used as a five-dimensional vector to represent the amino acid conservation score. Search query: the HMM profile of each protein is first subjected to Z-score normalization by row and then subjected to averaging by column, which produces a 20-dimensional vector for each amino acid, and the 20-dimensional vector is further condensed into a 4-dimensional vector according to the polarity, charge, hydrophobicity and GP amino acid categories.

2. The method of predicting phase separation driving residues according to claim 1, wherein, Step S1 specifically comprises the following steps: Obtaining experimentally verified phase separation proteins from PhaSepDB, LLPSDB, DrLLPS and PhaSePro databases; Extracting human proteins by using the keyword "human" in the UniProt database and the screening standard "Reviewed"; Obtaining background proteins by excluding the experimentally verified phase separation proteins from the human proteins; The phase separation proteins are filtered by limiting the sequence similarity between any two proteins below a preset threshold, to obtain positive samples; A same number of proteins as the positive samples are randomly selected from the background proteins as negative samples; The pairs of positive samples and negative samples are divided into a training data set and an independent test data set.

3. The method of predicting phase separation driving residues according to claim 2, wherein, The amino acid composition is specifically: Each amino acid is divided into polar amino acids, charged amino acids, hydrophobic amino acids and GP amino acids, wherein the polar amino acids include N, Q, S and T, the charged amino acids include R, K, D and E, the hydrophobic amino acids include L, A, V, I, F, Y, M, H, W and C, and the GP amino acids include G and P; The percentage of all amino acids in each category of amino acid in the total number of amino acids and the percentage of consecutive two amino acids are calculated respectively, to generate a 20-dimensional vector representing the amino acid composition feature, representing the amino acid composition.

4. The method of predicting phase separation driving residues according to claim 3, wherein, The predicted functional site annotation is specifically: All potential phosphorylation sites, methylation sites, S-nitrosylation sites and palmitoylation sites of each protein are extracted by using the PTM predictor GPS; for the predicted mutation information, we use Rhapsody to perform saturation mutagenesis on each amino acid, and according to the output of Rhapsody, the percentages of deleterious and neutral mutations are calculated, and the length of the protein is used to represent the size of the protein, and each protein is represented by a 14-dimensional vector.

5. The method of predicting phase separation driving residues according to claim 4, wherein, The word vector is specifically: Each protein sequence is regarded as a sentence, and its subsequence is regarded as a word, and a distributed representation is constructed by using the word2vec method.

6. The method of predicting phase separating driver residues of claim 5, wherein, The protein function annotation information is specifically: Experimental verified PTM and mutation information are used to annotate each protein, four types of post-translational modification information are downloaded from the PhosphoSitePlus database, including phosphorylation, acetylation, ubiquitination and methylation, the PTM frequency of each modification is defined as the number of annotated sites divided by the length of the protein sequence, mutation information is extracted from the HuVarBase database, by searching the uniport ID of each protein, variants are obtained, including pathogenic variants and neutral variants; protein expression abundance information is extracted from PAXdb, the age of each queried protein is inferred using phylogenetic analysis, and is extracted from ProteinHistorian, wherein 'PPODv4_Jaccard_families' is selected as the protein family database, 'Wagnerparsimony' is selected as the ancestral reconstruction algorithm, and whether the target protein is enriched in essential genes and housekeeping genes is checked, the basic gene list of the genes is the intersection generated by whole genome single guide RNA screening and haploid gene capture screening, and the housekeeping gene list includes genes expressed in all tissues, and each protein is represented by an 11-dimensional vector.

7. The method of predicting phase separating driver residues according to claim 6, wherein, The network attribute is specifically: The PPI in the hippie database was used to establish a network, and three PS-related network features were calculated for each protein sequence to be searched, including the shortest distance between the protein sequence to be searched and its nearest PS, the average shortest distance between the protein sequence to be searched and known PSs, and the proportion of PSs near the protein sequence to be searched. Based on the complete PPI network, four general attributes were also calculated, including degree, interactivity, clustering coefficient, and average neighbor degree.

8. A system for predicting phase separation driving residues, characterized in that, The system applies the method for predicting phase separation driving residues according to any one of claims 1 to 7, comprising: a data module that collects experimentally verified phase separation proteins and background proteins, and establishes a training data set and an independent test data set according to the phase separation proteins and the background proteins; a model construction module that constructs a classification model, wherein the input of the classification model is a sequence attribute feature and a functional feature calculated according to a specific sequence in a protein, and the output of the classification model is a phase separation probability of the specific sequence; a training and testing module that trains the classification model using the training data set and tests the classification model using the independent test data set to obtain a final classification model; a calculation module that truncates a full-length protein sequence into fragments of a specific length one by one, and for each fragment, calculates a phase separation probability of a protein sequence missing the fragment using the final classification model to form a curve showing the influence of each amino acid on the phase separation ability of a protein; a driving residue positioning module that obtains the specific positions of driving residues according to regions in the curve that deviate greatly from the mean value of the curve.

Citation Information

Patent Citations

  • Prediction method and prediction device for membrane protein residue interaction relation

    CN106650309A

  • System and method for discovering drug active site of protein using pathogenic mutation

    KR102380934B1