Method and system for identifying liquid-liquid phase separation regulatory protein based on artificial intelligence technology
Through methods based on artificial intelligence technology, the data set is constructed and balanced, and the ESM-2 model and MLP classifier are used to solve the problem of time-consuming and costly identification of regulatory proteins in traditional biochemical methods, achieving efficient and accurate protein recognition.
Patent Information
- Application Number
- CN202510107352.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional biochemistry-based methods are time-consuming and costly in identifying liquid-liquid phase isolation of regulatory proteins, making it difficult to meet the needs of high-throughput research and practical applications.
Using an artificial intelligence technology-based method, the construction of benchmark data sets, data set balance, feature coding, and model training and evaluation is carried out, and the ESM-2 model and a multi-layer perceptron (MLP) classifier are used to achieve efficient identification of liquid-liquid phase isolation regulatory proteins.
This method significantly improves the research efficiency, reduces the research cost, and achieves high accuracy recognition of liquid-liquid phase isolation regulatory proteins, achieving an accuracy rate of 77.78%.
Smart Images

Figure CN120048360A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a method and system for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology. Background Art
[0002] Regulatory proteins play a central role in the process of liquid-liquid phase separation (LLPS). They are crucial for the formation, stability, and maintenance of the dynamic properties of LLPS, ensuring an appropriate phase separation response to cellular signals. Targeting these regulatory proteins is crucial for manipulating LLPS applications in biotechnology, materials science, and medicine. Therefore, identifying these regulatory proteins is particularly critical.
[0003] There are various methods for identifying regulatory proteins in liquid-liquid phase separation (LLPS), including conventional biochemical assays, mass spectrometry, immunofluorescence co-localization, and in situ hybridization techniques, etc. These methods have their own advantages and can analyze the functions and mechanisms of proteins in LLPS from different levels. For example, proteins such as TDP-43 and FUS have been revealed to play important roles in neurodegenerative diseases and cellular responses through these techniques.
[0004] However, traditional biochemistry-based methods for identifying these regulatory proteins are time-consuming and costly, and there may be certain errors and limitations in the experimental process. These methods usually require complex operation procedures and a large number of samples, making high-throughput screening and large-scale research difficult. Therefore, to address this challenge, it is particularly important to develop a powerful computational model to identify and predict LLPS-related regulatory proteins. By utilizing bioinformatics and machine learning techniques, potential phase separation proteins can be screened more efficiently from large-scale data, providing important theoretical support and experimental guidance for research, while reducing costs and improving research efficiency.
[0005] Through the above analysis, the problems and defects existing in the prior art are:
[0006] Traditional biochemistry-based methods for identifying these regulatory proteins are time-consuming and costly. Summary of the Invention
[0007] Aiming at the problems existing in the prior art, the present invention provides a method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology.
[0008] The present invention is implemented as follows. A method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology includes:
[0009] Step 1, constructing a benchmark data set;
[0010] Step 2, balancing the data set;
[0011] Step 3, feature encoding;
[0012] Step 4, model training and evaluation.
[0013] Furthermore, the construction of the benchmark dataset:
[0014] Collect 987 experimentally verified regulatory proteins involved in the LLPS process from the DrLLPS database to form a positive dataset; for the negative dataset, extract protein sequences from the UniProt database, including cytoskeletal proteins, membrane-associated proteins, highly structured proteins (such as globulins), proteins involved in forming stable structures, and single-chain proteins, which are assumed to be less likely to play a regulatory role in LLPS; use the CD-HIT tool, set the threshold for the positive dataset to 70%, and the threshold for the negative dataset to 40%; through this series of sorting and screening, finally obtain 913 positive protein sequences and 6584 negative protein sequences.
[0015] Furthermore, the dataset balancing:
[0016] Randomly select 113 positive sequences from the positive dataset as the test set, and the remaining 800 positive sequences are used for training; similarly, randomly select 184 negative sequences from the negative dataset as the test set, and the remaining 6400 negative sequences are used for training; the remaining negative sequences are randomly divided into eight subsets, each subset containing 800 sequences, to match the 800 positive sequences to ensure the complete balance of the training dataset; finally, obtain 8 balanced training datasets; then combine the selected 113 positive and 184 negative sequences to get a test dataset containing 297 protein sequences; due to the priority given to the balanced training dataset, a slightly higher proportion of negative sequences is finally obtained in the test dataset.
[0017] Furthermore, the feature encoding:
[0018] Use the ESM-2 model to encode the protein sequences to generate embedding vectors, and this model can capture the sequence, structure, and physicochemical properties of proteins.
[0019] Furthermore, the model training and evaluation:
[0020] Model evaluation:
[0021] Use the sensitivity (Sn), specificity (Sp), accuracy (ACC), and area under the curve (AUC) metrics; sensitivity measures the ability of the model to correctly identify regulatory proteins in LLPS, specificity evaluates its accuracy in distinguishing non-regulatory proteins, and accuracy provides an overall performance assessment of the model on the two classes of regulatory and non-regulatory proteins; AUC is derived from the receiver operating characteristic (ROC) curve and is particularly important in binary classification tasks as it reflects the model's ability to separate positive and negative classes; the formulas are as follows:
[0022]
[0023] Where TP represents that a regulatory protein is correctly identified as a regulatory protein; TN represents that a non-regulatory protein is correctly identified as a non-regulatory protein; FP represents that a non-regulatory protein is misclassified as a regulatory protein, and FN represents that a regulatory protein is mislabeled as a non-regulatory protein;
[0024] Model selection:
[0025] XG-Boost (XGB)
[0026] XGB uses gradient boosting decision trees for predictive modeling, and the final prediction result is the sum of the predictions of all individual trees; it can optimize the objective function, which combines the loss function l and the regularization term Ω to prevent overfitting and improve generalization ability; the general form of the objective function can be expressed as:
[0027]
[0028] Where n represents the number of samples, y t and y p correspond to the true value and the predicted value of the data sample respectively; the term l represents the loss function, and K and f k represent the number of trees and the k-th tree respectively; Ω(f k ) is a regularization term that helps to mitigate overfitting;
[0029] Convolutional neural network:
[0030] The convolutional neural network includes two one-dimensional convolutional layers and four fully connected layers, and its structure can be represented by the following formula:
[0031] (I * K)(i, j) = ∑m∑n(i + m, j + n).K(m, n)
[0032] Where I is the input, K is the convolutional kernel, and * represents the convolution operation; (i, j) represents the spatial position in the output, and (m, n) represents the iteration in the spatial dimension; the pooling layer minimizes the computational burden and mitigates the impact of overfitting by reducing the spatial dimension of the feature map; the final fully connected layer aggregates the learned features for the classification task;
[0033] Multi-Layer Perceptron (MLP)
[0034] The MLP can be expressed as:
[0035] y p = g(W (l) . f(W (l-1) . f(…f(W (1) . x + b (1) )…)+b (l-1) )+b (l) )
[0036] where x represents the input data, W l is the weight matrix of the l-th layer, b l is the bias function of the l-th layer, f l is the activation function of the l-th layer, g is the activation function of the output layer, and y p is the output;
[0037] Model training.
[0038] Furthermore, the model training:
[0039] The protein feature representations extracted from each balanced dataset by the protein language model ESM2-36 are used as the input for model training; three classifiers, XGB, MLP, and CNN, are used to train the model respectively, and 10-fold cross-validation is performed on each balanced dataset to ensure the robustness of training and validation; during the training process, a custom callback is used to monitor the validation accuracy, and the weights of the best-performing model are retained for testing; through this process, a total of 8 individual models are trained; to combine the prediction results of these models, the majority voting method of the output results is used for ensemble prediction, and if the number of votes in favor and against is the same, the final decision is made by calculating the average probability and using a threshold of 0.5; finally, the ensemble method is validated on the test dataset.
[0040] Another object of the present invention is to provide a system for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology, including:
[0041] A construction module for constructing a benchmark dataset;
[0042] A balancing module for dataset balancing;
[0043] An encoding module for feature encoding;
[0044] A training module for model training and evaluation.
[0045] Another object of the present invention is to provide a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology.
[0046] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology.
[0047] Another object of the present invention is to provide an information data processing terminal for implementing the system for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology.
[0048] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0049] Existing methods for identifying liquid-liquid phase separation (LLPS) regulatory proteins mainly rely on traditional biochemical wet experiments. Although these experiments are accurate, they are often time-consuming and costly, and it is difficult to meet the needs of high-throughput research and practical applications. The limitations of traditional methods have prompted scientists to seek more efficient and cost-effective alternatives to quickly screen potential LLPS regulatory proteins and provide a basis for further research.
[0050] The present invention first proposes a method for identifying LLPS regulatory proteins based on artificial intelligence technology, filling the research gap in related fields. Different from traditional wet experiments, this method relies on deep learning technology in computer science, uses a high-quality protein sequence dataset and an advanced pre-trained protein language model ESM2-36 to extract feature information from the sequences, and realizes the efficient identification of LLPS regulatory proteins. This innovative method has greatly improved the research efficiency and reduced the research cost.
[0051] The present invention constructs a robust dataset containing 913 positive protein sequences and 6584 negative protein sequences. The positive sequences are obtained by screening experimentally verified LLPS regulatory proteins from the DrLLPS database; the negative sequences are selected from the UniProt database for protein categories that have no direct relationship with LLPS, such as cytoskeletal proteins and membrane-associated proteins. The CD-HIT tool is used to perform redundancy removal on the positive and negative datasets respectively to ensure the representativeness and high quality of the data, providing a solid data foundation for subsequent analysis.
[0052] To address the problem of the imbalance in the number of positive and negative samples, the present invention has carried out a meticulous balancing process on the dataset. 113 sequences were randomly selected from the positive dataset and 184 sequences were randomly selected from the negative dataset to construct an independent test set containing 297 protein sequences. The remaining 6400 negative sequences were randomly divided into eight subsets, each subset containing 800 sequences, which were matched with 800 positive sequences, ultimately generating 8 balanced training datasets. This process effectively avoids classification errors caused by data bias and improves the robustness of the model.
[0053] The pre-trained protein language model ESM2-36 was used to encode the features of the protein sequences, extracting their sequence, structure, and physicochemical feature information. The present invention adopted a multi-layer perceptron (MLP) classifier for model training and performed 10-fold cross-validation on each balanced dataset to optimize the model performance. The training process was monitored through custom callbacks, and the optimal model weights were retained for testing. Finally, the prediction results of all models were subjected to majority voting through an ensemble learning method, further improving the accuracy and stability of the prediction.
[0054] The experimental results on the test dataset show that the proposed classifier achieved an accuracy of 77.78%, demonstrating excellent performance. This method is not only efficient and accurate but also significantly reduces the time and economic costs of identifying LLPS regulatory proteins. The present invention provides a novel solution for the intelligent identification of LLPS regulatory proteins, opening up a new direction for research and application in the field of bioinformatics. At the same time, the source code and dataset used in the present invention have been publicly available on a public platform, ensuring the transparency and reproducibility of the method, and having broad research and application prospects.
[0055] Facilitate the discovery of drug targets and the development of new drugs. Liquid-liquid phase separation plays an important role in the occurrence and development of diseases such as cancer and neurodegenerative diseases. Identifying LLPS regulatory proteins based on AI technology can help pharmaceutical companies and research institutions discover new drug targets, especially in the treatment of refractory diseases such as Alzheimer's disease, Parkinson's disease, and cancer. This technology provides new ideas and methods for the development of targeted drugs, helping researchers identify potential regulatory proteins as drug targets. With the development of AI identification technology, it may greatly improve the efficiency of drug development in the future and accelerate the process of new drug development, thus bringing huge market competitiveness and commercial profits to pharmaceutical companies.
[0056] For example, Alzheimer's disease is a neurodegenerative disease mainly characterized by cognitive impairment. Its typical feature is the abnormal accumulation of amyloid-β proteins, which form "protein gels" or insoluble aggregates through liquid-liquid phase separation. Liquid-liquid phase separation regulatory proteins may affect their aggregation, expansion, and pathological changes within cells by regulating the phase separation process of these proteins. If regulatory proteins can stabilize or alter the dynamics of these molecular droplets, they may be able to intervene in the formation of amyloid plaques, thereby affecting the occurrence and development of the disease. Therefore, small molecule drugs can be designed targeting liquid-liquid phase separation regulatory proteins, and then the LLPS process can be regulated to restore the normal function of proteins, reduce their aggregability, and prevent neuronal damage.
[0057] In the field of liquid-liquid phase separation, traditional techniques generally rely on laboratory biochemical techniques such as immunoprecipitation and microscopy. However, these methods are often limited by experimental conditions, the experience of operators, and equipment, etc. They are often time-consuming and costly, and it is difficult to meet the needs of high-throughput research and practical applications.
[0058] The technical solution of the present invention breaks through this barrier. It uses artificial intelligence technology to replace traditional experimental methods. By means of computer modeling of protein sequences, data analysis, etc., it overcomes the problem that experimental data cannot be processed on a large scale, and successfully combines AI technology with the field of liquid-liquid phase separation. The successful application of this technology demonstrates the integration of cross-field technologies and shows the great potential of artificial intelligence technology in the field of life sciences. Brief Description of the Drawings
[0059] Figure 1 It is a flowchart of the method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology provided by an embodiment of the present invention.
[0060] Figure 2 It is a block diagram of the system structure for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology provided by an embodiment of the present invention.
[0061] Figure 3 It is a graph showing the average cross-validation sensitivity, specificity, and accuracy of the model trained on datasets 1 to 8 provided by an embodiment of the present invention, with error bars indicating the maximum and minimum values.
[0062] Figure 4 It is a graph showing the average test sensitivity, specificity, and accuracy of the model trained on datasets 1 to 8 provided by an embodiment of the present invention, with error bars indicating the maximum and minimum values.
[0063] Figure 5 It is a graph showing the cross-validation AUC distribution of the model trained on datasets 1 to 8 provided by an embodiment of the present invention, with bar graphs indicating the maximum and minimum values.
[0064] Figure 6 It is the test AUC distribution of the model trained on datasets 1 to 8 provided by the embodiments of the present invention. The bar chart shows the maximum and minimum values.
[0065] Figure 7 It is a comparison chart of sensitivity, specificity, and accuracy using XGB, MLP, and CNN models respectively provided by the embodiments of the present invention. Detailed implementation manners
[0066] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0067] As Figure 1 shown, a method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology provided by the embodiments of the present invention includes the following steps:
[0068] S101, construction of a benchmark dataset;
[0069] S102, dataset balancing;
[0070] S103, feature encoding;
[0071] S104, model training and evaluation.
[0072] The identification of liquid-liquid phase separation regulatory proteins depends on the construction of a high-quality dataset. The positive dataset is collected by collecting experimentally verified regulatory protein sequences from the DrLLPS database. These proteins are involved in the LLPS process and have clear experimental support. The negative dataset is derived from the UniProt database, and protein categories that have no obvious relationship with LLPS are screened out, including cytoskeletal proteins, membrane-associated proteins, highly structured proteins, etc. The functions of these proteins are usually related to the formation of stable structures and are assumed not to be involved in LLPS regulation. To avoid data redundancy and noise, the CD-HIT tool is used to remove redundancy from the dataset, setting the threshold for the positive dataset to 70% and the threshold for the negative dataset to 40%. Through this process, a benchmark dataset of 913 positive sequences and 6584 negative sequences is constructed.
[0073] Due to the significant difference in the number of positive and negative samples, it may lead to bias in the model training process. To address this issue, undersampling or oversampling techniques are used to balance the dataset. Specific methods include data augmentation for positive samples, using techniques such as data transformation and perturbation to expand the sample size; and random sampling for negative samples, retaining a representative subset to reduce data imbalance. Through data balancing, the robustness of the model can be improved, and the error in the classification process can be reduced.
[0074] Feature extraction of protein sequences is an important step in model training. Multiple bioinformatics methods are used to encode the features of the sequences, such as amino acid composition (AAC), amino acid pair binary pattern (DPC), and physicochemical properties (such as hydrophobicity, polarity, etc.). These features can reflect the functional and structural characteristics of protein sequences. By combining features of different dimensions, the attributes of proteins can be comprehensively described, providing support for subsequent model training.
[0075] Build an artificial intelligence-based classification model, such as support vector machine (SVM), random forest (RF), or deep learning model (such as convolutional neural network CNN). By inputting the feature-encoded data into the model, training and parameter optimization are carried out. The performance of the model is evaluated through cross-validation, and metrics such as accuracy, sensitivity, specificity, ROC curve, and AUC value are used to measure the classification ability of the model. During multiple training and validation processes, the hyperparameters of the model are optimized, and the model with the best performance is selected for the final identification task.
[0076] To further improve the performance of the model, attempts are made to introduce pre-trained models or multi-task learning frameworks. Through transfer learning, protein sequence features are combined with other biological information (such as domains, functional annotations, etc.) to enhance the generalization ability of the model. In addition, using ensemble learning methods, the prediction results of multiple models are fused to further improve the accuracy and stability of the prediction.
[0077] After the model training is completed, it is applied to predict the LLPS regulatory function of unknown proteins. Combining experimental data to verify the accuracy of the prediction results provides an effective tool for protein function research related to liquid-liquid phase separation. At the same time, this method can also be extended to the prediction of protein functions in other biological processes, with broad application prospects.
[0078] Construction of the benchmark dataset provided by the embodiments of the present invention:
[0079] Collect 987 experimentally verified regulatory proteins involved in the LLPS process from the DrLLPS database to form a positive dataset; for the negative dataset, extract protein sequences from the UniProt database, including cytoskeletal proteins, membrane-associated proteins, highly structured proteins (such as globulins), proteins involved in forming stable structures, and single-chain proteins. These sequences are assumed to be less likely to play a regulatory role in LLPS; using the CD-HIT tool, the threshold for the positive dataset is set at 70%, and the threshold for the negative dataset is set at 40%; through this series of sorting and screening, 913 positive protein sequences and 6,584 negative protein sequences are finally obtained.
[0080] Dataset balance provided by the embodiments of the present invention:
[0081] Randomly select 113 positive sequences from the positive dataset as the test set, and the remaining 800 positive sequences are used for training; similarly, randomly select 184 negative sequences from the negative dataset as the test set, and the remaining 6,400 negative sequences are used for training; the remaining negative sequences are randomly divided into eight subsets, each subset containing 800 sequences, to match the 800 positive sequences and ensure the complete balance of the training dataset; finally, 8 balanced training datasets are obtained; then the selected 113 positive and 184 negative sequences are combined to obtain a test dataset containing 297 protein sequences; due to the priority given to the balanced training dataset, a slightly higher proportion of negative sequences is finally obtained in the test dataset.
[0082] Feature encoding provided by the embodiments of the present invention:
[0083] Use the ESM-2 model to encode the protein sequences to generate embedding vectors. This model can capture the sequence, structure, and physicochemical properties of proteins.
[0084] Model training and evaluation provided by the embodiments of the present invention:
[0085] Model evaluation:
[0086] Use sensitivity (Sn), specificity (Sp), accuracy (ACC), and area under the curve (AUC) metrics; sensitivity measures the ability of the model to correctly identify regulatory proteins in LLPS, specificity evaluates its accuracy in distinguishing non-regulatory proteins, and accuracy provides an overall performance evaluation of the model in the two categories of regulatory proteins and non-regulatory proteins; AUC is derived from the receiver operating characteristic (ROC) curve and is particularly important in binary classification tasks because it reflects the ability of the model to separate positive and negative classes; the formulas are as follows:
[0087]
[0088] Among them, TP represents that the regulatory protein is correctly recognized as a regulatory protein; TN represents that the non-regulatory protein is correctly recognized as a non-regulatory protein; FP represents that the non-regulatory protein is misclassified as a regulatory protein, and FN represents that the regulatory protein is mislabeled as a non-regulatory protein;
[0089] Model selection:
[0090] XG-Boost (XGB)
[0091] XGB uses gradient boosting decision trees for predictive modeling, and the final prediction result is the sum of the predictions of all individual trees; it can optimize the objective function, which combines the loss function l and the regularization term Ω to prevent overfitting and improve generalization ability; the general form of the objective function can be expressed as:
[0092]
[0093] where n represents the number of samples, y t and y p correspond to the true value and the predicted value of the data sample respectively; the term l represents the loss function, and K and f k represent the number of trees and the k-th tree respectively; Ω(f k ) is a regularization term that helps to mitigate overfitting;
[0094] Convolutional neural network:
[0095] The convolutional neural network includes two one-dimensional convolutional layers and four fully connected layers, and its structure can be expressed by the following formula:
[0096] (I * K)(i, j) = ∑m∑n(i + m, j + n).K(m, n)
[0097] where I is the input, K is the convolutional kernel, and * represents the convolution operation; (i, j) represents the spatial position in the output, and (m, n) represents the iteration in the spatial dimension; the pooling layer minimizes the computational burden and mitigates the impact of overfitting by reducing the spatial dimension of the feature map; the final fully connected layer aggregates the learned features for the classification task;
[0098] Multi-layer perceptron (MLP)
[0099] MLP can be expressed as:
[0100] y p = g(W (l) .f(W (l-1) .f(…f(W (1) .x + b (1) )…) + b (l-1) ) + b (l) )
[0101] Among them, x represents the input data, and W l is the weight matrix of the l-th layer, and b l is the bias function of the l-th layer, f l is the activation function of the l-th layer, g is the activation function of the output layer, and y p is the output;
[0102] Model training.
[0103] Model training provided by the embodiments of the present invention:
[0104] The protein feature representations extracted from each balanced dataset by the protein language model ESM2-36 are used as the input for model training; three classifiers, namely XGB, MLP, and CNN, are used to train the model respectively, and 10-fold cross-validation is performed on each balanced dataset to ensure the robustness of training and validation; during the training process, custom callbacks are used to monitor the validation accuracy, and the best-performing model weights are retained for testing; through this process, a total of 8 individual models are trained; to combine the prediction results of these models, the majority voting method of the output results is used for ensemble prediction. If the number of votes in favor and against is the same, the final decision is made by calculating the average probability and using a threshold of 0.5; finally, the ensemble method is verified on the test dataset.
[0105] As Figure 2 shown, a system for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology provided by the embodiments of the present invention includes:
[0106] A construction module for constructing a benchmark dataset;
[0107] A balancing module for balancing the dataset;
[0108] An encoding module for feature encoding;
[0109] A training module for model training and evaluation.
[0110] Another object of the present invention is to provide a computer device, which includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology.
[0111] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology.
[0112] Another object of the present invention is to provide an information data processing terminal for implementing the system for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology.
[0113] Specific implementation of the present invention:
[0114] Technical solution:
[0115] The present invention aims to construct an intelligent recognition model for LLPS regulatory proteins. This research includes four steps: (1) construction of a benchmark dataset, (2) dataset balancing, (3) feature encoding, and (4) model training and evaluation.
[0116] 1 Construction of a benchmark dataset:
[0117] A reliable dataset is the basis for a high-performance prediction model. Based on this understanding, we collected 987 experimentally verified regulatory proteins involved in the LLPS process from the DrLLPS database to form our positive dataset. For the negative dataset, we extracted protein sequences from the UniProt database, including cytoskeletal proteins, membrane-associated proteins, highly structured proteins (such as globulins), proteins involved in forming stable structures, and single-chain proteins. These sequences were assumed to be less likely to play a regulatory role in LLPS. To eliminate redundant sequences that might increase the computational load and time, we used the CD-HIT tool, with a threshold of 70% for the positive dataset and 40% for the negative dataset. Through this series of sorting and screening, we finally obtained 913 positive protein sequences and 6584 negative protein sequences, laying a solid foundation for subsequent model training and prediction.
[0118] 2 Dataset balancing:
[0119] To ensure the balance of the training dataset, we randomly selected 113 positive sequences from the positive dataset as the test set, and the remaining 800 positive sequences were used for training. Similarly, we randomly selected 184 negative sequences from the negative dataset as the test set, and the remaining 6400 negative sequences were used for training. These remaining negative sequences were randomly divided into eight subsets, each subset containing 800 sequences, to match the 800 positive sequences and ensure the complete balance of the training dataset. Finally, we obtained 8 balanced training datasets. Then, the selected 113 positive and 184 negative sequences were combined to obtain a test dataset containing 297 protein sequences. Due to the priority given to the balanced training dataset, we finally obtained a slightly higher proportion of negative sequences in the test dataset.
[0120] 3 Feature encoding:
[0121] Encode the protein sequence using the ESM-2 model to generate an embedding vector. This model can capture the sequence, structure, and physicochemical properties of proteins. This representation is more biologically meaningful than traditional embeddings based on natural language processing (NLP) such as word2vec.
[0122] 4 Model training and evaluation:
[0123] Model evaluation:
[0124] To evaluate the performance of the model, we used sensitivity (Sn), specificity (Sp), accuracy (ACC), and area under the curve (AUC) metrics. Sensitivity measures the ability of the model to correctly identify regulatory proteins in LLPS, specificity evaluates its accuracy in distinguishing non-regulatory proteins, and accuracy provides an overall performance assessment of the model on the two classes of regulatory and non-regulatory proteins. AUC is derived from the receiver operating characteristic (ROC) curve and is particularly important in binary classification tasks as it reflects the model's ability to separate positive and negative classes. The formulas are as follows:
[0125]
[0126] Where TP represents a regulatory protein correctly identified as a regulatory protein. TN represents a non-regulatory protein correctly identified as a non-regulatory protein. FP represents a non-regulatory protein misclassified as a regulatory protein, and FN represents a regulatory protein mislabeled as a non-regulatory protein.
[0127] Model selection:
[0128] XG-Boost (XGB)
[0129] XGB uses gradient-boosted decision trees for predictive modeling, and the final prediction result is the sum of the predictions of all individual trees. It can optimize the objective function, which combines the loss function l and the regularization term Ω to prevent overfitting and improve generalization ability. The general form of the objective function can be expressed as:
[0130]
[0131] Where n represents the number of samples, y t and y p correspond to the true value and the predicted value of the data sample, respectively. The term l represents the loss function, and K and f k represent the number of trees and the k-th tree, respectively. Ω(f k ) is a regularization term that helps to reduce overfitting.
[0132] Convolutional neural network (CNN)
[0133] Convolutional neural network (CNN) is a deep learning neural network architecture, renowned for its efficiency in capturing the spatial hierarchical structure of features in protein sequences. The CNN architecture consists of layers that use convolutional filters to capture patterns in the input data. In this study, the convolutional neural network includes two one-dimensional convolutional layers and four fully connected layers, and its structure can be represented by the following formula:
[0134] (I * K)(i, j) = ∑m∑n(i + m, j + n).K(m, n)
[0135] Where I is the input, K is the convolutional kernel, and * represents the convolution operation. (i, j) represents the spatial position in the output, and (m, n) represents the iteration in the spatial dimension. The pooling layer minimizes the computational burden and reduces the impact of overfitting by reducing the spatial dimension of the feature map. The final fully connected layer aggregates the learned features for classification tasks.
[0136] Multi-layer perceptron (MLP)
[0137] The multi-layer perceptron (MLP) classifier is a commonly used feed-forward neural network suitable for supervised learning tasks such as classification and regression. It consists of an input layer, multiple hidden layers, and an output layer, with each neuron fully connected to the neurons in the adjacent layer. The MLP can be expressed as:
[0138] y p = g(W (l) . f(W (l-1) . f(…f(W (1) . x + b (1) )…) + b (l-1) ) + b (l) )
[0139] Where x represents the input data, W l is the weight matrix of the l-th layer, b l is the bias function of the l-th layer, f l is the activation function of the l-th layer, g is the activation function of the output layer, and y p is the output.
[0140] In this study, the architecture of the MLP includes an input layer with 2560 neurons, followed by hidden layers containing 128, 64, 32, and 16 neurons, each using the ReLU activation function. The output layer for binary classification contains a single neuron using the sigmoid activation function. The process of training this MLP includes minimizing the binary cross-entropy loss through the backpropagation algorithm, with the optimizer being Adam and the learning rate set to 0.001.
[0141] Model training:
[0142] The protein feature representations extracted from each balanced dataset by the protein language model ESM2-36 were used as the input for model training. Three classifiers, namely XGB, MLP, and CNN, were used to train the model respectively, and 10-fold cross-validation was performed on each balanced dataset to ensure the robustness of training and validation. During the training process, a custom callback was adopted to monitor the validation accuracy, and the best-performing model weights were retained for testing. Through this process, a total of 8 individual models were trained. To combine the prediction results of these models, we used the majority voting method of the output results for ensemble prediction. If the number of votes in favor and against was the same, the final decision was made by calculating the average probability and adopting a threshold of 0.5. Finally, the ensemble method was verified on the test dataset.
[0143] MLP-based Model Construction and Performance Evaluation
[0144] The performance of the MLP model was evaluated through a comprehensive training and validation process. The 10-fold cross-validation method was adopted to train and validate the model on eight balanced training datasets, ensuring the robustness and reliability of the model. During the training process, a custom callback function was used to monitor the validation accuracy, and the best model weights were retained for testing. In addition, the test dataset was also evaluated, and the results are as follows:
[0145] For the model trained on Dataset 1, the sensitivity values ranged from 64.86% to 77.78% across the 1 to 10 folds, with an average sensitivity of 71.78%; the specificity varied between 72.37% and 89.04%, with an average of 81.18%; the accuracy fluctuated between 71.88% and 81.25%, with an average accuracy of 76.50%. For the test data, the sensitivity of the model was between 61.06% and 67.26%, with an average of 63.98%; the specificity varied between 75.54% and 80.98%, with an average of 77.72%; the accuracy fluctuated between 71.38% and 74.07%, with an average accuracy of 72.49%. The validation AUC ranged from 0.79 to 0.87, with an average of 0.83; the test AUC ranged from 0.80 to 0.82, with an average of 0.81.
[0146] The model trained on Dataset 2 has a sensitivity ranging from 68.92% to 88.10% across 1 to 10 folds, with an average sensitivity of 74.55%; the specificity fluctuates between 67.11% and 83.56%, averaging 77.73%; the accuracy varies between 70% and 80%, averaging 76.25%. For the test data, the sensitivity is between 63.72% and 84.07%, with an average of 69.82%; the specificity fluctuates between 60.87% and 78.80%, averaging 73.64%; the accuracy varies between 69.70% and 75.76%, averaging 72.19%. The validation AUC ranges from 0.78 to 0.85, averaging 0.83; the test AUC ranges from 0.79 to 0.82, averaging 0.80.
[0147] The model trained on Dataset 3 has a sensitivity ranging from 66.75% to 83.33% across 1 to 10 folds, with an average sensitivity of 74.70%; the specificity varies from 72.50% to 80.49%, averaging 76.90%; the accuracy fluctuates between 73.12% and 81.88%, averaging 75.82%. For the test data, the sensitivity ranges from 61.06% to 78.76%, with an average of 68.50%; the specificity varies from 67.93% to 79.89%, averaging 74.62%; the accuracy ranges from 69.02% to 77.78%, averaging 72.29%. The validation AUC ranges from 0.78 to 0.87, averaging 0.81; the test AUC ranges from 0.79 to 0.83, averaging 0.81.
[0148] The model trained on Dataset 4 has a sensitivity varying between 67.82% and 83.33%, with an average of 76.46%; the specificity ranges from 69.74% to 86.84%, averaging 77.27%; the accuracy fluctuates between 70% and 81.88%, averaging 76.88%. For the test data, the sensitivity varies from 65.49% to 76.11%, with an average of 69.74%; the specificity ranges from 69.02% to 75%, averaging 72.56%; the accuracy fluctuates between 69.70% and 74.41%, averaging 71.48%. The validation AUC ranges from 0.77 to 0.86, averaging 0.82; the test AUC ranges from 0.81 to 0.82, averaging 0.81.
[0149] The models trained on dataset 5 have sensitivities ranging from 66.67% to 85.71%, with an average sensitivity of 75.40%; specificities ranging from 67.47% to 89.47%, with an average of 77.73%; and accuracies varying from 72.50% to 83.12%, with an average of 76.50%. For the test data, the sensitivities range from 61.95% to 81.42%, with an average of 69.74%; the specificities range from 62.50% to 81.42%, with an average of 72.83%; and the accuracies range from 69.70% to 74.07%, with an average of 71.65%. The validation AUCs range from 0.78 to 0.90, with an average of 0.82; the test AUCs range from 0.79 to 0.83, with an average of 0.81.
[0150] The models trained on dataset 6 have sensitivities ranging from 63.22% to 77.38%, with an average sensitivity of 71.04%; specificities varying from 75% to 83.13%, with an average of 79.84%; and accuracies ranging from 70% to 78.75%, with an average of 75.44%. For the test data, the sensitivities range from 61.95% to 73.34%, with an average of 65.22%; the specificities range from 70.65% to 82.07%, with an average of 77.45%; and the accuracies fluctuate from 70.37% to 75.08%, with an average of 72.79%. The validation AUCs range from 0.78 to 0.86, with an average of 0.82; the test AUCs range from 0.79 to 0.82, with an average of 0.81.
[0151] The models trained on dataset 7 have sensitivities ranging from 63.22% to 80.77%, with an average sensitivity of 72.89%; specificities ranging from 71.95% to 86.05%, with an average of 78.73%; and accuracies ranging from 71.88% to 80%, with an average of 75.75%. For the test data, the sensitivities range from 60.18% to 76.11%, with an average of 66.02%; the specificities range from 70.65% to 80.43%, with an average of 76.14%; and the accuracies range from 68.01% to 75.08%, with an average of 72.29%. The validation AUCs range from 0.80 to 0.85, with an average of 0.82; the test AUCs range from 0.79 to 0.82, with an average of 0.81.
[0152] Finally, for the model trained on dataset 8, the sensitivity ranged from 63.22% to 80.77%, with an average of 69.79%; the specificity varied from 71.95% to 80.05%, with an average of 79.37%; the accuracy ranged from 71.88% to 80.00%, with an average of 74.56%. For the test data, the sensitivity ranged from 60.18% to 76.11%, with an average of 62.66%; the specificity ranged from 70.65% to 80.43%, with an average of 75.49%; the accuracy ranged from 68.01% to 75.08%, with an average of 70.60%. The validation AUC ranged from 0.80 to 0.85, with an average of 0.81; the test AUC ranged from 0.79 to 0.82, with an average of 0.80.
[0153] The above results are shown in ( Figures 3 to 6 , Tables 1 to 8).
[0154] After completing the cross-validation for all datasets, the folds with the best accuracy on the test datasets in each dataset were selected. The prediction results of each model were aggregated, and the final ensemble prediction result was determined based on the majority voting method. When the number of positive and negative votes was the same, the final decision was made by calculating the average probability and using a threshold of 0.5. This method ensured that each model played an equal role in the decision-making process and made full use of the negative sample dataset in the decision. Finally, the ensemble model was evaluated on the test dataset, and the following performance metrics were calculated: sensitivity, specificity, accuracy, and AUC. The sensitivity of the ensemble model was 74.34%, the specificity was 79.89%, and the accuracy was 77.78%. In addition, the AUC of the ensemble model was 0.83.
[0155] Comparison of the performance of different classifiers
[0156] To determine the effectiveness of the model, multiple classifiers were evaluated, including XGB, MLP, and CNN. The performance metrics of these classifiers are as follows: the sensitivity of XGB was 76.99%, the specificity was 71.74%, and the accuracy was 73.74%. Similarly, the sensitivity of MLP was 74.34%, the specificity was 79.89%, and the accuracy was 77.78%, while the sensitivity of CNN was 65.49%, the specificity was 74.46%, and the accuracy was 71.04%. The comparison of these results is shown in Figure 7 .
[0157] Analysis shows that the MLP classifier performs best in terms of both specificity and accuracy. Therefore, the MLP-based model was selected as the final method for accurately identifying regulatory proteins. Although the sensitivity of the XGB classifier is slightly higher, its specificity and accuracy are relatively low. Therefore, the MLP-based model was finally selected.
[0158] Figure 7The sensitivity, specificity, and accuracy were compared using the XGB, MLP, and CNN models respectively.
[0159] Table 1: Sensitivity, Specificity, and Accuracy of Dataset 1
[0160]
[0161]
[0162] Table 2: Sensitivity, Specificity, and Accuracy of Dataset 2
[0163]
[0164] Table 3: Sensitivity, Specificity, and Accuracy of Dataset 3
[0165]
[0166] Table 4: Sensitivity, Specificity, and Accuracy of Dataset 4
[0167]
[0168] Table 5: Sensitivity, Specificity, and Accuracy of Dataset 5
[0169]
[0170] Table 6: Sensitivity, Specificity, and Accuracy of Dataset 6
[0171]
[0172]
[0173] Table 7: Sensitivity, Specificity, and Accuracy of Dataset 7
[0174]
[0175] Table 8: Sensitivity, Specificity, and Accuracy of Dataset 8
[0176]
[0177] In this study, an ensemble classifier based on a multi-layer perceptron (MLP) was trained, and this classifier achieved an accuracy of 77.78% on the test dataset. By leveraging bioinformatics and machine learning techniques, potential phase separation proteins can be screened more efficiently from large-scale data, providing important theoretical support and experimental guidance for the research.
[0178] In this study, an ensemble classifier based on a multi-layer perceptron (MLP) was trained, and this classifier achieved an accuracy of 77.78% on the test dataset. To determine the effectiveness of the model, we evaluated multiple classifiers, including XGB, MLP, and CNN. The performance metrics of these classifiers are as follows: the sensitivity of XGB is 76.99%, the specificity is 71.74%, and the accuracy is 73.74%. Similarly, the sensitivity of MLP is 74.34%, the specificity is 79.89%, and the accuracy is 77.78%, while the sensitivity of CNN is 65.49%, the specificity is 74.46%, and the accuracy is 71.04%. The comparison of these results is shown in Figure 7 .
[0179] Our analysis shows that the MLP classifier performs best in terms of both specificity and accuracy. Therefore, we finally selected the MLP-based model.
[0180] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and their modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware.
[0181] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.
Claims
1. A method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology, characterized in that: The following steps are involved: Construction of benchmark datasets; Balancing of data sets; Characteristic encoding of protein sequences; Model training and evaluation.
2. The method according to claim 1, characterized in that The construction of the benchmark dataset includes: Positive protein sequences were collected from the DrLLPS database and processed for redundancy removal using the CD-HIT tool, with the threshold set at 70%; Negative protein sequences, including cytoskeletal proteins, membrane-associated proteins, and highly structural proteins, were extracted from the UniProt database and de-redundanted using the CD-HIT tool, with the threshold set at 40%; Finally, positive protein sequences and negative protein sequences are generated.
3. The method according to claim 1, characterized in that The balancing process of the data set includes: Randomly extract some sequences from the positive data set to form the test set, and the rest as the training set; Some sequences are randomly selected from the negative dataset to form the test set, and the remaining sequences are divided into several subsets; The positive training set is paired with each negative subset to form multiple balanced training data sets.
4. The method according to claim 1, characterized in that The feature encoding of protein sequences uses the ESM-2 model to generate feature vectors by embedding protein sequences.
5. The method according to claim 1, characterized in that Model training and evaluation include: Use multiple balanced training datasets to train XG-Boost, convolutional neural network, and multi-layer perceptron classifiers respectively; Each model was cross-validated 10-fold and the best model weights were retained for testing; The model performance is verified on the test dataset.
6. The method according to claim 5, characterized in that The XG-Boost model achieves prediction by gradient boosting decision trees, and its objective function includes a loss function and a regularization term.
7. The method according to claim 5, characterized in that The convolutional neural network consists of two one-dimensional convolutional layers and four fully connected layers, and extracts and classifies protein sequence features through convolution and pooling operations.
8. The method according to claim 5, characterized in that By making a majority vote on the prediction results of multiple models, an ensemble prediction result is finally generated.
9. A system for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology for implementing the method for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology as described in any one of claims 1 to 8, characterized in that: The system for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology includes: Building module, used for benchmark dataset construction; Balance module, used for data set balancing; Encoding module, used for feature encoding; Training module, used for model training and evaluation.
10. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the system for identifying liquid-liquid phase separation regulatory proteins based on artificial intelligence technology as described in claim 9.