Protein soluble expression level prediction method, system, equipment and medium

Through the integration of multi-source data sets and parallel feature extraction, and combining the ProtSATT architecture with self-attention and cross-attention mechanism, the existing prediction methods are solved inefficient and limited accuracy, and efficient and accurate prediction of protein soluble expression levels is achieved.

CN120496638APending Publication Date: 2025-08-15NANJING AGRICULTURAL UNIVERSITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510581320.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing methods for predicting protein soluble expression levels are inefficient, costly and limited prediction accuracy, unable to provide continuous solubility tendency scores, and it is difficult for a single embedded source to fully characterize the evolution-structure-function multimodal characteristics of sequences, resulting in insufficient ability to generalize across species.

Method used

The multi-source biological dataset integration optimization was adopted to construct a high-quality training verification dataset. The sequence features were extracted in parallel by three large-language models of UniRep, ESM-2 and ProtT5, and linear position coding was added. Combined with the ProtSATT architecture of self-attention and cross-attention mechanism, a dual-task prediction model was built to output continuous solubility tendency scores and binary classification results.

Benefits of technology

It improves prediction accuracy and stability, improves data utilization and model generalization capabilities, and achieves more efficient and accurate protein soluble prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496638A_ABST
    Figure CN120496638A_ABST
Patent Text Reader

Abstract

The invention discloses a protein soluble expression level prediction method, system, equipment and medium, and relates to the technical field of bioinformatics and artificial intelligence, and the method comprises the steps: integrating and optimizing a multi-source biological data set, and constructing a high-quality training verification data set; based on a protein large language model, parallel extraction sequence features are adopted to add linear position codes; a deep learning model based on a self-attention mechanism and a cross attention mechanism is adopted, a dichotomy normal form prediction framework is optimized, a double-task prediction model is constructed, and classification and regression double-task prediction is completed; according to the method disclosed by the invention, an efficient, accurate and universal soluble prediction tool ProtSATT is constructed by fusing a multi-source protein large language model and an innovative deep learning architecture; the ProtSATT innovatively breaks through a traditional dichotomy prediction framework, continuous solubility tendency score prediction is achieved, and the prediction precision and stability are effectively improved in combination with a self-attention mechanism and a cross attention mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics and artificial intelligence technology, and in particular to a method, system, device and medium for predicting protein soluble expression levels. Background Art

[0002] Traditional approaches rely primarily on empirical optimization strategies based on experiments, including using low-strength promoter systems to regulate transcription rates, using low-temperature culture conditions to slow misfolding kinetics, and optimizing culture medium composition through factorial design. While these approaches have achieved some success in early studies, their core drawback is the intensive trial-and-error wet experimentation required to optimize conditions, resulting in lengthy and costly R&D cycles. With the rapid advancement of bioinformatics, sequence-based machine learning prediction models are emerging as an alternative. Traditional machine learning architectures, by incorporating handcrafted feature descriptors such as amino acid composition, hydropathicity index, and isoelectric point, have shown initial promise in solubility prediction. For example, early models such as Protein-Sol achieved approximately 70% accuracy on limited datasets by integrating physicochemical features with sequence properties. However, these models are limited in their ability to process high-dimensional omics data and fail to effectively capture key structural determinants such as long-range residue interactions and tertiary folding patterns.

[0003] In recent years, deep learning technology has promoted a paradigm shift in this field. Models based on convolutional neural networks and recurrent neural networks have significantly improved prediction accuracy by automatically extracting local patterns in sequences. Even more groundbreaking is that the introduction of protein language models (PLMs) has achieved implicit encoding of evolutionary conservation and structural characteristics. However, existing deep learning frameworks still have significant technical bottlenecks: most models adopt a binary classification paradigm (soluble / insoluble) and cannot provide continuous solubility propensity scores, which limits their application value in targeted optimization of protein engineering; a single embedding source is difficult to fully characterize the evolutionary-structural-functional multimodal characteristics of the sequence, resulting in insufficient cross-species generalization capabilities; the model architecture lacks modeling depth for residual contact maps and folding dynamics, making it difficult to meet prediction needs in complex biological contexts.

[0004] The core flaws of the current technology system are concentrated in three aspects: the one-sidedness of feature representation, the limitations of the task framework, and the lack of adaptability of the model architecture. Existing methods often rely on a single protein language model to extract features, but the encoding biases of different protein language models lead to incomplete information coverage. The binary classification framework simplifies solubility into a binary output, which cannot meet the quantitative gradient analysis requirements of solubility in antibody engineering. Traditional architectures also lack the synergistic utilization of attention mechanisms and residual learning, resulting in inefficient feature extraction.

[0005] The present invention proposes a systematic solution to the above problems: (1) By strengthening the multidimensional representation through parallel feature extraction and linear position encoding of UniRep, ESM-2 and ProtT5, the information bias bottleneck of a single model is broken, and the amino acid sequence information is decoupled from the semantic embedding, so that the model can simultaneously capture the sequence semantics and spatial topological constraints; (2) a dual-task prediction framework is constructed to synchronously output continuous solubility scores and classification labels with a shared coding layer, which retains the quantitative analysis capability; (3) a ProtSATT architecture based on self-attention and cross-attention is designed, which realizes deep feature refinement through stacked self-attention layers with residual connections, and then uses cross-attention blocks to complete the nonlinear fusion of multimodal features.

[0006] In addition, the improved Cos activation function enhances nonlinear representation through periodic mapping, which improves prediction stability in TR tasks compared to the standard Swish function. After cross-dataset verification, the present invention is superior to the existing technical system in feature completeness, task adaptability and architectural robustness, providing a new methodological paradigm for protein solubility prediction. Summary of the Invention

[0007] In view of the above-mentioned problems, the present invention is proposed.

[0008] Therefore, the technical problem solved by the present invention is: the existing methods for predicting protein soluble expression levels are low in efficiency, high in cost, and have limited prediction accuracy, and how to predict protein solubility more efficiently and accurately.

[0009] To solve the above technical problems, the present invention provides the following technical solutions: a method for predicting protein soluble expression levels, including the integration and optimization of multi-source biological data sets to construct a high-quality training and verification data set; based on a large protein language model, parallel extraction of sequence features is used to add linear position encoding; a deep learning model based on a self-attention mechanism and a cross-attention mechanism is used to optimize the binary classification paradigm prediction framework, construct a dual-task prediction model, and complete classification and regression dual-task prediction; the training and verification data set includes data cleaning, de-redundancy, and normalization operations based on four major data sets to balance the sample ratio between different data sets; parallel extraction of sequence features includes using three large protein language models in parallel to extract sequence features of different dimensions; optimization of the binary classification paradigm prediction framework includes designing a dual prediction framework with both classification and regression tasks, and outputting a continuous solubility propensity score and a binary classification result; construction of a dual-task prediction model includes optimizing the binary classification paradigm prediction framework, constructing a dual-task prediction model, performing classification and regression predictions, and outputting a continuous solubility propensity score and a binary classification result.

[0010] As a preferred embodiment of the method for predicting protein soluble expression levels of the present invention, the construction of a high-quality training and validation dataset includes obtaining original protein sequences and solubility labels from four datasets: eSOL, SC, E.Coli, and TR.

[0011] As a preferred embodiment of the method for predicting protein soluble expression levels described in the present invention, the parallel extraction of sequence features to add linear position coding includes embedding the protein sequence using three large protein language models, UniRep, ESM-2, and ProtT5, and adding linear position coding information to each embedding vector.

[0012] As a preferred embodiment of the method for predicting the protein soluble expression level described in the present invention, the construction of the dual-task prediction model includes constructing a deep learning model ProtSATT based on the attention mechanism, adding linear position encoding information to the input, optimizing the physical position dimension, and constructing a deep feature extraction task.

[0013] As a preferred solution of the method for predicting the protein soluble expression level described in the present invention, the task of constructing a deep feature extraction includes stacking self-attention layers, converting the protein sequence after embedding encoding by three large protein language models into a hidden vector, adding a variable number of stacked self-attention layers of a defect mechanism, performing deep feature extraction, and constructing a weighted fusion feature vector basis.

[0014] As a preferred solution of the method for predicting the protein soluble expression level described in the present invention, the construction of the weighted fusion feature vector basis includes: using a cross-attention mechanism combined with a residual block, reshaping the output vector after depth extraction into a two-dimensional vector, generating three major matrices through a linear layer, and based on the Softmax activation function and the Swish activation function, combining the residual mechanism to perform weighted addition and fusion of the feature vectors.

[0015] As a preferred embodiment of the method for predicting the protein soluble expression level of the present invention, the completion of the classification and regression dual-task prediction includes splicing and fusing feature vectors, reducing the dimension of the fused feature vectors based on the Swish activation function, and outputting the prediction results.

[0016] Another object of the present invention is to provide a prediction system for protein soluble expression levels, which can solve the problems of low data utilization and insufficient model generalization ability in current technologies for predicting protein soluble expression levels by combining a large protein language model and a deep learning model.

[0017] As a preferred solution of the protein soluble expression level prediction system described in the present invention, it includes: a preprocessing module, a deep feature extraction module, a feature fusion module and an output module; the preprocessing module is used to integrate and optimize high-quality data sets and add linear position encoding information to the three inputs of the model; the deep feature extraction module is used to use a self-attention mechanism to perform deep feature extraction on the protein sequence processed by the preprocessing module; the feature fusion module is used to use a cross-attention mechanism to perform pairwise cross-fusion on the three inputs; the output module is used to splice the six hidden vectors that have undergone cross-fusion processing, process them through an activation function and a fully connected layer, and output the prediction result.

[0018] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a method for predicting protein soluble expression level.

[0019] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for predicting protein soluble expression levels.

[0020] Beneficial effects of the present invention: The method for predicting protein soluble expression levels adopts a transfer learning strategy, designs a dual prediction framework that combines classification and regression tasks, proposes the ProtSATT model based on the attention mechanism, combines self-attention and cross-attention mechanisms, realizes deep extraction and fusion of multi-source features, and verifies it in independent test sets and across tasks. At the same time, ablation experiments are used to verify the necessity of model components and optimize hyperparameter configuration. The present invention achieves better results in terms of data utilization, model performance, and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 This is an overall flow chart of the method for predicting protein soluble expression levels provided in the first embodiment of the present invention.

[0023] Figure 2 This is a bar chart comparing the performance of the deep feature extraction module of the protein soluble expression level prediction method provided in the second embodiment of the present invention.

[0024] Figure 3This is a histogram comparing model input patterns of the method for predicting protein soluble expression levels provided in the second embodiment of the present invention.

[0025] Figure 4 This is an overall flow chart of the system for predicting protein soluble expression levels provided in the third embodiment of the present invention. DETAILED DESCRIPTION

[0026] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0027] Example 1, with reference to Figure 1 , as one embodiment of the present invention, provides a method for predicting the soluble expression level of a protein, comprising:

[0028] S1: Integration and optimization of multi-source biological datasets to construct high-quality training and validation datasets.

[0029] Furthermore, the construction of high-quality training and validation datasets includes obtaining original protein sequences and solubility labels from four datasets: eSOL, SC, E.Coli, and TR.

[0030] It should be noted that eSOL represents a database of Escherichia coli protein solubility data synthesized by a cell-free expression system; SC represents Saccharomyces cerevisiae, a selected collection of non-redundant protein sequences annotated with UniProt identifiers; E.Coli represents a dataset of soluble expression of Escherichia coli proteins, which are divided into six categories according to the expression levels on polyacrylamide gel electrophoresis gels and then merged into a binary classification task; TR represents a dataset that quantifies the conformational performance of riboswitches by switching factors.

[0031] It should be noted that the preferred solution for constructing a high-quality training and validation dataset includes quality screening of the eSOL dataset and dividing it into a training subset and an independent test set in a 3:1 ratio; optimizing the Saccharomyces cerevisiae dataset to obtain non-redundant protein sequences with experimentally verified solubility indicators; dividing the Escherichia coli protein expression dataset into six different expression categories according to expression levels, and merging some data to form low-expression data and high-expression data; and quantitatively characterizing the tetracycline riboswitch dataset through the switching efficiency coefficient.

[0032] It should also be noted that a transfer learning strategy is adopted to use a pre-trained large protein language model to extract multi-dimensional features, alleviating the bottleneck of insufficient protein soluble data; through the general feature representation ability of the pre-trained model, the generalization performance of the model on small data sets is enhanced, and the comprehensiveness of feature extraction is improved.

[0033] S2: Based on the protein language model, parallel extraction of sequence features is used to add linear position encoding.

[0034] Furthermore, the parallel extraction of sequence features and the addition of linear position coding include embedding protein sequences using three large protein language models, UniRep, ESM-2, and ProtT5, and adding linear position coding information to each embedding vector.

[0035] It should be noted that UniRep (unified representation of protein sequences) represents the generation of universal embeddings for protein sequences through unsupervised representation learning. The UniRep model adopts a multi-layer long short-term memory (mLSTM) architecture to capture evolutionary and structural patterns by predicting the next amino acid in the sequence, and derives a fixed-length global representation vector of 1900*1; ESM-2 represents a Transformer-based protein language model, which aims to infer evolutionary, structural and functional patterns from protein sequences. It uses a masked language modeling method for training to obtain a fixed vector of 1280*1 as the feature information of the protein; ProtT5 is a protein sequence pre-training model based on the T5 architecture, which belongs to the protTrans model family. Its core goal is to capture complex patterns in protein sequences through large-scale unsupervised learning, support transfer learning for various downstream tasks, encode protein sequences into corresponding embedding information, and obtain a fixed vector representation with a dimension of 1024*1 by taking the average of amino acids.

[0036] It should be noted that the embedding process includes processing the protein sequence to complete the corresponding regression or classification task, and obtaining the primary feature one-dimensional vector with the shape of 1900*1, 1280*1, and 1024*1.

[0037] It should also be noted that the position information encoding formula is expressed as:

[0038]

[0039] Here, X represents a fixed-dimensional vector embedded in a large protein language model, vector_length represents the vector length, and X' represents a vector with added linear encoding information, which enriches the input information in the physical location dimension and provides a more sufficient information basis for subsequent processing.

[0040] S3: Adopt a deep learning model based on self-attention mechanism and cross-attention mechanism, optimize the binary classification paradigm prediction framework, build a dual-task prediction model, and complete classification and regression dual-task prediction.

[0041] Furthermore, building a dual-task prediction model includes building a deep learning model ProtSATT based on the attention mechanism, adding linear position encoding information to the input, optimizing the physical position dimension, and building a deep feature extraction task.

[0042] It should be noted that the construction of the deep learning model ProtSATT includes using different loss functions according to different prediction tasks when training the model. When training the binary classification task model, the binary cross entropy loss function is used, which is expressed as:

[0043]

[0044] in, represents the average difference between the true label and the predicted probability, y i Indicates the true classification label with a value of 0 or 1. Indicates the probability that the i-th sample belongs to the positive class, and N is the number of samples. When training the model for the regression task, the mean square error loss function is used, which is expressed as:

[0045]

[0046] Among them, y i is the true value; is the model prediction value; N is the number of samples.

[0047] It should be noted that the task of constructing deep feature extraction includes stacking self-attention layers, converting the protein sequence after embedding encoding by three large protein language models into hidden vectors, adding self-attention layers with a variable number of stacking layers of incomplete mechanisms, performing deep feature extraction, and constructing the basis of weighted fusion feature vectors.

[0048] It should be noted that the basis for constructing a weighted fusion feature vector includes using a cross-attention mechanism combined with a residual block to reshape the output vector after depth extraction into a two-dimensional vector, generating three major matrices through a linear layer, and performing weighted addition and fusion of feature vectors based on the Softmax activation function and the Swish activation function combined with the residual mechanism.

[0049] It should be noted that the Softmax activation function is expressed as:

[0050]

[0051] Among them, i represents the sequence number of the element, x iRepresents the value of the i-th element; the Swish activation function is expressed as:

[0052]

[0053] Where x is the input value and β is a customizable value.

[0054] It should also be noted that completing the dual-task prediction of classification and regression includes concatenating and fusing feature vectors, reducing the dimension of the fused feature vectors based on the Swish activation function, and outputting the prediction results.

[0055] It should also be noted that based on the coefficient of determination (R 2 ), Accuracy, Precision and Recall are used to evaluate the prediction results. The coefficient of determination formula is expressed as:

[0056]

[0057] Among them, y i represents the true value; represents the predicted value; Represents the average value of the true value, and the accuracy formula is expressed as:

[0058]

[0059] The accuracy formula is expressed as:

[0060]

[0061] The recall formula is expressed as:

[0062]

[0063] Among them, TP represents the true positive of soluble protein, FP represents the false positive of soluble protein, TN represents the true negative of insoluble protein, and FN represents the false negative of insoluble protein.

[0064] Example 2, reference Figure 2-Figure 3 , which is an embodiment of the present invention, provides a method for predicting the soluble expression level of a protein. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0065] First, this experiment conducts an ablation experiment on the model input information. A three-input model is constructed. The three feature information is input into the model at the same time and the same processing method is adopted. The deep feature extraction is performed through the module with the self-attention mechanism as the core, and the feature fusion is performed through the cross-attention mechanism. Finally, the output is completed. The corresponding model input ablation experiment is performed on the four data sets. The results are as follows Figure 3 shown.

[0066] from Figure 3 The comprehensive results in

[15] show that the three-input model demonstrates superior performance compared to the two-input or single-input configurations. The two-input model loses some feature information when reducing the input representation, while the single-input model has two key limitations: severe information loss due to the lack of complementary features, and an inability to utilize cross-fusion modules, which weakens feature extraction capabilities.

[0067] This experimental evidence confirms that the three-input ensemble systematically outperforms the single-input baseline model and all pairwise combinations. When using only a single-input model, the results obtained by using the embedding information of ProtT5 are better than those of other models. In the two-input model, the combination of ESM-2 and ProtT5 produces the best performance. Based on the contribution of each representation in feature extraction, ProtT5's contribution is greater than ESM-2, and UniRep's individual contribution is the weakest.

[0068] Then, for model structure ablation experiments and hyperparameter optimization, systematically removing any core architectural component will continuously degrade the model's performance. Specifically, removing the deep feature extraction module will lead to the most severe performance degradation in all tasks, while removing the feature fusion module will cause significant but relatively small performance degradation. Figure 2 As shown, both ablation cases show a significant drop in performance compared to the full architecture, confirming that the deep feature extraction module is the main driver of the model performance, while the feature fusion module provides important cross-modal integration capabilities.

[0069] In addition, this experiment conducts extensive ablation experiments on other hyperparameter configurations, as shown in Table 1:

[0070] Table 1 Comparison data of other hyperparameter configurations

[0071]

[0072]

[0073] Regarding the residual ratio in the cross-fusion module, skip connections not only preserve the original feature information with a certain weight, but also effectively alleviate the vanishing gradient problem, helping the model further improve feature extraction performance. Regarding the activation function of the output module, experiments were conducted on Softmax, Swish, and a modified cosine function. The experiments showed that the custom cosine activation function not only provides enhanced nonlinear mapping capabilities for the neural network, but also better scales information within the appropriate range of 0 to 1, achieving a 1.2% to 2.8% improvement over the standard activation function. A suitable dropout rate can optimally address underfitting and overfitting, particularly improving performance by 10% in small data sets.

[0074] Specifically, on the eSOL test set, the accuracy is improved by 2%; on the SC test set, R 2 The performance was improved by 0.015; on the E. coli (E.Coli) dataset, the average accuracy was improved by 1.81%; and on the TR independent test set, the Spearman correlation coefficient reached 0.66.

[0075] The eSOL dataset is related to protein soluble expression prediction, where the label represents the solubility probability, ranging from 0 to 1. Samples with a probability higher than 0.5 are classified as soluble, while samples with a probability lower than 0.5 are considered insoluble. The ProtSATT model handles both regression tasks (predicting solubility probability) and classification tasks (binary classification of solubility). Experimental results show that the performance of the ProtSATT model is better than the current state-of-the-art model in both tasks. The GATSol method achieved an R of 0.517 on this dataset. 2 value and an accuracy of 0.791. Under the same training and testing conditions, as shown in Table 2:

[0076] Table 2 Comparison of data of different models

[0077]

[0078] ProtSATT achieved an R of 0.534 2 The accuracy of ProtSATT is 0.811, which surpasses the existing benchmark model. The precision of ProtSATT is slightly lower than that of the comparison model (1.8% lower than GraphSol and 0.9% lower than GATSol), and the recall rate is significantly higher, reaching 82.2%, which is 12% higher than GraphSol and 7.7% higher than GATSol. This shows that ProtSATT can identify more positive samples while maintaining competitive precision, thereby enabling more comprehensive detection of highly soluble sequences. Overall, the results highlight the balanced and excellent overall performance of ProtSATT in the solubility prediction task.

[0079] To validate the generalizability of the ProtSATT model, the Saccharomyces cerevisiae (SC) dataset was used as an independent test set for further evaluation. The original SC dataset contained 447 soluble protein sequences. After rigorously removing internal redundancy and overlap with the training set, an optimized dataset containing 108 non-redundant proteins, designated SC.108, was obtained. Using the same model weights trained on the eSOL dataset, the ProtSATT model was tested on the SC.108 independent test set, demonstrating excellent generalization and versatility.

[0080] Table 3 shows that the ProtSATT model achieved an R of 0.439. 2 The value is 0.067 higher than the GraphSol model and 0.015 higher than the GATSol model, which highlights that the ProtSATT model has strong generalization ability and versatility in predicting protein solubility in different biological backgrounds even when trained on only a single dataset.

[0081] Table 3 Comparison of coefficients of determination

[0082] Solubility Predictor Coefficient of determination DeepSol 0.090 Protein-Sol 0.281 GraphSol 0.358 GraphSol(ensemble) 0.372 GATSol 0.424 ProtSatt(ours) 0.439

[0083] The E.Coli dataset contains 4281 protein sequences. This task was transformed into a binary classification problem with a 1:1 ratio between the two classes. After applying the same data processing method and ten-fold cross-validation to the model, the following results are shown in Table 4:

[0084] Table 4 Average accuracy comparison table

[0085] Model Average accuracy (%) RF 61.76 DT 59.52 LR 61.07 Support Vector Machine 62.02 MPEPE 69.81 ProtSatt(ours) 71.62

[0086] Finally, an average accuracy of 71.62% was achieved, which is 1.81% higher than MPEPE. MPEPE uses codon encoding, which contains richer sequence-level features compared to amino acid sequences. Models such as UniRep, ESM2, and ProtT5 are based on amino acid sequences for feature extraction and structure prediction. They are pre-trained on large-scale datasets, which can capture the comprehensive features and structural information of proteins, and therefore have significant advantages. In contrast, MPEPE's codon embedding is only trained on a limited dataset of a few thousand entries in the training set. The embedding method relies on a simple unique encoding rather than large-scale pre-training using big data, resulting in significant limitations in data representation and generalization capabilities compared to the above models.

[0087] In order to further verify the generalization ability, structural rationality and versatility of the ProtSATT model, and to prove that it can not only handle protein solubility prediction tasks, but also perform well in other prediction tasks related to protein sequences, another dataset unrelated to protein solubility was selected for training and testing: the tetracycline riboswitch dataset, which is used to predict the switch factors representing the differential activity of riboswitches in the presence or absence of tetracycline. The tetracycline riboswitch dataset is divided into training, validation and test sets with ratios of 0.7, 0.15 and 0.15, respectively. The ProtSATT model not only performs extremely well in protein solubility prediction, but also achieves excellent results in other tasks related to protein sequences, proving that the ProtSATT model has strong learning ability in predicting protein-related properties from sequences, even when trained on small sample training sets. These findings further verify that the model has strong deep feature extraction capabilities for protein sequences.

[0088] Example 3, reference Figure 4 , which is an embodiment of the present invention, provides a prediction system for protein soluble expression level, including a preprocessing module 100, a deep feature extraction module 200, a feature fusion module 300 and an output module 400.

[0089] Among them, X1: the preprocessing module 100 includes a data collection submodule 101, a position encoding submodule 102, a feature dimension alignment submodule 103, and a residual initialization submodule 104.

[0090] It should be noted that the data collection submodule 101 is used to collect data from four benchmark data sets covering different biological backgrounds; the position encoding submodule 102 is used to add linear position encoding information to the model input to make the input information richer in the physical position dimension; the feature dimension alignment submodule 103 adjusts the embedding vectors of different dimensions output by the three large protein language models to the same dimension; the residual initialization submodule 104 adds an initial residual connection to the input vector to prevent the gradient of the deep network from disappearing.

[0091] It should also be noted that the data collection submodule 101 and the position encoding submodule 102 receive the embedding vectors of the three large protein language models and position encode the enhanced spatial information, align the feature dimensions based on the feature dimension alignment submodule 103, and output to the deep feature extraction module 200 after the initialized residual connection of the residual initialization submodule 104.

[0092] X2: The deep feature extraction module 200 includes a linear transformation submodule 201 and a variable stacked self-attention submodule 202.

[0093] It should be noted that the linear conversion submodule 201 is used to convert the three input information processed by the preprocessing module 100 into three hidden vectors through a linear layer; the variable stacking self-attention submodule 202 uses a variable number of stacking layers of self-attention layers with an added residual mechanism to perform deep feature extraction on the hidden vectors to obtain three deep feature vectors.

[0094] It should also be noted that the input preprocessed vector is reduced in dimension to reduce complexity through the linear transformation submodule 201, and then the variable stacking self-attention submodule 202 stacks the self-attention to improve the high-order semantic features and transmits it to the feature fusion module 300.

[0095] X3: The feature fusion module 300 includes a cross-attention pairing submodule 301 and a multimodal residual fusion submodule 302.

[0096] It should be noted that the cross-attention pairing submodule 301 is used to construct cross-attention for the three inputs in pairs and calculate the cross-model feature similarity; the multimodal residual fusion submodule 302 is used to introduce a dynamic residual ratio and weightedly fuse the original features with the cross-attention features.

[0097] It should also be noted that after the feature fusion module 300 receives the three-way deep features transmitted by the variable stacking self-attention sub-module 202, it performs cross-attention calculation of feature correlation through the cross-attention pairing sub-module 301, and dynamically fuses the complementary information through the multimodal residual fusion sub-module 302, and then reduces the dimension and deduplicates the output to the output module 400.

[0098] Among them, X4: the output module 400 includes a multi-source feature splicing submodule 401, an activation function submodule 402, and a dual-task output submodule 403.

[0099] It should be noted that the multi-source feature splicing submodule 401 is used to splice the six hidden vectors that have undergone cross-fusion processing; the activation function submodule 402 uses the Swish activation function to process the spliced vectors to increase the expressive power of the model; the dual-task output submodule 403 is used to output classification and regression results in parallel in the fully connected layer to achieve end-to-end multi-task learning.

[0100] It should also be noted that after the multi-source feature splicing submodule 401 splices and fuses the features, the activation function submodule 402 performs nonlinear mapping and transmits them to the dual-task output submodule 403 to synchronously output the classification label and regression probability.

[0101] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0102] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0103] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0104] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc. It should be noted that the above embodiments are merely illustrative of the technical solutions of the present invention and are not intended to be limiting. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications should be encompassed by the claims of the present invention.

Claims

1. A method for predicting protein soluble expression levels, characterized in that: include: Integration and optimization of multi-source biological datasets to build high-quality training and validation datasets; Based on the protein language model, the linear position encoding is added by parallel extraction of sequence features; Adopting a deep learning model based on self-attention mechanism and cross-attention mechanism, optimizing the binary classification paradigm prediction framework, building a dual-task prediction model, and completing classification and regression dual-task prediction; The training and validation datasets include data cleaning, redundancy removal, and normalization based on the four datasets to balance the sample ratios between different datasets; Parallel extraction of sequence features involves using three large protein language models in parallel to extract sequence features of different dimensions; Optimizing the binary classification paradigm prediction framework involves designing a dual prediction framework that combines classification and regression tasks, outputting continuous solubility propensity scores and binary classification results; Building a dual-task prediction model includes optimizing the binary classification paradigm prediction framework, building a dual-task prediction model, performing classification and regression predictions, and outputting continuous solubility propensity scores and binary classification results.

2. The method for predicting protein soluble expression level according to claim 1, wherein: The construction of high-quality training and verification data sets includes: The original protein sequences and solubility labels were obtained from four datasets: eSOL, SC, E.Coli and TR.

3. The method for predicting protein soluble expression level according to claim 1 or 2, wherein: The parallel extraction of sequence features and adding linear position coding includes: Three large protein language models, UniRep, ESM-2 and ProtT5, are used to embed protein sequences, and linear position encoding information is added to each embedding vector.

4. The method for predicting protein soluble expression level according to claim 3, wherein: The construction of the dual-task prediction model includes: Based on the attention mechanism, a deep learning model ProtSATT is constructed, linear position encoding information is added to the input, the physical position dimension is optimized, and a deep feature extraction task is constructed.

5. The method for predicting protein soluble expression level according to any one of claims 1, 2 or 4, wherein: The construction of the deep feature extraction task includes: The self-attention layers are stacked to convert the protein sequences after embedding encoding by three large protein language models into hidden vectors. The self-attention layers with a variable number of stacking layers of the incomplete mechanism are added to perform deep feature extraction and build the basis of weighted fusion feature vectors.

6. The method for predicting protein soluble expression level according to claim 5, wherein: The basis for constructing the weighted fusion feature vector includes: The cross-attention mechanism is combined with the residual block to reshape the output vector after depth extraction into a two-dimensional vector. Three matrices are generated through the linear layer. Based on the Softmax activation function and the Swish activation function, the residual mechanism is combined to perform weighted addition and fusion of the feature vectors.

7. The method for predicting the soluble expression level of a protein according to any one of claims 1, 2, 4 or 6, wherein: The completion of the classification and regression dual-task prediction includes: The fused feature vectors are concatenated and dimensionality reduced based on the Swish activation function to output the prediction results.

8. A system for predicting protein soluble expression levels, characterized by: It includes a pre-processing module (100), a depth feature extraction module (200), a feature fusion module (300) and an output module (400); The pre-processing module (100) is used to integrate and optimize high-quality data sets and add linear position encoding information to the three inputs of the model; The deep feature extraction module (200) is used to perform deep feature extraction on the protein sequence processed by the preprocessing module using a self-attention mechanism; The feature fusion module (300) is used to perform pairwise cross fusion on the three inputs using a cross attention mechanism; The output module (400) is used to splice the six hidden vectors that have undergone cross-fusion processing, process them through activation functions and fully connected layers, and output prediction results.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for predicting the soluble expression level of a protein according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for predicting the soluble expression level of a protein according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Large-model-driven intelligent biological research method and system

    CN120805532A