Methods of predicting protein solubility, neural network training methods, apparatuses, devices, and media
Patent Information
- Application Number
- CN202510908055.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-07-02
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-07-02
AI Technical Summary
在传统方法中,蛋白质的溶解性多数通过湿实验进行检测,完成该过程所需的时间和材料成本较高
[0012] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, causes the processor to perform the methods described above.
Smart Images

Figure CN120748480B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese patent application No. 202410878381.7, filed on July 2, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of bioinformatics, specifically to the fields of artificial intelligence and biological computing, and particularly to a method for predicting protein solubility, a neural network training method, apparatus, device, and medium. Background Technology
[0004] Solubility is a key biophysical property of proteins, crucial for assessing their viability in biochemical engineering. Protein solubility depends on the interaction of various external physical conditions (such as pH and temperature) with intrinsic factors (such as the protein's amino acid composition and structure). Traditionally, protein solubility is mostly determined through wet experiments, a process that is time-consuming and materially expensive. Therefore, there is an urgent need for a dry experimental method to predict and analyze the solubility of potential proteins, thereby reducing the costs associated with wet experiments. Summary of the Invention
[0005] It would be beneficial to provide a mechanism to alleviate, reduce, or even eliminate one or more of the aforementioned problems.
[0006] According to one aspect of this disclosure, a method for predicting protein solubility using a neural network is provided, wherein the trained neural network includes a first graph neural network, a second graph neural network, and a prediction subnetwork. The method includes: acquiring a target protein sequence comprising multiple amino acid residues; encoding the target protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the multiple amino acid residues; acquiring a position-specific score matrix of the target protein sequence, the position-specific score matrix including position-specific score vectors for each of the multiple amino acid residues; determining the contact probability of multiple amino acid residue pairs in the target protein sequence; processing the encoding feature vectors for each of the multiple amino acid residues and the contact probabilities of the multiple amino acid residue pairs using the first graph neural network to obtain a first comprehensive feature vector for each of the multiple amino acid residues; processing the position-specific score vectors for each of the multiple amino acid residues and the contact probabilities of the multiple amino acid residue pairs using the second graph neural network to obtain a second comprehensive feature vector for each of the multiple amino acid residues; and inputting the first comprehensive feature vector and the second comprehensive feature vector for each of the multiple amino acid residues into the prediction subnetwork to obtain a prediction result of the solubility of the target protein sequence.
[0007] According to another aspect of this disclosure, a neural network training method is provided, wherein the initial neural network includes a first initial graph neural network, a second initial graph neural network, and an initial prediction subnetwork. The method includes: acquiring a sample protein sequence and the true solubility of the sample protein sequence, the sample protein sequence comprising multiple amino acid residues; encoding the sample protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the multiple amino acid residues; acquiring a position-specific score matrix of the sample protein sequence, the position-specific score matrix comprising position-specific score vectors for each of the multiple amino acid residues; determining the contact probability of multiple amino acid residue pairs in the sample protein sequence; and using the first initial graph neural network... The initial graph neural network processes the encoding feature vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs to obtain the first comprehensive feature vectors of multiple amino acid residues. The second initial graph neural network processes the position-specific score vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs to obtain the second comprehensive feature vectors of multiple amino acid residues. The first and second comprehensive feature vectors of multiple amino acid residues are input into the initial prediction sub-network to obtain the prediction result of the protein solubility of the sample protein sequence. Based on the prediction result and the actual solubility, the parameters of the initial neural network are adjusted to obtain the trained neural network.
[0008] According to another aspect of this disclosure, an apparatus for predicting protein solubility using a neural network is provided, wherein the trained neural network includes a first graph neural network, a second graph neural network, and a prediction subnetwork. The apparatus includes: a first acquisition unit configured to acquire a target protein sequence comprising a plurality of amino acid residues; a first encoding unit configured to encode the target protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the plurality of amino acid residues; a second acquisition unit configured to acquire a position-specific score matrix of the target protein sequence, the position-specific score matrix comprising position-specific score vectors for each of the plurality of amino acid residues; and a first determination unit configured to determine the target protein... The sequence comprises: a first processing unit configured to process the encoded feature vectors of each amino acid residue and the contact probabilities of the amino acid residue pairs using the first graph neural network to obtain a first comprehensive feature vector for each amino acid residue; a second processing unit configured to process the position-specific score vectors of each amino acid residue and the contact probabilities of the amino acid residue pairs using the second graph neural network to obtain a second comprehensive feature vector for each amino acid residue; and a first prediction unit configured to input the first and second comprehensive feature vectors of each amino acid residue into the prediction subnetwork to obtain a prediction result of the solubility of the target protein sequence.
[0009] According to another aspect of this disclosure, a neural network training apparatus is provided, wherein the initial neural network includes a first initial graph neural network, a second initial graph neural network, and an initial prediction subnetwork. The apparatus includes: a second acquisition unit configured to acquire a sample protein sequence and the true solubility of the sample protein sequence, the sample protein sequence comprising a plurality of amino acid residues; a second encoding unit configured to encode the sample protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the plurality of amino acid residues; a third determination unit configured to determine replacement score vectors for each of the plurality of amino acid residues, the replacement score vectors including a tendency score indicating that a corresponding amino acid residue is replaced at the position of the corresponding amino acid residue with a plurality of preset amino acids; and a fourth determination unit configured to determine the connection of multiple amino acid residue pairs in the sample protein sequence. The initial graph neural network is configured to: a first initial graph neural network for processing the coding feature vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs to obtain a first comprehensive feature vector for each amino acid residue; a second initial graph neural network for processing the substitution score vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs to obtain a second comprehensive feature vector for each amino acid residue; a third initial graph neural network for processing the substitution score vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs to obtain a second comprehensive feature vector for each amino acid residue; a fourth initial graph neural network for processing the substitution score vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs to obtain a second comprehensive feature vector for each amino acid residue; a fifth initial graph neural network for inputting the first comprehensive feature vectors and the second comprehensive feature vectors of multiple amino acid residues into the initial prediction sub-network to obtain a prediction result of the protein solubility of the sample protein sequence; and a parameter tuning unit for adjusting the parameters of the initial neural network based on the prediction result and the actual solubility to obtain a trained neural network.
[0010] According to another aspect of this disclosure, a computer device is provided, comprising: at least one processor; and a memory having a computer program stored thereon, wherein, when executed by the processor, the computer program causes the processor to perform the method described above.
[0011] According to another aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to perform the above-described method.
[0012] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, causes the processor to perform the methods described above.
[0013] According to another aspect of this disclosure, a trained neural network obtained by the above-described neural network training method is provided.
[0014] According to one or more embodiments of this disclosure, within the framework of a graph neural network, the encoded feature vector obtained by encoding a target protein sequence using a protein pre-trained model and the contact probabilities of multiple amino acid residue pairs in the target protein sequence are combined. Furthermore, the substitution score vectors of each amino acid residue in the target protein sequence and the contact probabilities of multiple amino acid residue pairs are combined. This allows the neural network to utilize richer protein information, thereby improving the diversity and comprehensiveness of feature representation. Moreover, it enables the prediction of protein solubility without acquiring or predicting the spatial structure information of the target protein sequence. By setting up two graph neural networks to process the above two combinations respectively, the information related to protein solubility in the encoded feature vectors and substitution score vectors obtained through different methods can be fully extracted, and interference between the two can be avoided. This improves the predictive ability of the obtained first and second comprehensive feature vectors, ultimately resulting in a more accurate prediction of the solubility of the target protein sequence.
[0015] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0016] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0017] Figure 1 A flowchart illustrating a method for predicting protein solubility using a neural network according to an exemplary embodiment of the present disclosure is shown. Figure 2 A flowchart is shown illustrating a process of processing the encoded feature vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs using a first graph neural network according to an exemplary embodiment of the present disclosure. Figure 3 A flowchart illustrating a process for updating node features in first graph data using a first graph neural network according to an exemplary embodiment of the present disclosure is shown. Figure 4 A flowchart is shown illustrating a process of processing the substitution score vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs using a second graph neural network according to an exemplary embodiment of the present disclosure. Figure 5A flowchart illustrating a process for updating node features in second graph data using a second graph neural network according to an exemplary embodiment of the present disclosure is shown. Figure 6 A flowchart is shown illustrating a process of processing a first integrated feature vector and a second integrated feature of each of multiple amino acid residues using a predictive subnetwork, according to an exemplary embodiment of the present disclosure. Figure 7 A schematic diagram of a neural network according to an exemplary embodiment of the present disclosure is shown; Figure 8 A flowchart of a neural network training method according to an exemplary embodiment of the present disclosure is shown; Figure 9 A structural block diagram of an apparatus for predicting protein solubility using a neural network according to an exemplary embodiment of the present disclosure is shown. Figure 10 A structural block diagram of a neural network training apparatus according to exemplary embodiments of the present disclosure is shown; and Figure 11 This is a block diagram illustrating an exemplary computer device that can be applied to exemplary embodiments. Detailed Implementation
[0018] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0019] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0020] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0021] Currently, the relevant technologies mainly rely on traditional laboratories to predict and verify the solubility properties of proteins, and there is a lack of more efficient methods that apply biotechnology to predict protein solubility.
[0022] To address the aforementioned issues, this disclosure, within the framework of a graph neural network, combines the encoded feature vector obtained by encoding the target protein sequence using a protein pre-trained model with the contact probabilities of multiple amino acid residue pairs in the target protein sequence. It also combines the substitution score vectors of each amino acid residue in the target protein sequence with the contact probabilities of multiple amino acid residue pairs. This allows the neural network to utilize richer protein information, thereby improving the diversity and comprehensiveness of feature representation. Furthermore, it enables the prediction of protein solubility without acquiring or predicting the spatial structure information of the target protein sequence. By setting up two graph neural networks to process the above two combinations separately, the information related to protein solubility in the encoded feature vectors and substitution score vectors obtained through different methods can be fully extracted, and interference between the two can be avoided. This improves the predictive power of the obtained first and second comprehensive feature vectors, ultimately resulting in a more accurate prediction of the solubility of the target protein sequence.
[0023] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0024] According to one aspect of this disclosure, a method for predicting protein solubility using a neural network is provided. Figure 1 A flowchart illustrating a method for predicting protein solubility using a neural network according to an exemplary embodiment of the present disclosure is shown. Figure 1As shown, method 100 includes: step S101, obtaining a target protein sequence, the target protein sequence including multiple amino acid residues; step S102, encoding the target protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the multiple amino acid residues; step S103, determining substitution score vectors for each of the multiple amino acid residues, the substitution score vectors including a tendency score indicating that the corresponding amino acid residue is replaced by multiple preset amino acids at the position of the corresponding amino acid residue; step S104, determining the contact probability of multiple amino acid residue pairs in the target protein sequence; step S105, processing the encoding feature vectors for each of the multiple amino acid residues and the contact probability of multiple amino acid residue pairs using a first graph neural network to obtain a first comprehensive feature vector for each of the multiple amino acid residues; step S106, processing the substitution score vectors for each of the multiple amino acid residues and the contact probability of multiple amino acid residue pairs using a second graph neural network to obtain a second comprehensive feature vector for each of the multiple amino acid residues; and step S107, inputting the first comprehensive feature vector and the second comprehensive feature vector for each of the multiple amino acid residues into a prediction sub-network to obtain a prediction result for the solubility of the target protein sequence.
[0025] Therefore, within the framework of graph neural networks, the encoded feature vector obtained by encoding the target protein sequence using a protein pre-trained model is combined with the contact probabilities of multiple amino acid residue pairs in the target protein sequence. Furthermore, the substitution score vectors of each amino acid residue in the target protein sequence are combined with the contact probabilities of multiple amino acid residue pairs. This allows the neural network to utilize richer protein information, thereby improving the diversity and comprehensiveness of feature representation. Moreover, it can predict protein solubility without acquiring or predicting the spatial structure information of the target protein sequence. By setting up two graph neural networks to process the above two combinations respectively, the information related to protein solubility in the encoded feature vectors and substitution score vectors obtained through different methods can be fully extracted, and interference between the two can be avoided. This improves the predictive ability of the obtained first and second comprehensive feature vectors, ultimately resulting in a more accurate prediction of the solubility of the target protein sequence.
[0026] In this disclosure, although the various operations are depicted in the accompanying drawings in a specific order, this should not be construed as requiring that these operations must be performed in the specific order shown or in a sequential order. For example, the order of steps S102, S103 and S104 can be interchanged or performed simultaneously. Similarly, the order of steps S105 and S106 can be interchanged or performed simultaneously.
[0027] In some embodiments, in step S101, the sequence sources of the target protein include, but are not limited to: obtaining the amino acid sequence information of the protein through peptide mass spectrometry and tandem mass spectrometry analysis; transcribing complementary deoxyribonucleic acid (cDNA) from messenger ribonucleic acid (mRNA) and then performing deoxyribonucleic acid (DNA) sequencing to deduce the amino acid sequence of the protein; direct synthesis using synthetic biology; artificially designed protein amino acid sequences; computer-designed protein amino acid sequences; NCBI database; EXProt database; UniProtKB database; and other common bioinformatics databases. The target protein sequence contains multiple amino acid residues, which contain information related to protein properties, such as protein solubility information.
[0028] In some embodiments, in step S102, multiple amino acid residues included in the target protein sequence can be directly input into the protein pre-training model to obtain the encoding result output by the protein pre-training model, i.e., the encoding feature vector of each of the multiple amino acid residues.
[0029] Pre-trained models can be obtained by learning language features in advance on a large amount of unlabeled data, and then fine-tuning them on a small amount of labeled data. This training method can improve the performance of downstream tasks. The protein pre-trained model used in step S102 can be obtained by training on a large corpus of protein or nucleic acid sequences.
[0030] According to some embodiments, the protein pre-training model can be the ProtT5-XL model. The ProtT5-XL model, pre-trained on a large-scale protein sequence dataset, can capture the sequence, structural, and functional information of proteins, providing rich and accurate feature representations. Using the ProtT5-XL model as the protein pre-training model can significantly improve the quality of feature extraction, ultimately enhancing the accuracy of protein solubility prediction.
[0031] In one exemplary embodiment, the ProtT5-XL model can be first trained on the BFD dataset and then fine-tuned using the UniRef50 dataset. This model contains 24 attention layers, each consisting of 32 self-attention heads and 1024 neurons. The output of the last layer can be selected as the encoding feature vector for multiple amino acid residues. Each encoding feature vector has a length of 1024, thus the encoding feature vectors for multiple amino acid residues can form a matrix of size [size missing]. L A matrix of ×1024, where L This represents the number of amino acid residues in the target protein sequence.
[0032] In some embodiments, the protein pre-training model may also be ESM-1b, UniRef, ProteinBert, TAPE, ProtGPT2, ProtTXL, ProtBert, ProtXLNet, ProtAlbert, ProtElectra, ProtT5-XXL, Ankh, or other models, without limitation.
[0033] Observations revealed that the propensity score for amino acid residues to be replaced with other amino acids at specific positions is helpful in predicting protein solubility. By determining the substitution score vectors for multiple amino acid residues in a target protein sequence (including the propensity score for each amino acid residue to be replaced with multiple preset amino acids at its position), and processing the substitution score vectors and the basic probabilities of multiple amino acid residue pairs using a graph neural network, a feature vector with strong predictive power for protein solubility can be obtained.
[0034] In some embodiments, step S103, determining the replacement score vector for each of the multiple amino acid residues, may include: searching a protein database for known protein sequences similar to the target protein sequence to obtain multiple sequence alignment results; for each of the multiple amino acid residues, calculating the target frequency of multiple preset amino acids appearing at the position of the amino acid residue based on the multiple sequence alignment results; and determining the tendency score for the amino acid residue to be replaced by multiple preset amino acids at the position of the amino acid residue based on the target frequency and the background frequency of each of the multiple preset amino acids.
[0035] Background frequency is the probability of an amino acid appearing at all possible positions without any specific context or constraints. It represents the global frequency of each amino acid in a protein sequence and can be statistically obtained from large-scale protein databases. Various predefined amino acids can include natural amino acids as well as non-natural amino acids, such as synthetic amino acids, amino acid analogs that function similarly to natural amino acids, and amino acid mimics. Natural amino acids can include, but are not limited to, 20 common natural amino acids, such as A (alanine), V (valine), L (leucine), I (isoleucine), F (phenylalanine), W (tryptophan), M (methionine), P (proline), G (glycine), T (threonine), S (serine), C (cysteine), N (aspartic acid), Q (glutamine), Y (tyrosine), K (lysine), R (arginine), H (histidine), D (aspartic acid), E (glutamate), B (aspartic acid or aspartic acid), Z (glutamine or glutamate), and J (leucine or isoleucine). It can also include two less common amino acids: pyrolysine and selenocysteine. A propensity score can be represented, for example, as: S ij = log ( P ij / P j ),in S ij Indicates the first i The position corresponding to each amino acid residue is replaced with an amino acid. j The score, P ij In the multiple sequence alignment results, at the first... i An amino acid appears at the position corresponding to each amino acid residue. j frequency, P j It is an amino acid j The background frequency.
[0036] In one exemplary embodiment, a PSI-BLAST tool (e.g., version v2.4.0) can be used to search a protein database (e.g., UniRef90) using the target protein sequence as the query sequence, and multiple sequence alignment results can be obtained after three iterations. The substitution score vectors of multiple amino acid residues can form a matrix of size [size missing]. L A 20 × 10 matrix, where L The number of amino acid residues in the target protein sequence is used, and 20 standard amino acids are used as multiple preset amino acids.
[0037] Contact probability is the likelihood that an amino acid residue pair in a protein sequence will form a contact. In some embodiments, the formation of a contact between two amino acid residue pairs can be defined as the distance between the atoms of two amino acid residues in a protein (e.g., Cβ atoms) being less than a preset distance threshold (e.g., 8 Å).
[0038] Observations revealed that by acquiring the contact probabilities of multiple amino acid residue pairs and combining the encoding feature vectors or replacement score vectors of multiple amino acid residues in the target protein sequence, and then processing them with a graph neural network, a feature vector with strong predictive power for protein solubility can be obtained without acquiring or predicting the spatial structure information of the target protein sequence, thereby significantly reducing the cost of protein solubility prediction.
[0039] In some embodiments, step S104, determining the contact probability of multiple amino acid residue pairs in the target protein sequence, may include: identifying multiple amino acid residue pairs in the target protein sequence; for each amino acid residue pair, obtaining a high-dimensional feature vector between the first and second amino acid residues included in the amino acid residue pair; and simultaneously inputting the high-dimensional feature vectors of the multiple amino acid residue pairs into a trained neural network for predicting contact probabilities to obtain the contact probability of the multiple amino acid residue pairs. In this way, the neural network for predicting contact probabilities can consider information from other amino acids when predicting the contact probability between each amino acid residue pair, thereby improving the accuracy of the predicted contact probability.
[0040] According to some embodiments, multiple amino acid residue pairs may include amino acid residue pairs consisting of an optional first amino acid residue and a second amino acid residue in the target protein sequence. That is, the contact probability between any two amino acid residues in the target protein sequence can be determined in step S104. By determining the contact probabilities between all possible amino acid residue pairs, more comprehensive protein structure information can be provided, reducing information omissions and adapting to the needs of complex proteins. Furthermore, graph neural networks rely on the relationships between nodes (amino acid residues) and edges (amino acid residue pairs) for feature learning and prediction. By determining the contact probabilities of all amino acid residue pairs, more accurate edge topology and edge features can be provided, enabling the graph neural network to make fuller use of this information, thereby improving the performance of the neural network and ultimately contributing to improving the accuracy and reliability of protein solubility prediction.
[0041] In some embodiments, the high-dimensional feature vector between the first and second amino acid residues (i.e., the high-dimensional feature vector of the amino acid residue pair) may include any combination of the following features: the one-hot encoding of each of the first and second amino acid residues, the attention scores of the first and second amino acid residues, and / or the structural information of the target protein sequence. In an exemplary embodiment, the attention score may come from an attention map generated by a protein pre-training model based on the target protein sequence; the structural information may include the probabilities of 3-state secondary structures (SS3) and 8-state secondary structures (SS8), solvent-accessible surface area (ASA), hemispherical exposure (HSE), and the dihedral angle ψ of the protein backbone. θ, τ and / or other protein structure information.
[0042] In some embodiments, the above features can be combined using splicing, element-level operations (e.g., element addition, subtraction, multiplication) or other methods to obtain a high-dimensional feature vector of amino acid residue pairs.
[0043] In some embodiments, the neural network used to predict the probability of contact may employ a residual network (e.g., ResNet). This neural network may be trained using labeled data describing whether or not contact occurs based on amino acid residues.
[0044] In some embodiments, in step S105, the encoded feature vectors of each of the multiple amino acid residues obtained in step S102 and the contact probabilities of the multiple amino acid residue pairs obtained in step S104 can be input into the first graph neural network to obtain the first comprehensive feature vectors of each of the multiple amino acid residues output by the first graph neural network.
[0045] Graph neural networks are deep learning models based on graph-structured data, used to process structured data of nodes and edges. In some embodiments, the first graph neural network and the second graph neural network, which will be described in detail below, can adopt graph convolutional networks (GCN), graph attention networks (GAT), graph isomorphism networks (GIN), diffusion convolutional neural networks (DCNN), or other graph neural network structures.
[0046] According to some embodiments, the first and second graph neural networks can be based on GraphSAGE. GraphSAGE is a sampling-based graph neural network that achieves efficient learning and representation of node features by sampling and aggregating neighboring nodes. These characteristics of GraphSAGE give it significant advantages in processing large-scale graph data, especially when processing protein sequences to perform protein solubility prediction tasks, where GraphSAGE can balance efficiency and accuracy.
[0047] Figure 2 A flowchart illustrating a process, according to an exemplary embodiment of the present disclosure, of processing the encoded feature vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs using a first graph neural network. Process 200 can be used to implement step S105 of the method 100 described above. According to some embodiments, such as... Figure 2 As shown, process 200 may include: step S201, using the encoding feature vectors of multiple amino acid residues as node features and the contact probability of multiple amino acid residue pairs as edge features to construct first graph data; and step S202, using the first graph neural network to update the node features in the first graph data based on the node features and edge features in the first graph data to obtain the first comprehensive feature vectors of multiple amino acid residues.
[0048] Therefore, by using the above method, the sequence information of proteins and the contact information of amino acid residue pairs can be utilized more effectively within the framework of graph neural networks, thereby obtaining a first comprehensive feature vector with stronger predictive ability and improving the accuracy and robustness of subsequent protein solubility prediction.
[0049] In step S201, multiple amino acid residues included in the target protein sequence can be identified as multiple nodes in the first graph data, and the topological relationship of the first graph data can be determined based on multiple amino acid residue pairs. In some embodiments, edges can be set between two nodes corresponding to the first and second amino acid residues included in each of the multiple amino acid residue pairs determined in step S104, thereby obtaining multiple edges in the first graph data. In some embodiments, the multiple amino acid residue pairs include amino acid residue pairs composed of optional first and second amino acid residues in the target protein sequence, then edges can be set between any two nodes in the first graph data, that is, the first graph data is a complete graph.
[0050] In some embodiments, at least one amino acid residue pair with a contact relationship can be determined based on the contact probability of multiple amino acid residue pairs and a preset probability threshold, and the topological relationship of the first graph data can be determined based on at least one amino acid residue pair. For example, an edge can be set between the two nodes corresponding to the first and second amino acid residues included in the amino acid residue pair with a contact probability greater than the preset probability threshold. In this way, it is possible to avoid establishing edges between nodes corresponding to amino acid residues that have no contact relationship, thereby effectively reducing the number of edges, reducing the demand for computing and storage resources, and improving the inference efficiency of the first graph neural network.
[0051] According to some embodiments, the first graph neural network may include three first graph convolutional layers. By configuring three graph convolutional layers in the first graph neural network to process the encoded feature vectors and contact probabilities of multiple amino acid residues, optimal prediction performance can be achieved while maintaining reasonable computational resource consumption.
[0052] In some embodiments, in step S202, the node features in the first graph data can be updated using three first graph convolutional layers iteratively, and the updated features of multiple amino acid residues output by the last first graph convolutional layer can be obtained as the first comprehensive feature vector.
[0053] Figure 3 A flowchart illustrating a process for updating node features in first graph data using a first graph neural network according to an exemplary embodiment of the present disclosure is shown. Process 300 can be used to implement step S202 in process 200 described above. In some embodiments, such as Figure 3 As shown, process 300 may include: step S301, updating the node features in the first graph data by iteratively using three first graph convolutional layers, and obtaining the first intermediate feature vectors of each of the multiple amino acid residues output by the three first graph convolutional layers respectively; and step S302, fusing the first intermediate feature vectors of each of the multiple amino acid residues output by the three first graph convolutional layers respectively to obtain the first comprehensive feature vector of each of the multiple amino acid residues.
[0054] Therefore, through iterative updates of multi-layer graph convolutions, each convolutional layer can collect information from a wider neighborhood, thus aggregating more and more contextual information layer by layer. More specifically, the first layer mainly captures the local features of other amino acid residues that are close to the amino acid residues, the second layer begins to fuse a larger range of neighborhood information, and the third layer captures more global structural information. Finally, by fusing features from multiple layers, the neural network can comprehensively utilize the feature representations of each layer, thereby capturing the complex relationships between amino acid residues and obtaining feature vectors with stronger predictive power.
[0055] In an exemplary embodiment, in step S301, the output size of each first graph convolutional layer is L A matrix of size 128, where L 128 represents the number of amino acid residues in the target protein sequence, and 128 represents the hidden dimension of the first intermediate feature vector.
[0056] In some embodiments, in step S302, the three first intermediate feature vectors output by the three first graph convolutional layers of each amino acid residue can be fused by splicing, weighted summation, mean pooling, max pooling, attention mechanism, direct summation, processing with a multilayer perceptron or other neural network, or any combination of the above methods, to obtain the first comprehensive feature vector of the amino acid residue.
[0057] According to some embodiments, the fusion in step S302 may include merging the first intermediate feature vectors of the plurality of amino acid residues output by the three first graph convolutional layers using a residual network structure.
[0058] Therefore, by using the above method, the first intermediate feature vectors at different levels can be more effectively integrated, enhancing the expressive power of the first comprehensive feature vector obtained after integration, thereby improving the performance of the neural network in protein solubility prediction.
[0059] In an exemplary embodiment, in step S302, for each amino acid residue, the three first intermediate feature vectors (of length 128) output by the three first graph convolutional layers of that amino acid residue can be merged using a residual network structure to obtain the first comprehensive feature vector (of length 384) of that amino acid residue. In other words, the first comprehensive feature vectors of multiple amino acid residues can ultimately form a structure of size [missing information]. L A 384 × matrix.
[0060] Figure 4 A flowchart illustrating a process for processing substitution score vectors of multiple amino acid residues and contact probabilities of multiple amino acid residue pairs using a second graph neural network according to an exemplary embodiment of the present disclosure. Process 400 can be used to implement step S106 in method 100 described above. According to some embodiments, such as Figure 4 As shown, process 400 may include: step S401, using the replacement score vectors of multiple amino acid residues as node features and the contact probabilities of multiple amino acid residue pairs as edge features to construct second graph data; and step S402, using the second graph neural network to update the node features in the second graph data based on the node features and edge features in the second graph data to obtain the second comprehensive feature vectors of the multiple amino acid residues.
[0061] Therefore, by using the above method, it is possible to more effectively utilize the tendency scores of amino acid residues in proteins being replaced by other amino acids and the contact information of amino acid residue pairs within the framework of graph neural networks, thereby obtaining a second comprehensive feature vector with stronger predictive ability and improving the accuracy and robustness of subsequent protein solubility prediction.
[0062] In some embodiments, the second graph data may have a similar or identical topological structure to the first graph data. In step S401, multiple amino acid residues included in the target protein sequence can be identified as multiple nodes in the second graph data, and the topological relationship of the second graph data can be determined based on multiple amino acid residue pairs. In some embodiments, edges can be set between two nodes corresponding to the first and second amino acid residues included in each of the multiple amino acid residue pairs determined in step S104, thereby obtaining multiple edges in the second graph data. In some embodiments, the multiple amino acid residue pairs include amino acid residue pairs composed of optional first and second amino acid residues in the target protein sequence, in which case edges can be set between any two nodes in the second graph data, that is, the second graph data is a complete graph.
[0063] In some embodiments, at least one amino acid residue pair with a contact relationship can be determined based on the contact probability of multiple amino acid residue pairs and a preset probability threshold, and the topological relationship of the second graph data can be determined based on at least one amino acid residue pair. For example, an edge can be set between the two nodes corresponding to the first and second amino acid residues included in the amino acid residue pair with a contact probability greater than the preset probability threshold. In this way, it is possible to avoid establishing edges between nodes corresponding to amino acid residues that have no contact relationship, thereby effectively reducing the number of edges, reducing the demand for computing and storage resources, and improving the inference efficiency of the second graph neural network.
[0064] According to some embodiments, the second graph neural network may include two second graph convolutional layers. By configuring two graph convolutional layers in the second graph neural network to process the substitution fraction vector and contact probability of multiple amino acid residues, optimal prediction performance can be achieved while maintaining reasonable computational resource consumption.
[0065] In some embodiments, in step S402, the node features in the second graph data can be updated using two second graph convolutional layers iteratively, and the updated features of multiple amino acid residues output by the last second graph convolutional layer can be obtained as the second comprehensive feature vector.
[0066] Figure 5A flowchart illustrating a process for updating node features in second graph data using a second graph neural network according to an exemplary embodiment of the present disclosure is shown. Process 500 can be used to implement step S402 in process 400 described above. In some embodiments, such as Figure 5 As shown, process 500 may include: step S501, updating the node features in the second graph data by iteratively using two second graph convolutional layers, and obtaining the second intermediate feature vectors of each of the multiple amino acid residues output by the two second graph convolutional layers respectively; and step S502, fusing the second intermediate feature vectors of each of the multiple amino acid residues output by the two second graph convolutional layers respectively to obtain the second comprehensive feature vector of each of the multiple amino acid residues.
[0067] Therefore, through iterative updates of multi-layer graph convolutions, each convolutional layer can collect information from a wider neighborhood, thus aggregating more and more contextual information layer by layer. More specifically, the first layer mainly captures the local features of other amino acid residues that are close to the amino acid residues, the second layer begins to fuse a larger range of neighborhood information, and finally, by fusing features from multiple layers, the neural network can comprehensively utilize the feature representations of each layer, thereby capturing the complex relationships between amino acid residues and obtaining feature vectors with stronger predictive power.
[0068] In an exemplary embodiment, in step S501, the output size of each second graph convolutional layer is... L A 32 × 10⁻³² matrix, where L 32 represents the number of amino acid residues in the target protein sequence, and 32 represents the hidden dimension of the second intermediate feature vector.
[0069] In some embodiments, in step S502, the two second intermediate feature vectors output by the two second graph convolutional layers of each amino acid residue can be fused to obtain the second comprehensive feature vector of the amino acid residue by splicing, weighted summation, mean pooling, max pooling, attention mechanism, direct summation, processing with multilayer perceptron or other neural networks, or any combination of the above methods.
[0070] According to some embodiments, the fusion in step S502 may include merging the second intermediate feature vectors of the plurality of amino acid residues output by the two second graph convolutional layers using a residual network structure.
[0071] Therefore, by using the above method, the second intermediate feature vectors at different levels can be integrated more effectively, enhancing the expressive power of the second comprehensive feature vector obtained after integration, thereby improving the performance of the neural network in protein solubility prediction.
[0072] In an exemplary embodiment, in step S502, for each amino acid residue, the two second intermediate feature vectors (length 32) output by the two second graph convolutional layers of that amino acid residue can be merged using a residual network structure to obtain a second comprehensive feature vector (length 64) for that amino acid residue. In other words, the final second comprehensive feature vectors of multiple amino acid residues can form a structure with a size of... L A 64×64 matrix.
[0073] Figure 6 A flowchart illustrating a process, according to an exemplary embodiment of the present disclosure, of processing a first integrated feature vector and a second integrated feature vector of each of multiple amino acid residues using a predictive subnetwork. Process 600 can, for example, be used to implement step S107 of method 100 described above. According to some embodiments, such as... Figure 6 As shown, process 600 may include: step S601, pooling the first comprehensive feature vectors of multiple amino acid residues to obtain a first pooled feature vector of the target protein sequence; step S602, pooling the second comprehensive feature vectors of multiple amino acid residues to obtain a second pooled feature vector of the target protein sequence; step S603, fusing the first pooled feature vector and the second pooled feature vector to obtain a target feature vector of the target protein sequence; and step S604, processing the target feature vector using a fully connected layer and an output layer to obtain a prediction result.
[0074] Therefore, by using the above method, the comprehensive feature vector of multiple amino acid residues can be simplified into a fixed-length feature vector that characterizes the target protein sequence, making subsequent processing more efficient. Furthermore, it can effectively fuse the first and second comprehensive feature vectors, which represent different feature sources (encoding feature vectors and replacement score vectors), so that the neural network can make full use of this information to obtain more accurate protein solubility prediction results.
[0075] In some embodiments, in step S601, the first integrated feature vectors of multiple amino acid residues can be pooled along the direction of the amino acid residues to obtain the first pooled feature vector of the target protein sequence. In an exemplary embodiment, the matrix (of size ) formed by the first integrated feature vectors of multiple amino acid residues can be pooled. L A first pooling feature vector of size 1×128 is obtained by performing an average pooling operation along the direction of amino acid residues (×128).
[0076] In some embodiments, in step S602, the second integrated feature vectors of multiple amino acid residues can be pooled along the direction of the amino acid residues to obtain the second pooled feature vector of the target protein sequence. In an exemplary embodiment, the matrix (of size ) formed by the second integrated feature vectors of multiple amino acid residues can be... L A second pooling feature vector of size 1×64 is obtained by performing an average pooling operation along the direction of amino acid residues (×64).
[0077] It is understood that this disclosure does not limit the execution order between steps S601 and S602. The execution order of these two steps can be interchanged, or the two steps can be executed simultaneously.
[0078] In some embodiments, in step S603, the first pooling feature vector obtained in steps S601 and S602 can be fused with the second pooling feature vector by splicing, weighted summation, mean pooling, max pooling, attention mechanism, direct summation, processing with a multilayer perceptron or other neural network, or any combination of the above methods, to obtain the target feature vector of the target protein sequence.
[0079] According to some embodiments, the fusion in step S603 may include merging the first pooling feature vector and the second pooling feature vector using a residual network structure.
[0080] Therefore, by using the above method, the first pooling feature vector and the second pooling feature vector based on different protein information sources can be more effectively fused, enhancing the expressive power of the target feature vector obtained after fusion, thereby improving the performance of the neural network in protein solubility prediction.
[0081] In an exemplary embodiment, in step S603, a residual network can be used to merge the first pooling feature vector of size 1×384 with the second pooling feature vector of size 1×64 to obtain a target feature vector of size 1×448, which serves as a comprehensive representation of the entire target protein sequence.
[0082] In some embodiments, in step S604, the target feature vector can be input into a prediction head including a fully connected layer and an output layer to obtain the prediction result of the solubility of the target protein sequence output by the prediction head. It is understood that the method of this disclosure can be used for both qualitative prediction of whether a target protein sequence can be dissolved and quantitative prediction of the solubility of a target protein sequence. Different prediction methods can be implemented by using different prediction heads (e.g., regression heads, classification heads).
[0083] In an exemplary embodiment, in step S604, the target feature vector obtained in step S603 can be input into two fully connected layers to obtain a feature vector of size 1×128, and this feature vector is then sent to the output layer to obtain the final prediction result. Softmax can be used as the classification function of the output layer to predict the protein solubility probability. A protein solubility probability value of 0.5 can be used as a threshold to classify protein solubility; that is, when the protein solubility probability is higher than 0.5, the target protein sequence can be classified as a soluble protein, and the output result can be marked as 1; otherwise, it can be classified as an insoluble protein, and the output result can be marked as 0.
[0084] In some embodiments, both the neural network and the protein pre-trained model used in method 100 are trained. The neural network may be trained, for example, using a neural network training method provided in this disclosure (e.g., method 800, which will be described below).
[0085] According to some embodiments, the trained neural network is obtained through end-to-end training using a sample protein sequence and the actual solubility of the sample protein sequence. The encoded feature vectors of multiple sample amino acid residues corresponding to the sample protein sequence (encoded using a protein pre-training model), the substitution score vectors of multiple sample amino acid residues, and the contact probabilities of multiple sample amino acid residue pairs in the sample protein sequence can be input into a first initial graph neural network and a second initial graph neural network to be trained. The sample prediction results output by the initial prediction subnetwork in the initial neural network are then obtained. Finally, the initial neural network is trained end-to-end based on the sample prediction results and the actual solubility to obtain the trained neural network.
[0086] In some specific embodiments, the trained neural network is obtained through end-to-end training using a sample protein sequence and the actual solubility of the sample protein sequence. The training includes: acquiring the sample protein sequence and the actual solubility of the sample protein sequence, wherein the sample protein sequence comprises multiple amino acid residues; encoding the sample protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the multiple amino acid residues; determining substitution score vectors for each of the multiple amino acid residues, wherein the substitution score vectors include a tendency score indicating that the corresponding amino acid residue is replaced by multiple preset amino acids at the corresponding amino acid residue position; determining the contact probability of multiple amino acid residue pairs in the sample protein sequence; and using the first initial... The initial graph neural network processes the encoding feature vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs to obtain a first comprehensive feature vector for each of the plurality of amino acid residues; the second initial graph neural network processes the substitution score vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs to obtain a second comprehensive feature vector for each of the plurality of amino acid residues; the first and second comprehensive feature vectors of each of the plurality of amino acid residues are input into the initial prediction sub-network to obtain a prediction result of the protein solubility of the sample protein sequence; and based on the prediction result and the actual solubility, the parameters of the initial neural network are adjusted to obtain a trained neural network.
[0087] Figure 7 A schematic diagram of a neural network according to an exemplary embodiment of the present disclosure is shown. (As...) Figure 7 As shown, the neural network 700 may include: a first graph convolutional network 720, a second graph convolutional network 730, and a prediction sub-network 740. The first graph convolutional network 720 is configured to receive encoding feature vectors 702 for each of multiple amino acid residues and contact probabilities 704 for multiple amino acid residue pairs, and outputs a first comprehensive feature vector for each of the multiple amino acid residues. The second graph convolutional network 730 is configured to receive substitution score vectors 708 for each of the multiple amino acid residues and contact probabilities 704 for multiple amino acid residue pairs, and outputs a second comprehensive feature vector for each of the multiple amino acid residues. The prediction sub-network 740 is configured to receive the first and second comprehensive feature vectors for each of the multiple amino acid residues, and outputs a prediction result 712 of the solubility of the target protein sequence.
[0088] In some embodiments, first graph data 706 can be constructed based on the encoding feature vectors 702 of each amino acid residue and the contact probabilities 704 of multiple amino acid residue pairs. The first graph convolutional network 720 may include three first graph convolutional layers 722, configured to iteratively update the node features in the first graph data and output first intermediate feature vectors of each amino acid residue. The first graph convolutional network 720 may also include a first fusion module 724, configured to fuse the first intermediate feature vectors of each amino acid residue output by the three first graph convolutional layers to obtain a first comprehensive feature vector of each amino acid residue.
[0089] In some embodiments, second graph data 710 can be constructed based on the substitution score vectors 708 of each amino acid residue and the contact probabilities 704 of multiple amino acid residue pairs. The second graph convolutional network 730 may include two second graph convolutional layers 732, configured to iteratively update the node features in the second graph data and output second intermediate feature vectors of each amino acid residue. The second graph convolutional network 730 may also include a second fusion module 734, configured to fuse the second intermediate feature vectors of each amino acid residue output by the two second graph convolutional layers to obtain a second comprehensive feature vector of each amino acid residue.
[0090] In some embodiments, the prediction subnetwork 740 may include: a first pooling module 742 configured to pool the first comprehensive feature vectors of multiple amino acid residues to obtain a first pooled feature vector of the target protein sequence; a second pooling module 744 configured to pool the second comprehensive feature vectors of multiple amino acid residues to obtain a second pooled feature vector of the target protein sequence; a third fusion module 746 configured to fuse the first pooling feature vector and the second pooling feature vector to obtain a target feature vector of the target protein sequence; and a fully connected layer and an output layer 748 configured to process the target feature vector to obtain a prediction result 712.
[0091] In this disclosure, the terms neural network, neural network model, deep learning model, model or other similar terms may be used interchangeably and, unless otherwise specified, may refer to the trained neural network used in method 100 or the initial neural network trained in method 800, which will be described below.
[0092] According to another aspect of this disclosure, a neural network training method is provided. Figure 8A flowchart of a neural network training method according to an exemplary embodiment of the present disclosure is shown. The method 800 includes: step S801, obtaining a sample protein sequence and the true solubility of the sample protein sequence, the sample protein sequence comprising multiple amino acid residues; step S802, encoding the sample protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the multiple amino acid residues; step S803, determining replacement score vectors for each of the multiple amino acid residues, the replacement score vectors including a tendency score indicating that the corresponding amino acid residue is replaced by multiple preset amino acids at the corresponding amino acid residue position; step S804, determining the contact probability of multiple amino acid residue pairs in the sample protein sequence; step S805, using a first initial graph neural network to train each of the multiple amino acid residue pairs... Step S806: Process the encoding feature vector of each amino acid residue and the contact probability of multiple amino acid residue pairs to obtain the first comprehensive feature vector of each amino acid residue; Step S807: Input the first comprehensive feature vector and the second comprehensive feature vector of each amino acid residue into the initial prediction sub-network to obtain the prediction result of the protein solubility of the sample protein sequence; and Step S808: Based on the prediction result and the actual solubility, adjust the parameters of the initial neural network to obtain the trained neural network.
[0093] Therefore, within the framework of graph neural networks, the encoded feature vector obtained by encoding the protein sequence using a protein pre-trained model is combined with the contact probabilities of multiple amino acid residue pairs in the protein sequence. Furthermore, the substitution score vectors of each amino acid residue in the protein sequence are combined with the contact probabilities of multiple amino acid residue pairs. This allows the neural network to utilize richer protein information, thereby improving the diversity and comprehensiveness of feature representation. Moreover, it can predict protein solubility without acquiring or predicting the spatial structure information of the protein sequence. By setting up two graph neural networks to process the above two combinations respectively, the information related to protein solubility in the encoded feature vectors and substitution score vectors obtained through different methods can be fully extracted, and interference between the two can be avoided. This improves the predictive ability of the obtained first and second comprehensive feature vectors, ultimately enabling the trained neural network to output more accurate predictions of protein sequence solubility.
[0094] It is understood that the operations and effects of steps S801-S807 on the sample protein sequence in method 800 can be referred to the description of steps S101-S107 in method 100 above. The implementation details of steps S101 and S107, as well as the procedures for implementing steps S101-S107, are applicable to steps S801-S807 (only the target protein sequence needs to be replaced with the sample protein sequence). In some embodiments, the neural network trained using method 800 can be used in method 100 to predict protein solubility.
[0095] Furthermore, the order of steps S802, S803, and S804 can be interchanged or performed simultaneously. For example, the order of steps S805 and S806 can be interchanged or performed simultaneously.
[0096] In some embodiments, in step S801, sample protein sequences and their true solubility can be obtained from a predetermined training set. In an exemplary embodiment, the PSI: Biology dataset provided by the NetSolP tool can be used as the training set. This dataset has 11,226 original protein samples. After removing samples with conflicting labels for the same protein sequence, the final dataset used to train the model contains 11,110 samples. The number of soluble and insoluble protein samples in the dataset are 7,402 and 3,708, respectively, and the classification labels for soluble and insoluble proteins are 1 and 0, respectively. By training the initial neural network using the sample protein sequences and true solubility (classification labels) obtained from the above training set, the trained neural network can be equipped with the ability to output classification results indicating whether a protein sequence is soluble.
[0097] In some embodiments, the solubility of sample protein sequences obtained by other methods can be used as annotation data for the true solubility of the sample protein sequences. By training an initial neural network using such sample data, the trained neural network can be equipped with the ability to output regression results describing protein solubility.
[0098] In some embodiments, in step S802, the protein pre-training model can use the ProtT5-XL pre-training model, which encodes each amino acid residue in the sample protein sequence, with each encoded amino acid residue corresponding to a vector of length 1024, i.e., a vector of length 1024. L The protein sequence will encode a protein of size . LThe ProtT5-XL model is a 1024 × 10 matrix. Pre-trained on a large-scale protein sequence dataset, it captures sequence, structural, and functional information of proteins, providing rich and accurate feature representations. Using the ProtT5-XL model as a protein pre-training model significantly improves the quality of feature extraction, ultimately enhancing the accuracy of protein solubility prediction using trained neural networks.
[0099] In some embodiments, step S803, determining the replacement score vector for each of the multiple amino acid residues, may include: searching a protein database for known protein sequences similar to the sample protein sequence to obtain multiple sequence alignment results; for each of the multiple amino acid residues, calculating the target frequency of multiple preset amino acids appearing at the position of the amino acid residue based on the multiple sequence alignment results; and determining the tendency score for the amino acid residue to be replaced by multiple preset amino acids at the position of the amino acid residue based on the target frequency and the background frequency of each of the multiple preset amino acids.
[0100] In one exemplary embodiment, a PSI-BLAST tool (e.g., version v2.4.0) can be used to search a protein database (e.g., UniRef90) using the sample protein sequence as the query sequence, and multiple sequence alignment results can be obtained after three iterations. The substitution score vectors of multiple amino acid residues can form a matrix of size [size missing]. L A 20 × 10 matrix, where L The number of amino acid residues in the sample protein sequence is used, and 20 standard amino acids are used as multiple preset amino acids.
[0101] Contact probability is the likelihood that an amino acid residue pair in a protein sequence will form a contact. In some embodiments, the formation of a contact between two amino acid residue pairs can be defined as the distance between the atoms of two amino acid residues in a protein (e.g., Cβ atoms) being less than a preset distance threshold (e.g., 8 Å).
[0102] In some embodiments, step S804, determining the contact probability of multiple amino acid residue pairs in a sample protein sequence, may include: identifying multiple amino acid residue pairs in the sample protein sequence; for each amino acid residue pair, obtaining a high-dimensional feature vector between the first and second amino acid residues included in that amino acid residue pair; and simultaneously inputting the high-dimensional feature vectors of the multiple amino acid residue pairs into a trained neural network for predicting contact probabilities to obtain the contact probability of the multiple amino acid residue pairs. In this way, the neural network for predicting contact probabilities can consider information from other amino acids when predicting the contact probability between each amino acid residue pair, thereby improving the accuracy of the predicted contact probability.
[0103] In some embodiments, multiple amino acid residue pairs may include amino acid residue pairs consisting of a first amino acid residue and a second amino acid residue selected from the sample protein sequence. That is, the contact probability between any two amino acid residues in the sample protein sequence can be determined in step S804. By determining the contact probabilities between all possible amino acid residue pairs, more comprehensive protein structure information can be provided, reducing information omissions and adapting to the needs of complex proteins. Furthermore, graph neural networks rely on the relationships between nodes (amino acid residues) and edges (amino acid residue pairs) for feature learning and prediction. By determining the contact probabilities of all amino acid residue pairs, more accurate edge topology and edge features can be provided, enabling the graph neural network to make fuller use of this information, thereby improving model performance and ultimately contributing to improving the accuracy and reliability of protein solubility prediction using trained neural networks.
[0104] In some embodiments, the high-dimensional feature vector between the first and second amino acid residues (i.e., the high-dimensional feature vector of the amino acid residue pair) may include any combination of the following features: the one-hot encoding of each of the first and second amino acid residues, the attention scores of the first and second amino acid residues, and / or the structural information of the target protein sequence. In an exemplary embodiment, the attention score may come from an attention map generated by a protein pre-training model based on the target protein sequence; the structural information may include the probabilities of 3-state secondary structures (SS3) and 8-state secondary structures (SS8), solvent-accessible surface area (ASA), hemispherical exposure (HSE), and the dihedral angle ψ of the protein backbone. θ, τ and / or other protein structure information.
[0105] In some embodiments, the above features can be combined using splicing, element-level operations (e.g., element addition, subtraction, multiplication) or other methods to obtain a high-dimensional feature vector of amino acid residue pairs.
[0106] In some embodiments, the neural network used to predict the probability of contact may employ a residual network (e.g., ResNet). This neural network may be trained using labeled data describing whether or not contact occurs based on amino acid residues.
[0107] In some embodiments, in step S805, the encoded feature vectors of each of the multiple amino acid residues obtained in step S802 and the contact probabilities of the multiple amino acid residue pairs obtained in step S804 can be input into the first initial graph neural network to obtain the first comprehensive feature vectors of each of the multiple amino acid residues output by the first initial graph neural network.
[0108] According to some embodiments, step S805, processing the encoding feature vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs using a first initial graph neural network, may include (not shown in the figure): step S8051, using the encoding feature vectors of multiple amino acid residues as node features and the contact probabilities of multiple amino acid residue pairs as edge features to construct first graph data; and step S8052, using the first initial graph neural network to update the node features in the first graph data based on the node features and edge features in the first graph data to obtain the first comprehensive feature vectors of multiple amino acid residues.
[0109] Therefore, by using the above method, the sequence information of proteins and the contact information of amino acid residue pairs can be utilized more effectively within the framework of graph neural networks, thereby obtaining a first comprehensive feature vector with stronger predictive ability and improving the accuracy and robustness of protein solubility prediction using trained neural networks.
[0110] In step S8051, multiple amino acid residues included in the sample protein sequence can be identified as multiple nodes in the first graph data, and the topological relationship of the first graph data can be determined based on multiple amino acid residue pairs. In some embodiments, edges can be set between two nodes corresponding to the first and second amino acid residues included in each of the multiple amino acid residue pairs determined in step S804, thereby obtaining multiple edges in the first graph data. In some embodiments, the multiple amino acid residue pairs include amino acid residue pairs composed of randomly selected first and second amino acid residues in the sample protein sequence, in which case edges can be set between any two nodes in the first graph data, that is, the first graph data is a complete graph.
[0111] In some embodiments, at least one amino acid residue pair with a contact relationship can be determined based on the contact probability of multiple amino acid residue pairs and a preset probability threshold, and the topological relationship of the first graph data can be determined based on at least one amino acid residue pair. For example, an edge can be set between the two nodes corresponding to the first and second amino acid residues included in the amino acid residue pair with a contact probability greater than the preset probability threshold. In this way, it is possible to avoid establishing edges between nodes corresponding to amino acid residues that have no contact relationship, thereby effectively reducing the number of edges, reducing the demand for computing and storage resources, and improving the inference efficiency of the first graph neural network.
[0112] According to some embodiments, the first initial graph neural network may include three first graph convolutional layers. By configuring three graph convolutional layers in the first graph neural network to process the encoded feature vectors and contact probabilities of multiple amino acid residues, optimal prediction performance can be achieved while maintaining reasonable computational resource consumption.
[0113] In some embodiments, in step S202, the node features in the first graph data can be updated using three first graph convolutional layers iteratively, and the updated features of multiple amino acid residues output by the last first graph convolutional layer can be obtained as the first comprehensive feature vector.
[0114] In some embodiments, in step S8052, the node features in the first graph data can be updated using three first graph convolutional layers iteratively, and the updated features of multiple amino acid residues output by the last first graph convolutional layer can be obtained as the first comprehensive feature vector.
[0115] In some embodiments, step S8052, updating the node features in the first graph data based on the node features and edge features in the first graph data using a first initial graph neural network to obtain the first comprehensive feature vectors of each of the multiple amino acid residues, may include (not shown in the figure): step S80521, updating the node features in the first graph data iteratively using three first graph convolutional layers, and obtaining the first intermediate feature vectors of each of the multiple amino acid residues output by the three first graph convolutional layers respectively; and step S80522, fusing the first intermediate feature vectors of each of the multiple amino acid residues output by the three first graph convolutional layers respectively to obtain the first comprehensive feature vectors of each of the multiple amino acid residues.
[0116] Therefore, through iterative updates of multi-layer graph convolutions, each convolutional layer can collect information from a wider neighborhood, thus aggregating more and more contextual information layer by layer. More specifically, the first layer mainly captures the local features of other amino acid residues that are close to the amino acid residues, the second layer begins to fuse a larger range of neighborhood information, and the third layer captures more global structural information. Finally, by fusing features from multiple layers, the model can comprehensively utilize the feature representations of each layer, thereby capturing the complex relationships between amino acid residues and obtaining feature vectors with stronger predictive power.
[0117] In an exemplary embodiment, in step S80521, the output size of each first graph convolutional layer is L A matrix of size 128, where L 128 represents the number of amino acid residues in the sample protein sequence, and 128 represents the hidden dimension of the first intermediate feature vector.
[0118] In some embodiments, in step S80522, the three first intermediate feature vectors output by the three first graph convolutional layers of each amino acid residue can be fused by splicing, weighted summation, mean pooling, max pooling, attention mechanism, direct summation, processing with multilayer perceptron or other neural networks, or any combination of the above methods, to obtain the first comprehensive feature vector of the amino acid residue.
[0119] According to some embodiments, the fusion in step S80522 may include merging the first intermediate feature vectors of the plurality of amino acid residues output by the three first graph convolutional layers using a residual network structure.
[0120] Therefore, by using the above method, the first intermediate feature vectors at different levels can be more effectively integrated, enhancing the expressive power of the first comprehensive feature vector obtained after integration, thereby improving the performance of the neural network in protein solubility prediction.
[0121] In an exemplary embodiment, in step S80522, for each amino acid residue, the three first intermediate feature vectors (of length 128) output by the three first graph convolutional layers of that amino acid residue can be merged using a residual network structure to obtain the first comprehensive feature vector (of length 384) of that amino acid residue. In other words, the first comprehensive feature vectors of multiple amino acid residues can ultimately form a structure with a size of L A 384 × matrix.
[0122] According to some embodiments, step S806, processing the substitution score vectors of each of the multiple amino acid residues and the contact probabilities of the multiple amino acid residue pairs using a second initial graph neural network to obtain the second comprehensive feature vectors of each of the multiple amino acid residues, may include (not shown in the figure): step S8061, using the substitution score vectors of each of the multiple amino acid residues as node features and the contact probabilities of the multiple amino acid residue pairs as edge features to construct second graph data; and step S8062, using the second graph neural network to update the node features in the second graph data based on the node features and edge features in the second graph data to obtain the second comprehensive feature vectors of each of the multiple amino acid residues.
[0123] Therefore, by using the above method, it is possible to more effectively utilize the tendency scores of amino acid residues in proteins being replaced by other amino acids and the contact information of amino acid residue pairs within the framework of graph neural networks, thereby obtaining a second comprehensive feature vector with stronger predictive ability, which improves the accuracy and robustness of subsequent protein solubility prediction using trained neural networks.
[0124] In some embodiments, the second graph data may have a similar or identical topological structure to the first graph data. In step S8061, multiple amino acid residues included in the sample protein sequence can be identified as multiple nodes in the second graph data, and the topological relationships of the second graph data can be determined based on multiple amino acid residue pairs. In some embodiments, edges can be set between two nodes corresponding to the first and second amino acid residues included in each of the multiple amino acid residue pairs determined in step S804, thereby obtaining multiple edges in the second graph data. In some embodiments, the multiple amino acid residue pairs include amino acid residue pairs composed of optional first and second amino acid residues in the sample protein sequence, in which case edges can be set between any two nodes in the second graph data; that is, the second graph data is a complete graph.
[0125] In some embodiments, at least one amino acid residue pair with a contact relationship can be determined based on the contact probability of multiple amino acid residue pairs and a preset probability threshold, and the topological relationship of the second graph data can be determined based on at least one amino acid residue pair. For example, an edge can be set between the two nodes corresponding to the first and second amino acid residues included in the amino acid residue pair with a contact probability greater than the preset probability threshold. In this way, it is possible to avoid establishing edges between nodes corresponding to amino acid residues that have no contact relationship, thereby effectively reducing the number of edges, reducing the demand for computing and storage resources, and improving the inference efficiency of the second graph neural network.
[0126] According to some embodiments, the second graph neural network may include two second graph convolutional layers. By configuring two graph convolutional layers in the second graph neural network to process the substitution fraction vector and contact probability of multiple amino acid residues, optimal prediction performance can be achieved while maintaining reasonable computational resource consumption.
[0127] In some embodiments, in step S8062, the node features in the second graph data can be updated using two second graph convolutional layers, and the updated features of multiple amino acid residues output by the last second graph convolutional layer can be obtained as the second comprehensive feature vector.
[0128] In some embodiments, step S8062, updating the node features in the second graph data based on the node features and edge features in the second graph data using the second graph neural network to obtain the second comprehensive feature vector of each of the plurality of amino acid residues, may include (not shown in the figure): step S80621, updating the node features in the second graph data iteratively using two second graph convolutional layers, and obtaining the second intermediate feature vectors of each of the plurality of amino acid residues output by the two second graph convolutional layers respectively; and step S80622, fusing the second intermediate feature vectors of each of the plurality of amino acid residues output by the two second graph convolutional layers respectively to obtain the second comprehensive feature vector of each of the plurality of amino acid residues.
[0129] Therefore, through iterative updates of multi-layer graph convolutions, each convolutional layer can collect information from a wider neighborhood, thus aggregating more and more contextual information layer by layer. More specifically, the first layer mainly captures the local features of other amino acid residues that are close to the amino acid residues, the second layer begins to fuse a larger range of neighborhood information, and finally, by fusing features from multiple layers, the model can comprehensively utilize the feature representations of each layer, thereby capturing the complex relationships between amino acid residues and obtaining a feature vector with stronger predictive power.
[0130] In an exemplary embodiment, in step S80621, the output size of each second graph convolutional layer is... L A 32 × 10⁻³² matrix, where L 32 represents the number of amino acid residues in the sample protein sequence, and 32 represents the hidden dimension of the second intermediate feature vector.
[0131] In some embodiments, in step S80622, the two second intermediate feature vectors output by the two second graph convolutional layers of each amino acid residue can be fused to obtain the second comprehensive feature vector of the amino acid residue by splicing, weighted summation, mean pooling, max pooling, attention mechanism, direct summation, processing with multilayer perceptron or other neural networks, or any combination of the above methods.
[0132] According to some embodiments, the fusion in step S80622 may include merging the second intermediate feature vectors of the plurality of amino acid residues output by the two second graph convolutional layers using a residual network structure.
[0133] Therefore, by using the above method, the second intermediate feature vectors at different levels can be more effectively fused, enhancing the expressive power of the second comprehensive feature vector obtained after fusion, thereby improving the performance of the trained neural network in protein solubility prediction.
[0134] In an exemplary embodiment, in step S80622, for each amino acid residue, the two second intermediate feature vectors (length 32) output by the two second graph convolutional layers of that amino acid residue can be merged using a residual network structure to obtain a second comprehensive feature vector (length 64) for that amino acid residue. In other words, the final second comprehensive feature vectors of multiple amino acid residues can form a structure with a size of... L A 64×64 matrix.
[0135] According to some embodiments, step S807, inputting the first and second integrated feature vectors of each of the multiple amino acid residues into the initial prediction sub-network to obtain the prediction result of the protein solubility of the sample protein sequence, may include (not shown in the figure): step S8071, pooling the first integrated feature vectors of each of the multiple amino acid residues to obtain the first pooled feature vector of the sample protein sequence; step S8072, pooling the second integrated feature vectors of each of the multiple amino acid residues to obtain the second pooled feature vector of the sample protein sequence; step S8073, fusing the first pooled feature vector and the second pooled feature vector to obtain the target feature vector of the sample protein sequence; and step S8074, processing the target feature vector using a fully connected layer and an output layer to obtain the prediction result.
[0136] Therefore, by using the above method, the comprehensive feature vector of multiple amino acid residues can be simplified into a fixed-length feature vector that characterizes the protein sequence of the sample, making subsequent processing more efficient. Furthermore, it can effectively fuse the first and second comprehensive feature vectors, which represent different feature sources (encoded feature vectors and replacement score vectors), ultimately enabling the trained neural network to make full use of this information to obtain more accurate protein solubility prediction results.
[0137] In some embodiments, in step S8071, the first integrated feature vectors of multiple amino acid residues can be pooled along the direction of the amino acid residues to obtain a first pooled feature vector of the sample protein sequence. In an exemplary embodiment, the matrix (of size ) formed by the first integrated feature vectors of multiple amino acid residues can be pooled. L A first pooling feature vector of size 1×128 is obtained by performing an average pooling operation along the direction of amino acid residues (×128).
[0138] In some embodiments, in step S8072, the second integrated feature vectors of multiple amino acid residues can be pooled along the direction of the amino acid residues to obtain a second pooled feature vector of the sample protein sequence. In an exemplary embodiment, the matrix (of size ) formed by the second integrated feature vectors of multiple amino acid residues can be pooled. LA second pooling feature vector of size 1×64 is obtained by performing an average pooling operation along the direction of amino acid residues (×64).
[0139] It is understood that this disclosure does not limit the execution order between steps S8071 and S8072. The execution order of these two steps can be interchanged, or the two steps can be executed simultaneously.
[0140] In some embodiments, in step S8073, the first pooling feature vector obtained in steps S8071 and S8072 can be fused with the second pooling feature vector by splicing, weighted summation, mean pooling, max pooling, attention mechanism, direct summation, processing with a multilayer perceptron or other neural network, or any combination of the above methods, to obtain the target feature vector of the sample protein sequence.
[0141] According to some embodiments, the fusion in step S8073 may include merging the first pooling feature vector and the second pooling feature vector using a residual network structure.
[0142] Therefore, by using the above method, the first pooling feature vector and the second pooling feature vector based on different protein information sources can be more effectively fused, enhancing the expressive power of the target feature vector obtained after fusion, thereby improving the performance of the trained neural network in protein solubility prediction.
[0143] In an exemplary embodiment, in step S8073, a residual network can be used to merge the first pooling feature vector of size 1×384 with the second pooling feature vector of size 1×64 to obtain a target feature vector of size 1×448, which serves as a comprehensive representation of the entire sample protein sequence.
[0144] In some embodiments, in step S8074, the target feature vector can be input into a prediction head including a fully connected layer and an output layer to obtain the prediction result of the solubility of the sample protein sequence output by the prediction head. It is understood that the method of this disclosure can be used for both qualitative prediction of whether a sample protein sequence is soluble and quantitative prediction of the solubility of a sample protein sequence. Different prediction heads (e.g., regression heads, classification heads) and different true solubilities (e.g., classification labels indicating whether a sample protein sequence can dissolve and regression labels describing the solubility of a sample protein sequence) can be used to train the initial neural network to achieve the aforementioned different prediction methods.
[0145] In an exemplary embodiment, in step S8074, the target feature vector obtained in step S8073 can be input into two fully connected layers to obtain a feature vector of size 1×128, and this feature vector is then sent to the output layer to obtain the final prediction result. Softmax can be used as the classification function of the output layer to predict the protein solubility probability. A threshold of 0.5 can be used to classify the protein; that is, when the protein solubility probability is higher than 0.5, the sample protein sequence can be classified as a soluble protein, and the output result can be marked as 1; otherwise, it can be classified as an insoluble protein, and the output result can be marked as 0.
[0146] In some embodiments, in step S807, a loss value can be calculated based on the prediction result, the true solubility, and a predetermined loss function. The learnable parameters of each sub-network (first initial graph convolutional network, second initial graph convolutional network, and initial prediction sub-network) in the initial neural network are then adjusted based on the loss value to obtain a trained neural network. In an exemplary embodiment, the loss function can use binary cross-entropy loss. During training, the neural network is adjusted and optimized according to the loss value until the neural network performance no longer shows a significant improvement.
[0147] In an exemplary embodiment, the initial neural network can employ any combination of the following hyperparameters to achieve better training results and prediction performance: using the Adam optimizer; using a rectified linear unit (ReLU) as the activation function, employing an L2 regularization term with a coefficient set to 1e-03; setting the dropout rate to 0.3; setting the learning rate to 0.01; using the average aggregation method to aggregate neighbor node information in the first and second initial graph convolutional networks; setting the batch size to 512; and setting the training epoch to 30, in which all data will be trained once according to the selected batch size during each training epoch.
[0148] In some embodiments, to test the generalization performance of the trained neural network, the following three independent test sets can be selected to validate the model, and samples in all test sets with a global protein sequence similarity higher than 25% with the training set are removed. The global sequence similarity comparison tool used is 32-bit USEARCH v11.0.667. The three independent test sets are: the NESG dataset, the S. Cerevisiaes dataset, and the eSol dataset. These three independent test sets have the following characteristics: 1. NESG Dataset: The North East Structural Consortium (NESG) expressed 1323 proteins in E. coli using a standardized production process and scored their solubility. Using this data and comparing it with the training set, 1319 samples were ultimately retained, including 838 soluble proteins and 481 insoluble proteins.
[0149] 2. *S. Cerevisiaes* dataset: Contains 109 protein samples expressed in *Saccharomyces cerevisiae*. Under the same external conditions, to minimize environmental impact, solubility was measured using a cell-free expression system called PURE. After filtering this dataset against the training set, 103 samples were retained for regression testing, with solubility values as input to the model. Using the same threshold as the eSol dataset, this dataset was classified, ultimately yielding 72 samples for classification testing, including 53 soluble proteins and 19 insoluble proteins.
[0150] 3. eSol Dataset: Whole *E. coli* proteins were synthesized individually using an in vitro reconstructed translation system, and protein aggregation tendencies were analyzed, ultimately resulting in the successful quantification of 3173 proteins. Solubility was defined as the ratio of protein content in the supernatant to the total protein content in a physicochemical experiment, using a threshold of 30% and 70% to define insoluble and soluble proteins. After aligning the eSol dataset with the training set and filtering, 2943 proteins were retained for regression applications. Following the defined thresholds, 2067 proteins were obtained for classification applications, with 928 soluble proteins and 1139 insoluble proteins.
[0151] To better verify the accuracy of the solubility data, the following performance metrics were selected to validate the solubility prediction performance of the neural network: True Positive (TP), False Positive (FP), True Negative (TN), False Negative (FN), Accuracy (ACC), Receiver Operating Characteristic (ROC), Area Under Curve (AUC), Matthews Correlation Coefficient (MCC), Precision, Recall, F1 Score, and Pearson Correlation Coefficient (R). The specific definitions are as follows: Accuracy (ACC) is the most basic classification metric, representing the percentage of samples that correctly predict the outcome out of the total sample. The formula for ACC is ACC = (TP+TN) / (P+N), and its value ranges from 0 to 1. Generally, higher accuracy indicates better model performance.
[0152] The area under the ROC curve, also known as AUC, is a popular parameter-independent metric used to describe binary classifiers. AUC is the area under the Receiver Operational Feature (ROC) curve, with values ranging from 0 to 1. A higher AUC indicates better classifier performance.
[0153] The Matthews correlation coefficient (MCC) is a comprehensive performance metric used to evaluate binary classification models, particularly suitable for handling imbalanced datasets. The formula for calculating MCC is MCC = (TP×TN-FP×FN) / (sqrt((TP+FP)×(TP+FN)×(TN+FP)×(TN+FN))). The value of MCC ranges from -1 to +1, where +1 represents perfect prediction, 0 represents random prediction, and -1 represents a prediction that is completely inconsistent with actual observations.
[0154] Precision refers to the proportion of correctly predicted positive samples out of all predicted positive samples. The formula for precision is: precision = TP / (TP+FP), where TP (False Negative) represents the positive samples classified as positive, and FP (False Positive) represents the negative samples classified as positive. The value of precision ranges from 0 to 1; higher precision means fewer false positives (FP) generated by the model.
[0155] Recall is the proportion of correctly predicted positive samples out of the total number of positive samples. The formula for recall is: recall = TP / (TP+FN), where FN (False Negative) represents the positive samples that were classified as negative. The value of recall ranges from 0 to 1; a higher recall means that the model misses fewer positive examples.
[0156] The F1 score is the harmonic mean of precision and recall, calculated as F1 = 2 × precision × recall / (precision + recall). The F1 score ranges from 0 to 1; a higher F1 score indicates a better balance between precision and recall.
[0157] Pearson correlation coefficient: In statistics, the Pearson correlation coefficient, also known as the Pearson product-moment correlation coefficient (PPMCC or PCCs), is used to measure the correlation (linear correlation) between two variables X and Y. Its value ranges from -1 to 1. A coefficient of 1 means that X and Y can be well described by a linear equation, all data points lie on a straight line, and Y increases as X increases. A coefficient of -1 means that all data points lie on a straight line, and Y decreases as X increases. A coefficient of 0 means that there is no linear relationship between the two variables.
[0158] To better demonstrate the predictive accuracy of the proposed neural network-based protein solubility prediction method, the following comparative experimental results are provided, comparing the proposed method with other protein solubility prediction methods. It should be noted that although the PSI:Biology dataset is preferably used as the training set for the proposed method, for fair comparison, NetSolP will be used as the training set for five-fold cross-validation to compare its performance with other methods in the following experiments, while the PSI:Biology dataset will be used as the test set. Furthermore, the labels were balanced in the experiments, and each training and test dataset was designed to ensure that sequences with a global sequence identity >25% were not shared. Global sequence identity was determined using ggsearch36.
[0159] In addition, to better verify the prediction accuracy of the method for predicting protein solubility using neural networks proposed in this disclosure, two new methods are established: Method A is based on Method 100 (Method 800) proposed in this disclosure, but removes the second graph neural network part, that is, only the first graph neural network is used to process the encoding feature vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs, and only the first comprehensive feature vector is input into the prediction sub-network for protein solubility prediction. Method B is based on Method A, but replaces the encoding feature vectors of multiple amino acid residues processed by the first graph neural network with the fusion result of the encoding feature vectors of multiple amino acid residues and the replacement score vector, and remains unchanged.
[0160] The performance of different methods on this test set is shown in Table 1 below: Table 1. Performance of different methods on the PSI:Biology dataset
[0161] As shown in Table 1, the method proposed in this disclosure significantly outperforms other methods in accuracy (ACC), classification ability (AUC), Matthews correlation coefficient (MCC), F1 score, and recall on the PSI:Biology dataset. Since SWI, SoluProt, NetSolP, and DeepSoluE all use models trained on the PSI:Biology dataset, they are not compared here.
[0162] To further evaluate the differences between the proposed method and other methods, three independent test sets—NESG, S. Cerevisiae, and eSol—were selected to compare the performance of the proposed method with other methods. It should be noted that for most models, the value recommended by the model authors was used as the threshold for determining protein solubility. Since Camsol is not a binary classification model, we used 1 as the threshold, and 0.5 was used for all other methods. The comparison results are as follows: Table 2 Performance of different methods on the independent test set NESG
[0163] The results show that the method proposed in this disclosure significantly outperforms other methods in accuracy (ACC), classification ability (AUC), Matthews correlation coefficient (MCC), F1 score, and recall on the independent test set NESG. DeepSoluE's predictions on the NESG dataset contained 7 NaN values.
[0164] Table 3. Performance of different methods on the independent test set S. Cerevisiae
[0165] The results show that the method proposed in this disclosure significantly outperforms other methods in accuracy (ACC), classification ability (AUC), F1 score, and Pearson correlation coefficient (R coefficient) on the independent test set *S. cerevisiae*. It is worth noting that the proteins in the *S. cerevisiae* dataset are expressed in *Saccharomyces cerevisiae*, while other datasets, including the training set, are expressed in *E. coli*. This demonstrates that the method proposed in this disclosure still exhibits good generalization performance across expression systems.
[0166] Table 4 Performance of different methods on the independent test set eSol
[0167] The results show that the method proposed in this disclosure significantly outperforms other methods in accuracy (ACC), classification ability (AUC), Matthews correlation coefficient (MCC), F1 score, and Pearson correlation coefficient (R coefficient) on the independent test set eSol, indicating that the method proposed in this disclosure can simultaneously address classification and regression problems. Protein-Sol, GraphSol, and HybridGCN all use the eSol dataset as their training set, and therefore are not compared here.
[0168] The above comparison results further demonstrate that, within the framework of graph neural networks, combining the encoded feature vector obtained by encoding the target protein sequence using a protein pre-trained model with the contact probabilities of multiple amino acid residue pairs in the target protein sequence, and combining the substitution score vectors of each amino acid residue in the target protein sequence with the contact probabilities of multiple amino acid residue pairs, allows the neural network to utilize richer protein information, thereby improving the diversity and comprehensiveness of feature representation. Furthermore, it enables the prediction of protein solubility without acquiring or predicting the spatial structure information of the target protein sequence. By setting up two graph neural networks to process the above two combinations separately, the information related to protein solubility in the encoded feature vectors and substitution score vectors obtained through different methods can be fully extracted, and interference between the two can be avoided. This improves the predictive power of the obtained first and second comprehensive feature vectors, ultimately resulting in a more accurate prediction of the solubility of the target protein sequence.
[0169] According to another aspect of this disclosure, an apparatus for predicting protein solubility using a neural network is provided. Figure 9 A structural block diagram of an apparatus for predicting protein solubility using a neural network according to an exemplary embodiment of the present disclosure is shown. Figure 9As shown, the device 900 includes: a first acquisition unit 910 configured to acquire a target protein sequence, the target protein sequence including multiple amino acid residues; and a first encoding unit 920 configured to encode the target protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the multiple amino acid residues. A first determining unit 930 is configured to determine the substitution score vectors of multiple amino acid residues, the substitution score vectors including a tendency score indicating that the corresponding amino acid residue is replaced by multiple preset amino acids at the corresponding amino acid residue position; a second determining unit 940 is configured to determine the contact probability of multiple amino acid residue pairs in the target protein sequence; a first processing unit 950 is configured to process the encoding feature vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs using a first graph neural network to obtain a first comprehensive feature vector of multiple amino acid residues; a second processing unit 960 is configured to process the substitution score vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs using a second graph neural network to obtain a second comprehensive feature vector of multiple amino acid residues; and a first prediction unit 970 is configured to input the first comprehensive feature vector and the second comprehensive feature vector of multiple amino acid residues into a prediction subnetwork to obtain a prediction result of the solubility of the target protein sequence.
[0170] It is understood that the operation of units 910-970 in device 900 can refer to the description of steps S101-S107 in method 100 above, and will not be repeated here.
[0171] According to another aspect of this disclosure, a neural network training apparatus is provided. Figure 10 A structural block diagram of a neural network training apparatus according to an exemplary embodiment of the present disclosure is shown. Figure 10As shown, the device 1000 includes: a second acquisition unit 1010 configured to acquire a sample protein sequence and the true solubility of the sample protein sequence, the sample protein sequence comprising multiple amino acid residues; a second encoding unit 1020 configured to encode the sample protein sequence using a protein pre-training model to obtain encoding feature vectors for each of the multiple amino acid residues; a third determination unit 1030 configured to determine substitution score vectors for each of the multiple amino acid residues, the substitution score vectors including a tendency score indicating that the corresponding amino acid residue is replaced by multiple preset amino acids at the position of the corresponding amino acid residue; a fourth determination unit 1040 configured to determine the contact probability of multiple amino acid residue pairs in the sample protein sequence; and a third processing unit 1050 configured to utilize a first initial graph neural network. The network processes the encoding feature vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs to obtain the first comprehensive feature vectors of multiple amino acid residues; the fourth processing unit 1060 is configured to process the substitution score vectors of multiple amino acid residues and the contact probabilities of multiple amino acid residue pairs using the second initial graph neural network to obtain the second comprehensive feature vectors of multiple amino acid residues; the second prediction unit 1007 is configured to input the first and second comprehensive feature vectors of multiple amino acid residues into the initial prediction sub-network to obtain the prediction result of the protein solubility of the sample protein sequence; and the parameter tuning unit 1080 is configured to adjust the parameters of the initial neural network based on the prediction result and the actual solubility to obtain the trained neural network.
[0172] It is understood that the operation of units 1010-1080 in device 1000 can refer to the description of steps S801-S808 in method 800 above, and will not be repeated here.
[0173] According to one aspect of this disclosure, a computer device is provided, including a memory, a processor, and a computer program stored in the memory. The processor is configured to execute the computer program to implement the steps of any of the method embodiments described above.
[0174] According to one aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the method embodiments described above.
[0175] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of any of the method embodiments described above.
[0176] In the following text, combined with Figure 11Illustrative examples describing such computer devices, non-transitory computer-readable storage media, and computer program products.
[0177] Figure 11 An example configuration of computer device 1100 that can be used to implement the methods described herein is shown. The aforementioned devices 900 and 1000 may also be implemented wholly or at least partially by computer device 1100 or similar devices or systems.
[0178] Computer device 1100 can be a variety of different types of devices. Examples of computer device 1100 include, but are not limited to: desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablets, cellular or other wireless phones (e.g., smartphones), notebook computers, mobile stations), wearable devices (e.g., glasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to a display device, game consoles), televisions or other display devices, automotive computers, and so on.
[0179] Computer device 1100 may include at least one processor 1102, memory 1104, multiple communication interfaces 1106, display device 1108, other input / output (I / O) devices 1110, and one or more mass storage devices 1112 capable of communicating with each other, such as via system bus 1114 or other suitable connections.
[0180] Processor 1102 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 1102 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Among other capabilities, processor 1102 may be configured to acquire and execute computer-readable instructions stored in memory 1104, mass storage device 1112, or other computer-readable media, such as program code of operating system 1116, program code of application program 1118, program code of other program 1120, etc.
[0181] Memory 1104 and mass storage device 1112 are examples of computer-readable storage media for storing instructions executed by processor 1102 to perform the various functions described above. For example, memory 1104 may generally include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, mass storage device 1112 may generally include hard disk drives, solid-state drives, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Both memory 1104 and mass storage device 1112 may be collectively referred to herein as memory or computer-readable storage media, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code, which may be executed by processor 1102 as a specific machine configured to perform the operations and functions described in the examples herein.
[0182] Multiple programs can be stored on mass storage device 1112. These programs include operating system 1116, one or more applications 1118, other programs 1120, and program data 1122, and they can be loaded into memory 1104 for execution.
[0183] Although Figure 11 The modules 1116, 1118, 1120, and 1122, or portions thereof, are illustrated as being stored in memory 1104 of computer device 1100, but modules 1116, 1118, 1120, and 1122, or portions thereof, may be implemented using any form of computer-readable medium accessible by computer device 1100. As used herein, “computer-readable medium” includes at least two types of computer-readable media: computer-readable storage media and communication media.
[0184] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD, or other optical storage devices, magnetic cassettes, magnetic tapes, disk storage devices or other magnetic storage devices, or any other non-transmission medium that can be used to store information for access by a computer device. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms. Computer-readable storage media as defined herein do not include communication media.
[0185] One or more communication interfaces 1106 are used for exchanging data with other devices, such as via a network, direct connection, etc. Such communication interfaces can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), wired or wireless (such as an IEEE 802.11 wireless LAN (WLAN)) interface, Wi-MAX interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth. TM Interfaces, near field communication (NFC) interfaces, etc. Communication interface 1106 can facilitate communication across various network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, etc. Communication interface 1106 can also provide communication with external storage devices (not shown) such as storage arrays, network-attached storage, storage area networks, etc.
[0186] In some examples, a display device 1108, such as a monitor, may be included for displaying information and images to the user. Other I / O devices 1110 may be devices that receive various inputs from the user and provide various outputs to the user, and may include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, and so on.
[0187] The technologies described herein can be supported by these various configurations of computer device 1100, and are not limited to specific examples of the technologies described herein. For example, the functionality can also be implemented wholly or partially on a “cloud” using a distributed system. A cloud includes and / or represents a platform for resources. The platform abstracts the underlying functionality of the cloud’s hardware (e.g., servers) and software resources. Resources may include applications and / or data that can be used when performing computational processing on a server remote from computer device 1100. Resources may also include services provided via the Internet and / or via subscriber networks such as cellular or Wi-Fi networks. The platform can abstract resources and functionality to connect computer device 1100 to other computer devices. Therefore, the implementation of the functionality described herein can be distributed throughout the cloud. For example, the functionality may be implemented partly on computer device 1100 and partly through a platform that abstracts the functionality of the cloud.
[0188] Although this disclosure has been described and illustrated in detail in the accompanying drawings and the foregoing description, such description and illustration should be considered illustrative and suggestive, not restrictive; this disclosure is not limited to the disclosed embodiments. By studying the drawings, the disclosure, and the appended claims, those skilled in the art will be able to understand and implement variations of the disclosed embodiments in practice with respect to the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps not listed, the indefinite article "a" or "an" does not exclude a plurality, the term "a plurality" means two or more, and the term "based on" should be interpreted as "at least partially based on". The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be beneficial.
Claims
1. A method for predicting protein solubility using a neural network, wherein, The trained neural network includes a first graph neural network, a second graph neural network, and a prediction subnetwork, and the method includes: Obtain the target protein sequence, wherein the target protein sequence comprises multiple amino acid residues; The target protein sequence is encoded using a protein pre-training model to obtain the encoding feature vectors of each of the multiple amino acid residues; Determine the substitution score vector for each of the plurality of amino acid residues, wherein the substitution score vector includes a tendency score indicating that the corresponding amino acid residue is replaced at the position of the corresponding amino acid residue with a plurality of preset amino acids; Determine the contact probability of multiple amino acid residue pairs in the target protein sequence; The first graph neural network is used to process the encoding feature vectors of each of the multiple amino acid residues and the contact probabilities of the multiple amino acid residue pairs to obtain the first comprehensive feature vector of each of the multiple amino acid residues. The substitution score vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs are processed using the second graph neural network to obtain the second comprehensive feature vectors of each of the plurality of amino acid residues; and The first and second integrated feature vectors of each of the multiple amino acid residues are input into the prediction sub-network to obtain the prediction result of the solubility of the target protein sequence.
2. The method according to claim 1, wherein, The first graph neural network is used to process the encoding feature vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs to obtain the first comprehensive feature vector of each of the plurality of amino acid residues, including: The encoding feature vectors of each of the multiple amino acid residues are used as node features, and the contact probabilities of the multiple amino acid residue pairs are used as edge features to construct the first graph data; and Using the first graph neural network, the node features in the first graph data are updated based on the node features and edge features in the first graph data to obtain the first comprehensive feature vector of each of the multiple amino acid residues.
3. The method according to claim 2, wherein, The first graph neural network includes three first graph convolutional layers. Using the first graph neural network, the node features in the first graph data are updated based on the node features and edge features in the first graph data to obtain the first comprehensive feature vectors for each of the multiple amino acid residues, including: The node features in the first graph data are updated iteratively using the three first graph convolutional layers, and the first intermediate feature vectors of the plurality of amino acid residues output by the three first graph convolutional layers are obtained respectively; and The first intermediate feature vectors of each of the multiple amino acid residues output by the three first graph convolutional layers are fused to obtain the first comprehensive feature vector of each of the multiple amino acid residues.
4. The method according to claim 3, wherein, The fusion includes merging the first intermediate feature vectors of the multiple amino acid residues output by the three first graph convolutional layers using a residual network structure.
5. The method according to claim 1, wherein, The substitution score vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs are processed using the second graph neural network to obtain the second comprehensive feature vectors of each of the plurality of amino acid residues, including: The substitution score vectors of each of the multiple amino acid residues are used as node features, and the contact probabilities of the multiple amino acid residue pairs are used as edge features to construct the second graph data; and Using the second graph neural network, the node features in the second graph data are updated based on the node features and edge features in the second graph data to obtain the second comprehensive feature vector of each of the multiple amino acid residues.
6. The method according to claim 5, wherein, The second graph neural network includes two second graph convolutional layers. Using the second graph neural network, the node features in the second graph data are updated based on the node features and edge features in the second graph data to obtain the second comprehensive feature vectors for each of the multiple amino acid residues, including: The node features in the second graph data are updated iteratively using the two second graph convolutional layers, and the second intermediate feature vectors of the multiple amino acid residues output by the two second graph convolutional layers are obtained respectively; and The second intermediate feature vectors of the multiple amino acid residues output by the two second graph convolutional layers are fused to obtain the second comprehensive feature vector of each of the multiple amino acid residues.
7. The method according to claim 6, wherein, The fusion includes merging the second intermediate feature vectors of the multiple amino acid residues output by the two second graph convolutional layers using a residual network structure.
8. The method according to claim 1, wherein, The first and second integrated feature vectors of each of the multiple amino acid residues are input into the prediction sub-network to obtain the prediction results of the solubility of the target protein sequence, including: The first integrated feature vectors of each of the multiple amino acid residues are pooled to obtain the first pooled feature vector of the target protein sequence; The second integrated feature vectors of each of the multiple amino acid residues are pooled to obtain the second pooled feature vector of the target protein sequence; The first pooling feature vector and the second pooling feature vector are fused to obtain the target feature vector of the target protein sequence; and The target feature vector is processed using a fully connected layer and an output layer to obtain the prediction result.
9. The method according to claim 8, wherein, The fusion includes merging the first pooling feature vector and the second pooling feature vector using a residual network structure.
10. The method according to any one of claims 1-9, wherein, The protein pre-training model is the ProtT5-XL model.
11. The method according to any one of claims 1-9, wherein, The first graph neural network and the second graph neural network are based on GraphSAGE.
12. The method according to any one of claims 1-9, wherein, The plurality of amino acid residue pairs include amino acid residue pairs consisting of a first amino acid residue and a second amino acid residue selected from the target protein sequence.
13. The method according to any one of claims 1-9, wherein, The trained neural network is obtained through end-to-end training using sample protein sequences and the actual solubility of the sample protein sequences.
14. A neural network training method, wherein, The initial neural network includes a first initial graph neural network, a second initial graph neural network, and an initial prediction subnetwork, and the method includes: Obtain the sample protein sequence and the true solubility of the sample protein sequence, wherein the sample protein sequence comprises multiple amino acid residues; The protein sequence of the sample is encoded using a protein pre-training model to obtain the encoding feature vectors of each of the multiple amino acid residues; Determine the substitution score vector for each of the plurality of amino acid residues, wherein the substitution score vector includes a tendency score indicating that the corresponding amino acid residue is replaced at the position of the corresponding amino acid residue with a plurality of preset amino acids; Determine the contact probability of multiple amino acid residue pairs in the protein sequence of the sample; The first initial graph neural network is used to process the encoding feature vectors of each of the multiple amino acid residues and the contact probabilities of the multiple amino acid residue pairs to obtain the first comprehensive feature vector of each of the multiple amino acid residues. The replacement score vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs are processed using the second initial graph neural network to obtain the second comprehensive feature vectors of each of the plurality of amino acid residues. The first and second integrated feature vectors of each of the multiple amino acid residues are input into the initial prediction sub-network to obtain the prediction result of the protein solubility of the sample protein sequence; and Based on the prediction results and the actual solubility, the parameters of the initial neural network are adjusted to obtain a trained neural network.
15. A device for predicting protein solubility using a neural network, wherein, The trained neural network includes a first graph neural network, a second graph neural network, and a prediction subnetwork, and the device includes: The first acquisition unit is configured to acquire a target protein sequence, the target protein sequence comprising a plurality of amino acid residues; The first coding unit is configured to encode the target protein sequence using a protein pre-training model to obtain the coding feature vectors of each of the plurality of amino acid residues; The first determining unit is configured to determine the substitution score vector for each of the plurality of amino acid residues, the substitution score vector including a tendency score indicating that the corresponding amino acid residue is replaced by a plurality of preset amino acids at the position of the corresponding amino acid residue; The second determining unit is configured to determine the contact probability of multiple amino acid residue pairs in the target protein sequence; The first processing unit is configured to process the encoded feature vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs using the first graph neural network to obtain the first comprehensive feature vector of each of the plurality of amino acid residues. The second processing unit is configured to process the substitution score vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs using the second graph neural network to obtain a second comprehensive feature vector for each of the plurality of amino acid residues; and The first prediction unit is configured to input the first and second integrated feature vectors of each of the plurality of amino acid residues into the prediction sub-network to obtain a prediction result of the solubility of the target protein sequence.
16. A neural network training device, wherein, The initial neural network includes a first initial graph neural network, a second initial graph neural network, and an initial prediction subnetwork. The device includes: The second acquisition unit is configured to acquire a sample protein sequence and the true solubility of the sample protein sequence, the sample protein sequence comprising multiple amino acid residues; The second coding unit is configured to encode the sample protein sequence using a protein pre-training model to obtain the coding feature vectors of each of the multiple amino acid residues. The third determining unit is configured to determine the substitution score vector for each of the plurality of amino acid residues, the substitution score vector including a tendency score indicating that the corresponding amino acid residue is replaced by a plurality of preset amino acids at the position of the corresponding amino acid residue; The fourth determining unit is configured to determine the contact probability of multiple amino acid residue pairs in the sample protein sequence; The third processing unit is configured to process the encoded feature vectors of each of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs using the first initial graph neural network to obtain the first comprehensive feature vector of each of the plurality of amino acid residues. The fourth processing unit is configured to process the substitution score vectors of the plurality of amino acid residues and the contact probabilities of the plurality of amino acid residue pairs using the second initial graph neural network to obtain the second comprehensive feature vectors of the plurality of amino acid residues. The second prediction unit is configured to input the first and second integrated feature vectors of each of the plurality of amino acid residues into the initial prediction sub-network to obtain a prediction result of the protein solubility of the sample protein sequence; and The parameter tuning unit is configured to adjust the parameters of the initial neural network based on the prediction results and the actual solubility to obtain a trained neural network.
17. A computer device, comprising: At least one processor; as well as The memory, on which computer programs are stored, When the computer program is executed by the processor, it causes the processor to perform the method described in any one of claims 1-14.
18. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to perform the method of any one of claims 1-14.
19. A computer program product comprising a computer program that, when executed by a processor, causes the processor to perform the method of any one of claims 1-14.
20. A trained neural network obtained by the neural network training method of claim 14.
Citation Information
Patent Citations
Codon optimization
CN112513989A
Protein phosphorylation site prediction method based on inner product self-attention neural network
CN113096722A