A method, system, device, and medium for predicting protein interactions
By combining the three-dimensional structure and amino acid sequence information of the protein, a comprehensive feature vector is generated for prediction, which solves the problem that existing methods cannot effectively utilize the three-dimensional structure information and improves the accuracy of protein interaction prediction.
Patent Information
- Application Number
- CN202210967699.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-08-12
AI Technical Summary
Existing protein interaction prediction methods cannot fully utilize the three-dimensional structural information of proteins, and sequence-based methods cannot effectively extract interaction characteristics, resulting in insufficient prediction accuracy.
By obtaining the three-dimensional structure and amino acid sequence of the protein, the three-dimensional structure is scaled and a shape vector is generated using voxel representation, and the characteristic vector is generated by combining the amino acid sequence to form a comprehensive feature vector for prediction.
This method can make full use of the multi-dimensional information of proteins to improve the accuracy and coverage of protein interaction predictions.
Smart Images

Figure CN115410644B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of proteins, and in particular to a method, system, device, and storage medium for predicting protein-protein interactions. Background Art
[0002] Proteins are biological macromolecules composed of one or more long chains of α-amino acid residues, and have four levels of structure: primary structure, secondary structure, tertiary structure, and quaternary structure. Among them, the tertiary structure and quaternary structure have complex spatial structures and are functional. When a protein performs its biochemical function, it always forms a protein complex through non-covalent bonds. The interactions (protein-protein interaction, PPI) between these different proteins constitute a major part of the cell biochemical reaction network and play a very important role in most biochemical functions.
[0003] According to the characteristics input into the model, PPI prediction methods are mainly divided into two categories: sequence-based and protein three-dimensional structure-based. Among them, the sequence-based methods mainly include: Conjoint triads (CT) and Auto-Covariance (AC) methods.
[0004] However, the techniques for identifying protein interactions based on traditional experiments have the disadvantages of being time-consuming, having limited coverage, and being expensive. In recent years, researchers have developed some methods for identifying protein interactions using CNN or LSTM and protein amino acid sequences. However, these methods generally have some disadvantages: the vectorized encoding method of protein amino acid sequences cannot fully extract interaction features, ignores the complementary information between multiple amino acid sequence encodings and classifiers, and at the same time there is also a problem that the three-dimensional structure information is not fully utilized for prediction. Summary of the Invention
[0005] In view of this, in order to overcome at least one aspect of the above problems, an embodiment of the present invention provides a method for predicting protein-protein interactions, including the following steps:
[0006] S1, obtaining the three-dimensional structure and amino acid sequence of the protein to be predicted;
[0007] S2, placing the three-dimensional structure of the protein in a three-dimensional space with a voxel of l*l*l, wherein the centroid of the three-dimensional structure of the protein coincides with the origin of the three-dimensional space, and scaling all coordinates of the three-dimensional structure of the protein;
[0008] S3. Obtain an \(l\times l\times l\) shape vector according to whether the three-dimensional structure of the protein passes through the voxels in the three-dimensional space, where if the three-dimensional structure of the protein passes through a voxel, the value is 1, otherwise it is 0; and obtain an \(l\times l\times l\) first feature vector and an \(l\times l\times l\) second feature vector according to the position, hydrophobicity, and side chain net charge index of each amino acid in the three-dimensional structure of the protein in the three-dimensional space;
[0009] S4. Generate a third feature vector using the amino acid sequence;
[0010] S5. Concatenate the shape vector, the first feature vector, the second feature vector, and the third feature vector to obtain a final feature vector;
[0011] S6. Input the final feature vector into the trained prediction model for prediction to obtain the interaction of the protein to be predicted.
[0012] In some embodiments, generating a third feature vector using the amino acid sequence further includes obtaining the value of each element in the vector based on the following formula:
[0013]
[0014] where \(P\) represents the amino acid sequence of the protein, with a length of \(N\) r , \(m\) is the number of the \(m\)th amino acid in the sequence, \(n\) is the number of the physicochemical property, \(lag\) is the value of the lag, and \(P\) m,n is the physicochemical property with the number \(n\) of the \(m\)th amino acid, thereby converting the amino acid sequence into a vector of \((n\times lag)\) dimensions.
[0015] In some embodiments, scaling all the coordinates of the three-dimensional structure of the protein further includes:
[0016] Using to scale all the coordinates of the three-dimensional structure of the protein, where the \(R\) max is the radius of the largest inscribed sphere in the three-dimensional space.
[0017] In some embodiments, it further includes:
[0018] Construct a prediction model using a connection layer, a linear layer, \(H\) Transformer modules, and a normalization layer;
[0019] Construct a training set, and perform the processing of steps S2 - S5 on each sample in the training set to obtain the final feature vector corresponding to each sample;
[0020] Train the prediction model using the final feature vector corresponding to each sample.
[0021] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a prediction system for protein interaction, including:
[0022] An acquisition module configured to acquire the three-dimensional structure and amino acid sequence of the protein to be predicted;
[0023] A scaling module configured to place the three-dimensional structure of the protein in a three-dimensional space with a voxel of l*l*l, wherein the centroid of the three-dimensional structure of the protein coincides with the origin of the three-dimensional space, and scale all coordinates of the three-dimensional structure of the protein;
[0024] A first generation module configured to obtain an l*l*l shape vector according to whether the three-dimensional structure of the protein passes through the voxels in the three-dimensional space, wherein if the three-dimensional structure of the protein passes through the voxel, the value is 1, otherwise it is 0; and obtain an l*l*l first characteristic vector and an l*l*l second characteristic vector according to the position, hydrophobicity and side chain net charge index of each amino acid in the three-dimensional structure of the protein in the three-dimensional space;
[0025] A second generation module configured to generate a third feature vector using the amino acid sequence;
[0026] A connection module configured to connect the shape vector, the first feature vector, the second feature vector and the third feature vector to obtain a final feature vector;
[0027] A prediction module configured to input the final feature vector into a trained prediction model for prediction to obtain the interaction of the protein to be predicted.
[0028] In some embodiments, the second generation module is further configured to:
[0029]
[0030] Where P represents the amino acid sequence of the protein, with a length of N r , m is the number of the m-th amino acid in the sequence, n is the number of the physicochemical property, lag is the value of the lag, and P m,n is the physicochemical property of the m-th amino acid with the number n, thereby converting the amino acid sequence into a vector of (n×lag) dimensions.
[0031] In some embodiments, the scaling module is further configured to:
[0032] Use to scale all coordinates of the three-dimensional structure of the protein, wherein the R maxis the radius of the largest inscribed sphere in the three-dimensional space.
[0033] In some embodiments, it further includes a neural network module configured to:
[0034] Construct a prediction model using a connection layer, a linear layer, H transformer modules, and a normalization layer;
[0035] Construct a training set, and process each sample in the training set through steps S2 - S5 to obtain the final feature vector corresponding to each sample;
[0036] Train the prediction model using the final feature vector corresponding to each sample.
[0037] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a computer device, including:
[0038] At least one processor; and
[0039] A memory storing a computer program that can run on the processor, wherein when the processor executes the program, it executes the steps of any one of the protein interaction prediction methods described above.
[0040] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it executes the steps of any one of the protein interaction prediction methods described above.
[0041] One of the beneficial technical effects of the present invention is as follows: The solution proposed by the present invention can make full use of information in different dimensions of proteins, namely one-dimensional sequence information and three-dimensional structure information, for a model of predicting protein interactions. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.
[0043] Figure 1 It is a flowchart showing the protein interaction prediction method provided by the embodiment of the present invention;
[0044] Figure 2 The R provided by the embodiment of the present invention maxSchematic diagram of the relationship between and V;
[0045] Figure 3 Schematic diagram of the 20 amino acids provided in the embodiments of the present invention divided into five major categories;
[0046] Figure 4 Schematic diagram of the six physicochemical properties of the 20 amino acids provided in the embodiments of the present invention;
[0047] Figure 5 Schematic diagram of the feature extraction process provided in the embodiments of the present invention;
[0048] Figure 6 Schematic diagram of the prediction model provided in the embodiments of the present invention;
[0049] Figure 7 Schematic diagram of the structure of the prediction system for protein-protein interaction provided in the embodiments of the present invention;
[0050] Figure 8 Schematic diagram of the structure of the computer device provided in the embodiments of the present invention;
[0051] Figure 9 Schematic diagram of the structure of the computer-readable storage medium provided in the embodiments of the present invention. Detailed implementation manners
[0052] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention in detail with reference to specific embodiments and the accompanying drawings.
[0053] It should be noted that all the expressions using "first" and "second" in the embodiments of the present invention are used to distinguish two entities or parameters with the same name but different identities. It can be seen that "first" and "second" are only for the convenience of expression and should not be construed as limitations on the embodiments of the present invention. This will not be elaborated one by one in the subsequent embodiments.
[0054] According to one aspect of the present invention, embodiments of the present invention propose a method for predicting protein-protein interaction, as Figure 1 shown, which may include the steps:
[0055] S1, obtaining the three-dimensional structure and amino acid sequence of the protein to be predicted;
[0056] S2, placing the three-dimensional structure of the protein in a three-dimensional space with a voxel of l*l*l, wherein the centroid of the three-dimensional structure of the protein coincides with the origin of the three-dimensional space, and scaling all the coordinates of the three-dimensional structure of the protein;
[0057] S3. Obtain an \(l\times l\times l\) shape vector based on whether the three-dimensional structure of the protein passes through the voxels in the three-dimensional space, where if the three-dimensional structure of the protein passes through a voxel, the value is 1, otherwise it is 0; and obtain an \(l\times l\times l\) first feature vector and an \(l\times l\times l\) second feature vector based on the position, hydrophobicity, and side chain net charge index of each amino acid in the three-dimensional structure of the protein in the three-dimensional space;
[0058] S4. Generate a third feature vector using the amino acid sequence;
[0059] S5. Concatenate the shape vector, the first feature vector, the second feature vector, and the third feature vector to obtain a final feature vector;
[0060] S6. Input the final feature vector into the trained prediction model for prediction to obtain the interaction of the protein to be predicted.
[0061] The solution proposed by the present invention can make full use of information in different dimensions of proteins: a model for predicting protein interactions using one-dimensional sequence information and three-dimensional structure information.
[0062] In some embodiments, further comprising scaling all coordinates of the three-dimensional structure of the protein:
[0063] Use to scale all coordinates of the three-dimensional structure of the protein, where the \(R\) max is the radius of the largest inscribed sphere in the three-dimensional space.
[0064] Specifically, for three-dimensional structure information, perform the following operations on the three-dimensional coordinates of the protein structure: Move the centroid of the protein structure to the origin (the center of the cube \(V\)), scale all coordinate values, and the multiplier
[0065] where in this work \(R\) max = 40, \(l = 32\), and the relationship between \(R\) max and \(V\) is as Figure 2 shown.
[0066] In some embodiments, in step S3, an l×l×l shape vector is obtained according to whether the three-dimensional structure of the protein passes through the voxels in the three-dimensional space, where if the three-dimensional structure of the protein passes through a voxel, the value is 1, otherwise it is 0; and a first l×l×l characteristic vector and a second l×l×l characteristic vector are obtained according to the position, hydrophobicity, and side-chain net charge index of each amino acid in the three-dimensional structure of the protein in the three-dimensional space. Specifically, using volumetric representation, the three-dimensional coordinates of the protein are represented as voxel representation (grid density is l×l×l) in a fixed three-dimensional space. If the backbone of the protein passes through a voxel, the value is 1, otherwise the value is 0. However, the voxel representation only includes the shape in the three-dimensional structure of the protein. In addition to the shape of the three-dimensional structure of the protein, the hydrophobicity and side-chain net charge index of the amino acids in the protein also fundamentally affect the interaction between proteins. Therefore, this work will also extract this part of the properties and use them as inputs to the model at the same time. Similarly using the above volumetric representation, the hydrophobicity and side-chain net charge index of the amino acids in the protein are respectively represented by voxels in the three-dimensional cubic space V.
[0067] In this way, the three-dimensional structure information generates three l×l×l vectors through volumetric representation.
[0068] In some embodiments, generating a third feature vector using the amino acid sequence further includes obtaining the value of each element in the vector based on the following formula:
[0069]
[0070] where P represents the amino acid sequence of the protein, with a length of N r , m is the number of the m-th amino acid in the sequence, n is the number of the physicochemical property, lag is the value of the lag, and P m,n is the physicochemical property with the number n of the m-th amino acid, thereby converting the amino acid sequence into a vector with a dimension of (n×lag).
[0071] Specifically, as Figure 3 shown, there are 20 kinds of amino acids that make up proteins in the human body, which can be divided into 5 categories according to the properties of the side-chain R groups: amino acids with positively charged R groups, amino acids with negatively charged R groups, amino acids with polar uncharged R groups, amino acids with non-polar hydrophobic R groups in the side chain, and three special categories of amino acids. As Figure 4As shown, amino acids have six physicochemical properties, where H is hydrophobicity, VSC is the volume of side chains, P1 is polarity, P2 is polarizability, SASA is solvent accessible surface area, and NCISC is the net charge index of side chains.
[0072] The amino acid sequence information of a protein is represented as a vector of a certain dimension using the above formula. In this work, n = 6 and lag = 30. Thus, when calculating AC in the above formula lag,n n takes values from 1 to 6, and lag takes values from 1 to 30. That is, when n = 1, lag takes values from 1 to 30; when n = 2, lag takes values from 1 to 30, and so on, thereby obtaining a third feature vector with a dimension of 6 * 30.
[0073] In some embodiments, it further includes:
[0074] Construct a prediction model using a connection layer, a linear layer, H transformer modules, and a normalization layer;
[0075] Construct a training set, and process each sample in the training set through steps S2 - S5 to obtain the final feature vector corresponding to each sample;
[0076] Train the prediction model using the final feature vector corresponding to each sample.
[0077] Specifically, the PPI dataset can be selected to perform data feature extraction, extracting one-dimensional sequence and three-dimensional structure information.
[0078] PPI dataset: Pan’s PPI dataset
Pan, X.-Y., Zhang, Y.-N. & Shen, H.-B. Large-scale prediction of human protein–protein interactions from amino acid sequence based on latent topic features. J. Proteome Res. 9, 4992–5001 (2010).
[0079] The positive samples are from the June 2007 version of the HPRD dataset (human protein references database), which contains 38,788 experimentally verified protein-protein interaction pairs from 9,630 different human proteins. After removing self-interactions and duplicate data, there are finally 36,630 positive samples. The negative samples contain 36,480 non-interacting protein-protein pairs. Considering that not all sequences have corresponding protein three-dimensional structure data in the PDB dataset, the final dataset contains 25,493 protein-protein pairs, including 18,025 positive samples and 7,468 negative samples.
[0080] Protein three-dimensional structure dataset: Protein three-dimensional structure information is obtained from the.pdb files provided on the RCSB Protein Data Bank (PDB, http: / / www.rcsb.org / pdb / ) website. The original files contain the atoms in the protein and their corresponding three-dimensional coordinates.
[0081] Since this is a binary classification problem, the output of the model can be classified into the following four categories:
[0082] ● FN: False Negative, which is judged as a negative sample but is actually a positive sample.
[0083] ● FP: False Positive, which is judged as a positive sample but is actually a negative sample.
[0084] ● TN: True Negative, which is judged as a negative sample and is actually a negative sample.
[0085] ● TP: True Positive, which is judged as a positive sample and is actually a positive sample.
[0086] Therefore, accuracy, precision, recall, etc. can be used as evaluation criteria for model performance.
[0087] accuracy = (TP + TN) / (P + N): Accuracy, that is, the number of truly correct and truly incorrect in the retrieval results divided by the total number of samples;
[0088] sensitive = TP / P, Sensitivity, which represents the proportion of correctly classified positive examples among all positive examples, measuring the classifier's ability to identify positive examples;
[0089] specificity = TN / N, Specificity, which represents the proportion of correctly classified negative examples among all negative examples, measuring the classifier's ability to identify negative examples;
[0090] Precision = TP / (TP + FP): Precision, that is, among the results returned after retrieval, the proportion of the truly correct ones in the entire result;
[0091] Recall = TP / (TP + FN): Recall rate, that is, among the retrieval results, the proportion of the truly correct ones in the entire dataset (retrieved and not retrieved) that are truly correct.
[0092] As Figure 5 shown, when performing data feature extraction, since the prediction model uses one-dimensional sequence and three-dimensional structure information as input, the feature extraction includes two major steps: the three-dimensional structure information generates three l*l*l vector representations through volumetric representation, and the one-dimensional amino acid sequence information generates a 6*30 vector representation through the autocovariance method. Both are subjected to feature extraction through the pre-trained ResNet50, and finally a feature of length 2048 is obtained.
[0093] As Figure 6 shown, the present invention constructs a transformer prediction model based on the self-attention mechanism, uses the sequence information feature and three-dimensional structure feature obtained in the previous step as the input of the model, now concatenates the two together, then passes through a linear layer and H transformer modules, and finally predicts the label of PPI through a sigmoid layer. If the value is greater than 0.5, the label is positive (meaning there is an interaction between proteins), and if the value is less than 0.5, the label is negative (that is, there is no interaction between proteins).
[0094] The solution proposed by the present invention can simultaneously use one-dimensional sequence information and three-dimensional structure information to predict protein-protein interactions, utilizes more dimensional information, and can improve the accuracy of PPI prediction.
[0095] Based on the same inventive concept, according to another aspect of the present invention, an embodiment of the present invention also provides a protein-protein interaction prediction system 400, as Figure 7 shown, including:
[0096] An acquisition module 401, configured to acquire the three-dimensional structure and amino acid sequence of the protein to be predicted;
[0097] A scaling module 402, configured to place the three-dimensional structure of the protein in a three-dimensional space with a voxel of l*l*l, wherein the centroid of the three-dimensional structure of the protein coincides with the origin of the three-dimensional space, and scale all coordinates of the three-dimensional structure of the protein;
[0098] The first generation module 403 is configured to obtain an l×l×l shape vector according to whether the three-dimensional structure of the protein passes through the voxels in the three-dimensional space, where if the three-dimensional structure of the protein passes through a voxel, the value is 1, otherwise it is 0; and obtain an l×l×l first feature vector and an l×l×l second feature vector according to the positions, hydrophobicity, and side chain net charge indices of each amino acid in the three-dimensional structure of the protein;
[0099] The second generation module 404 is configured to generate a third feature vector by using the amino acid sequence;
[0100] The connection module 405 is configured to connect the shape vector, the first feature vector, the second feature vector, and the third feature vector to obtain a final feature vector;
[0101] The prediction module 406 is configured to input the final feature vector into a trained prediction model for prediction to obtain the interaction of the protein to be predicted.
[0102] In some embodiments, the second generation module 404 is further configured to:
[0103]
[0104] where P represents the amino acid sequence of the protein, with a length of N r , m is the number of the m-th amino acid in the sequence, n is the number of the physicochemical property, lag is the value of the lag, and P m,n is the physicochemical property with the number n of the m-th amino acid, thereby converting the amino acid sequence into a vector of (n×lag) dimensions.
[0105] In some embodiments, the scaling module 402 is further configured to:
[0106] Use to scale all the coordinates of the three-dimensional structure of the protein, where the R max is the radius of the largest inscribed sphere in the three-dimensional space.
[0107] In some embodiments, it further includes a neural network module, which is configured to:
[0108] Construct a prediction model by using a connection layer, a linear layer, H transformer modules, and a normalization layer;
[0109] Construct a training set, and perform the processing of steps S2-S5 on each sample in the training set to obtain the final feature vector corresponding to each sample;
[0110] Use the final feature vector corresponding to each sample to train the prediction model.
[0111] Based on the same inventive concept, according to another aspect of the present invention, as Figure 7 shown, an embodiment of the present invention further provides a computer device 501, including:
[0112] At least one processor 520; and
[0113] A memory 510, where the memory 510 stores a computer program 511 that can run on the processor, and when the processor 520 executes the program, it performs the steps of any of the above protein interaction prediction methods.
[0114] Based on the same inventive concept, according to another aspect of the present invention, as Figure 8 shown, an embodiment of the present invention further provides a computer-readable storage medium 601, where the computer-readable storage medium 601 stores a computer program 610, and when the computer program 610 is executed by a processor, it performs the steps of any of the above protein interaction prediction methods.
[0115] Finally, it should be noted that those of ordinary skill in the art can understand that all or part of the processes in the above embodiment methods can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above various methods.
[0116] In addition, it should be understood that the computer-readable storage medium herein (for example, a memory) can be a volatile memory or a non-volatile memory, or can include both a volatile memory and a non-volatile memory.
[0117] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, a general description of the functions of various illustrative components, blocks, modules, circuits, and steps has been presented. Whether this function is implemented as software or hardware depends on the specific application and the design constraints imposed on the overall system. The functions that those skilled in the art can implement in various ways for each specific application, but such implementation decisions should not be construed as causing a departure from the scope of the disclosure of the embodiments of the present invention.
[0118] The above are exemplary embodiments disclosed by the present invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments disclosed by the present invention as defined by the claims. The functions, steps, and / or actions of the method claims according to the disclosed embodiments herein do not need to be performed in any specific order. In addition, although the elements disclosed in the embodiments of the present invention can be described or claimed in individual form, they can also be understood as plural unless explicitly limited to the singular form.
[0119] It should be understood that, as used herein, unless the context clearly supports the exception, the singular form "a" is also intended to include the plural form. It should also be understood that the "and / or" used herein refers to any and all possible combinations including one or more of the associated listed items.
[0120] The serial numbers of the above-disclosed embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.
[0121] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk, an optical disc, etc.
[0122] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the embodiments disclosed by the present invention (including the claims) is limited to these examples; under the concept of the embodiments of the present invention, the technical features between the above embodiments or different embodiments can also be combined, and there are many other variations in different aspects of the above embodiments of the present invention, which are not provided in detail for the sake of brevity. Therefore, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.
Claims
1. A method for predicting protein interactions, characterized in that, It includes the following steps: S1. Obtain the three-dimensional structure and amino acid sequence of the protein to be predicted; S2, place the three-dimensional structure of the protein in a three-dimensional space with a voxel size of l×l×l , where the centroid of the three-dimensional structure of the protein coincides with the origin of the three-dimensional space, and scale all coordinates of the three-dimensional structure of the protein; S3, obtain the shape vector l×l×l according to whether the three-dimensional structure of the protein passes through the voxels in the three-dimensional space, where if the three-dimensional structure of the protein passes through the voxel, the value is 1, otherwise it is 0; and obtain the first eigenvector l×l×l and the second eigenvector l×l×l according to the position, hydrophobicity and side chain net charge index of each amino acid in the three-dimensional structure of the protein; S4. Generate a third feature vector using the amino acid sequence; S5. Concatenate the shape vector, the first feature vector, the second feature vector, and the third feature vector to obtain the final feature vector; S6. Input the final feature vector into the trained prediction model for prediction to obtain the interaction of the protein to be predicted; Generating a third feature vector using the amino acid sequence further includes obtaining each element value in the vector based on the following formula: Among them, P represents the amino acid sequence of the protein, with a length of , m is the number of the m-th amino acid in the sequence, n is the number of the physicochemical property, lag is the value of lag, is the physicochemical property with the number n of the m-th amino acid, thereby converting the amino acid sequence into a vector of (n×lag) dimensions.
2. The method according to claim 1, wherein Scaling all coordinates of the three-dimensional structure of the protein further includes: Scale all coordinates of the three-dimensional structure of the protein using λ = , where the is the radius of the largest inscribed sphere in the three-dimensional space.
3. The method according to claim 1, characterized in that, It also includes: Construct a prediction model using a connection layer, a linear layer, H transformer modules, and a normalization layer; Construct a training set, and perform the processing of steps S2 - S5 on each sample in the training set to obtain the final feature vector corresponding to each sample; Train the prediction model using the final feature vector corresponding to each sample.
4. A prediction system for protein-protein interactions, characterized in that, It includes: An acquisition module configured to obtain the three-dimensional structure and amino acid sequence of the protein to be predicted; Scaling module, configured to place the three-dimensional structure of the protein in a three-dimensional space with a voxel of l×l×l , wherein the centroid of the three-dimensional structure of the protein coincides with the origin of the three-dimensional space, and scales all coordinates of the three-dimensional structure of the protein; A first generation module configured to obtain a shape vector of l×l×l based on whether the three-dimensional structure of the protein passes through a voxel in the three-dimensional space, where if the three-dimensional structure of the protein passes through the voxel, the value is 1, otherwise 0; and obtain a first feature vector of l×l×l and a second feature vector of l×l×l according to the position, hydrophobicity, and side chain net charge index of each amino acid in the three-dimensional structure of the protein in the three-dimensional space; l×l×l wherein if the three-dimensional structure of the protein passes through the voxel, the value is 1, otherwise 0; and obtain a first feature vector of l×l×l and a second feature vector of l×l×l according to the position of each amino acid in the three-dimensional structure of the protein in the three-dimensional space, hydrophobicity, and side chain net charge index; l×l×l a first feature vector of l×l×l a second feature vector of; A second generation module configured to generate a third feature vector using the amino acid sequence; A connection module configured to concatenate the shape vector, the first feature vector, the second feature vector, and the third feature vector to obtain the final feature vector; A prediction module configured to input the final feature vector into the trained prediction model for prediction to obtain the interaction of the protein to be predicted; The second generation module is further configured to: Among them, P represents the amino acid sequence of the protein, with a length of , m is the number of the m-th amino acid in the sequence, n is the number of the physicochemical property, lag is the value of the lag, is the physicochemical property with the number n of the m-th amino acid, thereby converting the amino acid sequence into a vector of (n × lag) dimensions.
5. The system according to claim 4, wherein The scaling module is further configured to: Using λ = scale all coordinates of the three-dimensional structure of the protein, where the is the radius of the largest inscribed sphere in the three-dimensional space.
6. The system according to claim 4, wherein It also includes a neural network module configured to: Construct a prediction model using a connection layer, a linear layer, H transformer modules, and a normalization layer; Construct a training set, and perform the processing of steps S2 - S5 on each sample in the training set to obtain the final feature vector corresponding to each sample; Train the prediction model using the final feature vector corresponding to each sample.
7. A computer device, comprising: At least one processor; And A memory storing a computer program that can run on the processor, wherein when the processor executes the program, it executes the steps of the method according to any one of claims 1 - 3.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it executes the steps of the method according to any one of claims 1 - 3.
Citation Information
Patent Citations
Prokaryotic acetylation site prediction method based on information fusion and deep learning
CN111063393A
Drug-protein interaction prediction model based on convolutional neural network
CN113593633A