Protein allosteric site recognition system based on deep learning
By adopting an improved hollow convolutional neural network and frequency domain channel attention mechanism in the protein allosteric site recognition system, combined with the XGBoost model, the data scarcity and complex dependency recognition problem of identifying protein allosteric sites in the prior art is solved, and higher recognition accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510250499.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art has problems such as scarce data, time-consuming and labor-intensive identification of protein allosteric sites, and the difficulty of traditional machine learning models to capture the complex dependence of protein structure, resulting in high false positives or false negatives in the identification results.
A deep learning-based protein allosteric site recognition system is proposed, using an improved hollow convolutional neural network (DCNN) combined with the frequency domain channel attention mechanism to enhance the capture of periodic and long-distance-dependent features, and an XGBoost model is introduced to reduce the risk of overfitting.
Through this system, the characterization ability of allosteric site characteristics is significantly improved, the accuracy and reliability of identifying allosteric sites are improved, and the false positive and false negative rates are reduced.
Smart Images

Figure CN120089210A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and biomedicine, and in particular to a protein allosteric site recognition system based on deep learning. Background Art
[0002] In modern drug development, identifying allosteric sites on proteins is of great significance for the development of new drugs with high efficiency and low toxicity, especially in the field of targeted therapy and precision medicine. However, allosteric sites are not only found on special proteins, but also have the potential to be widely distributed in proteins, which means that almost all proteins may have allosteric sites. However, there are currently very few proteins with confirmed allosteric sites, and the scarcity of data has severely limited the development of allosteric drugs. Therefore, the identification of allosteric sites on proteins has become a research topic that needs to be solved urgently.
[0003] At present, there are still some challenges in the identification of allosteric sites of proteins. Existing methods mainly rely on traditional experimental methods and machine learning models. However, the use of traditional experimental methods to find allosteric sites requires a lot of manpower, material resources and time, and the results are minimal. At the same time, traditional machine learning models are difficult to effectively capture the complex local and global dependencies in protein structures, and the recognition results may have a high false positive or false negative rate. In addition, allosteric sites are often associated with long-range structural regions. Most methods only focus on the local characteristics of protein allosteric sites, ignoring the long-distance dependencies and periodic characteristics that may be involved in allosteric sites, which further limits the accurate identification of allosteric sites. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and proposes a protein allosteric site identification system based on deep learning. The improved atrous convolutional neural network is adopted to break through the limitations of traditional methods in capturing complex dependencies in protein structure, enhance the capture of periodic and long-distance dependency features, and introduce the XGBoost model to reduce the risk of overfitting, while improving the fitting ability of complex data, further improving the accuracy and reliability of identifying allosteric sites.
[0005] To achieve the above object, the technical solution provided by the present invention is: a protein allosteric site recognition system based on deep learning, comprising:
[0006] The data import module is used to load the PDB file of proteins with allosteric effects and the CSV file of their allosteric site location information, and perform preprocessing to obtain the spatial structure data of the protein;
[0007] The feature extraction module uses the DSSP tool, residue contact network, and Feature software to extract secondary structure features, contact network features, and physicochemical features of the residue microenvironment from the spatial structure data of proteins, and fuses the secondary structure features, contact network features, and physicochemical features of the residue microenvironment to generate a feature matrix;
[0008] The training module is based on the improved dilated convolutional neural network DCNN to learn the local and global feature distributions of each allosteric protein from the fused feature matrix, capture the feature information of allosteric sites, learn the constraint conditions of allosteric sites and the complex dependencies in space, and finally obtain a trained improved model; this improved model improves the feature capture module and prediction module of the dilated convolutional neural network DCNN; the improvement of the feature capture module is: introducing three layers of parallel dilated convolutions to extract features with large, medium, and small receptive fields respectively, and at the same time combining the frequency domain channel attention mechanism to enhance the capture of periodic and long-range dependence features through the conversion of frequency domain features and global pooling; the improvement of the prediction module is: introducing the XGBoost model, using the method of cascading multiple decision trees, and improving the prediction accuracy and generalization ability of the model by weighting and regularizing the extracted key features;
[0009] The recognition module uses the trained improved model to identify allosteric sites on proteins according to the PDB files of proteins and provide recognition explanations.
[0010] Furthermore, the data import module includes a data loading module and a data preprocessing module, where:
[0011] The data loading module loads the PDB file of the protein with allosteric effect and the CSV file of the allosteric site position information of the corresponding protein from the local through the Bio.PDB and pandas toolkits in the Python environment, and parses the PDB file to obtain the three-dimensional structure information of the protein;
[0012] The data preprocessing module uses the Bio.PDB and pandas toolkits to retain the spatial structure information of the peptide chain where the allosteric site is located according to the allosteric site position information provided by the CSV file, and generates the spatial structure data of the protein related only to the allosteric site for subsequent feature extraction and model training.
[0013] Furthermore, the feature extraction module performs the following operations:
[0014] 1) Extract secondary structure features based on the DSSP tool:
[0015] The DSSP tool parses the spatial structure data of a protein and, based on secondary structure prediction methods, identifies the secondary structure type of each residue, its corresponding sequence position, and chain identifier. First, the spatial structure data of the protein is input into the DSSP tool to generate a DSSP file containing the secondary structure features of the protein. Then, by parsing the spatial structure data of the protein and the secondary structure features in the DSSP file, the secondary structure features of the protein pocket region are extracted, including the secondary structure type of the residue and the relative solvent accessibility RSA. Subsequently, the Bio.PDB.DSSP module in Biopython is used to parse the DSSP file, convert the secondary structure type of each residue into a one-hot encoding, and calculate and normalize its relative solvent accessibility RSA for use in the construction of subsequent feature matrices. Finally, the generated DSSP file will be retained for further extraction of the secondary structure features in the residue microenvironment by the Feature software;
[0016] 2) Extract the contact network features of atoms, i.e., short-range interaction features, based on the residue contact network:
[0017] First, a contact network of the protein is generated using Biopython and NetworkX tools; the spatial coordinates of each residue are parsed from the PDB file of the protein, the alpha carbon atom of each residue is selected as the representative atom, and based on the distance relationship in three-dimensional space, according to the set threshold where is a unit describing the atomic spacing, find the contact residues, and construct a contact network of alpha carbon atom - alpha carbon atom:
[0018]
[0019] In the formula, d k,c (R) is the residue contact density of residue R at position k with threshold c, len(P) is the sequence length of protein P, and C k,c (R) is the total number of contacts of residue R at position k with threshold c;
[0020] Next, according to the spatial coordinates of the alpha carbon atom of residue R, the contacts of residue R are divided into the upper hemisphere and the lower hemisphere, the number of contacts in the upper hemisphere and the lower hemisphere of each residue at different thresholds is counted, and the exposure ratio is calculated:
[0021]
[0022] D k,c (R) = 1 - U k,c (R)
[0023] In the formula, U k,c (R) is the upper hemisphere exposure ratio of residue R at position k with threshold c, and D k,c(R) is the exposure ratio of the lower hemisphere of the k - residue R with c as the threshold, C u,k,c (R) is the number of contacts in the upper hemisphere of the k - residue R with c as the threshold, C k,c (R) is the total number of contacts in the upper and lower hemispheres of the k - residue R with c as the threshold;
[0024] In the contact network, further analyze the local network properties of each residue, and calculate the Clustering Coefficient and Betweenness Centrality of each residue:
[0025]
[0026] In the formula, C c (R) is the clustering coefficient of residue R, T(R) is the number of triangles passing through residue R in the network, deg(R) is the degree of residue R in the network, is the theoretical maximum number of triangles for residue R;
[0027]
[0028] In the formula, C B (R) is the betweenness centrality of residue R, V is the set of nodes, s and t are any two nodes in the network, σ(s,t) is the total number of shortest paths from node s to node t, and σ(s,t|R) is the number of shortest paths passing through node s, node t and residue R; in this network, it is defined that if σ(s,t) = 0, then
[0029] Through the residue contact network, extract the residue short - range interaction characteristics in the protein contact network; the residue contact density reveals the local contact situation of residues, the hemisphere exposure ratio reflects the exposure characteristics of residues in space, and the clustering coefficient and betweenness centrality characterize the local and global importance of residues in the contact network from the perspective of network properties;
[0030] 3) Extract the physicochemical characteristics of the residue microenvironment based on Feature software:
[0031] Perform microenvironment sampling through the featurize module of Feature software, analyze the protein structure at the atomic level, perform microenvironment sampling on the region centered on the alpha - carbon atom of each residue. In the microenvironment sampling, define the spatial range of the microenvironment as a spherical shell with a radius of and ;
[0032] Using the Atomselect module of Feature software, select the alpha carbon atoms of the target residues and perform microenvironment characterization on each residue one by one. The content of the characterization includes the physical, chemical, and structural properties within the sphere and spherical shell:
[0033] Extract the quantities of various elements within the spherical shell or sphere according to the elemental level, including: the quantity of any element ELEMENT_IS_ANY, the quantity of carbon element ELEMENT_IS_C, the quantity of nitrogen element ELEMENT_IS_N, the quantity of oxygen element ELEMENT_IS_O, the quantity of sulfur element ELEMENT_IS_S, and the quantity of other elements excluding carbon, nitrogen, oxygen, and sulfur ELEMENT_IS_OTHER;
[0034] Extract the quantities of various atoms within the spherical shell or sphere according to the atomic level, including:
[0035] The quantity of carbon atoms connected to oxygen atoms ATOM_TYPE_IS_C, the quantity of terminal carbon atoms on the side chain ATOM_TYPE_IS_CT, the quantity of carbon atoms connected to amino carbon atoms ATOM_TYPE_IS_CA, the quantity of nitrogen atoms on the amino group ATOM_TYPE_IS_N, the quantity of atoms with the structural atom identifier N2 in the PDB file ATOM_TYPE_IS_N2, the quantity of atoms with the structural atom identifier N3 in the PDB file ATOM_TYPE_IS_N3, the quantity of nitrogen atoms connected to the alpha carbon atom ATOM_TYPE_IS_NA, the quantity of double-bonded oxygen atoms connected to the alpha carbon atom ATOM_TYPE_IS_O, the quantity of carboxyl double-bonded oxygen atoms of the alpha carbon atom closest to the alpha carbon atom on the side chain ATOM_TYPE_IS_O2, the quantity of hydroxyl hydrogen atoms connected to the alpha carbon atom ATOM_TYPE_IS_OH, the quantity of sulfur atoms ATOM_TYPE_IS_S, the quantity of hydrogen atoms on sulfur ATOM_TYPE_IS_SH, and the quantity of all atoms other than the above atoms ATOM_TYPE_IS_OTHER;
[0036] Extract the various charge quantities within the spherical shell or sphere according to the residue level, including: partial atomic charge PARTIAL_CHARGE, negative charge NEG_CHARGE, positive charge POS_CHARGE, total charge considering histidine CHARGE_WITH_HIS, and total charge without considering histidine CHARGE;
[0037] Extract the quantities of various residues within the spherical shell or sphere according to the atomic level, including:
[0038] The number of alanine (RESIDUE_NAME_IS_ALA), the number of arginine (RESIDUE_NAME_IS_ARG), the number of asparagine (RESIDUE_NAME_IS_ASN), the number of aspartic acid (RESIDUE_NAME_IS_ASP), the number of cysteine (RESIDUE_NAME_IS_CYS), the number of glutamine (RESIDUE_NAME_IS_GLN), the number of glutamic acid (RESIDUE_NAME_IS_GLU), the number of glycine (RESIDUE_NAME_IS_GLY), the number of histidine (RESIDUE_NAME_IS_HIS), the number of isoleucine (RESIDUE_NAME_IS_ILE), the number of leucine (RESIDUE_NAME_IS_LEU), the number of lysine (RESIDUE_NAME_IS_LYS), the number of methionine (RESIDUE_NAME_IS_MET), the number of phenylalanine (RESIDUE_NAME_IS_PHE), the number of proline (RESIDUE_NAME_IS_PRO), the number of serine (RESIDUE_NAME_IS_SER), the number of threonine (RESIDUE_NAME_IS_THR), the number of tryptophan (RESIDUE_NAME_IS_TRP), the number of tyrosine (RESIDUE_NAME_IS_TYR), the number of valine (RESIDUE_NAME_IS_VAL), the number of residues whose names do not belong to all of the above residue names (RESIDUE_NAME_IS_OTHER);
[0039] According to the various residue and atomic properties in the microenvironment, the overall physical and chemical properties within the sphere or spherical shell are statistically analyzed, including:
[0040] The number of hydrophobic residues (RESIDUE_CLASS1_IS_HYDROPHOBIC), the number of charged residues (RESIDUE_CLASS1_IS_CHARGED), the number of polar residues (RESIDUE_CLASS1_IS_POLAR), the number of residues that do not belong to hydrophobic, polar, and uncharged residues
[0041] RESIDUE_CLASS1_IS_UNKNOWN, the number of residues with non - polarity RESIDUE_CLASS2_IS_NONPOLAR, the number of hydrophobic residues that distinguish RESIDUE_CLASS2_IS_POLAR from CLASS1, basic residues RESIDUE_CLASS2_IS_BASIC, the number of acidic residues RESIDUE_CLASS2_IS_ACIDIC, the number of residues that do not belong to the non - polar class and have no acid - base properties RESIDUE_CLASS2_IS_UNKNOWN;
[0042] Combined with the secondary structure features provided by the DSSP tool, count the number of various secondary structures contained in the microenvironment, including:
[0043] The number of 3-turn α-helices of secondary structure type SECONDARY_STRUCTURE1_IS_3HELIX, the number of 4-turn α-helices of secondary structure type SECONDARY_STRUCTURE1_IS_4HELIX, the number of 5-turn α-helices of secondary structure type SECONDARY_STRUCTURE1_IS_5HELIX, the number of short segments connecting α-helices and β-helices of secondary structure type SECONDARY_STRUCTURE1_IS_BRIDGE, the number of β-strands and anti-β-sheets of secondary structure type SECONDARY_STRUCTURE1_IS_STRAN, the number of turns of secondary structure type SECONDARY_STRUCTURE1_IS_TURN, the number of regular bends of secondary structure type SECONDARY_STRUCTURE1_IS_BEND, the number of secondary structures of secondary structure type that do not belong to strand, helix, bend, bridge, turn, heterocycle SECONDARY_STRUCTURE1_IS_COIL, the number of heteroatoms SECONDARY_STRUCTURE1_IS_HET, the number of secondary structures that cannot be recognized by DSSP SECONDARY_STRUCTURE1_IS_UNKNOWN, the number of α-helices of secondary structure type SECONDARY_STRUCTURE2_IS_HELIX, the number of β-helices including BRIDGE, BEND, TURN of secondary structure type SECONDARY_STRUCTURE2_IS_BETA, the number of random bends of secondary structure type SECONDARY_STRUCTURE2_IS_COIL, the number of heterocyclic structures of secondary structure type SECONDARY_STRUCTURE2_IS_HET, the number of secondary structures that do not belong to HELIX, BETA, COIL, HET SECONDARY_STRUCTURE2_IS_UNKNOWN;
[0044] Finally, count the special structures and functional groups within the spherical shell or sphere, including: hydroxyl group HYDROXYL, amide AMIDE, amine AMINE, carbonyl CARBONYL, ring system RING_SYSTEM, peptide PEPTIDE; and physical properties at the atomic level: van der Waals volume VDW_VOLUME, hydrophobicity HYDROPHOBICITY, mobility MOBILITY, solvent accessibility SOLVENT_ACCESSIBILITY;
[0045] The Feature software correlates protein function with its structure by analyzing the structures of functionally similar proteins; the core idea of the Feature software is to analyze protein structure at the atomic level, sampling the small spherical volumes around each atom or a specified set of points, and these volumes are called microenvironments; the microenvironment is represented by a real vector of a feature vector, and the feature vector contains the physicochemical feature information within the sphere or spherical shell; by extracting the microenvironment information, the complex three-dimensional structure data of proteins can be transformed into numerical features that can be used in machine learning models;
[0046] 4) Use the pandas toolkit in the Python environment to concatenate the secondary structure features, contact network features, and residue microenvironment physicochemical features obtained in steps 1), 2), and 3) to obtain the feature matrix of the protein; through the feature extraction module, 258 feature descriptions are generated for each residue of the protein.
[0047] Furthermore, the training module performs the following operations:
[0048] 1) Input the feature matrix into the embedding layer:
[0049] The shape of the input feature matrix X of each protein is (L, 258, 1), where L is the number of residues of the protein, 258 represents the number of features extracted for each residue, and 1 represents single-channel input; for the 258 features of each residue, two-layer fully connected neural network is used for embedding processing to reduce the dimension and introduce non-linearity; the input dimension of the first-layer fully connected neural network is 258, and the output dimension is 128. The input dimension of the second-layer fully connected neural network is 128, and the output dimension is 64. Both two-layer fully connected neural networks use the ReLU function as the activation function, and the shape of the embedded feature matrix is (L, 64);
[0050] 2) Extract the local and global features of the embedded feature matrix through the dilated convolution layer:
[0051] Apply a dilated convolutional neural network DCNN composed of three parallel dilated convolutional layers to the embedded feature matrix to extract local and global features. The shape of the input feature matrix is (L, 64, 1), and the mathematical expression of the dilated convolution layer is shown as follows:
[0052]
[0053] Where, O(i,j′) is the value of the output feature matrix after dilated convolution at position (i,j′), where i and j′ represent the i-th row and j′-th column of the output feature matrix, W(m,n) is the weight of the convolutional kernel at position (m,n), where m and n represent the m-th row and n-th column of the convolutional kernel, I(i+m·r,j′+n·r) is the value of the input feature matrix at position (i+m·r,j′+n·r), r is the dilation rate, M c and N c represent the size of the convolutional kernel; Three parallel convolutional layers with different dilation rates are used, and the dilation rates r are 4, 2, and 1 respectively. The dilated convolutional layer with r = 4 is responsible for extracting features with large receptive fields, and the dilated convolutional layer with r = 2 is responsible for extracting features with medium receptive fields. When r = 1, the dilated convolution degenerates into ordinary convolution. When r > 1, gaps are inserted between the elements of the convolutional kernel; The size of the convolutional kernel of each convolutional layer is (3,1), the stride is 1, and the padding method is the same padding method that keeps the output feature map size consistent with the input size. The number of output channels of each convolutional layer is 64, and the activation function is the ReLU function;
[0054] The outputs of the three parallel dilated convolutions are concatenated in the channel dimension to form a joint feature matrix X concat , whose shape is (L,192,1). The concatenated feature matrix is dimensionally reduced through a 1×1 convolutional layer to obtain a global feature matrix X gap , whose shape is (L,64);
[0055] 3) Use the frequency-domain channel attention mechanism in the feature capture module to capture the implicit periodicity and long-range dependence in the global feature matrix X gap :
[0056] Apply the fast Fourier transform to the global feature matrix X gap to convert it from the spatial domain to the frequency domain to obtain a frequency-domain feature vector F:
[0057]
[0058] Where, F(u,v) is the value of the frequency-domain feature matrix at position (u,v), where u and v represent the u-th row and v-th column of the frequency-domain feature matrix, X gap (x,y) is the value of the global feature matrix at position (x,y), where x and y represent the x-th row and y-th column of the global feature matrix, H and W are the height and width of the feature matrix, and j is the imaginary unit;
[0059] Extract the amplitude spectrum A from the frequency-domain feature F:
[0060] A(u,v) = |F(u,v)|
[0061] where \(A(u, v)\) is the value of the frequency-domain amplitude spectrum at position \((u, v)\);
[0062] Perform global pooling on the frequency-domain amplitude spectrum \(A\) to obtain the frequency-domain feature \(S\) at the channel level:
[0063]
[0064] 4) Input the frequency-domain feature \(S\) into a two-layer fully connected neural network to generate the attention weight \(Attention(S)\), and apply the generated channel attention weight \(Attention(S)\) to the original feature:
[0065] Attention(S)=\(\sigma(W\) 2 \(\cdot ReLU(W\) 1 \(\cdot S))\)
[0066] F' = Attention(S) \(\cdot F\)
[0067] where \(Attention(S)\) is the attention weight of the frequency-domain feature \(S\), \(W\) 1 and \(W\) 2 are the weight matrices of the two-layer fully connected neural network, \(ReLU\) is the activation function, and \(\sigma\) is the Sigmoid function; the input dimension of the first-layer fully connected neural network is 64, the output dimension is 32, and the activation function is the \(ReLU\) function, which is used to introduce non-linearity. The input dimension of the second-layer fully connected neural network is 32, the output dimension is 64, and the activation function is the Sigmoid function, which is used to limit the output within the range of \([0, 1]\). \(F'\) is the weighted channel feature, and its shape is \((L, 64)\);
[0068] 5) Use a frequency-domain variational encoder to perform feature compression and latent variable modeling on the weighted channel feature \(F'\) to generate the latent variable matrix \(Z\):
[0069] The encoder of the frequency-domain variational autoencoder uses a two-layer fully-connected neural network. The input dimension of the first-layer fully-connected neural network is 64, the output dimension is 32, and the activation function is the ReLU function. The input dimension of the second-layer fully-connected neural network is 32, the output dimension is 16, and the activation function is the Linear function. The encoder of the frequency-domain variational autoencoder performs a non-linear transformation on the channel feature F', extracts its potential low-dimensional representation, generates the mean and variance after passing through the layers of the encoder of the frequency-domain variational autoencoder, and then obtains the latent variable matrix Z through reparameterization, with its shape being (L, 16). The decoder of the frequency-domain variational autoencoder also uses a two-layer fully-connected neural network. The input dimension of the first-layer fully-connected neural network is 16, the output dimension is 32, and the activation function is the ReLU function. The input dimension of the second-layer fully-connected neural network is 32, the output dimension is 64, and the activation function is the Linear function. The loss function of the encoder and decoder of the frequency-domain variational autoencoder consists of the reconstruction loss and the KL divergence loss:
[0070] L total = L recon + 0.1·L KL
[0071] In the formula, L total represents the total loss function, L recon represents the difference between F' and the reconstructed feature matrix generated by the decoder, which is calculated using the mean squared error, and L KL represents the KL divergence loss;
[0072] 6) Residue-level identification:
[0073] The latent variable matrix Z generated by the frequency-domain variational autoencoder is input into the XGBoost model as the feature of each residue for residue-level identification. The identification target value of the XGBoost model is the binary label vector Y with the same length as the number of amino acids of the currently trained protein θ , where the value range of Y θ is in {0, 1}. When Y θ = 0, it means that the θ-th residue is a non-allosteric site, and when Y θ = 1, it means that the θ-th residue is an allosteric site. The parameter settings of the XGBoost model are: the maximum tree depth is 6, the learning rate is 0.1, the number of trees is 100, the subsample ratio is 0.8, and the regularization parameters are γ = 0.1 and λ = 1. Finally, the prediction module determines whether a residue is an allosteric site based on the probability value output by the XGBoost model. If the probability that a certain residue is predicted to be an allosteric site is greater than 0.5, the prediction module will identify it as an allosteric site.
[0074] Furthermore, the identification module performs the following operations:
[0075] 1) Read the PDB file of the input protein and parse it using the Bio.PDB toolkit in the Python environment to extract the three-dimensional structure information of the protein; store the parsed three-dimensional structure information of the protein in the variable protein_structure;
[0076] 2) Use the data import module to retain the spatial structure information of the peptide chains related to the allosteric site in the protein_structure obtained in step 1) and generate the spatial structure data of the protein related only to the allosteric site; store the processed spatial structure data of the protein in the variable processed_protein_data as the input for feature extraction;
[0077] 3) Use the feature extraction module to extract features from the spatial structure data of the protein stored in the variable processed_protein_data in step 2), call the DSSP tool, residue contact network, and Feature software to extract secondary structure features, contact network features, and physicochemical features of the residue microenvironment, and store them in the variables dssp_feature, residue_connect_feature, and mirco_env_feature respectively. Use the pandas toolkit to concatenate dssp_feature, residue_connect_feature, and mirco_env_feature to generate the feature matrix feature_matrix of the protein;
[0078] 4) Input the generated feature matrix feature_matrix into the trained improved model for prediction. The model combines local and global feature distributions, captures the spatial dependence and constraints of residues, and obtains the probability P that each residue is an allosteric site. The value range of P is in [0,1], indicating the possibility that the residue is an allosteric site. Store the allosteric site probability of each residue in the prediction probability list prediction_probability_list;
[0079] 5) For each residue in the prediction probability list prediction_probability_list with the allosteric site probability P, a set probability threshold T = 0.5 is used to classify each residue: If the probability P of a certain residue is ≥ T, it is classified as an allosteric site with a label of 1; if the probability P of a certain residue is > T, it is classified as a non-allosteric site with a label of 0; the classified labels are stored in the prediction label list prediction_list; finally, based on the probability value of the residue in the prediction probability list prediction_probability_list and the label in the prediction label list prediction_list, it is explained whether it is an allosteric site.
[0080] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0081] 1. Through the systematic feature extraction of the DSSP tool, residue contact network, and Feature software, secondary structure features, contact network features, and physicochemical features of the residue microenvironment are obtained from the protein structure, significantly enhancing the characterization ability of allosteric site features.
[0082] 2. By analyzing the multi-dimensional feature matrix of the protein through the improved dilated convolutional network, dilated convolutional layers with different dilation rates are introduced to extract features of large, medium, and small receptive fields respectively, enhancing the ability to capture structural information at different scales.
[0083] 3. By introducing the frequency domain channel attention mechanism, the capture of periodic and long-range dependence features is enhanced. This improvement not only improves the comprehensiveness of feature capture, enabling it to learn more complex spatial dependence relationships, but also enhances the model's ability to identify long-range structural associations.
[0084] 4. The multi-layer structure design of the frequency domain variational encoder is conducive to compressing weighted features and latent variable modeling, enabling better extraction of potential low-dimensional representations, and through the optimization of reconstruction loss and KL divergence loss, improving the effectiveness of feature representation and reducing the impact of data noise.
[0085] 5. By introducing the XGBoost model for final prediction, using the method of cascading multiple decision trees, the key features extracted are weighted and regularized, providing accurate and reliable allosteric site recognition results, and improving the prediction accuracy and generalization ability of the model.
[0086] 6. The accurate identification of allosteric sites of proteins by the present invention can not only provide key target information for new drug research and development, but also provide important theoretical basis for understanding the allosteric mechanism of proteins and designing allosteric modulators, promoting the process of targeted drug development. Brief Description of the Drawings
[0087] Figure 1 It is a schematic diagram of the relationships among the various modules of the system of the present invention.
[0088] Figure 2 It is a flowchart of the training and recognition of the system of the present invention.
[0089] Figure 3 It is a schematic diagram of the structure of the improved dilated convolutional neural network used in the system of the present invention. Specific implementation manners
[0090] The present invention will be further described below in conjunction with specific embodiments.
[0091] This embodiment discloses a protein allosteric site recognition system based on deep learning, which is a protein allosteric site recognition system based on deep learning developed using the Python language and capable of running on Windows devices. The relationships among the various modules of the system are as Figure 1 shown, and the processes of system training and recognition are as Figure 2 shown.
[0092] It includes:
[0093] A data import module, which is used to load the PDB file of the protein with allosteric effect and the CSV file of its allosteric site position information, and perform preprocessing to obtain the spatial structure data of the protein;
[0094] A feature extraction module, which uses the DSSP tool, residue contact network and Feature software to extract secondary structure features, contact network features and physicochemical features of the residue microenvironment from the spatial structure data of the protein, and fuses the secondary structure features, contact network features and physicochemical features of the residue microenvironment to generate a feature matrix;
[0095] A training module, which learns the local and global feature distributions of each allosteric protein from the feature matrix generated by fusion based on the improved dilated convolutional neural network DCNN, captures the feature information of the allosteric site, learns the constraint conditions of the allosteric site therein and the complex dependencies in space, and finally obtains a trained improved model (i.e., the optimal model); as Figure 3 shown, this improved model improves the feature capture module and prediction module of the dilated convolutional neural network DCNN; the improvement of the feature capture module is: introducing three layers of parallel dilated convolutions to extract features with large, medium and small receptive fields respectively, and at the same time combining the frequency domain channel attention mechanism to enhance the capture of periodic and long-range dependence features through the conversion of frequency domain features and global pooling; the improvement of the prediction module is: introducing the XGBoost model, using the method of cascading multiple decision trees, and improving the prediction accuracy and generalization ability of the model by weighting and regularizing the extracted key features;
[0096] The recognition module uses the trained improved model to identify allosteric sites on a protein according to the PDB file of the protein and provide recognition explanations.
[0097] Specifically, the data import module includes a data loading module and a data preprocessing module, where:
[0098] The data loading module loads the PDB file of the protein with allosteric effects and the CSV file of the allosteric site position information of the corresponding protein from the local through the Bio.PDB and pandas toolkits in the Python environment, and parses the PDB file to obtain the three-dimensional structure information of the protein;
[0099] The data preprocessing module uses the Bio.PDB and pandas toolkits to retain the spatial structure information of the peptide chain where the allosteric site is located according to the allosteric site position information provided by the CSV file, and generates the spatial structure data of the protein related only to the allosteric site for subsequent feature extraction and model training.
[0100] Specifically, the feature extraction module performs the following operations:
[0101] 1) Extract secondary structure features based on the DSSP tool:
[0102] The DSSP tool parses the spatial structure data of the protein and, based on the secondary structure prediction method, identifies the secondary structure type of each residue and its corresponding sequence position and chain identifier; First, input the spatial structure data of the protein into the DSSP tool to generate a DSSP file containing the secondary structure features of the protein. Then, by parsing the spatial structure data of the protein and the secondary structure features in the DSSP file, extract the secondary structure features of the protein pocket region, including the secondary structure type of the residue and the relative solvent accessibility RSA. Subsequently, use the Bio.PDB.DSSP module in Biopython to parse the DSSP file, convert the secondary structure type of each residue into a one-hot encoding, and calculate and normalize its relative solvent accessibility RSA for subsequent construction of the feature matrix. Finally, the generated DSSP file will be retained for further extraction of the secondary structure features in the residue microenvironment by the Feature software;
[0103] 2) Extract the contact network features of atoms, i.e., short-range interaction features, based on the residue contact network:
[0104] First, generate the contact network of the protein through the Biopython and NetworkX tools; Parse the spatial coordinates of each residue from the PDB file of the protein, select the alpha carbon atom of each residue as the representative atom, and based on the distance relationship in three-dimensional space, according to the set threshold, respectively wherein is a unit for describing the atomic spacing. Search for contact residues and construct a contact network of alpha-carbon atom - alpha-carbon atom:
[0105]
[0106] In the formula, d k,c (R) is the residue contact density of residue R at position k with c as the threshold, len(P) is the sequence length of protein P, and C k,c (R) is the total number of contacts of residue R at position k with c as the threshold;
[0107] Next, according to the spatial coordinates of the alpha-carbon atom of residue R, divide the contacts of residue R into the upper hemisphere and the lower hemisphere, count the number of contacts in the upper and lower hemispheres of each residue at different thresholds, and calculate the exposure ratio:
[0108]
[0109] D k,c (R) = 1 - U k,c (R)
[0110] In the formula, U k,c (R) is the upper hemisphere exposure ratio of residue R at position k with c as the threshold, D k,c (R) is the lower hemisphere exposure ratio of residue R at position k with c as the threshold, C u,k,c (R) is the number of contacts in the upper hemisphere of residue R at position k with c as the threshold, and C k,c (R) is the total number of contacts in the upper and lower hemispheres of residue R at position k with c as the threshold;
[0111] In the contact network, further analyze the local network properties of each residue, and calculate the clustering coefficient and betweenness centrality of each residue:
[0112]
[0113] In the formula, C c (R) is the clustering coefficient of residue R, T(R) is the number of triangles passing through residue R in the network, and deg(R) is the degree of residue R in the network, is the theoretical maximum number of triangles for residue R;
[0114]
[0115] In the formula, C B(R) is the betweenness centrality of residue R, V is the set of nodes, s and t are any two nodes in the network, σ(s,t) is the total number of shortest paths from node s to node t, and σ(s,t|R) is the number of shortest paths passing through node s, node t and residue R; in this network, it is defined that if σ(s,t) = 0, then
[0116] Through the residue contact network, extract the residue short-range interaction characteristics in the protein contact network; the residue contact density reveals the local contact situation of residues, the hemisphere exposure ratio reflects the exposure characteristics of residues in space, and the clustering coefficient and betweenness centrality characterize the local and global importance of residues in the contact network from the perspective of network properties;
[0117] 3) Extract the physicochemical characteristics of the residue microenvironment based on the Feature software:
[0118] Perform microenvironment sampling through the featurize module of the Feature software, analyze the protein structure at the atomic level, perform microenvironment sampling on the region centered on the alpha carbon atom of each residue, and in the microenvironment sampling, define the spatial range of the microenvironment as a spherical shell with a radius of and ;
[0119] Use the Atomselect module of the Feature software to select the alpha carbon atoms of the target residues, and perform microenvironment characterization on each residue one by one. The content of the characterization includes the physical, chemical and structural characteristics within the sphere and the spherical shell:
[0120] Extract the quantity of various elements within the spherical shell or the sphere according to the element level, including: the quantity of any element ELEMENT_IS_ANY, the quantity of carbon element ELEMENT_IS_C, the quantity of nitrogen element ELEMENT_IS_N, the quantity of oxygen element ELEMENT_IS_O, the quantity of sulfur element ELEMENT_IS_S, and the quantity of other elements except carbon, nitrogen, oxygen and sulfur ELEMENT_IS_OTHER;
[0121] Extract the quantity of various atoms within the spherical shell or the sphere according to the atomic level, including:
[0122] The number of carbon atoms connected to oxygen atoms ATOM_TYPE_IS_C, the number of terminal carbon atoms on the side chain ATOM_TYPE_IS_CT, the number of carbon atoms connected to the amino carbon atom ATOM_TYPE_IS_CA, the number of nitrogen atoms on the amino group ATOM_TYPE_IS_N, the number of atoms with the structure atom identifier N2 in the PDB file ATOM_TYPE_IS_N2, the number of atoms with the structure atom identifier N3 in the PDB file ATOM_TYPE_IS_N3, the number of nitrogen atoms connected to the α-carbon atom ATOM_TYPE_IS_NA, the number of double-bonded oxygen atoms connected to the α-carbon atom ATOM_TYPE_IS_O, the number of carboxyl double-bonded oxygen atoms of the α-carbon atom closest to the α-carbon atom on the side chain ATOM_TYPE_IS_O2, the number of hydroxyl hydrogen atoms connected to the α-carbon atom ATOM_TYPE_IS_OH, the number of sulfur atoms ATOM_TYPE_IS_S, the number of hydrogen atoms on sulfur ATOM_TYPE_IS_SH, the number of all atoms other than the above atoms ATOM_TYPE_IS_OTHER;
[0123] Extract various charge amounts within the spherical shell or sphere at the residue level, including: partial atomic charge PARTIAL_CHARGE, negative charge NEG_CHARGE, positive charge POS_CHARGE, total charge considering histidine CHARGE_WITH_HIS, total charge without considering histidine CHARGE;
[0124] Extract the number of various residues within the spherical shell or sphere at the atomic level, including:
[0125] The number of alanine (RESIDUE_NAME_IS_ALA), the number of arginine (RESIDUE_NAME_IS_ARG), the number of asparagine (RESIDUE_NAME_IS_ASN), the number of aspartic acid (RESIDUE_NAME_IS_ASP), the number of cysteine (RESIDUE_NAME_IS_CYS), the number of glutamine (RESIDUE_NAME_IS_GLN), the number of glutamic acid (RESIDUE_NAME_IS_GLU), the number of glycine (RESIDUE_NAME_IS_GLY), the number of histidine (RESIDUE_NAME_IS_HIS), the number of isoleucine (RESIDUE_NAME_IS_ILE), the number of leucine (RESIDUE_NAME_IS_LEU), the number of lysine (RESIDUE_NAME_IS_LYS), the number of methionine (RESIDUE_NAME_IS_MET), the number of phenylalanine (RESIDUE_NAME_IS_PHE), the number of proline (RESIDUE_NAME_IS_PRO), the number of serine (RESIDUE_NAME_IS_SER), the number of threonine (RESIDUE_NAME_IS_THR), the number of tryptophan (RESIDUE_NAME_IS_TRP), the number of tyrosine (RESIDUE_NAME_IS_TYR), the number of valine (RESIDUE_NAME_IS_VAL), the number of residues whose names do not belong to all of the above residue names (RESIDUE_NAME_IS_OTHER);
[0126] According to the characteristics of various residues and atoms in the microenvironment, the overall physical and chemical properties within the sphere or spherical shell are statistically analyzed, including:
[0127] The number of hydrophobic residues (RESIDUE_CLASS1_IS_HYDROPHOBIC), the number of charged residues (RESIDUE_CLASS1_IS_CHARGED), the number of polar residues (RESIDUE_CLASS1_IS_POLAR), the number of residues that are neither hydrophobic, polar nor charged
[0128] RESIDUE_CLASS1_IS_UNKNOWN, the number of residues with non-polar, RESIDUE_CLASS2_IS_NONPOLAR, the number of hydrophobic residues that distinguish RESIDUE_CLASS2_IS_POLAR from CLASS1, basic residues RESIDUE_CLASS2_IS_BASIC, the number of acidic residues RESIDUE_CLASS2_IS_ACIDIC, the number of residues that do not belong to the non-polar class and are neither acidic nor basic RESIDUE_CLASS2_IS_UNKNOWN;
[0129] Combined with the secondary structure features provided by the DSSP tool, count the number of various secondary structures contained in the microenvironment, including:
[0130] The number of 3-turn α-helices of the secondary structure type SECONDARY_STRUCTURE1_IS_3HELIX, the number of 4-turn α-helices of the secondary structure type SECONDARY_STRUCTURE1_IS_4HELIX, the number of 5-turn α-helices of the secondary structure type SECONDARY_STRUCTURE1_IS_5HELIX, the number of short segments connecting α-helices and β-helices of the secondary structure type SECONDARY_STRUCTURE1_IS_BRIDGE, the number of β-strands and anti-β-sheets of the secondary structure type SECONDARY_STRUCTURE1_IS_STRAN, the number of turns of the secondary structure type SECONDARY_STRUCTURE1_IS_TURN, the number of regular bends of the secondary structure type SECONDARY_STRUCTURE1_IS_BEND, the number of secondary structures of the secondary structure type that do not belong to strand, helix, bend, bridge, turn, or heterocycle SECONDARY_STRUCTURE1_IS_COIL, the number of heteroatoms SECONDARY_STRUCTURE1_IS_HET, the number of secondary structures that cannot be recognized by DSSP SECONDARY_STRUCTURE1_IS_UNKNOWN, the number of α-helices of the secondary structure type SECONDARY_STRUCTURE2_IS_HELIX, the number of β-helices including BRIDGE, BEND, and TURN of the secondary structure type SECONDARY_STRUCTURE2_IS_BETA, the number of irregular bends of the secondary structure type SECONDARY_STRUCTURE2_IS_COIL, the number of heterocyclic structures of the secondary structure type SECONDARY_STRUCTURE2_IS_HET, the number of secondary structures that do not belong to HELIX, BETA, COIL, or HET SECONDARY_STRUCTURE2_IS_UNKNOWN;
[0131] Finally, count the special structures and functional groups within the spherical shell or sphere, including: hydroxyl group HYDROXYL, amide AMIDE, amine AMINE, carbonyl CARBONYL, ring system RING_SYSTEM, peptide PEPTIDE; and physical properties at the atomic level: van der Waals volume VDW_VOLUME, hydrophobicity HYDROPHOBICITY, mobility MOBILITY, solvent accessibility SOLVENT_ACCESSIBILITY;
[0132] The Feature software associates protein functions with their structures by analyzing the structures of functionally similar proteins. The core idea of the Feature software is to analyze protein structures at the atomic level, sampling the small spherical volumes around each atom or a specified set of points, which are called microenvironments. The microenvironment is represented by a real-valued vector of a feature vector that contains physicochemical feature information within the sphere or spherical shell. By extracting microenvironment information, the complex three-dimensional structure data of proteins can be transformed into numerical features that can be used in machine learning models.
[0133] 4) Use the pandas toolkit in the Python environment to concatenate the secondary structure features, contact network features, and residue microenvironment physicochemical features obtained in steps 1), 2), and 3) to obtain a feature matrix of proteins. Through the feature extraction module, 258 feature descriptions are generated for each residue of the protein.
[0134] Specifically, as Figure 3 shown, the training module performs the following operations:
[0135] 1) Input the feature matrix into the embedding layer:
[0136] The shape of the input feature matrix X of each protein is (L, 258, 1), where L is the number of residues of the protein, 258 represents the number of features extracted for each residue, and 1 represents single-channel input. For the 258 features of each residue, a two-layer fully connected neural network is used for embedding processing to reduce the dimension and introduce non-linearity. The input dimension of the first layer of the fully connected neural network is 258, and the output dimension is 128. The input dimension of the second layer of the fully connected neural network is 128, and the output dimension is 64. Both layers of the fully connected neural network use the ReLU function as the activation function. The shape of the embedded feature matrix is (L, 64).
[0137] 2) Extract local and global features of the embedded feature matrix through the dilated convolutional layer:
[0138] Apply a dilated convolutional neural network DCNN composed of three parallel dilated convolutional layers to the embedded feature matrix to extract local and global features. The shape of the input feature matrix is (L, 64, 1). The mathematical expression of the dilated convolutional layer is as follows:
[0139]
[0140] Wherein, O(i, j′) is the value of the output feature matrix after dilated convolution at position (i, j′), where i and j′ represent the i-th row and j′-th column of the output feature matrix, W(m, n) is the weight of the convolutional kernel at position (m, n), where m and n represent the m-th row and n-th column of the convolutional kernel, I(i + m·r, j′ + n·r) is the value of the input feature matrix at position (i + m·r, j′ + n·r), r is the dilation rate, M c and N c represent the size of the convolutional kernel; Three parallel convolutional layers with different dilation rates are used, and the dilation rates r are 4, 2, and 1 respectively. The dilated convolutional layer with r = 4 is responsible for extracting features with a large receptive field, and the dilated convolutional layer with r = 2 is responsible for extracting features with a medium receptive field. When r = 1, the dilated convolution degenerates into ordinary convolution. When r > 1, gaps are inserted between the elements of the convolutional kernel; The size of the convolutional kernel for each convolutional layer is (3, 1), the stride is 1, and the padding method is the padding method (same) that keeps the output feature map size consistent with the input size. The number of output channels for each convolutional layer is 64, and the activation function is the ReLU function;
[0141] The outputs of the three parallel dilated convolutions are concatenated in the channel dimension to form a joint feature matrix X concat , whose shape is (L, 192, 1). The concatenated feature matrix is dimensionally reduced through a 1×1 convolutional layer to obtain a global feature matrix X gap , whose shape is (L, 64);
[0142] 3) Use the frequency-domain channel attention mechanism in the feature capture module to capture the implicit periodicity and long-range dependence in the global feature matrix X gap :
[0143] Apply the fast Fourier transform to the global feature matrix X gap to convert it from the spatial domain to the frequency domain, obtaining a frequency-domain feature vector F:
[0144]
[0145] Wherein, F(u, v) is the value of the frequency-domain feature matrix at position (u, v), where u and v represent the u-th row and v-th column of the frequency-domain feature matrix, X gap (x, y) is the value of the global feature matrix at position (x, y), where x and y represent the x-th row and y-th column of the global feature matrix, H and W are the height and width of the feature matrix, and j is the imaginary unit;
[0146] Extract the amplitude spectrum A from the frequency-domain feature F:
[0147] A(u, v) = |F(u, v)|
[0148] where \(A(u, v)\) is the value of the frequency-domain amplitude spectrum at position \((u, v)\);
[0149] Perform global pooling on the frequency-domain amplitude spectrum \(A\) to obtain the channel-level frequency-domain feature \(S\):
[0150]
[0151] 4) Input the frequency-domain feature \(S\) into a two-layer fully connected neural network to generate the attention weight \(Attention(S)\), and apply the generated channel attention weight \(Attention(S)\) to the original feature:
[0152] \(Attention(S)=\sigma(W 2 \cdot ReLU(W 1 \cdot S))
[0153] F' = Attention(S)\cdot F
[0154] where \(Attention(S)\) is the attention weight of the frequency-domain feature \(S\), \(W 1 and \(W 2 are the weight matrices of the two-layer fully connected neural network, \(ReLU\) is the activation function, and \(\sigma\) is the Sigmoid function; the input dimension of the first-layer fully connected neural network is 64, the output dimension is 32, and the activation function is the \(ReLU\) function, which is used to introduce non-linearity. The input dimension of the second-layer fully connected neural network is 32, the output dimension is 64, and the activation function is the Sigmoid function, which is used to limit the output within the range of \([0, 1]\). \(F'\) is the weighted channel feature, and its shape is \((L, 64)\);
[0155] 5) Use a frequency-domain variational autoencoder to perform feature compression and latent variable modeling on the weighted channel feature \(F'\) to generate the latent variable matrix \(Z\):
[0156] The encoder of the frequency-domain variational encoder uses a two-layer fully-connected neural network. The input dimension of the first-layer fully-connected neural network is 64, the output dimension is 32, and the activation function is the ReLU function. The input dimension of the second-layer fully-connected neural network is 32, the output dimension is 16, and the activation function is the Linear function. The encoder of the frequency-domain variational encoder performs a non-linear transformation on the channel feature F', extracts its potential low-dimensional representation, generates the mean and variance after passing through each layer of the encoder of the frequency-domain variational encoder, and then obtains the latent variable matrix Z through reparameterization, whose shape is (L, 16). The decoder also uses a two-layer fully-connected neural network. The input dimension of the first-layer fully-connected neural network is 16, the output dimension is 32, and the activation function is the ReLU function. The input dimension of the second-layer fully-connected neural network is 32, the output dimension is 64, and the activation function is the Linear function. The loss function of the encoder and decoder of the frequency-domain variational encoder consists of the reconstruction loss and the KL divergence loss:
[0157] L total = L recon + 0.1·L KL
[0158] In the formula, L total represents the total loss function, L recon represents the difference between F' and the reconstructed feature matrix generated by the decoder, which is calculated using the mean squared error, and L KL represents the KL divergence loss;
[0159] 6) Residue-level identification:
[0160] The latent variable matrix Z generated by the frequency-domain variational encoder is input into the XGBoost model as the feature of each residue for residue-level identification. The identification target value of the XGBoost model is the binary label vector Y with the same length as the number of amino acids in the currently trained protein θ , where the value range of Y θ is in {0, 1}. When Y θ = 0, it means that the θ-th residue is a non-allosteric site, and when Y θ = 1, it means that the θ-th residue is an allosteric site. The parameter settings of the XGBoost model are: the maximum tree depth is 6, the learning rate is 0.1, the number of trees is 100, the subsample ratio is 0.8, and the regularization parameters are γ = 0.1 and λ = 1. Finally, the prediction module determines whether a residue is an allosteric site according to the probability value output by the XGBoost model. If the probability that a certain residue is predicted to be an allosteric site is greater than 0.5, the prediction module will identify it as an allosteric site.
[0161] Specifically, the identification module performs the following operations:
[0162] 1) Read the PDB file of the input protein and parse it using the Bio.PDB toolkit in the Python environment to extract the three-dimensional structure information of the protein; store the three-dimensional structure information of the parsed protein in the variable protein_structure;
[0163] 2) Use the data import module to retain the spatial structure information of the peptide chains related to the allosteric site in the protein_structure obtained in step 1) and generate the spatial structure data of the protein related only to the allosteric site; store the processed spatial structure data of the protein in the variable processed_protein_data as the input for feature extraction;
[0164] 3) Use the feature extraction module to extract features from the spatial structure data of the protein stored in the variable processed_protein_data in step 2), call the DSSP tool, residue contact network, and Feature software to extract secondary structure features, contact network features, and physicochemical features of the residue microenvironment, and store them in the variables dssp_feature, residue_connect_feature, and mirco_env_feature respectively. Use the pandas toolkit to concatenate dssp_feature, residue_connect_feature, and mirco_env_feature to generate the feature matrix feature_matrix of the protein;
[0165] 4) Input the generated feature matrix feature_matrix into the trained improved model for prediction. The model combines local and global feature distributions, captures the spatial dependencies and constraints of residues, and obtains the probability P that each residue is an allosteric site. The value range of P is [0, 1], indicating the likelihood that the residue is an allosteric site. Store the allosteric site probabilities of each residue in the prediction probability list prediction_probability_list;
[0166] 5) For each residue in the prediction probability list prediction_probability_list with the allosteric site probability P, a set probability threshold T = 0.5 is used to classify each residue: If the probability P of a certain residue is P ≥ T, it is classified as an allosteric site with a label of 1; if the probability P of a certain residue is P < T, it is classified as a non-allosteric site with a label of 0; the classified labels are stored in the prediction label list prediction_list; finally, based on the probability value of the residue in the prediction probability list prediction_probability_list and the label in the prediction label list prediction_list, it is explained whether it is an allosteric site.
[0167] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A protein allosteric site recognition system based on deep learning, characterized in that: include: The data import module is used to load the PDB file of proteins with allosteric effects and the CSV file of their allosteric site location information, and perform preprocessing to obtain the spatial structure data of the protein; The feature extraction module uses the DSSP tool, residue contact network and Feature software to extract secondary structure features, contact network features and residue microenvironment physicochemical features from the spatial structure data of proteins, and fuses the secondary structure features, contact network features and residue microenvironment physicochemical features to generate a feature matrix; The training module, based on the improved atrous convolutional neural network (DCNN), learns the local and global feature distribution of each allosteric protein from the fused feature matrix, captures the feature information of the allosteric site, learns the constraints of the allosteric site and the complex spatial dependencies, and finally obtains a trained improved model; the improved model is an improvement on the feature capture module and prediction module of the atrous convolutional neural network (DCNN); The improvements to the feature capture module are: introducing three layers of parallel dilated convolution to extract features of large, medium, and small receptive fields respectively, and combining the frequency domain channel attention mechanism to enhance the capture of periodic and long-distance dependency features through frequency domain feature conversion and global pooling; The improvements to the prediction module are: introducing the XGBoost model, using the method of connecting multiple decision trees in series, and weighting and regularizing the extracted key features to improve the prediction accuracy and generalization ability of the model; The recognition module uses the trained improved model to identify the allosteric sites on the protein according to the PDB file of the protein and provides identification explanations.
2. A protein allosteric site recognition system based on deep learning according to claim 1, characterized in that: The data import module includes a data loading module and a data preprocessing module, wherein: The data loading module uses Bio.PDB and pandas toolkits in the Python environment to load the PDB file of the protein with allosteric effect and the CSV file of the allosteric site position information of the corresponding protein from the local, and parses the PDB file to obtain the three-dimensional structure information of the protein; The data preprocessing module uses Bio.PDB and pandas toolkits to retain the spatial structural information of the peptide chain where the allosteric site is located in the three-dimensional structural information according to the allosteric site position information provided by the CSV file, and generates spatial structural data of proteins only related to the allosteric site for subsequent feature extraction and model training.
3. A protein allosteric site recognition system based on deep learning according to claim 2, characterized in that: The feature extraction module performs the following operations: 1) Extraction of secondary structure features based on DSSP tools: The DSSP tool analyzes the spatial structure data of proteins and identifies the secondary structure type of each residue and its corresponding sequence position and chain identifier based on the secondary structure prediction method. First, the spatial structure data of the protein is input into the DSSP tool to generate a DSSP file containing the secondary structure characteristics of the protein. Then, by analyzing the spatial structure data of the protein and the secondary structure characteristics in the DSSP file, the secondary structure characteristics of the protein pocket region are extracted, including the secondary structure type of the residue and the relative solvent accessibility RSA. Subsequently, the Bio.PDB.DSSP module in Biopython is used to parse the DSSP file, convert the secondary structure type of each residue into a one-hot encoding, and calculate and normalize its relative solvent accessibility RSA for the subsequent construction of the feature matrix. Finally, the generated DSSP file will be retained so that the secondary structure characteristics in the residue microenvironment can be further extracted through the Feature software. 2) Extract the atomic contact network characteristics based on the residue contact network, that is, the close-range interaction characteristics: First, the contact network of the protein was generated by Biopython and NetworkX tools. The spatial coordinates of each residue were parsed from the PDB file of the protein, and the α carbon atom of each residue was selected as the representative atom. Based on the distance relationship in three-dimensional space, the contact network was divided into in It is a unit that describes the distance between atoms. It searches for contact residues and builds a contact network between α carbon atoms and α carbon atoms: Where, d k,c (R) is the residue contact density of the k-th residue R with c as the threshold, len(P) is the sequence length of protein P, C k,c (R) is the total number of contacts when the residue R at position k is set as the threshold value; Next, according to the spatial coordinates of the α carbon atom of residue R, the residue R contacts are divided into the upper hemisphere and the lower hemisphere. The number of upper and lower hemisphere contacts of each residue at different thresholds is counted, and the exposure ratio is calculated: D k,c (R)=1-U k,c (R) Where U k,c (R) is the upper hemisphere exposure ratio of residue R at position k with c as the threshold, D k,c (R) is the exposure ratio of the lower hemisphere when the residue R at position k is set as the threshold, C u,k,c (R) is the number of contacts in the upper hemisphere when the residue R at position k is set as the threshold, C k,c (R) is the total number of contacts in the upper and lower hemispheres when the residue R at position k is set as the threshold; In the contact network, the local network properties of each residue are further analyzed, and the clustering coefficient and betweenness centrality of each residue are calculated: In the formula, C c (R) is the clustering coefficient of residue R, T(R) is the number of triangles in the network passing through residue R, deg(R) is the degree of residue R in the network, is the theoretical maximum number of triangles for R residues; In the formula, C B (R) is the betweenness centrality of residue R, V is the set of nodes, s and t are any two nodes in the network, σ(s,t) is the total number of shortest paths from node s to node t, and σ(s,t|R) is the number of shortest paths through node s, node t and residue R; in this network, it is defined that if σ(s,t) = 0, then Through the residue contact network, the characteristics of the close-range interactions of residues in the protein contact network are extracted; the residue contact density reveals the local contact of the residues, the hemispheric exposure ratio reflects the exposure characteristics of the residues in space, and the clustering coefficient and betweenness centrality characterize the local and global importance of the residues in the contact network from the perspective of network properties; 3) Extraction of the physicochemical characteristics of the residue microenvironment based on Feature software: Microenvironment sampling is performed by the featurize module of the Feature software. The protein structure is analyzed at the atomic level. Microenvironment sampling is performed on the area centered on the α carbon atom of each residue. In the microenvironment sampling, the spatial range of the microenvironment is defined as a radius of and spherical shell; Use the Atomselect module of the Feature software to select the alpha carbon atom of the target residue and characterize the microenvironment of each residue one by one. The characterization content includes the physical, chemical and structural properties within the sphere and spherical shell: Extract the number of various elements in the spherical shell or sphere according to the element level, including: the number of any element ELEMENT_IS_ANY, the number of carbon elements ELEMENT_IS_C, the number of nitrogen elements ELEMENT_IS_N, the number of oxygen elements ELEMENT_IS_O, the number of sulfur elements ELEMENT_IS_S, and the number of other elements except carbon, nitrogen, oxygen and sulfur ELEMENT_IS_OTHER; Extract the number of various atoms in a shell or sphere at the atomic level, including: The number of carbon atoms connected to oxygen atoms ATOM_TYPE_IS_C, the number of terminal carbon atoms on the side chain ATOM_TYPE_IS_CT, the number of carbon atoms connected to the amino carbon atom ATOM_TYPE_IS_CA, the number of nitrogen atoms on the amino group ATOM_TYPE_IS_N, the number of atoms with the structural atom identifier N2 in the PDB file ATOM_TYPE_IS_N2, the number of atoms with the structural atom identifier N3 in the PDB file ATOM_TYPE_IS_N3, the number of nitrogen atoms connected to the alpha carbon atom ATOM_TYPE_IS_NA, the number of double-bonded oxygen atoms connected to the α carbon atom ATOM_TYPE_IS_O, the number of double-bonded oxygen atoms of the carboxyl group of the α carbon atom closest to the α carbon atom on the side chain ATOM_TYPE_IS_O2, the number of hydroxyl hydrogen atoms connected to the α carbon atom ATOM_TYPE_IS_OH, the number of sulfur atoms ATOM_TYPE_IS_S, the number of hydrogen atoms on sulfur ATOM_TYPE_IS_SH, the number of all atoms except the above atoms ATOM_TYPE_IS_OTHER; Extract various charges in the spherical shell or sphere according to the residue level, including: partial charge of the atom PARTIAL_CHARGE, negative charge NEG_CHARGE, positive charge POS_CHARGE, total charge considering histidine CHARGE_WITH_HIS, total charge excluding histidine CHARGE; Extract the number of various residues in the shell or sphere at the atomic level, including: The number of alanine RESIDUE_NAME_IS_ALA, the number of arginine RESIDUE_NAME_IS_ARG, the number of asparagine RESIDUE_NAME_IS_ASN, the number of aspartic acid RESIDUE_NAME_IS_ASP, the number of cysteine RESIDUE_NAME_IS_CYS, the number of glutamine RESIDUE_NAME_IS_GLN, the number of glutamic acid RESIDUE_NAME_IS_GLU, the number of glycine RESIDUE_NAME_IS_GLY, the number of histidine RESIDUE_NAME_IS_HIS, the number of isoleucine RESIDUE_NAME_IS_ILE, the number of leucine RESIDUE_NA ME_IS_LEU, the number of lysines RESIDUE_NAME_IS_LYS, the number of methionines RESIDUE_NAME_IS_MET, the number of phenylalanines RESIDUE_NAME_IS_PHE, the number of prolines RESIDUE_NAME_IS_PRO, the number of serines RESIDUE_NAME_IS_SER, the number of threonines RESIDUE_NAME_IS_THR, the number of tryptophans RESIDUE_NAME_IS_TRP, the number of tyrosines RESIDUE_NAME_IS_TYR, the number of valines RESIDUE_NAME_IS_VAL, the residue name does not belong to any of the above residue names RESIDUE_NAME_IS_OTHER; Based on the various residues and atomic properties in the microenvironment, the overall physicochemical properties within the sphere or shell are calculated, including: The number of residues that are hydrophobic RESIDUE_CLASS1_IS_HYDROPHOBIC, the number of residues that are charged RESIDUE_CLASS1_IS_CHARGED, the number of residues that are polar RESIDUE_CLASS1_IS_POLAR, the number of residues that are neither hydrophobic nor polar and are uncharged RESIDUE_CLASS1_IS_UNKNOWN, the number of residues with non-polarity RESIDUE_CLASS2_IS_NONPOLAR, the number of residues with hydrophobicity RESIDUE_CLASS2_IS_POLAR that are different from CLASS1, the number of residues with basicity RESIDUE_CLASS2_IS_BASIC, the number of residues with acidity RESIDUE_CLASS2_IS_ACIDIC, the number of residues that do not belong to the non-polar class and are not acidic or basic RESIDUE_CLASS2_IS_UNKNOWN; Combined with the secondary structure features provided by the DSSP tool, the number of various secondary structures contained in the microenvironment is counted, including: The number of secondary structures with three turns of alpha helix SECONDARY_STRUCTURE1_IS_3HELIX, the number of secondary structures with four turns of alpha helix SECONDARY_STRUCTURE1_IS_4HELIX, the number of secondary structures with five turns of alpha helix SECONDARY_STRUCTURE1_IS_5HELIX, the number of secondary structures with short segments connecting alpha helix and beta helix SECONDARY_STRUCTURE1_IS_BRIDGE, the number of secondary structures with beta-strand and anti-beta fold SECONDARY_STRUCTURE1_IS_STRAN, the number of secondary structures with turns SECONDARY_STRUCTURE1_IS_TURN, the number of secondary structures with regular bends SECONDARY_STRUCTURE1_IS_BEND, the secondary structures are not strand, helix, bend, bridge, turn, heterocycle The number of SECONDARY_STRUCTURE1_IS_COIL, the number of heteroatoms SECONDARY_STRUCTURE1_IS_HET, the number of secondary structures that cannot be recognized by DSSP SECONDARY_STRUCTURE1_IS_UNKNOWN, the number of secondary structures of type α helix SECONDARY_STRUCTURE2_IS_HELIX, the number of secondary structures of type β helix including BRIDGE, BEND, and TURN SECONDARY_STRUCTURE2_IS_BETA, the number of secondary structures of type random bend SECONDARY_STRUCTURE2_IS_COIL, the number of secondary structures of type heterocyclic structures SECONDARY_STRUCTURE2_IS_HET, the number of secondary structures that are not HELIX, BETA, COIL, or HET SECONDARY_STRUCTURE2_IS_UNKNOWN; Finally, the special structures and functional groups in the spherical shell or sphere are counted, including: HYDROXYL, AMIDE, AMINE, CARBONYL, RING_SYSTEM, PEPTIDE; and physical properties at the atomic level: van der Waals volume VDW_VOLUME, HYDROPHOBICITY, MOBILITY, SOLVENT_ACCESSIBILITY; Feature software associates protein function with its structure by analyzing the structure of proteins with similar functions. The core idea of Feature software is to analyze protein structure at the atomic level and sample the small spherical volume around each atom or a specified set of points. These volumes are called microenvironments. Microenvironments are represented by a real vector of feature vectors, which contain the physicochemical characteristic information within the sphere or shell. By extracting microenvironment information, complex three-dimensional structural data of proteins can be converted into numerical features that can be used in machine learning models. 4) Using the pandas toolkit in the Python environment, the secondary structure features, contact network features and residue microenvironment physicochemical features obtained in steps 1), 2) and 3) were spliced together to obtain the protein feature matrix; through the feature extraction module, 258 feature descriptions were generated for each residue of the protein.
4. A protein allosteric site recognition system based on deep learning according to claim 3, characterized in that: The training module performs the following operations: 1) Input the feature matrix into the embedding layer: The shape of the input feature matrix X of each protein is (L, 258, 1), where L is the number of residues in the protein, 258 represents the number of features extracted from each residue, and 1 represents single-channel input; for the 258 features of each residue, two layers of fully connected neural networks are used for embedding processing to reduce the dimension and introduce nonlinearity; the input dimension of the first layer of fully connected neural network is 258, and the output dimension is 128, and the input dimension of the second layer of fully connected neural network is 128, and the output dimension is 64. Both layers of fully connected neural networks use the ReLU function as the activation function, and the shape of the feature matrix after embedding is (L, 64); 2) Extract local and global features of the embedded feature matrix through the hole convolution layer: A dilated convolutional neural network (DCNN) consisting of three layers of parallel dilated convolutions is applied to the embedded feature matrix to extract local and global features. The shape of the input feature matrix is (L, 64, 1). The mathematical expression of the dilated convolution layer is as follows: Where O(i,j′) is the value of the output feature matrix at position (i,j′) after the dilated convolution, where i and j′ represent the i-th row and j′th column of the output feature matrix, W(m,n) is the weight of the convolution kernel at position (m,n), where m and n represent the m-th row and n-th column of the convolution kernel, I(i+m·r,j′+n·r) is the value of the input feature matrix at position (i+m·r,j′+n·r), r is the dilation rate, M c and N c Indicates the size of the convolution kernel; uses three parallel convolutional layers with different dilation rates, the dilation rates r are 4, 2, and 1 respectively. The dilated convolution layer with r = 4 is responsible for extracting features of large receptive fields, and the dilated convolution layer with r = 2 is responsible for extracting features of medium receptive fields. When r = 1, the dilated convolution degenerates into ordinary convolution. When r > 1, gaps are inserted between the elements of the convolution kernel. The convolution kernel size of each convolution layer is (3, 1), the stride is 1, and the padding method is the same as the padding method that keeps the output feature map size consistent with the input size. The number of output channels of each convolution layer is 64, and the activation function is the ReLU function. The outputs of the three layers of parallel dilated convolution are concatenated in the channel dimension to form a joint feature matrix X concat , whose shape is (L, 192, 1), and the concatenated feature matrix is reduced in dimension through a 1×1 convolutional layer to obtain the global feature matrix X gap , whose shape is (L,64); 3) In the feature capture module, the frequency domain channel attention mechanism is used to capture the global feature matrix X gap Implicit periodicity and long-range dependencies in : For the global feature matrix X gap Apply fast Fourier transform to convert it from spatial domain to frequency domain and obtain frequency domain feature vector F: In the formula, F(u,v) is the value of the frequency domain feature matrix at position (u,v), where u and v represent the uth row and vth column of the frequency domain feature matrix, and X gap (x, y) is the value of the global feature matrix at position (x, y), where x and y represent the xth row and yth column of the global feature matrix, H and W are the height and width of the feature matrix, and j is an imaginary unit; Extract the amplitude spectrum A from the frequency domain feature F: A(u,v)=|F(u,v)| Where A(u,v) is the value of the frequency domain amplitude spectrum at position (u,v); Perform global pooling on the frequency domain amplitude spectrum A to obtain the channel-level frequency domain features S: 4) Input the frequency domain feature S into the two-layer fully connected neural network to generate the attention weight Attention(S), and apply the generated channel attention weight Attention(S) to the original feature: Attention(S)=σ(W2·ReLU(W1·S)) F'=Attention(S)·F In the formula, Attention(S) is the attention weight of the frequency domain feature S, W1 and W2 are the weight matrices of the two-layer fully connected neural network, ReLU is the activation function, and σ is the Sigmoid function; the input dimension of the first layer of the fully connected neural network is 64, the output dimension is 32, and the activation function is the ReLU function, which is used to introduce nonlinearity. The input dimension of the second layer of the fully connected neural network is 32, the output dimension is 64, and the activation function is the Sigmoid function, which is used to limit the output to the range of [0,1]. F' is the weighted channel feature, and its shape is (L,64); 5) Use the frequency domain variational encoder to perform feature compression and latent variable modeling on the weighted channel feature F', and generate the latent variable matrix Z: The encoder of the frequency domain variational encoder uses a two-layer fully connected neural network. The input dimension of the first layer of the fully connected neural network is 64, the output dimension is 32, and the activation function is the ReLU function. The input dimension of the second layer of the fully connected neural network is 32, the output dimension is 16, and the activation function is the Linear function. The encoder of the frequency domain variational encoder performs a nonlinear transformation on the channel feature F' to extract its potential low-dimensional representation. After being processed by each layer of the encoder of the frequency domain variational encoder, the mean and variance are generated, and then the latent variable matrix Z is obtained by reparameterization, and its shape is (L, 16). The decoder of the frequency domain variational encoder also uses a two-layer fully connected neural network. The input dimension of the first layer of the fully connected neural network is 16, the output dimension is 32, and the activation function is the ReLU function. The input dimension of the second layer of the fully connected neural network is 32, the output dimension is 64, and the activation function is the Linear function. The loss function of the encoder and decoder of the frequency domain variational encoder consists of reconstruction loss and KL divergence loss: L total =L recon +0.1 L KL Where, L total Represents the total loss function, L recon represents the difference between F' and the reconstructed feature matrix generated by the decoder, calculated using the mean square error, L KL represents KL divergence loss; 6) Residue level identification: The latent variable matrix Z generated by the frequency domain variational encoder is input into the XGBoost model as the feature of each residue for residue-level recognition; the recognition target value of the XGBoost model is a binary label vector Y with a length equal to the number of amino acids in the currently trained protein θ , where Y θ The value range of Y is in {0,1}. θ = 0, indicating that the θth residue is a non-allosteric site. θ =1 indicates that the θth residue is an allosteric site; the parameter settings of the XGBoost model are: the maximum tree depth is 6, the learning rate is 0.1, the number of trees is 100, the subsample ratio is 0.8, and the regularization parameters are γ=0.1 and λ=1; finally, the prediction module determines whether the residue is an allosteric site based on the probability value output by the XGBoost model. If the probability of a residue being predicted as an allosteric site is greater than 0.5, the prediction module will identify it as an allosteric site.
5. A protein allosteric site recognition system based on deep learning according to claim 4, characterized in that: The identification module performs the following operations: 1) Read the PDB file of the input protein and parse it using the Bio.PDB toolkit in the Python environment to extract the three-dimensional structure information of the protein; store the parsed three-dimensional structure information of the protein in the variable protein_structure; 2) Using the data import module, retain the peptide chain spatial structure information related to the allosteric site in the protein_structure obtained in step 1), and generate the spatial structure data of the protein only related to the allosteric site; store the processed protein spatial structure data in the variable processed_protein_data as the input of feature extraction; 3) Use the feature extraction module to extract features from the spatial structure data of the protein stored in the variable processed_protein_data in step 2), call the DSSP tool, residue contact network, and Feature software to extract secondary structure features, contact network features, and residue microenvironment physicochemical features, which are stored in variables dssp_feature, variable residue_connect_feature, and variable mirco_env_feature, respectively, and use the pandas toolkit to splice dssp_feature, residue_connect_feature, and mirco_env_feature to generate the protein feature matrix feature_matrix; 4) Input the generated feature matrix feature_matrix into the trained improved model for prediction. The model combines local and global feature distributions, captures the spatial dependence and constraints of residues, and obtains the probability P of each residue being an allosteric site. The value range of P is [0,1], indicating the possibility that the residue is an allosteric site. The allosteric site probability of each residue is stored in the prediction probability list prediction_probability_list; 5) For each residue in the prediction probability list prediction_probability_list, the probability threshold T=0.5 is set to classify each residue: if the probability P of a residue is ≥T, it is classified as an allosteric site with a label of 1; if the probability P of a residue is <T, it is classified as a non-allosteric site with a label of 0; the classified labels are stored in the prediction label list prediction_list; finally, whether the residue is an allosteric site is interpreted based on the probability value of the residue in the prediction probability list prediction_probability_list and the label in the prediction label list prediction_list.
Citation Information
Patent Citations
Protein interaction site identification method
CN110265085A
Protein secondary structure prediction method based on multi-scale convolution attention neural network
CN112767997A
Protein-protein interaction site prediction method based on deep learning and XGBoost
CN113611360A
Method for identifying protein allosteric modulator based on deep learning and computational simulation
CN115938488A
Determining variant pathogenicity based on images
CN118648063A