Heterodimer interchain residue contact prediction method
Through the method of combining deep learning with protein language model, using multiple feature fusion and attention mechanisms, the problem of low prediction accuracy of residue contact between heterodimer chains in the prior art is solved, and higher prediction accuracy and more accurate contact mode capture is achieved.
Patent Information
- Application Number
- CN202510230527.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing heterodimer inter-chain residue contact prediction methods have low accuracy and face problems such as MSAs pairing, noise processing, and model integration.
A method of combining deep learning with protein language models is used to predict residue contact between heterodimer chains through multiple feature fusion and attention mechanisms such as channel attention and spatial attention. The specific steps include screening heterodimer structure information from the protein database, extracting monomer structure and sequence information, calculating the Euro-type distance, extracting multi-sequence alignment information and row attention characteristics using tools such as HHblits and ESM-MSA, fusing the feature matrix, and inputting the depth residual network and KAN convolutional network for prediction.
It significantly improves the accuracy of residue contact prediction between heterodimer chains, which is better than existing methods (such as DeepComplex, GLINTER, CDPred, DeepInter), and can more accurately capture the contact patterns between protein chains.
Smart Images

Figure CN120072030A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to protein-protein interactions, and in particular to a method for predicting inter-chain residue contacts of heterodimers. Background Art
[0002] Protein heterodimers are complexes composed of two different types of protein monomers (varying in length and structure). Heterodimers usually consist of protein chains with different functional and structural characteristics, and these different protein monomers are often functionally complementary. Their interaction is crucial for achieving specific biological functions. For example, many key biomolecules such as receptors, enzymes, and transcription factors exist in the form of heterodimers, and they regulate cell signal transduction, metabolic processes, and gene expression by forming heterodimer complexes. The formation of heterodimers is usually driven by protein-protein interactions (PPIs), and these interactions often have high specificity and selectivity. The interaction between protein chains not only depends on electrostatic forces, hydrogen bonds, and hydrophobic interactions, but also involves the precise docking of molecular structures. Therefore, understanding and predicting the residue contacts between heterodimer chains is of great significance for in-depth understanding of the structure-function relationship of protein complexes. Compared with homodimers, predicting inter-chain contacts of heterodimers is more difficult and faces greater challenges because there is more heterogeneity between the two chains of heterodimers, and the differences in their structural and functional characteristics make the interaction pattern more complex.
[0003] Judging from the currently published methods for predicting inter-chain residue contacts of heterodimers, the overall prediction accuracy is relatively low, and the accuracy is lower than 50% for most prediction samples. The reasons are as follows: First, it is difficult to construct high-quality combined MSAs. Especially when the sequence differences of heterodimers between different species are large, it is often difficult to find enough homologous sequences. When the quality of MSAs is poor and the sequence pair matching is inaccurate, the prediction results will be affected by noise. Second, the existing models are not well integrated. Different models are good at capturing different features, but there is a lack of integration of existing models. Generally speaking, although the existing methods for predicting inter-chain residue contacts of heterodimers have made progress to a certain extent, they still face problems such as MSAs pairing, noise processing, model integration, and relatively low overall prediction accuracy. Summary of the Invention
[0004] The present invention provides a method for predicting inter-chain residue contacts in heterodimers (AttCON-Hetero) based on deep learning and protein language models, which is used to predict the inter-chain residue contacts in heterodimers. The present invention processes the input features by means of multi-feature fusion. The network model combines the channel attention (ECA) and spatial attention (SA) modules and the KAN convolutional network module to achieve higher performance in predicting inter-chain residue contacts, providing very useful information for dimer structure prediction, function prediction, drug design, etc.
[0005] The technical solution of the present invention is: a method for predicting inter-chain residue contacts in heterodimers, comprising the following steps:
[0006] S1. Screen out the heterodimer structure information Heterodimer-PDB from the protein database;
[0007] S2. Divide the data obtained in step S1 into a training set, a validation set and a test set. According to the heterodimer structure information, preprocess the Heterodimer-PDB structure data, and extract the structure information MonomerA-PDB / MonomerB-PDB and sequence information MonomerA–FATSA / MonomerB-FASTA of the two monomers of the heterodimer. The sequence lengths are L A and L B , and directly concatenate them into a joint sequence FASTA with a length of L A +L B . Calculate the Euclidean distance between the amino acids of the two monomers of the heterodimer according to the heterodimer structure information Heterodimer-PDB, so as to obtain the distance contact between the heterodimers (Inter-Chain Contact, matrix size is L A ×L B );
[0008] S3. Use the two monomer sequence information MonomerA–FATSA / MonomerB-FASTA generated in step S2, and use the HHblits multiple sequence alignment search algorithm to search the Uniclust30 sequence database to obtain the multiple sequence alignment information MSAs of these two monomers respectively, and use the TaxID species classification number to match the MSAs of these two monomers to ensure that the selected homologous sequences come from the same species, and finally generate a joint multiple sequence alignment Paired MSAs; this method can integrate the information from the two monomers and capture the co-evolution signal between the two.
[0009] S4. Using the paired multiple sequence alignment information Paired MSAs generated in step S3, extract row attention features using two protein language models, ESM-MSA and ESM2, respectively; extract co-evolution information features in Paired MSAs using the CCMpred algorithm; calculate the distance map within the amino acid chain of each monomer using the structural information of the two monomers of the heterodimer in step S2; use the combined sequence FASTA obtained in step S2 and generate a position-specific scoring matrix (PSSM) using the PSI-BLAST tool.
[0010] S5. Fuse the seven types of features generated in step S4 to generate a fused feature matrix, enhancing the diversity of the input features.
[0011] S6. Input the features processed in step S5 into the model AttCON-Hetero composed of a deep residual network ResNet and a KAN (Kolmogorov-Arnold Networks) convolutional network. The model uses an efficient channel attention ECA and a spatial attention SA module to capture local and global dependence features of the feature matrix, and at the same time uses the KAN convolution that can adaptively adjust the convolution kernel weights to output the predicted contact probability matrix. The final output result is the probability matrix of amino acid contacts between the two chains of the heterodimer protein (matrix size is L A ×L B ).
[0012] S7. Based on the network output in step S6, use the inter-chain contact matrix Inter-ChainContact of the heterodimer in step S2 as the label for training. Subsequently, use the validation set in S2 to verify the performance of the model obtained in each batch of training, and save the model parameters with the best metrics on the validation set as the weights of the final classification model. Finally, verify the performance of the saved model on the test set in S2.
[0013] S1. Screen out heterodimers from the PDB database that meet the following conditions: (1) The aggregation state should be a heterodimer; (2) The biological assembly should only contain two single chains; (3) The resolution of the X-ray structure should be better than (4) The length of the single chain should be greater than 50, and the total length of the complex should be less than 500. (5) Cluster the heterodimers using a 40% sequence similarity threshold and select the heterodimer with the best resolution from each cluster.
[0014] As a further description of the above solution: S2 extracts the monomer structure information MonomerA-PDB / MonomerB-PDB and sequence information of the heterodimer, and calculates the contact matrix between pairwise amino acids of the two monomers of the heterodimer, including the following steps:
[0015] S2.1. Extract the information of heterodimer monomers and sequence information:
[0016] Extract the structural information and sequence information of the two chains of the heterodimer: For the eligible dimers, extract the three-dimensional structural information of the two chains from their PDB files, and save them as Monomer-PDB A and Monomer-PDB B respectively. At the same time, extract their amino acid sequence information MonomerA–FATSA and MonomerB-FASTA, and combine MonomerA–FATSA and MonomerB-FASTA into one sequence FASTA; during the extraction process, remove the residues with amino acid label UNK.
[0017] S2.2. Calculate the Euclidean distance and contact matrix between amino acid residues between the heterodimer chains:
[0018] For residues i and j (i from monomer A; j from monomer B), define two types of distance calculation rules:
[0019]
[0020] i∈MonomerA,j∈MonomerB (1)
[0021] In formula (1), is the closest heavy atom distance, where r a and r b represent the three-dimensional coordinates of the heavy atoms in residues i and j respectively, and ||r a -r b || is the Euclidean distance of the closest heavy atom between residue i and residue j.
[0022]
[0023] In formula (2), is the distance between the Cβ atoms in the middle of residues i and j (C β -distance), is the atomic coordinate of the C β atom of residue i; calculate the two types of contact matrices respectively according to formula (1) and formula (2). If or is less than or equal to , then the value of the corresponding position in the contact matrix is 1, otherwise it is 0, and save them as Inter-Chain Contact(Heavy Atom) and Inter-Chain Contact(C β ), and the matrix size is L A ×L B, the above matrix is used as the label of the model for comparing experimental results.
[0024] As a further description of the above solution: S3 specifically includes the following steps:
[0025] S3.1. Use HHblits to search the Uniclust30 database to obtain MSAs:
[0026] Using the monomer sequence information generated in step S2.1, use the HH-suite 3.30 tool to perform a preliminary search on the Uniclust30 database;
[0027] S3.2. Match the MSAs of the two monomers to generate a paired multiple sequence alignment (Paired MSAs):
[0028] Use the TaxID of each monomer's species classification to match the MSAs of monomer A and monomer B; ensure that the two monomers come from the same biological species or related homologous species; according to the TaxID information, integrate the matching results of monomer A and monomer B to generate a paired multiple sequence alignment Paired MSAs; this combined MSAs contains homologous sequence information from the two monomers, providing high-quality input data for subsequent feature extraction.
[0029] As a further description of the above solution: The detailed steps for S4 to extract multiple features using the paired multiple sequence alignment information (Paired MSAs) and other feature sources are as follows:
[0030] S4.1. Extract row attention features (F ESM-MSA and F ESM2 )
[0031] F ESM-MSA That is, the ESM-MSA row attention feature: Based on the paired multiple sequence alignment information (Paired MSAs) generated in step S3, use the ESM-MSA model (the model used in the present invention is esm_msa1_t12_100M_UR50S) to extract the row attention features in the sequence. The ESM-MSA model is specifically optimized for multiple sequence alignment, can capture the co-evolution signals and sequence differences in MSAs, and generate a high-quality row attention feature matrix (ESM-MSA Row Attention).
[0032] F ESM2That is, the ESM2 row attention feature: Using the joint sequence FASTA, the row attention feature is extracted through the ESM2 model (the model used in this invention is esm2_t6_8M_UR50D). The ESM2 model performs excellently in capturing the long-range dependencies and local properties of protein sequences, and the generated row attention feature matrix ESM2 RowAttention has strong generalization and prediction accuracy.
[0033] S4.2. Extract the co-evolution information features (F CCMpred , F DCA-APC and F DCA-DI )
[0034] F CCMpred That is, the CCMpred co-evolution feature: Based on Paired MSAs, the CCMpred algorithm is used to calculate the probability distribution of co-evolution between sequences, generating a diagonally symmetric co-evolution matrix;
[0035] F DCA-APC and F DCA-DI : Further based on the Direct Coupling Analysis (DCA) method, the DCA feature matrix (DCA-APC) with applied average product correction (APC) and the uncorrected DCA feature matrix (DCA-DI) are generated respectively. These features reveal the intensity and pattern of sequence co-evolution and are used to describe the interactions between residues;
[0036]
[0037] F DCA-DI = DCA(4)
[0038] The DCA in formula (3) is the result of direct coupling analysis. Row Mean, Col Mean, and Overall Mean are the row mean, column mean, and average of all elements of the DCA matrix respectively;
[0039] S4.3. Extract the intra-chain distance features of the two monomers (F intra )
[0040] Based on the two monomer structure information (MonomerA-PDB / MonomerB-PDB) generated in step S2, the Euclidean distance matrix (Intra-Chain Distance) between amino acids within the two monomers is calculated respectively; for each pair of residues (a, b) within the monomer chain, the distance can be expressed as:
[0041]
[0042] In formula (5), (x a , ya , z a ), and (x b , y b , z b ) are the three-dimensional coordinates of the a-th and b-th residues respectively; the intra-chain distance matrix generated by monomer A is D A , and the intra-chain distance matrix generated by monomer B is D B , which can be expressed as follows:
[0043]
[0044] S4.4. Generate the position-specific scoring matrix (F PSSM )
[0045] Using the combined sequence information (FASTA) in step S2, calculate the position-specific scoring matrix (PSSM) of the sequence through the PSI-BLAST tool; the PSSM captures the substitution patterns and conservation information of amino acids through multiple iterative searches of the UniProt database, and the generated matrix can be used for feature enhancement and model optimization;
[0046]
[0047] In formula (7), f ij is the frequency of amino acid j at position i, and b j is the background frequency.
[0048] As a further description of the above scheme: The detailed steps for generating the fusion features in S5 are as follows:
[0049] Concatenate the 7 features generated in step S4 along the channel dimension to generate the fusion feature F contact ;
[0050] F contact = [F ESM-MSA , F ESM2 , F CCMpred , F DCA-APC , F DCA-DI , F intra , F PSSM (8)
[0051] In formula (8), the dimension of F contact is (L A + L B ) × (L A + L B ) × C, where L A , L B are the sequence lengths of monomer A and monomer B respectively, and C represents the number of channels after fusion.
[0052] As a further description of the above solution: The construction of the network model based on deep learning and protein language model in S6 includes the following steps:
[0053] Input the enhanced feature matrix F generated in step S5 contact into the AttCON-Hetero model based on the deep residual network (ResNet) and KAN (Kolmogorov-Arnold Networks) convolution for feature learning. The dimension of the feature matrix is (L A +L B ) × (L A +L B ) × C; Use instance normalization to standardize the input feature matrix to reduce the distribution deviation of the input features; Compress the C feature channels through 2D convolution; At the same time, further extract the features of the channels through the Maxout layer to obtain X input ; Subsequently, input X input into the residual module. Each residual module consists of the following steps. Effectively extract local and global features through step-by-step standardization, convolution, and attention module, while maintaining the hierarchical consistency of the features;
[0054] First, standardize the input feature X input to unify the feature scale:
[0055]
[0056] Among them, Norm is the instance normalization operation. Subsequently, use 2D convolution to extract preliminary features:
[0057]
[0058] Among them, Conv2D is the 2D convolution operation. In order to capture multi-scale feature information, subsequently use three different sizes of convolution kernels to further extract features, and splice the three groups of features for fusion of different scale information:
[0059]
[0060] X multi =Contact(X 1 , X 2 , X 3 ) (14)
[0061] Among them, X 1 , X 2 , X 3is the result after extracting features with different convolutional kernels. Contact is the concatenation operation, and then X multi is normalized and two-dimensionally convolutionally compressed according to Formulas (9) and (10) to obtain Subsequently, the channel attention mechanism (Efficient Channel Attention, ECA) and the spatial attention mechanism (Spatial Attention, SA) are used for feature extraction:
[0062]
[0063] z = σ(Conv1D(y)) (16)
[0064]
[0065] f sa = Conv2D([AvgPool(X ECA ).;.MaxPool(X ECA )]) (18)
[0066] X SA = X ECA ⊙ σ(f sa ) (19)
[0067] Formulas (15) to (17) are the processing flow of the channel attention mechanism: GAP is global max pooling, which averages the spatial dimensions of each channel to generate a global feature representation; Conv1D is a one-dimensional convolutional operation; σ is the Sigmoid activation function; ⊙ represents element-wise multiplication. Formulas (18) to (19) are the spatial attention mechanism: AvgPool and MaxPool are average pooling and max pooling operations respectively; [·;·] represents the concatenation of the pooling results; X SA is the result after the spatial attention SA extracts features.
[0068]
[0069] Formula (20) adds the output of the attention mechanism and the input features through a residual connection to enhance the gradient flow and avoid the gradient vanishing problem;
[0070] Finally, it is the output layer of the model. X output is normalized according to Formula (9) to obtain X output_norm , and then the output is performed through two layers of KAN convolution. The KAN convolution ConvKAN adaptively adjusts the weights of the convolutional kernels, enabling the network to dynamically adjust its learning process according to the features of the input data, which has significant advantages for capturing the heterogeneity and complex contact patterns between protein chains;
[0071] X KAN = ConvKAN(X output_norm )(21)
[0072] ConvKAN is the KAN convolution, and X KAN is the result after KAN convolution processing. Use the KAN convolution layer and the Sigmoid classifier for the output of formula (21) to output the probability matrix (whether in contact), and the dimension of the output matrix is L A ×L B , representing the contact probability P contact :
[0073] P contact = Sigmoid(ConvKAN(X KAN ))(22)
[0074] As a further description of the above solution: The S7 model training process, model performance evaluation, and model testing steps are as follows:
[0075] S7.1 The use of the loss function and the tuning steps of the optimizer in model training
[0076] Take the contact prediction matrix output in step S6 as the output of the network, and use the inter-chain residue contact matrix (Inter-Chain Contact) obtained in step S2 as the label label (in the experiment, preferentially select Inter-Chain Contact (C β ) as the label. If there are problems in calculating some samples of Inter-Chain Contact (C β ), uniformly select Inter-Chain Contact (HeavyAtom) as the label) for model training; the training process is optimized through the Focal Loss function, and the AdamW optimizer is used in combination with the cosine annealing learning rate scheduler;
[0077] S7.2 Performance evaluation and model selection
[0078] After each batch of training, use the validation set provided in step S2 to evaluate the performance of the model, and evaluate the model performance by calculating the Top n accuracy, Top-L / k accuracy, and weighted evaluation index W score to evaluate the model performance, and finally select the best model parameters on the validation set.
[0079] The features of the present invention are as follows: Since there is more heterogeneity between the two chains of the heterodimer, the interaction mechanism is complex, and there is still great room for improvement in aspects such as feature integration and model optimization in the existing methods. Generally speaking, the current methods have low accuracy in predicting the inter-chain residue contacts of heterodimers. The present invention effectively integrates and fuses traditional co-evolution features, monomer structure information, and transformer row attention features generated by protein language models (ESM-MSA and ESM2). At the same time, the prediction model is optimized, and Efficient Channel Attention (ECA) and Spatial Attention (SA) attention modules are used to efficiently extract features, where ECA is used to weight channel features, and SA is used to capture spatial information and integrate channel features. Finally, the KAN convolutional network module is used to adaptively adjust the weights of the convolutional kernels, and efficiently output the inter-chain residue contact probability map. Compared with the existing prediction methods, more accurate inter-chain residue contact information can be obtained.
[0080] Compared with the prior art, the present invention has the following technical effects: By integrating the attention features generated by protein language models (ESM-MSA and ESM2), the present invention enhances the diversity of input data and the robustness of the model compared with traditional protein structure features (such as distance maps, co-evolution features, etc.). Subsequently, the network adopts efficient channel attention (ECA) and spatial attention (SA) modules and the KAN convolutional network module to effectively capture the local and global dependence features of heterodimers. Experiments show that the prediction accuracy of the present invention on the benchmark dataset is significantly better than that of the existing methods (such as DeepComplex, GLINTER, CDPred, DeepInter, etc.), and higher prediction performance can be achieved. The present invention can be widely applied to the fields of predicting inter-chain residue contacts of protein heterodimers and protein structure prediction, further promoting protein function research and protein drug development. Brief Description of the Drawings
[0081] Figure 1 is the flow chart of predicting the inter-chain residue contacts of the heterodimer of the proposed method AttCON-Hetero;
[0082] Figure 2 is the data processing flow chart of the residual module of the proposed method AttCON-Hetero;
[0083] Figure 3 is the data processing flow chart of the attention module of the proposed method AttCON-Hetero;
[0084] Figure 4 are the PDB quaternary structure diagrams of two heterodimer samples in the training set;
[0085] Figure 5It is a prediction map of inter-chain residue contacts of test set samples under different prediction methods. Detailed implementation manners
[0086] The present invention will be further described below in conjunction with the accompanying drawings and embodiments, but the content of the present invention is not limited to the described scope.
[0087] Embodiment 1
[0088] A method for predicting inter-chain residue contacts of a heterodimer, the flowchart is as Figure 1 shown, and specifically includes the following steps:
[0089] S1. Obtain heterodimer data from the PDB database. The three-dimensional structure schematic diagram of the heterodimer is as Figure 4 shown. In the figure, a and b are two different heterodimer samples. For example, the heterodimer in b is composed of two monomers in green and brown. The inter-chain residue contacts to be predicted are the contact information between the two monomers;
[0090] S2. Divide the heterodimer data that meets the conditions obtained from the PDB database into a training set, a validation set, and a test set. Table 1 shows the sample numbers of the training set, the validation set, and the test set. According to the information in these sets, obtain the PDB structures (MonomerA-PDB / MonomerB-PDB) and sequence information (MonomerA–FATSA / MonomerB-FASTA / FASTA) of the two monomers of each sample, as well as the contact matrix (Inter-ChainContact) of the inter-chain residues of each sample's heterodimer.
[0091] Table 1 Division of the data sets used in the experiment
[0092]
[0093]
[0094] S3. Using the monomer sequence information of all samples generated in step S2, use the HHblits multiple sequence alignment search algorithm to search the Uniclust30 sequence database to obtain the multiple sequence alignment information (MSAs) of all monomers, and use TaxID (species classification number) to match the MSAs of each heterodimer sample to ensure that the selected homologous sequences come from the same species and pair them according to the species information. Finally, a paired multiple sequence alignment (Paired MSAs) is generated for each sample.
[0095] S4. Using the joint multiple sequence alignment information (PairedMSAs) of each heterodimer sample generated in step S3, extract row attention features (ESM-MSAAttention, ESM2-Attention) using two protein language models (PLMs), ESM-MSA and ESM2, respectively; extract co-evolution information features (CCMpred, DCA-APC, DCA-DI) in the joint multiple sequence alignment information (Paired MSAs) using the CCMpred algorithm; calculate the intra-chain distance maps (Intra-ChainDistanceA and Intra-Chain DistanceB) within the amino acid chains of the monomers using the structural information (MonomerA-PDB / MonomerB-PDB) of the two monomers of the heterodimer in step S2; use the joint sequence information (FATSA) obtained in step S2 and generate a position-specific scoring matrix (PSSM) using the PSI-BLAST tool.
[0096] S5. Fuse the seven types of features of each sample generated in step S4 to generate a fused feature matrix, enhancing the diversity of the input features.
[0097] S6. Input the features of the training samples processed in step S5 into a model (AttCON-Hetero) composed of a deep residual network (ResNet, Figure 2 which is the network architecture diagram of the residual module) and a KAN convolutional network. This model adopts efficient channel attention (ECA) and spatial attention (SA) modules ( Figure 3 which is the network architecture of the attention module) to capture the local and global dependence features of the feature matrix, and at the same time uses the KAN convolution that can adaptively adjust the convolutional kernel weights to output the predicted contact probability matrix. The final output result is the probability matrix of pairwise amino acid contacts between the two chains of the heterodimer protein (the matrix size is L A ×L B ).
[0098] S7. Based on the network output in step S6, obtain the inter-chain residue contact prediction probability matrix. Use the inter-chain contact matrix (Inter-Chain Contact) of the heterodimers in the training samples obtained in step S2 as the label for training. Subsequently, use the validation set in S2 to verify the performance of the model obtained in each batch of training, and save the model parameters with the best metrics obtained on the validation set as the weights of the final classification model. Finally, verify the performance of the saved model on the test set in S2.
[0099] Furthermore, for the training set, validation set, and test set, extract the monomer structure information (MonomerA-PDB / MonomerB-PDB) and sequence information of the heterodimer, and calculate the Euclidean distance between the amino acids of the two monomers of the heterodimer, so as to obtain the contact label data (Inter-Chain Contact).
[0100] S2.1: Extract the three-dimensional structure information of two chains from the PDB files of the training set, validation set, and test set, and save them as Monomer-PDB A and Monomer-PDB B respectively. At the same time, extract their amino acid sequence information MonomerA–FATSA and MonomerB-FASTA, and combine MonomerA–FATSA and MonomerB-FASTA into one sequence FASTA. During the extraction process, remove the residues with amino acid label UNK.
[0101] S2.2: Use the heterodimer PDB files of the training set, validation set, and test set to calculate the Euclidean distance and contact matrix between the amino acid residues between the heterodimer chains:
[0102] For residues i and j (i from monomer A; j from monomer B), define two types of distance calculation rules:
[0103]
[0104] i∈MonomerA,j∈MonomerB(1)
[0105] In formula (1), is the closest heavy-atom distance, where r a and r b represent the three-dimensional coordinates of the heavy atoms in residues i and j respectively, and ||r a -r b || is the Euclidean distance of the closest heavy atom between residue i and residue j.
[0106]
[0107] In formula (2), is the distance between the Cβ atoms in the middle of residues i and j (C β -distance), is the C β atom coordinate of residue i; calculate the two types of contact matrices respectively according to formula (1) and formula (2). If or is less than or equal to If the value at the corresponding position of the contact matrix is 1, otherwise it is 0, which are saved as Inter-Chain Contact (Heavy Atom) and Inter-Chain Contact (C β ), and the matrix size is L A ×L B . The above matrices are used as the labels of the model for comparing experimental results.
[0108] Furthermore, use the HHblits multiple sequence alignment search algorithm on S3 to search the Uniclust30 sequence database to obtain the multiple sequence alignment information (MSAs) of two monomers for each dataset sample, and use TaxID (species classification number) to match the MSAs of these two monomers to generate a paired multiple sequence alignment (Paired MSAs), including the following steps:
[0109] S3.1. Use HHblits to search the Uniclust30 database to obtain MSAs:
[0110] Using the monomer sequence information generated in step S2.1, use the HH-suite 3.30 tool to perform a preliminary search on the Uniclust30 database. Uniclust30 is a database composed of UniProtKB sequences, and all sequences are clustered based on 30% pairwise similarity, which can capture rich co-evolution signals. By setting appropriate search parameters (in this invention, -diffinf-id 99 -cov 50 -n 3 are used), the MSAs information of two monomers are generated respectively.
[0111] S3.2. Match the MSAs of two monomers to generate a paired multiple sequence alignment (Paired MSAs):
[0112] Use the species classification number (TaxID) of each monomer to match the MSAs of monomer A and monomer B. TaxID is used to identify the species to which each monomer belongs. Through species information, ensure that the sequences of the two come from the same biological species or related homologous species. According to the TaxID information, integrate the matching results of monomer A and monomer B to generate a paired multiple sequence alignment (Paired MSAs). This paired MSA contains homologous sequence information from two monomers, providing high-quality input data for subsequent feature extraction.
[0113] Furthermore, the detailed steps of using paired multiple sequence alignment information (Paired MSAs) and other feature sources to extract various features for S4 are as follows:
[0114] S4.1. Extract row attention features (F ESM-MSA and FESM2 )
[0115] F ESM-MSA That is, the ESM-MSA row attention feature: Based on the paired multiple sequence alignment information (Paired MSAs) generated in step S3, use the ESM-MSA-1 model (a protein language model, in this invention, esm_msa1_t12_100M_UR50S is used) to extract the row attention features in the sequence. The ESM-MSA model is specifically optimized for multiple sequence alignment. By capturing the co-evolution signals and differences between sequences in the MSAs, it generates a high-quality row attention feature matrix (ESM-MSARowAttention).
[0116] F ESM2 That is, the ESM2 row attention feature: Using the combined sequence (FASTA), extract the row attention features through the ESM2 model (a protein language model, in this invention, esm2_t6_8M_UR50D is used). The ESM2 model performs excellently in capturing the long-range dependencies and local characteristics of protein sequences, and the generated row attention feature matrix (ESM2 Row Attention) has strong generalization and prediction accuracy.
[0117] S4.2. Extract co-evolution information features (F CCMpred 、F DCA-APC and F DCA-DI )
[0118] F CCMpred That is, the CCMpred co-evolution feature: Based on Paired MSAs, use the CCMpred algorithm to calculate the probability distribution of co-evolution between sequences, and generate a diagonally symmetric co-evolution matrix.
[0119] F DCA-APC and F DCA-DI : Further based on the Direct Coupling Analysis (DCA) method, use the CCMpred algorithm to generate the DCA feature matrix (DCA-APC) with application of average product correction (APC) and the uncorrected DCA feature matrix (DCA-DI) respectively. These features reveal the intensity and pattern of sequence co-evolution and are used to describe the interactions between residues.
[0120]
[0121] F DCA-DI =DCA (4)
[0122] In formula (3), DCA is the direct coupling analysis result, and Row Mean, Col Mean, and Overall Mean are the row mean, column mean, and the mean of all elements of the DCA matrix, respectively.
[0123] S4.3. Extract two intra-monomer chain distance features (F intra )
[0124] Based on the two monomer structure information (MonomerA-PDB / MonomerB-PDB) generated in step S2, calculate the Euclidean distance matrix (Intra-Chain Distance) between amino acids within each of the two monomers. For each pair of residues (a, b) within a monomer chain, the distance can be expressed as:
[0125]
[0126] In formula (5), (x a , y a , z a ) and (x b , y b , z b ) are the three-dimensional coordinates of the a-th and b-th residues, respectively. The intra-chain distance matrix generated by monomer A is D A , and the intra-chain distance matrix generated by monomer B is D B , which can be expressed as follows:
[0127]
[0128] S4.4. Generate a position-specific scoring matrix (F PSSM )
[0129] Using the combined sequence information (FASTA) in step S2, calculate the position-specific scoring matrix (PSSM) of the sequence through the PSI-BLAST tool. The PSSM captures the amino acid substitution patterns and conservation information by iteratively searching the UniProt database multiple times, and the generated matrix can be used for feature enhancement and model optimization.
[0130]
[0131] In formula (7), f ij is the frequency of amino acid j at position i, and b j is the background frequency.
[0132] Furthermore, the detailed steps for generating the fused features in S5 are as follows:
[0133] Concatenate the 7 features generated in step S4 along the channel dimension to generate the fused feature F contact .
[0134] F contact = [F ESM-MSA , F ESM2 , F CCMpred , F DCA-APC , F DCA-DI , F intra , F PSSM (8)
[0135] In formula (8), the dimension of F contact is (L A + L B ) × (L A + L B ), where L A , L B are the sequence lengths of monomer A and monomer B respectively, and C represents the number of channels after fusion. In the present invention, C is 308.
[0136] Furthermore, the network model based on deep learning and protein language model for S6 includes the following steps:
[0137] Input the feature matrix F contact generated in step S5 into the AttCON-Hetero model based on the deep residual network (ResNet) and KAN convolution for feature learning. The dimension of the feature matrix is (L A + L B ) × (L A + L B ), where L input . Subsequently, input X input into the residual module. Each residual module consists of the following steps, effectively extracting local and global features through step-by-step normalization, convolution, and attention module, while maintaining the hierarchical consistency of features.
[0138] First, perform normalization processing on the input. Normalize the input feature X input to unify the feature scale:
[0139]
[0140] where Norm is the instance normalization operation, and then use two-dimensional convolution to extract preliminary features:
[0141]
[0142] Among them, Conv2D is a two-dimensional convolution operation. In order to capture multi-scale feature information, convolution kernels of three different sizes are subsequently used to further extract features, and the three groups of features are concatenated to fuse different scale information:
[0143]
[0144] X multi = Contact(X 1 , X 2 , X 3 ) (14)
[0145] Among them, X 1 , X 2 , X 3 are the results after feature extraction by different convolution kernels. Contact is a concatenation operation. Subsequently, X multi is normalized and compressed by two-dimensional convolution with reference to formulas (9) and (10) to obtain Subsequently, an efficient channel attention (ECA) mechanism and a spatial attention (SA) mechanism are used for feature extraction:
[0146]
[0147] z = σ(Conv1D(y)) (16)
[0148]
[0149] f sa = Conv2D([AvgPool(X ECA ) ; MaxPool(X ECA )]) (18)
[0150] X SA = X ECA ⊙ σ(f sa ) (19)
[0151] Formulas (15) to (17) are the processing flow of the channel attention mechanism: GAP is global max pooling, which averages the spatial dimension of each channel to generate a global feature representation; Conv1D is a one-dimensional convolution operation; σ is the Sigmoid activation function; ⊙ represents element-wise multiplication. Formulas (18) to (19) are the spatial attention mechanism: AvgPool and MaxPool are average pooling and max pooling operations respectively; [· ; ·] represents the concatenation of pooling results; XSA It is the result after the spatial attention SA extracts features.
[0152]
[0153] Equation (20) adds the output of the attention mechanism and the input features through a residual connection to enhance the gradient flow and avoid the problem of gradient vanishing.
[0154] Finally, it is the output layer of the model. Standardize X output with reference to Equation (9) to obtain X output_norm , and then output through two layers of KAN convolution. The KAN convolution (Convolutional Kolmogorov - Arnold Network, ConvKAN) adaptively adjusts the weights of the convolution kernel, enabling the network to dynamically adjust its learning process according to the features of the input data, which has significant advantages for capturing the heterogeneity and complex contact patterns between protein chains.
[0155] X KAN = ConvKAN(X output_norm ) (21)
[0156] ConvKAN is the KAN convolution, and X KAN is the result after being processed by the KAN convolution. Use the KAN convolution layer and the Sigmoid classifier for the output of Equation (21) to output the probability matrix (whether in contact), and the dimension of the output matrix is L A ×L B , representing the contact probability P contact :
[0157] P contact = Sigmoid(ConvKAN(X KAN )) (22)
[0158] Furthermore, during the overall training process of the model, after calculating the loss, use the backpropagation algorithm for optimization, select the AdamW optimizer for optimization, and use the cosine annealing learning rate scheduler to dynamically adjust the learning rate.
[0159] (1) Loss function: FocalLoss, used to handle the imbalance of label data:
[0160] Focal Loss = -α(1 - p y ′) γ log p y ′ (23)
[0161] Among them, α is the loss weighting for different category samples, and p y' is the probability that the model predicts the contact category. γ is an adjustable parameter used to control the weight decay degree of easily classified samples (non-contact). In the present invention, α is 0.75 and γ is 2.
[0162] (2) Topn and Top-L / k prediction accuracies:
[0163]
[0164] Top n refers to the top n contacts with the highest probability values in the predicted inter-chain residue contact probability matrix. The Top n prediction accuracy is the ratio of true contacts among the top n contacts with the highest predicted probability values. In the present invention, n takes values of 1, 10, 25, and 50. Top-L / k refers to the top L / k contacts with the highest probability values in the predicted inter-chain residue contact probability matrix. The Top-L / k prediction accuracy is the ratio of true contacts among the top L (where L is the length of the monomer sequence with the smaller length among the two monomers of the heterodimer) contacts with the highest predicted probability values. In the present invention, k takes values of 1, 5, and 10.
[0165] (3) Weighted evaluation parameter W score :
[0166]
[0167] where α and β are the weight parameters of the Top n prediction accuracy and the Top-L / k prediction accuracy respectively. In the present invention, both α and β are set to 1.
[0168] Furthermore, in order to prove the effectiveness of the method proposed in the present invention, the present invention conducted a comparative experiment on the test data set in Table 1, and the experimental results are shown in Table 2.
[0169] Table 2 Results of the comparison between the present invention and existing advanced methods on the test data set
[0170]
[0171] Table 2 shows the performance of multiple models (DeepComplex, GLINTER, CDPred, DeepInter, and AttCON-Hetero) on the test data set in Table 1. The prediction results evaluate their performance on different metrics, including the prediction accuracies of "top1", "top10", "top25", "top50", "topL / 10", "topL / 5", "topL" and the weighted evaluation metric "W score ".
[0172] According to the table content, the following is a comparative analysis, with a focus on the performance of AttCON-Hetero (the present invention) on the test dataset: Top-1 prediction accuracy: The Top-1 prediction accuracy of AttCON-Hetero is 30.91%, significantly exceeding that of other models. Followed by CDpred, with a prediction accuracy of 27.27%. At the same time, AttCON-Hetero exceeds DeepInter by 12.73 percentage points in terms of Top-1 prediction accuracy. Top-10 prediction accuracy: The Top-10 prediction accuracy of AttCON-Hetero is 26.36%, also higher than that of all other models. CDPred ranks second, with a prediction accuracy of 21.09%. At the same time, AttCON-Hetero exceeds DeepInter by 8.36 percentage points in terms of Top-10 prediction accuracy. Top-25 prediction accuracy: The Top-25 prediction accuracy of AttCON-Hetero is 23.05%, also higher than that of all other models. CDPred ranks second, with a prediction accuracy of 18.55%. At the same time, AttCON-Hetero exceeds DeepInter by 5.96 percentage points in terms of Top-25 prediction accuracy. Top-50 prediction accuracy: The top-50 prediction accuracy of AttCON-Hetero is 20.55, also ranking first among all models. This shows that when AttCON-Hetero expands the prediction range, it can still maintain a high accuracy rate and can effectively rank in a larger prediction set.
[0173] Top-L / 10, Top-L / 5, Top-L prediction accuracy: In terms of the prediction accuracy indicators at these scales, AttCON-Hetero achieved 24.80%, 21.61%, and 15.66% respectively, exceeding CDPred by 5.49, 3.61, and 1.93 percentage points respectively. These results further indicate that AttCON-Hetero performs excellently in all accuracy ranking indicators, whether in a smaller range or a larger prediction range.
[0174] Weighted evaluation parameter W score : The weighted evaluation parameter W of AttCON-Hetero score is 162.94, significantly higher than 135.53 of CDPred and 119.51 of DeepInter. This means that AttCON-Hetero demonstrates stronger comprehensive capabilities and superiority in the comprehensive evaluation index.
[0175] Figure 5Shown is the predicted inter-chain residue contact probability map of the sample with PDB ID 7EW0 on the test set under different methods. The leftmost one is the true contact map, where the yellow-marked parts are contacts. The three maps on the right are the predicted contact probability maps of DeepInter, CDPred, and AttCON-Hetero respectively. The lighter the color, the greater the contact probability. It can be clearly seen from the figure that AttCON-Hetero is the closest to the true contact map, thus having the highest prediction accuracy. Table 2. Figure 5 Fully demonstrates that the present invention has extremely strong prediction ability for inter-chain residues in heterodimers.
[0176] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A method for predicting residue contacts between heterodimer chains, characterized in that: The following steps are involved: S1. Screen the heterodimer structure information Heterodimer-PDB from the protein database; S2, the data obtained in step S1 is divided into a training set, a validation set and a test set, and the Heterodimer-PDB structural data is preprocessed according to the heterodimer structural information to extract the structural information MonomerA-PDB / MonomerB-PDB and sequence information MonomerA–FATSA / MonomerB-FASTA of the two monomers of the heterodimer, with sequence lengths L A and L B , and directly splice into a line with a length of L A +L B The combined sequence FASTA of the heterodimer was used to calculate the Euclidean distance of amino acids between the two monomers of the heterodimer according to the heterodimer structure information Heterodimer-PDB, and the contact matrix between the heterodimers was calculated accordingly; S3, using the two monomer sequence information MonomerA-FATSA / MonomerB-FASTA generated in step S2, using the HHblits multiple sequence alignment search algorithm to search the Uniclust30 sequence database to obtain the multiple sequence alignment information MSAs of the two monomers respectively, and using the TaxID species classification number to match the MSAs of the two monomers, and finally generate a joint multiple sequence alignment Paired MSAs; S4, using the Paired MSAs generated in step S3, two protein language models ESM-MSA and ESM2 are used to extract row attention features respectively; The co-evolution information features in the paired MSAs were extracted using the CCMpred algorithm; the distance graphs within the amino acid chains of the two monomers of the heterodimer were calculated using the structural information of the two monomers in step S2; the position-specific scoring matrix (PSSM) was generated using the PSI-BLAST tool using the joint sequence FASTA obtained in step S2; S5, fusing the various types of features generated in step S4 to generate a fused feature matrix to improve the diversity of input features; S6, inputting the features processed in step S5 into a model AttCON-Hetero composed of a deep residual network ResNet and a KAN (Kolmogorov-Arnold Networks) convolutional network, wherein the model uses efficient channel attention ECA and spatial attention SA modules to capture local and global dependency features of the feature matrix, and uses a KAN convolution output with adaptively adjustable convolution kernel weights to predict the contact probability matrix, and the final output result is a probability matrix of amino acid contacts between the two chains of the heterodimer; S7, based on the network output of step S6, use the contact matrix Inter-ChainContact between heterodimers in step S2 as a label for training, then use the validation set of step S2 to verify the performance of the model obtained by each batch of training, and save the model parameters of the best index obtained on the validation set as the weight of the final classification model; Finally, the saved model performance is verified on the test set in step S2.
2. The method for predicting heterodimer interchain residue contacts according to claim 1, characterized in that: S1 screened heterodimers from the PDB database that met the following conditions: (1) the polymerization state should be a heterodimer; (2) the biological assembly contains only two single chains; (3) the resolution of the X-ray structure should be better than (4) The length of a single chain should be greater than 50, and the total length of the complex should be less than 500; (4) A sequence similarity threshold of 40% was used to cluster heterodimers, and the heterodimer with the best resolution was selected from each cluster.
3. The method for predicting heterodimer interchain residue contacts according to claim 1, characterized in that: S2 extracts the heterodimer monomer structure information MonomerA-PDB / MonomerB-PDB and sequence information and calculates the contact matrix between the two monomers of the heterodimer, including the following steps: S2.
1. Extract heterodimer monomer information and sequence information: Extract the structural and sequence information of the two chains of the heterodimer: For qualified dimers, extract the three-dimensional structural information of the two chains from their PDB files and save them as Monomer-PDB A and Monomer-PDB B respectively. At the same time, extract their amino acid sequence information Monomer A–FATSA and Monomer B-FASTA, and combine Monomer A–FATSA and Monomer B-FASTA into a sequence FASTA. During the extraction process, remove the residues with the amino acid number UNK. S2.
2. Calculate the Euclidean distance and contact matrix between amino acid residues in heterodimer chains: For residues i and j, i comes from monomer A and j comes from monomer B, two types of distance calculation rules are defined: i∈MonomerA,j∈MonomerB (1) In formula (1) is the heavy-atom distance to the nearest heavy atom, where r a and r b denote the three-dimensional coordinates of the heavy atoms in residues i and j, respectively, ‖r a -r b ‖ is the Euclidean distance between the nearest heavy atoms of residue i and residue j; In formula (2) is the distance between the β carbon atoms of residue i and residue j (C β -distance), is the C of residue i β Atomic coordinates; two types of contact matrices are calculated according to formula (1) and formula (2). If or Less than or equal to The value of the corresponding position in the contact matrix is 1, otherwise it is 0, and they are saved as Inter-Chain Contact (HeavyAtom) and Inter-Chain Contact (C β ), the matrix size is L A ×L B ,The above matrix is used as the label of the model to compare the experimental results.
4. The method for predicting heterodimer interchain residue contacts according to claim 1, characterized in that: S3 specifically includes the following steps: S3.
1. Use HHblits to search the Uniclust30 database to obtain MSAs: Using the monomer sequence information generated in step S2.1, the HH-suite 3.30 tool was used to perform a preliminary search on the Uniclust30 database to generate MSAs information for the two monomers; S3.
2. Match the MSAs of the two monomers to generate paired MSAs: The MSAs of monomer A and monomer B are matched using the species classification number TaxID of each monomer to ensure that the two monomers are from the same biological species or related homologous species; based on the TaxID information, the matching results of monomer A and monomer B are integrated to generate a joint multiple sequence alignment Paired MSAs; this joint MSAs contains homologous sequence information from the two monomers, providing high-quality input data for subsequent feature extraction.
5. The method for predicting heterodimer interchain residue contacts according to claim 1, characterized in that: The detailed steps of S4 using Paired MSAs and other feature sources to extract multiple features are as follows: S4.
1. Extracting row attention features F ESM-MSA and F ESM2 F ESM-MSA That is, ESM-MSA row attention features: Based on the Paired MSAs generated in step S3, the ESM-MSA model is used to extract the row attention features in the sequence and generate the row attention feature matrix ESM-MSARowAttention; F ESM2 That is, ESM2 row attention features: using the joint sequence FASTA, the row attention features are extracted through the ESM2 model, and the generated row attention feature matrix ESM2 RowAttention; S4.
2. Extraction of co-evolution information features F CCMpred 、F DCA-APC and F DCA-DI F CCMpred That is, CCMpred coevolution feature: Based on Paired MSAs, the CCMpred algorithm is used to calculate the probability distribution of coevolution between sequences and generate a diagonally symmetric coevolution matrix; F DCA-APC and F DCA-DI :Based on the direct coupling analysis DCA method, the CCMpred algorithm was used to generate the DCA feature matrix DCA-APC with average product correction APC and the uncorrected DCA feature matrix DCA-DI. These features reveal the strength and pattern of sequence co-evolution and are used to describe the interactions between residues. F DCA-DI =DCA (4) The DCA in formula (3) is the result of direct coupling analysis. Row Mean, Col Mean, and Overall Mean are the row mean, column mean, and average of all elements of the DCA matrix, respectively. S4.
3. Extract the distance feature F between two monomer chains intra Based on the two monomer structure information MonomerA-PDB / MonomerB-PDB generated in step S2, the Euclidean distance matrix Intra-Chain Distance between the amino acids in the two monomer chains is calculated respectively; for each pair of residues (a, b) in the monomer chain, the distance can be expressed as: In formula (5), (x a ,y a ,z a ) and (x b ,y b ,z b ) are the three-dimensional coordinates of the a-th and b-th residues respectively; the intrachain distance matrix generated by monomer A is D A , the intrachain distance matrix generated by monomer B is D B , which can be expressed as follows: S4.
4. Generate position-specific scoring matrix F PSSM Using the joint sequence FASTA in step S2, the position-specific scoring matrix PSSM of the sequence is calculated by the PSI-BLAST tool; PSSM captures the substitution pattern and conservation information of amino acids by searching the UniProt database multiple times, and the generated matrix can be used for feature enhancement and model optimization; In formula (7), f ij is the frequency of amino acid j at position i, b j is the background frequency.
6. The method for predicting heterodimer interchain residue contacts according to claim 1, characterized in that: The detailed steps of S5 to generate fusion features are as follows: The 7 features generated in step S4 are concatenated according to the channel dimension to generate the fusion feature F contact ; F contact =[F ESM-MSA ,F ESM2 ,F CCMpred ,F DCA-APC ,F DCA-DI ,F intra ,F PSSM ] (8) In formula (8), F contact The dimension is (L A +L B )×(L A +L B )×C, where L A , L B are the sequence lengths of monomer A and monomer B respectively, and C represents the number of channels after fusion.
7. The method for predicting heterodimer interchain residue contacts according to claim 1, characterized in that: S6 network model construction based on deep learning and protein language model includes the following steps: The feature matrix F generated in step S5 contact The input is fed into the AttCON-Hetero model based on deep residual network (ResNet) and KAN (Kolmogorov-Arnold Networks) convolution for feature learning. The dimension of the feature matrix is (L A +L B )×(L A +L B )×C; use instance normalization to standardize the input feature matrix to reduce the distribution deviation of the input features; compress the C feature channels through two-dimensional convolution; and further extract the channel features through the Maxout layer to obtain X input ; Then X input Input residual module, each residual module consists of the following steps, which effectively extracts local and global features through step-by-step normalization, convolution and attention modules while maintaining the consistency of feature hierarchy; First, the input feature X input Perform standardization to unify the feature scale: Among them, Norm is the instance normalization operation, and then two-dimensional convolution is used to extract preliminary features: Conv2D is a two-dimensional convolution operation. In order to capture multi-scale feature information, three convolution kernels of different sizes are used for further extraction, and the three sets of features are concatenated to fuse information of different scales: X multi =Contact(X1,X2,X3) (14) Among them, X1, X2, and X3 are the results of feature extraction by different convolution kernels, and Contact is a concatenation operation. multi Refer to formulas (9) and (10) to continue standardization and two-dimensional convolution compression to obtain Then use the channel attention mechanism ECA and the spatial attention mechanism SA for feature extraction: z=σ(Conv1D(y)) (16) f sa =Conv2D([AvgPool(X ECA ).;.MaxPool(X ECA )]) (18) X SA =X ECA ⊙σ(f sa ) (19) Formulas (15) to (17) are the processing flow of the channel attention mechanism: GAP is the global maximum pooling, which averages the spatial dimensions of each channel to generate a global feature representation; Conv1D is a one-dimensional convolution operation; σ is the Sigmoid activation function; ⊙ represents element-by-element multiplication; X ECA is the result of the extracted features of channel attention ECA; Formulas (18) to (19) are the spatial attention mechanism: AvgPool and MaxPool are average pooling and maximum pooling operations respectively; [·; ·] represents the concatenation of pooling results; X SA It is the result of feature extraction by spatial attention SA; Formula (20) adds the output of the attention mechanism to the input features through a residual connection to enhance the gradient flow and avoid the gradient vanishing problem; Finally, the output layer of the model is output Refer to formula (9) for standardization and get X output_norm , and then output through two layers of KAN convolution layers. KAN convolution ConvKAN adaptively adjusts the weights of the convolution kernel, so that the network can dynamically adjust its learning process according to the characteristics of the input data; X KAN =ConvKAN(X output_norm ) (21) ConvKAN is KAN convolution, X KAN is the result after KAN convolution processing; the output of formula (21) uses KAN convolution layer and Sigmoid classifier to output the probability matrix (whether it is contact), and the output matrix dimension is L A ×L B , represents the contact probability P of each pair of residues contact : P contact =Sigmoid(ConvKAN(X KAN )) (22)。 8. The method for predicting heterodimer interchain residue contacts according to claim 1, characterized in that: The steps of S7 model training process, model performance evaluation and model testing are as follows: S7.1 Use of loss function in model training and optimization steps of optimizer The contact prediction matrix output in step S6 is used as the output of the network, and the heterodimer interchain residue contact matrix obtained in step S2 is used as the label label for model training; The training process is optimized by the FocalLoss loss function, using the AdamW optimizer combined with cosine annealing learning rate scheduling; S7.2 Performance evaluation and model selection After each batch of training, the performance of the model is evaluated using the validation set provided in step S2 by calculating the Topn accuracy, Top-L / k accuracy, and the weighted evaluation index W score To evaluate the model performance, the model parameters with the best performance on the validation set are finally selected.
Citation Information
Patent Citations
Protein residue contact map prediction method
CN113257357A
Method for predicting distance between protein residues based on self-attention mechanism
CN114708903A
Method for predicting protein function based on transfer learning and three-channel combination GNN
CN118969060A
Deep learning systems and methods for predicting structural aspects of protein-related complexes
US20230154561A1
Cited By
Method for predicting protein structure, medium and electronic equipment
CN120853661A
Antibacterial peptide prediction method based on dual-channel sparse attention
CN121354654A