Protein compound model interface quality evaluation method based on multi-scale isotropic graph neural network

By fusing multi-scale information through a multi-scale equivariant graph neural network and using a prototype contrast learning strategy, the problems of inaccurate evaluation and insufficient generalization ability in existing methods are solved, and a more accurate and stable interface quality assessment of protein complex models is achieved.

CN120690276APending Publication Date: 2025-09-23ZHEJIANG UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510750162.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing methods for assessing the interface quality of protein complex models lack multi-scale information fusion, spatially equivariant modeling capabilities, and adaptability to unbalanced data, resulting in inaccurate assessment results and insufficient generalization capabilities, making it difficult to meet the needs of drug design and protein interaction mechanism research.

Method used

A multi-scale equivariant graph neural network is adopted, combining the geometric, physicochemical and evolutionary features at the surface, atomic and residue levels. Multi-scale information is deeply integrated through the equivariant graph neural network architecture, and the prototype comparative learning strategy is used to improve the recognition ability of high-quality conformations.

Benefits of technology

It significantly improves the accuracy and stability of the interface quality assessment of protein complex models, enhances the perception of spatial configuration, physical and chemical environment and evolutionary information, and improves the robustness and generalization of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690276A_ABST
    Figure CN120690276A_ABST
Patent Text Reader

Abstract

A protein complex model interface quality evaluation method based on a multi-scale isovariant graph neural network comprises the following steps: firstly, screening out a co-crystallized natural protein complex structure from a non-redundant protein interaction database PRISM, and generating a bait structure by using a HDock docking algorithm; the method comprises the following steps: firstly, extracting molecular surface interaction fingerprints, atomic-level features and residue-level features on the basis of each compound bait structure, obtaining graph representation of the compound bait structures, then fully capturing and fusing multi-scale information through a depth isotropic graph neural network, and finally obtaining an interface mass fraction through prototype comparison prediction. According to the method, the interface quality evaluation of the protein compound model can be accurately carried out, and the problems of low precision and poor generalization of the interface quality evaluation of the protein compound model are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and computer applications, and in particular to a method for evaluating the interface quality of a protein complex model based on a multi-scale equivariant graph neural network. Background Art

[0002] Proteins interact with other proteins through their interfaces to form complexes, playing an important role in life processes such as cell signaling, metabolic regulation, and immune response. Elucidating the structure of protein complexes is crucial for revealing the mechanisms of protein-protein interactions. Although experimental techniques such as X-ray crystallography, nuclear magnetic resonance spectroscopy, and cryo-electron microscopy can provide high-precision structural information, they are limited by factors such as high costs and long cycles. Therefore, computational methods have gradually become an important supplementary means. Although deep learning methods have made significant progress in predicting the structure of monomeric proteins, the accurate prediction of protein complex structures still faces challenges. Existing computational methods can generate a large number of complex models, but the quality is uneven, and the lack of an effective evaluation mechanism has become a major bottleneck limiting their application. Therefore, how to accurately evaluate and screen high-quality models that are close to the native conformation has become the key to improving the practicality of protein docking and structural modeling.

[0003] Traditional assessments of complex model interface quality typically rely on physical force fields, empirical scoring functions, or statistical potentials. These methods have limited generalization capabilities and struggle to capture complex nonlinear energy relationships. With the advancement of deep learning, assessment methods based on 3D convolutional networks and graph neural networks have gradually emerged. Researchers model the 3D structure of protein complexes at the atomic or residue scale and use neural networks to predict model quality. However, existing methods still have shortcomings: While 3D convolutional network-based methods can extract local geometric features through voxelization, they inevitably suffer from precision loss, and convolutional networks lack equivariance to spatial rotations, which affects robustness and generalization. While graph neural network-based methods break through the limitations of Cartesian grids and model topological structures based on node relationships, they cannot fully express the 3D geometric positions between nodes and lack equivariance to spatial transformations such as rotations and translations, resulting in inconsistent assessment results. Furthermore, most existing methods rely on feature representations at a single scale, making it difficult to effectively integrate the multi-level information of protein complex interfaces, including surface geometric complementarity, fine-grained interactions between atoms, and evolutionary conservation of residues. Single-scale feature modeling ignores the complementarity between features at different levels, limiting a comprehensive understanding of the overall structural complexity and biological functional relevance of the interface, which in turn affects the accurate assessment and determination of interface quality. Furthermore, due to biases introduced by existing protein-protein docking methods during the sampling process, the resulting complex conformational datasets are often unevenly distributed, with high-quality conformations far fewer than low-quality ones. This leads to overfitting to majority class samples during training, affecting the neural network model's ability to recognize high-quality conformations and weakening its discriminative performance on borderline samples.

[0004] With the deepening of protein interaction research, the importance of accurately assessing the interface quality of protein complexes in drug design and the study of protein-protein interaction mechanisms has become increasingly prominent. However, existing evaluation methods still have obvious limitations in terms of multi-scale information fusion, spatial equivariant modeling capabilities, and generalization capabilities, making it difficult to meet the needs of high accuracy and strong generalization in practical applications. Therefore, there is an urgent need to develop methods that have spatial equivariant modeling capabilities, can integrate multi-scale structural features, and have adaptability to imbalanced data, in order to achieve more accurate and universal protein complex model interface quality assessment and promote further development in fields such as structural bioinformatics and precision medicine. Summary of the Invention

[0005] In order to overcome the shortcomings of existing protein complex model interface quality assessment methods, the present invention proposes a protein complex model interface quality assessment method based on a multi-scale equivariant graph neural network. This method introduces geometric features, physicochemical features and evolutionary features at three scales: surface, atom and residue. The equivariant graph neural network architecture is used to deeply fuse multi-scale information while maintaining spatial transformation invariance, significantly enhancing the perception of the spatial configuration, physicochemical environment and evolutionary information of the protein complex interface. The prototype comparative learning strategy is used to cluster samples with similar semantics, thereby improving the recognition ability of high-quality conformations, thereby achieving more accurate and stable protein complex model interface quality assessment.

[0006] The technical solution adopted by the present invention to solve its technical problem is:

[0007] A method for assessing the interface quality of protein complex models based on a multi-scale equivariant graph neural network is proposed. First, redundant native protein complex crystal structures are screened out from the PRISM protein-protein interaction database. For each native protein complex structure, a decoy structure is generated using the HDock docking algorithm. Then, surface-level features, atomic-level features, and residue-level features are extracted from each protein complex decoy structure and input into an equivariant graph neural network that integrates multi-scale information. An embedded representation is then generated through a multi-head self-attention mechanism. Finally, the interface quality score of the protein complex decoy structure is obtained by calculating the difference in distance between the embedded structure and the positive and negative prototypes in high-dimensional space.

[0008] Furthermore, the method comprises the following steps:

[0009] 1) Construct a protein complex bait structure dataset.

[0010] 2) Prepare label data of the protein complex bait structure, use the DockQ tool to align the protein complex bait structure with its native structure, and divide all bait structures into two categories according to the calculated scores: incorrect conformation and near-native conformation, represented by 0 and 1.

[0011] 3) Extract the molecular surface interaction fingerprint of the bait structure of the protein complex.

[0012] 4) Extract atomic-level features of the protein complex bait structure: Use the RDKit tool to read the structure of the protein complex model and obtain atomic information, including atomic coordinates, atom type, formal charge, number of connected hydrogens, chirality, aromaticity, ring type, etc., and then obtain chemical bond information, including chemical bond type, bond length, etc.

[0013] 5) Extract residue-level features of the protein complex bait structure: Use tools such as PyRosetta to read the structure of the protein complex model and extract geometric features, physicochemical features, and evolutionary features.

[0014] 6) Convert the protein complex bait structure into a graph representation, and construct a heterogeneous graph composed of atoms and surface points, and a graph composed of residues.

[0015] 7) Build a multi-scale equivariant graph neural network, perform neural network model training and parameter optimization, and predict the interface quality of the bait structure.

[0016] Further, the process of 1) is:

[0017] 1.1) 2465 co-crystallized native structures of protein complexes were screened from the PRISM non-redundant protein-protein interaction database;

[0018] 1.2) Use TM-align to calculate the TM-score between each pair of binding interfaces of the native structures of all protein complexes;

[0019] 1.3) Perform hierarchical clustering based on the calculated TM-score matrix and divide the obtained protein complex natural structure set into training, validation and testing subsets;

[0020] 1.4) Each protein complex constituting protein-protein interactions in the subset was split into two monomers, and protein-protein docking was re-performed using HDOCK to generate 100 protein complex bait structures for each protein complex native structure.

[0021] Furthermore, the process of 2) is as follows:

[0022] 2.1) Defined as: When , the two residues are in contact. For each protein complex bait structure in the dataset, the DockQ tool is used to align it with the corresponding native structure. The number of contacts between the bait structure and the native structure residues is counted, and the proportion of the contacts to the total number of contacts in the native structure is calculated.

[0023] 2.2) Superimpose the decoy and native structures using the receptor (longer chain) as a benchmark and calculate the root mean square deviation (L_RMS) of the main chain atoms of the ligand (shorter chain);

[0024] 2.3) Relax the contact threshold to Perform structural superposition of the main chain atoms of the interface of the decoy structure and the interface of the native structure, and calculate the interface root mean square deviation i_RMS;

[0025] 2.4) Labels refer to the CAPRI classification standard. When the f_nat of the bait structure is less than 0.1, or the L_RMS is greater than And i_RMS is greater than When , the bait structure is defined as an incorrect conformation, and the remaining bait structures are defined as near-native conformations.

[0026] Furthermore, the process of 3) is as follows:

[0027] 3.1) Use MSMS tools to process the three-dimensional structure of the protein complex bait structure, with a radius of The probe ball rolls on the surface of the three-dimensional structure, converting the solvent exclusion surface into a mesh composed of triangular facets;

[0028] 3.2) Input the surface mesh generated in 2.1) into PDB2PQR and APBS tools, and obtain the surface electrostatic potential and charge distribution by solving the linear Poisson-Boltzmann equation;

[0029] 3.3) For each surface point, identify its nearest neighbor atoms and determine the amino acid residue to which the atom belongs. Based on the type of residue, consult the Kyte & Doolittle scale and assign a corresponding hydrophobicity index, with positive values ​​indicating hydrophobicity and negative values ​​indicating hydrophilicity.

[0030] 3.4) Use the knowledge-based hydrogen bond statistical potential to identify the hydrogen bond donor and acceptor positions on the molecular surface. The calculation formula is as follows:

[0031]

[0032] Where E(p) represents the hydrogen bond potential under the conformational parameter p, f protein (p) represents the probability density of the hydrogen bond geometric parameters p (angle, distance, etc.) observed in high-resolution protein crystal structures, and f random (p) represents the probability density of the geometric parameter p under a random distribution background; the difference between the experimentally observed hydrogen bond geometric distribution and the theoretical random distribution is expressed in the form of logarithmic probability. The higher the frequency, the more consistent it is with the real structure, the lower the energy, and the more stable it is.

[0033] 3.5) For each surface point, calculate its local curvature characteristics. First, calculate the maximum curvature k1 and minimum curvature k2 of the point, and then calculate the shape index according to the following formula:

[0034]

[0035] This index maps the local surface shape to the range [-1,1], representing the change from concave to convex.

[0036] Furthermore, the process of 5) is as follows:

[0037] 5.1) One-hot encode the amino acid sequence of each chain in the complex to generate a binary matrix of size L × 20, where L is the sequence length and the 20 dimensions correspond to the 20 standard amino acid types;

[0038] 5.2) Extract seven types of physicochemical properties corresponding to residue types, including isoelectric point, polarity, acidity and alkalinity, hydrogen bond acceptor, hydrogen bond donor, n-octanol-water partition coefficient, and topological polar surface area;

[0039] 5.3) Use the BLOSUM62 substitution scoring matrix to represent the conservation of substitutions between amino acids, encoding each residue as a 20-dimensional vector (i.e., the substitution score for other amino acid types), generating a feature matrix of shape L×20;

[0040] 5.4) Extract the main chain torsion angles φ and ψ and calculate the main chain atoms C i -N i+1 The bond lengths between bond angle;

[0041] 5.5) Use the DSSP tool to obtain secondary structure information and generate a binary matrix of size L × 3, where L is the sequence length and the three dimensions correspond to the helical, folded, and random coil states respectively;

[0042] 5.6) The solvent accessible surface area (SASA) and relative solvent accessible surface area (RASA) of each residue were calculated using the FreeSASA tool.

[0043] 5.7) Calculate Rosetta monomer energy terms, including the main chain dihedral angle preference term rama_prepro, the probability distribution term p_aa_pp of the main chain dihedral angles φ and ψ, the main chain dihedral angle ω term omega, and the side chain rotamer energy fa_dun;

[0044] 5.8) Calculate the distance matrix between residues, including C β -C β 、Tip-Tip、C α -Tip, Tip-C α Distances between atoms; extract inter-residue angles, including the distances along the C link connecting two residues. β The dihedral angle omega of the rotation of the virtual axis of the atom, and the C β The dihedral angles theta and phi of the orientation of atoms relative to a local reference frame centered on another residue;

[0045] 5.9) Calculate the Rosetta two-body energy terms, including main-chain long-range hydrogen bonds hbond_lr_bb, main-chain short-range hydrogen bonds hbond_sr_bb, main-chain-side-chain hydrogen bonds hbond_bb_sc, side-chain-side-chain hydrogen bonds hbond_sc, van der Waals attractive terms fa_atr, van der Waals repulsive terms fa_rep, isotropic solvation free energies fa_sol, anisotropic solvation free energies lk_ball_wtd, and Coulomb electrostatic potential fa_elec.

[0046] Furthermore, the process of step 6) is as follows:

[0047] 6.1) Construct a heterogeneous graph of atoms and surface points, expressed as follows:

[0048] G H =(v A ,ε Ab ,ε Ac ,v S ,ε S ,ε AS )

[0049] where v A Represents atoms, v S represents a surface point, ε Ab represents chemical bonds, ε Ac Indicates interatomic contact (distance less than ), ε S represents the edge of the surface triangle mesh, ε AS Represents the edge between a surface point and its nearest neighboring atoms. Surface node features include electrostatic potential, charge, hydrophobicity index, hydrogen bond potential, and shape index. Surface-surface edge features include edge length and edge type. Surface-atom edge features include edge length and edge type. Atomic node features include element type, formal charge, number of connected hydrogens, chirality, aromaticity, and ring type. Atomic chemical bond features include bond length and bond type. Interatomic contact edge features include edge length and edge type.

[0050] 6.2) Construct a residue-level graph representation in the following form:

[0051] G R =(v R ,ε R )

[0052] where v R represents the residue, ε R Represents the interaction between residues (residue C β The distance between atoms is less than ); node features include residue type, physicochemical properties, secondary structure, BLOSUM62 substitution matrix, solvent accessible surface area, Rosetta monomer energy term, bond length, bond angle, main chain dihedral angle; edge features include inter-residue distance, orientation angle, and Rosetta two-body energy term;

[0053] Furthermore, the process of 7) is as follows:

[0054] 7.1) The heterogeneous graph consisting of atoms and surface points and the residue graph are fed into two independent equivariant graph neural network modules. The representations of nodes and edges are updated through coordinate transformation and message passing. The equivariant graph neural network is updated as follows:

[0055]

[0056] where h i represents the feature of node i, x i represents the spatial coordinates of the node, ‖x i -x j ‖ 2 is the relative position information between nodes, e ij represents the edge feature, N(i) represents the neighbor set of node i, φ e 、φ h 、φ x is a learnable function (such as MLP), φ x The output of is a scalar, so the coordinates are updated in the same direction as x. i -x j Maintain consistency and ensure equal variability;

[0057] 7.2) The node representations output by the two equivariant graph neural network modules are stacked into a graph-level representation through global average pooling. Then, a multi-head self-attention mechanism is used to learn long-range interactions. The formula of the multi-head self-attention mechanism is as follows:

[0058] Q=XW Q

[0059] K=XW K

[0060] V=XW V

[0061]

[0062] MultiHead(Q,K,V)=Concat(head1,…,head h )W O

[0063] Where X represents the input sequence, Q, K, and V represent the query vector, key vector, and value vector, and W Q、W K 、W V 、W o represents the learnable parameter matrix, d k Indicates the dimensions of Q and K of each attention head, head i represents the output of the i-th attention head, and MultiHead represents the information of different subspaces captured in parallel;

[0064] 7.3) The representations obtained by the multi-head self-attention mechanism are globally average pooled and normalized. The distances between the negative prototype vector and the positive prototype vector (i.e., cluster centers) updated during training are calculated. The distance difference ranges from [-2, 2]. The larger the distance, the closer it is to the positive prototype, and the smaller the distance, the closer it is to the negative prototype. The score is normalized to the range of [0, 1] as the interface quality assessment score of the protein complex model.

[0065] 7.4) Use the AdamW optimizer to optimize the network weight parameters. The loss function is the sum of cross entropy loss, marginal ranking loss and supervised contrast loss, and the learning rate is 0.0001.

[0066] The main benefits of this invention are as follows: by introducing an equivariant graph neural network and multi-scale feature fusion, it effectively overcomes the shortcomings of existing methods, such as the lack of geometric equivariance and sensitivity to local perturbations, significantly enhancing the ability to perceive the spatial configuration of protein interfaces, the local physicochemical environment, and evolutionary information. Test results and ablation experiments on multiple widely used datasets demonstrate that, by introducing a prototype comparative learning strategy, this method outperforms existing state-of-the-art scoring methods in metrics such as CAPRI positive example recognition precision, recall, and ROC-AUC, significantly improving the robustness and generalization of the scoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is an overall flow chart of the interface quality assessment method for protein complex models based on multi-scale equivariant graph neural networks.

[0068] Figure 2 It is a schematic diagram of prototype contrastive learning. DETAILED DESCRIPTION

[0069] The present invention will be further described below with reference to the accompanying drawings.

[0070] Reference Figure 1 and Figure 2 A method for evaluating the interface quality of a protein complex model based on a multi-scale equivariant graph neural network, the method comprising the following steps:

[0071] 1) Construct a protein complex bait structure dataset. The process is as follows:

[0072] 1.1) 2465 co-crystallized native structures of protein complexes were screened from the PRISM non-redundant protein-protein interaction database;

[0073] 1.2) Use TM-align to calculate the TM-score between each pair of binding interfaces of the native structures of all protein complexes;

[0074] 1.3) Perform hierarchical clustering based on the calculated TM-score matrix and divide the obtained protein complex natural structure set into training, validation and testing subsets;

[0075] 1.4) Each protein complex constituting protein-protein interactions in the subset was split into two monomers, and protein-protein docking was re-performed using HDOCK to generate 100 protein complex bait structures for each protein complex native structure;

[0076] 2) Produce label data of protein complex bait structure, the process is as follows:

[0077] 2.1) Defined as: When the two residues are in contact, the two residues are in contact. For each protein complex bait structure in the data set, the DockQ tool is used to perform a structural alignment between it and the corresponding native structure, and the number of contacts between the bait structure and the native structure residue pairs is counted, and the proportion of the contacts to the total number of contacts in the native structure is calculated. nat ;

[0078] 2.2) Superimpose the decoy structure and the native structure based on the receptor (longer chain) and calculate the root mean square deviation L of the main chain atoms of the ligand (shorter chain) RMS ;

[0079] 2.3) Relax the contact threshold to The main chain atoms of the interface of the decoy structure and the interface of the native structure are superimposed and the interface root mean square deviation i is calculated. RMS ;

[0080] 2.4) Labeling refers to the CAPRI classification standard. When the bait structure is nat Less than 0.1, or L RMS Greater than And i RMS Greater than When , the bait structure is defined as an incorrect conformation, and the remaining bait structures are defined as near-native conformations;

[0081] 3) Extracting the molecular surface interaction fingerprint of the protein complex bait structure, the process is as follows:

[0082] 3.1) First, the solvent-excluded surface of the protein complex bait structure is triangulated using the MSMS tool and discretized into a triangular mesh consisting of surface point sets;

[0083] 3.2) Input the surface mesh into APBS to calculate the surface electrostatic potential and charge;

[0084] 3.3) Query the Kyte & Doolittle hydrophilicity scale to assign a hydrophobicity index based on the residues to which the nearest neighbor atoms of the surface points belong;

[0085] 3.4) Use hydrogen bond potential to calculate the positions of free electrons and hydrogen bond donors on the molecular surface;

[0086] 3.5) Calculate the shape index describing the local area around the surface point based on the maximum curvature and minimum curvature of the curve passing through the surface point;

[0087] 4) Extract the atomic-level features of the protein complex bait structure. Use the RDKit tool to read the structure of the protein complex model and obtain atomic information, including atomic coordinates, atom type, formal charge, number of connected hydrogens, chirality, aromaticity, ring type, etc., and then obtain chemical bond information, including chemical bond type, bond length, etc.

[0088] 5) Extract residue-level features of the protein complex bait structure by:

[0089] 5.1) One-hot encode the amino acid sequence of each chain in the complex to generate a binary matrix of size L × 20, where L is the sequence length and the 20 dimensions correspond to the 20 standard amino acid types;

[0090] 5.2) Extract seven types of physicochemical properties corresponding to residue types, including isoelectric point, polarity, acidity and alkalinity, hydrogen bond acceptor, hydrogen bond donor, n-octanol-water partition coefficient, and topological polar surface area;

[0091] 5.3) Use the BLOSUM62 substitution scoring matrix to represent the conservation of substitutions between amino acids, encoding each residue as a 20-dimensional vector (i.e., the substitution score for other amino acid types), generating a feature matrix of shape L×20;

[0092] 5.4) Extract the main chain torsion angles φ and ψ and calculate the main chain atoms C i -N i+1 The bond lengths between bond angle;

[0093] 5.5) Use the DSSP tool to obtain secondary structure information and generate a binary matrix of size L × 3, where L is the sequence length and the three dimensions correspond to the helical, folded, and random coil states respectively;

[0094] 5.6) The solvent accessible surface area (SASA) and relative solvent accessible surface area (RASA) of each residue were calculated using the FreeSASA tool.

[0095] 5.7) Calculate Rosetta monomer energy terms, including the main chain dihedral angle preference term rama_prepro, the probability distribution term p_aa_pp of the main chain dihedral angles φ and ψ, the main chain dihedral angle ω term omega, and the side chain rotamer energy fa_dun;

[0096] 5.8) Calculate the distance matrix between residues, including C β -C β 、Tip-Tip、Cα-Tip、Tip-C α Distances between atoms; extract inter-residue angles, including the distances along the C link connecting two residues. β The dihedral angle omega of the rotation of the virtual axis of the atom, and the C β The dihedral angles theta and phi of the orientation of atoms relative to a local reference frame centered on another residue;

[0097] 5.9) Calculate Rosetta two-body energy terms, including main chain-main chain long-range hydrogen bonds hbond_lr_bb, main chain-main chain short-range hydrogen bonds hbond_sr_bb, main chain-side chain hydrogen bonds hbond_bb_sc, side chain-side chain hydrogen bonds hbond_sc, van der Waals attraction term fa_atr, van der Waals repulsion term fa_rep, isotropic solvation free energy fa_sol, anisotropic solvation free energy lk_ball_wtd, and Coulomb electrostatic potential fa_elec;

[0098] 6) Convert the protein complex bait structure into a graphical representation by:

[0099] 6.1) Construct a heterogeneous graph consisting of atoms and surface points, expressed as follows:

[0100] G H =(v A ,ε Ab ,ε Ac ,v S ,ε S ,ε AS )

[0101] where v A Represents atoms, v S represents a surface point, ε Ab represents chemical bonds, ε Ac Indicates interatomic contact (distance less than ), εS represents the edge of the surface triangle mesh, ε AS Represents the edge formed by the surface point and the nearest neighbor atom. Surface node features include electrostatic potential, charge, hydrophobic index, hydrogen bond potential, and shape index. Surface-surface edge features include edge length and edge type. Surface-atom edge features include edge length and edge type. Atomic node features include element type, formal charge, number of connected hydrogens, chirality, aromaticity, and ring type. Atomic chemical bond features include bond length and bond type. Interatomic contact edge features include edge length and edge type.

[0102] 6.2) Construct a graph consisting of residues, represented as follows:

[0103] G R =(v R ,ε R )

[0104] where v R represents the residue, ε R Represents the interaction between residues (residue C β The distance between atoms is less than ); node features include residue type, physicochemical properties, secondary structure, BLOSUM62 substitution matrix, solvent accessible surface area, Rosetta monomer energy term, bond length, bond angle, main chain dihedral angle; edge features include inter-residue distance, orientation angle, and Rosetta two-body energy term;

[0105] 7) Build a multi-scale equivariant graph neural network, perform neural network model training and parameter optimization, and predict the interface quality of the decoy structure. The process is as follows:

[0106] 7.1) The heterogeneous graph consisting of atoms and surface points and the residue graph are fed into two independent equivariant graph neural network modules. The representations of nodes and edges are updated through coordinate transformation and message passing. The equivariant graph neural network is updated as follows:

[0107]

[0108] where h i represents the feature of node i, x i represents the spatial coordinates of the node, ‖x i -x j ‖ 2 is the relative position information between nodes, e ij represents the edge feature, N(i) represents the neighbor set of node i, φ e 、φ h 、φ x is a learnable function (such as MLP), φ x The output of is a scalar, so the coordinates are updated in the same direction as x. i -xj Maintain consistency and ensure equal variability;

[0109] 7.2) The node representations output by the two equivariant graph neural network modules are stacked into a graph-level representation through global average pooling. Then, a multi-head self-attention mechanism is used to learn long-range interactions. The formula of the multi-head self-attention mechanism is as follows:

[0110] Q=XW Q

[0111] K=XW K

[0112] V=XW V

[0113]

[0114] MultiHead(Q,K,V)=Concat(head1,…,head h )W O

[0115] Where X represents the input sequence, Q, K, and V represent the query vector, key vector, and value vector, and W Q 、W K 、W V 、W O represents the learnable parameter matrix, d k Indicates the dimensions of Q and K of each attention head, head i represents the output of the i-th attention head, and MultiHead represents the information of different subspaces captured in parallel;

[0116] 7.3) The representations obtained by the multi-head self-attention mechanism are globally averaged and normalized. The distances between the negative prototype vector and the positive prototype vector updated during training are calculated. The range of the distance difference is [-2, 2]. The larger the distance, the closer it is to the positive prototype, and the smaller the distance, the closer it is to the negative prototype. The score is normalized to the range of [0, 1] as the interface quality assessment score of the protein complex model. The schematic diagram of prototype comparative learning is attached. Figure 2 .

[0117] 7.4) Use the AdamW optimizer to optimize the network weight parameters. The loss function is the sum of cross entropy loss, marginal ranking loss and supervised contrast loss, and the learning rate is 0.0001.

[0118] This example uses the model T1206TS304_2 of the protein complex target T1206 with an amino acid sequence length of 474 predicted by the AlphaFold3 server in the CASP16 protein structure prediction competition as an implementation case. A method for evaluating the interface quality of protein complex models based on a multi-scale equivariant graph neural network includes the following steps:

[0119] 1) Construct a protein complex bait structure dataset. The process is as follows:

[0120] 1.1) 2465 co-crystallized native structures of protein complexes were screened from the PRISM non-redundant protein-protein interaction database;

[0121] 1.2) Use TM-align to calculate the TM-score between each pair of binding interfaces of the native structures of all protein complexes;

[0122] 1.3) Perform hierarchical clustering based on the calculated TM-score matrix and divide the obtained protein complex natural structure set into training, validation and testing subsets;

[0123] 1.4) Each protein complex constituting protein-protein interactions in the subset was split into two monomers, and protein-protein docking was re-performed using HDOCK to generate 100 protein complex bait structures for each protein complex native structure;

[0124] 2) Produce label data of protein complex bait structure, the process is as follows:

[0125] 2.1) Defined as: When the two residues are in contact, the protein complex bait structure is aligned with its native structure using the DockQ tool, and the ratio of the number of contacts in the bait structure T1206TS304_2 that are identical to those in the native structure to the total number of contacts in the native structure is calculated. nat is 0.907;

[0126] 2.2) The decoy structure and the native structure are superimposed on the receptor (longer chain) to calculate the root mean square deviation L of the main chain atoms of the ligand (shorter chain) RMS is 0.638;

[0127] 2.3) Relax the contact threshold to The main chain atoms of the interface of the decoy structure and the interface of the native structure are superimposed to calculate the interface root mean square deviation i RMS is 0.424;

[0128] 2.4) Labels refer to the CAPRI classification standard, T1206TS304_2 f nat Greater than 0.1, it is a near-native conformation, represented by 1;

[0129] 3) Extracting the molecular surface interaction fingerprint of the protein complex bait structure, the process is as follows:

[0130] 3.1) First, the solvent-excluded surface of the protein complex bait structure T1206TS304_2 was triangulated using the MSMS tool and discretized into a triangular mesh consisting of surface point sets;

[0131] 3.2) Input the surface mesh into APBS to calculate the surface electrostatic potential and charge;

[0132] 3.3) Query the Kyte & Doolittle hydrophilicity scale to assign a hydrophobicity index based on the residues to which the nearest neighbor atoms of the surface points belong;

[0133] 3.4) Use hydrogen bond potential to calculate the positions of free electrons and hydrogen bond donors on the molecular surface;

[0134] 3.5) Calculate the shape index describing the local area around the surface point based on the maximum curvature and minimum curvature of the curve passing through the surface point;

[0135] 4) Extract the atomic-level features of the protein complex bait structure. Use the RDKit tool to read the structure of the protein complex model and obtain atomic information, including atomic coordinates, atom type, formal charge, number of connected hydrogens, chirality, aromaticity, ring type, etc., and then obtain chemical bond information, including chemical bond type, bond length, etc.

[0136] 5) Extract residue-level features of the protein complex bait structure by:

[0137] 5.1) One-hot encode the amino acid sequence of each chain in the complex to generate a binary matrix of size 474 × 20, where 474 is the sequence length and the 20 dimensions correspond to the 20 standard amino acid types;

[0138] 5.2) Extract seven types of physicochemical properties corresponding to residue types, including isoelectric point, polarity, acidity and alkalinity, hydrogen bond acceptor, hydrogen bond donor, n-octanol-water partition coefficient, and topological polar surface area;

[0139] 5.3) Use the BLOSUM62 substitution scoring matrix to represent the conservation of substitutions between amino acids, encoding each residue as a 20-dimensional vector (i.e., the substitution score for other amino acid types), generating a feature matrix of shape 474×20;

[0140] 5.4) Extract the main chain torsion angles φ and ψ and calculate the main chain atoms C i -N i+1 The bond lengths between bond angle;

[0141] 5.5) Use the DSSP tool to obtain secondary structure information and generate a binary matrix of size 474 × 3, where 474 is the sequence length and the three dimensions correspond to the helical, folded, and random coil states respectively;

[0142] 5.6) The solvent accessible surface area (SASA) and relative solvent accessible surface area (RASA) of each residue were calculated using the FreeSASA tool.

[0143] 5.7) Calculate Rosetta monomer energy terms, including the main chain dihedral angle preference term rama_prepro, the probability distribution term p_aa_pp of the main chain dihedral angles φ and ψ, the main chain dihedral angle ω term omega, and the side chain rotamer energy fa_dun;

[0144] 5.8) Calculate the distance matrix between residues, including C β -C β 、Tip-Tip、C α -Tip, Tip-C α Distances between atoms; extract inter-residue angles, including the distances along the C link connecting two residues. β The dihedral angle omega of the rotation of the virtual axis of the atom, and the C β The dihedral angles theta and phi of the orientation of atoms relative to a local reference frame centered on another residue;

[0145] 5.9) Calculate Rosetta two-body energy terms, including main chain-main chain long-range hydrogen bonds hbond_lr_bb, main chain-main chain short-range hydrogen bonds hbond_sr_bb, main chain-side chain hydrogen bonds hbond_bb_sc, side chain-side chain hydrogen bonds hbond_sc, van der Waals attraction term fa_atr, van der Waals repulsion term fa_rep, isotropic solvation free energy fa_sol, anisotropic solvation free energy lk_ball_wtd, and Coulomb electrostatic potential fa_elec;

[0146] 6) Convert the protein complex bait structure into a graphical representation by:

[0147] 6.1) Construct a heterogeneous graph consisting of atoms and surface points, expressed as follows:

[0148] G H =(v A ,ε Ab ,ε Ac ,vS ,ε S ,ε AS )

[0149] where v A Represents atoms, v S represents a surface point, ε Ab represents chemical bonds, ε Ac Indicates interatomic contact (distance less than ), ε S represents the edge of the surface triangle mesh, ε AS Represents the edge between a surface point and its nearest neighboring atoms. Surface node features include electrostatic potential, charge, hydrophobicity index, hydrogen bond potential, and shape index. Surface-surface edge features include edge length and edge type. Surface-atom edge features include edge length and edge type. Atomic node features include element type, formal charge, number of connected hydrogens, chirality, aromaticity, and ring type. Atomic chemical bond features include bond length and bond type. Interatomic contact edge features include edge length and edge type.

[0150] 6.2) Construct a graph consisting of residues, represented as follows:

[0151] G R =(v R ,ε R )

[0152] where v R represents the residue, ε R Represents the interaction between residues (residue C β The distance between atoms is less than ); node features include residue type, physicochemical properties, secondary structure, BLOSUM62 substitution matrix, solvent accessible surface area, Rosetta monomer energy term, bond length, bond angle, main chain dihedral angle; edge features include inter-residue distance, orientation angle, and Rosetta two-body energy term;

[0153] 7) Build a multi-scale equivariant graph neural network, perform neural network model training and parameter optimization, and predict the interface quality of the decoy structure. The process is as follows:

[0154] 7.1) The heterogeneous graph consisting of atoms and surface points and the residue graph are fed into two independent equivariant graph neural network modules. The representations of nodes and edges are updated through coordinate transformation and message passing. The equivariant graph neural network is updated as follows:

[0155]

[0156] where h i represents the feature of node i, x i represents the spatial coordinates of the node, ‖x i -x j ‖2 is the relative position information between nodes, e ij represents the edge feature, N(i) represents the neighbor set of node i, φ e 、φ h 、φ x is a learnable function (such as MLP), φ x The output of is a scalar, so the coordinates are updated in the same direction as x. i -x j Maintain consistency and ensure equal variability;

[0157] 7.2) The node representations output by the two equivariant graph neural network modules are stacked into a graph-level representation through global average pooling. Then, a multi-head self-attention mechanism is used to learn long-range interactions. The formula of the multi-head self-attention mechanism is as follows:

[0158] Q=XW Q

[0159] K=XW K

[0160] V=XW V

[0161]

[0162] MultiHead(Q,K,V)=Concat(head1,…,head h )W O

[0163] Where X represents the input sequence, Q, K, and V represent the query vector, key vector, and value vector, and W Q 、W K 、W V 、W O represents the learnable parameter matrix, d k Indicates the dimensions of Q and K of each attention head, head i represents the output of the i-th attention head, and MultiHead represents the information of different subspaces captured in parallel;

[0164] 7.3) The representation obtained by the multi-head self-attention mechanism is globally average pooled and normalized. The difference in Euclidean distance between it and the negative prototype vector and the positive prototype vector is calculated as 1.541. The value 0.885, which is normalized to the range of [0,1], is used as the score for the interface quality assessment of the protein complex model.

[0165] The model T1206TS304_2 of the protein complex target T1206 with an amino acid sequence length of 474 predicted by the AlphaFold3 server in the CASP16 protein structure prediction competition is used as an implementation case. Its native contact retention ratio f nat is 0.907, and the ligand root mean square deviation L RMS is 0.638, and the interface root mean square deviation i RMS The value of the label is 1, the DockQ score is 0.943, and the interface quality score obtained by the above method is 0.885, which is correctly classified and close to the real DockQ score.

[0166] The above description is the result of an example given by the present invention. Obviously, the present invention is not only suitable for the above embodiment, but also can be implemented with various changes without departing from the basic spirit of the present invention and without exceeding the content involved in the essential content of the present invention.

Claims

1. A method for assessing the interface quality of a protein complex model based on a multi-scale equivariant graph neural network, characterized in that: First, redundant native protein complex crystal structures were screened out from the PRISM protein-protein interaction database, and a decoy structure was generated for each native protein complex structure using the HDock docking algorithm. Then, surface-level features, atomic-level features, and residue-level features were extracted based on each protein complex decoy structure, and input into an equivariant graph neural network that integrates multi-scale information. An embedded representation was then generated through a multi-head self-attention mechanism. Finally, the interface quality score of the protein complex decoy structure was obtained by calculating the difference in distance between the embedded structure and the positive and negative prototypes in high-dimensional space.

2. The method for evaluating the interface quality of a protein complex model based on a multi-scale equivariant graph neural network according to claim 1, wherein: The method comprises the following steps: 1) Construct a protein complex bait structure dataset; 2) Prepare label data of the protein complex bait structure, use the DockQ tool to align the protein complex bait structure with its native structure, and classify all bait structures into two categories: incorrect conformation and near-native conformation according to the calculated scores; 3) Extracting the molecular surface interaction fingerprint of the bait structure of the protein complex; 4) Extracting atomic-level features of the protein complex bait structure: Using the RDKit tool to read the structure of the protein complex model, atomic information is obtained, including atomic coordinates, atom type, formal charge, number of connected hydrogens, chirality, aromaticity, and ring type, followed by chemical bond information, including chemical bond type and bond length; 5) Extract residue-level features of the protein complex bait structure: Use tools such as PyRosetta to read the structure of the protein complex model and extract geometric features, physicochemical features, and evolutionary features; 6) Convert the protein complex decoy structure into a graph representation, constructing a heterogeneous graph consisting of atoms and surface points, and a graph consisting of residues; 7) Build a multi-scale equivariant graph neural network, perform neural network model training and parameter optimization, and predict the interface quality of the bait structure.

3. The method for evaluating the interface quality of a protein complex model based on a multi-scale equivariant graph neural network according to claim 2, wherein: The process of step 3) is as follows: 3.1) Use MSMS tools to process the three-dimensional structure of the protein complex bait structure, with a radius of The probe ball rolls on the surface of the three-dimensional structure, converting the solvent exclusion surface into a mesh composed of triangular facets; 3.2) Input the surface mesh generated in 3.1) into PDB2PQR and APBS tools, and obtain the surface electrostatic potential and charge distribution by solving the linear Poisson-Boltzmann equation; 3.3) For each surface point, identify its nearest neighbor atom and determine the amino acid residue to which the atom belongs. According to the type of the residue, consult the Kyte & Doolittle scale and assign the corresponding hydrophobicity index, where positive values ​​indicate hydrophobicity and negative values ​​indicate hydrophilicity; 3.4) Use the knowledge-based hydrogen bond statistical potential to identify the hydrogen bond donor and acceptor positions on the molecular surface. The calculation formula is as follows: Where E(p) represents the hydrogen bond potential under the conformational parameter p, f protein (p) represents the probability density of the hydrogen bond geometry parameter p observed in high-resolution protein crystal structures, and f random (p) represents the probability density of the geometric parameter p under a random distribution background; the difference between the experimentally observed hydrogen bond geometric distribution and the theoretical random distribution is expressed in the form of logarithmic probability. The higher the frequency, the more consistent it is with the real structure, the lower the energy, and the more stable it is. 3.5) For each surface point, calculate its local curvature feature, First calculate the maximum curvature k1 and minimum curvature k2 of the point, and then calculate the shape index according to the following formula: This index maps the local surface shape to the range [-1,1], representing the change from concave to convex.

4. The method for evaluating the interface quality of a protein complex model based on a multi-scale equivariant graph neural network according to claim 2, wherein: The process of step 5) is as follows: 5.1) One-hot encode the amino acid sequence of each chain in the complex to generate a binary matrix of size L × 20, where L is the sequence length and the 20 dimensions correspond to the 20 standard amino acid types; 5.2) Extract seven types of physicochemical properties corresponding to residue types, including isoelectric point, polarity, acidity and alkalinity, hydrogen bond acceptor, hydrogen bond donor, n-octanol-water partition coefficient, and topological polar surface area; 5.3) Use the BLOSUM62 substitution scoring matrix to represent the conservation of substitutions between amino acids, encoding each residue as a 20-dimensional vector (i.e., the substitution score for other amino acid types), generating a feature matrix of shape L×20; 5.4) Extract the main chain torsion angles φ and ψ and calculate the main chain atoms C i -N i+1 The bond lengths between bond angle; 5.5) Use the DSSP tool to obtain secondary structure information and generate a binary matrix of size L × 3, where L is the sequence length and the three dimensions correspond to the helical, folded, and random coil states respectively; 5.6) Use FreeSASA tool to calculate the solvent accessible surface area SASA and relative solvent accessible surface area RASA of each residue; 5.7) Calculate Rosetta monomer energy terms, including the main chain dihedral angle preference term rama_prepro, the probability distribution term p_aa_pp of the main chain dihedral angles φ and ψ, the main chain dihedral angle ω term omega, and the side chain rotamer energy fa_dun; 5.8) Calculate the distance matrix between residues, including C β -C β 、Tip-Tip、C α -Tip, Tip-C α Distances between atoms; extract inter-residue angles, including the distances along the C link connecting two residues. β The dihedral angle omega of the rotation of the virtual axis of the atom, and the C β The dihedral angles theta and phi of the orientation of atoms relative to a local reference frame centered on another residue; 5.9) Calculate the Rosetta two-body energy terms, including main-chain long-range hydrogen bonds hbond_lr_bb, main-chain short-range hydrogen bonds hbond_sr_bb, main-chain-side-chain hydrogen bonds hbond_bb_sc, side-chain-side-chain hydrogen bonds hbond_sc, van der Waals attractive terms fa_atr, van der Waals repulsive terms fa_rep, isotropic solvation free energies fa_sol, anisotropic solvation free energies lk_ball_wtd, and Coulomb electrostatic potential fa_elec.

5. The method for assessing the interface quality of a protein complex model based on a multi-scale equivariant graph neural network according to claim 2, wherein: The process of step 6) is as follows: 6.1) Construct a heterogeneous graph of atoms and surface points, expressed as follows: G H =(v A ,he Ab ,he Ac ,v S ,he S ,he AS ) where v A Represents atoms, v S represents a surface point, ε Ab represents chemical bonds, ε Ac Indicates interatomic contact, i.e., distances less than ε S represents the edge of the surface triangle mesh, ε AS represents the edge between a surface point and its nearest neighbor atom; 6.2) Construct a residue-level graph representation in the following form: G R =(v R ,he R ) where v R represents the residue, ε R Represents the interaction between residues, that is, residue C β The distance between atoms is less than 6. The method for assessing the interface quality of a protein complex model based on a multi-scale equivariant graph neural network according to claim 2, wherein: The process of step 7) is as follows: 7.1) The heterogeneous graph consisting of atoms and surface points and the residue graph are fed into two independent equivariant graph neural network modules. The representations of nodes and edges are updated through coordinate transformation and message passing. The equivariant graph neural network is updated as follows: where h i represents the feature of node i, x i represents the spatial coordinates of the node, ‖x i -x j ‖ 2 is the relative position information between nodes, e ij represents the edge feature, N(i) represents the neighbor set of node i, φ e 、φ h 、φ x is a learnable function, φ x The output of is a scalar, so the coordinates are updated in the same direction as x. i -x j Maintain consistency and ensure equal variability; 7.2) The node representations output by the two equivariant graph neural network modules are stacked into a graph-level representation through global average pooling. Then, a multi-head self-attention mechanism is used to learn long-range interactions. The formula of the multi-head self-attention mechanism is as follows: Q=XW Q K=XW K V=XW V MultiHead(Q,K,V)=Concat(head1,…,head h )W O Where X represents the input sequence, Q, K, and V represent the query vector, key vector, and value vector, and W Q 、W K 、W V 、W o represents the learnable parameter matrix, d k Indicates the dimensions of Q and K of each attention head, head i represents the output of the i-th attention head, and MultiHead represents the information of different subspaces captured in parallel; 7.3) The representations obtained by the multi-head self-attention mechanism are globally average pooled and normalized. The distances between the negative prototype vector and the positive prototype vector updated during training, i.e., the distances between cluster centers, are calculated. The range of the distance difference is [-2, 2]. The larger the distance, the closer it is to the positive prototype, and the smaller the distance, the closer it is to the negative prototype. The score is normalized to the range of [0, 1] as the interface quality assessment score of the protein complex model.

7. The method for evaluating the interface quality of a protein complex model based on a multi-scale equivariant graph neural network according to claim 2, wherein: The process of 1) is: 1.1) 2465 co-crystallized native structures of protein complexes were screened from the PRISM non-redundant protein-protein interaction database; 1.2) Use TM-align to calculate the TM-score between each pair of binding interfaces of the native structures of all protein complexes; 1.3) Perform hierarchical clustering based on the calculated TM-score matrix and divide the obtained protein complex natural structure set into training, validation and testing subsets; 1.4) Each protein complex constituting protein-protein interactions in the subset was split into two monomers, and protein-protein docking was re-performed using HDOCK to generate 100 protein complex bait structures for each protein complex native structure.

8. The method for evaluating the interface quality of a protein complex model based on a multi-scale equivariant graph neural network according to claim 2, wherein: The process of 2) is: 2.1) Defined as: When the two residues are in contact, for each protein complex bait structure in the data set, the DockQ tool is used to align it with the corresponding native structure, and the number of contacts between the bait structure and the native structure is counted, and the proportion of the contacts to the total number of contacts in the native structure is calculated. 2.2) Perform structural superposition of the decoy structure and the native structure based on the receptor and calculate the root mean square deviation (L_RMS) of the ligand main chain atoms; 2.3) Relax the contact threshold to Perform structural superposition of the main chain atoms of the interface of the decoy structure and the interface of the native structure, and calculate the interface root mean square deviation i_RMS; 2.4) Labels refer to the CAPRI classification standard. When the f_nat of the bait structure is less than 0.1, or the L_RMS is greater than And i_RMS is greater than When , the bait structure is defined as an incorrect conformation, and the remaining bait structures are defined as near-native conformations.

Citation Information

Cited By

  • Antigen-antibody docking and antibody generation method based on coarse-grained structure model

    CN121034391A

  • Compound bait data set construction method based on structure and sequence collaborative redundancy elimination

    CN121583345A

  • A method for multi-scale quality assessment of protein complex structure model

    CN122392646A

  • Multi-index consensus screening method for large-scale complex protein model pool

    CN122474119A