Method and device for predicting influence of mutation on protein-protein binding affinity
By using pre-trained graph neural network model and binding affinity change model, combined with the enhancement data set and the XGBoost model, the problem of the difficult impact of mutations on protein-protein binding affinity is solved, and high-precision prediction effect is achieved.
Patent Information
- Application Number
- CN202510464333.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to efficiently and simply predict the effect of mutations on protein-protein binding affinity, and traditional methods are costly and models are prone to overfitting.
The pre-trained graph neural network model and binding affinity change model are used to train the graph neural network model through the enhanced data set to learn the graph characteristics that characterize the interatomic interactions and protein-protein interface characteristics, and the XGBoost model is used to capture the relationship between graph characteristics and binding affinity changes.
High-precision prediction of the effect of mutations on protein-protein binding affinity is achieved, the prediction accuracy is improved, and overfitting is avoided.
Smart Images

Figure CN119993282A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bioinformatics, and in particular relates to a method and a device for predicting the influence of mutation on protein-protein binding affinity. Background Art
[0002] Protein-protein interactions (PPIs) play a vital role in biology. They are the basis of many basic biological processes in cells, including but not limited to cell signaling, immune response, cell cycle regulation, metabolic pathways, and apoptosis. An important basis for determining whether proteins can interact with each other is the prediction of protein-protein binding affinity. However, mutations have a great impact on protein-protein binding affinity. First, if the amino acid residues at the binding interface mutate, it may directly affect the interaction between proteins. For example, a mutation in a hydrogen bond donor or acceptor may destroy the original hydrogen bond, thereby reducing the binding affinity; second, a mutation in a hydrophobic amino acid on the interface (such as changing a hydrophobic residue to a polar residue) may destroy the hydrophobic core, thereby weakening the interaction; third, some mutations may lead to changes in the overall structure of the protein, so that the binding interface no longer maintains the original conformation, thereby affecting the binding affinity; fourth, mutations may cause changes in surface charge, thereby affecting protein-protein interactions. For example, mutating a positively charged lysine to a neutral alanine may reduce electrostatic interactions. In order to better understand the mechanism of protein interaction and provide valuable guidance for drug design and disease treatment, how to predict the effect of mutations on protein-protein binding affinity in an efficient and simple way has attracted increasing attention.
[0003] Traditional binding affinity prediction methods rely on complex feature engineering and a large amount of experimental data. These methods are not only costly but also prone to overfitting. Deep learning models are currently used for binding affinity prediction, but they do not predict the impact of mutations on protein-protein binding affinity. When constructing training data sets, they only use random masking methods or simple data enhancement methods for data enhancement, and do not fully consider the relative position relationship and geometric characteristics between residues in the protein-protein interaction interface. There are obvious defects, resulting in inaccurate prediction results. Summary of the invention
[0004] The purpose of the present invention is to provide a method, device, equipment and storage medium for predicting the effect of mutation on protein-protein binding affinity, which can predict the effect of mutation on protein-protein binding affinity with high prediction accuracy.
[0005] The first aspect of the present invention discloses a method for predicting the effect of a mutation on protein-protein binding affinity, comprising: Inputting the PDB files of the complex molecules before and after mutation into a pre-trained graph neural network model to obtain graph features, and inputting the graph features into a pre-trained binding affinity change model to obtain binding affinity change results; The pre-trained graph neural network model is obtained by training the graph neural network model with an enhanced data set, wherein the enhanced data set is obtained by performing data enhancement on the training data set based on the dihedral angle relationship between the residues at the protein-protein interaction interface and the main chain, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics.
[0006] In some embodiments, the step of training the graph neural network model includes: Get composite data; Changing the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface of the complex molecule in the complex data to obtain first enhanced data; perturbing the coordinates of the complex molecules in the complex data to obtain second enhanced data; combining the composite data, the first enhanced data, and the second enhanced data into a training data set; The graph neural network model is trained using the training data set.
[0007] In some embodiments, changing the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface of the complex molecule in the complex data to obtain the first enhanced data comprises: For each complex molecule in the complex data, side chain residues are randomly selected and side chain angles are randomly sampled according to Ramachandran distribution to perturb the structure of the complex molecule; the von Mises probability density function is used as the kernel function, and the probability of the rotational configuration of the complex molecule is estimated according to the Bayesian rule to obtain the probability density; random sampling is performed according to the probability density to generate complex molecules with different rotational configurations to obtain the first enhanced data.
[0008] In some embodiments, the probability density is expressed as: , in, Rotational configuration The probability of not being affected by the main chain, The number of amino acids from the N atom of the nth amino acid residue to the n+1th amino acid residue in the main chain The dihedral angles between atoms, The nth amino acid residue in the main chain The dihedral angle between the atom and the N atom of the n+1th amino acid residue, is the von Mises probability density function, For the rotational configuration Other possible rotation configurations are related.
[0009] In some embodiments, obtaining the second enhanced data by perturbing the coordinates of the complex molecules in the complex data comprises: For residues located at the protein-protein interface, the bond angle is randomly adjusted within the preset angle range. When the covalent bond is a single bond, the bond length is randomly adjusted within the first preset range; when the covalent bond is a double bond, the bond length is randomly adjusted within the second preset range.
[0010] In some embodiments, the binding affinity change model is an XGBoost model, and the step of training the binding affinity change model comprises: The PDB files of the complex molecules before and after mutation are used as the input of the pre-trained graph neural network model to obtain the graph features of the complex molecules before and after mutation; All of the graph features are combined into a training set to train the binding affinity change model.
[0011] The second aspect of the present invention discloses a device for predicting the effect of mutation on protein-protein binding affinity, comprising: A graph feature module, for inputting the PDB files of the complex molecules before and after mutation into a pre-trained graph neural network model to obtain graph features, wherein the pre-trained graph neural network model is obtained by training the graph neural network model with an enhanced data set, wherein the enhanced data set is obtained by performing data enhancement on the training data set based on the dihedral angle relationship between the residues and the main chain of the protein-protein interaction interface, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics; The prediction module is used to input the graph features into a pre-trained binding affinity change model to obtain a binding affinity change result.
[0012] In some embodiments, a training module is further included, wherein the training module includes a data acquisition unit, a data enhancement unit and a training unit; The data acquisition unit is used to acquire composite data; The data enhancement unit is used to change the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface of the complex molecules in the complex data to obtain first enhanced data; and to perturb the coordinates of the complex molecules in the complex data to obtain second enhanced data; The training unit is used to combine the composite data, the first enhanced data, and the second enhanced data into a training data set, and use the training data set to train the graph neural network model.
[0013] The third aspect of the present invention discloses an electronic device, comprising a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the method for predicting the effect of mutations on protein-protein binding affinity disclosed in the first aspect.
[0014] The fourth aspect of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the method for predicting the effect of mutations on protein-protein binding affinity disclosed in the first aspect.
[0015] The beneficial effect of the present invention is that by fully considering the dihedral angle between the residues and the main chain of the protein-protein interaction interface, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics, the data of the complex is enhanced, the graph neural network model is trained, and the spectral features that characterize the atomic interactions and protein-protein interface characteristics are learned. Then, the binding affinity change model is used to capture the relationship between these spectral features and the binding affinity change, which can predict the effect of mutations on protein-protein binding affinity, and the graph neural network model is fully trained to improve the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings herein show specific examples of the technical solutions described in the present invention, and together with the specific implementation methods, constitute a part of the specification, and are used to explain the technical solutions, principles and effects of the present invention.
[0017] Unless otherwise specified or defined, the same reference numerals in different drawings represent the same or similar technical features, and the same or similar technical features may also be represented by different reference numerals.
[0018] Figure 1 It is a flow chart of a method for predicting the effect of a mutation on protein-protein binding affinity disclosed in an embodiment of the present invention; Figure 2 is a flowchart of a training graph neural network model according to an embodiment of the present invention; Figure 3 is a schematic diagram of the structure of a device for predicting the effect of mutations on protein-protein binding affinity according to an embodiment of the present invention; Figure 4 It is a structural schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] Unless otherwise specified or defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art. In combination with the technical solution of the present invention in a realistic scenario, all technical and scientific terms used herein may also have meanings corresponding to the purpose of implementing the technical solution of the present invention. The "first, second..." used herein is only used to distinguish the names and does not represent a specific quantity or order. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0020] It should be noted that when a component is considered to be "fixed to" another component, it can be directly fixed to the other component or there can be a central component; when an component is considered to be "connected to" another component, it can be directly connected to the other component or there can be a central component at the same time; when an component is considered to be "installed on" another component, it can be directly installed on the other component or there can be a central component at the same time. When an component is considered to be "set on" another component, it can be directly set on the other component or there can be a central component at the same time.
[0021] Unless otherwise specified or defined, the "said" and "the" used in this document refer to the technical features or technical contents mentioned or described before the corresponding position, and the technical features or technical contents may be the same as or similar to the technical features or technical contents mentioned therein. In addition, the terms "including" and "having" and any variations thereof used in this document are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices.
[0022] The present invention proposes a method for predicting the effect of mutation on protein-protein binding affinity. The PDB files of the complex molecules before and after mutation are input into a pre-trained graph neural network model (GNN), and the spectral features that can characterize the interactions between atoms and the protein-protein interface are directly learned without feature engineering. Then, the binding affinity change model is used to capture the relationship between the input spectral features and the output change in binding affinity during mutation. By predicting the effect of mutation on binding affinity, it is helpful to predict the effect of different amino acid substitutions on the binding site of drugs and proteins, so as to design drug molecules with higher affinity, which has very broad practical significance in improving the efficacy and safety of drugs and accelerating the development process of new drugs.
[0023] The protein described in the present invention is a broad protein, including polypeptides. Protein-protein interaction includes protein-protein interaction, protein-polypeptide interaction, and polypeptide-polypeptide interaction.
[0024] The embodiment of the present invention discloses a method for predicting the effect of mutation on protein-protein binding affinity, which can be implemented by computer programming. The execution subject of the method can be an electronic device such as a computer, a laptop, a tablet computer, or a control chip embedded in an electronic device, and the present invention is not limited to this.
[0025] In order to facilitate the understanding of the present invention, specific embodiments of the present invention will be described in more detail below with reference to the accompanying drawings.
[0026] like Figure 1 As shown, the method comprises the following steps: Step S100: inputting the PDB files of the complex molecules before and after the mutation into a pre-trained graph neural network model to obtain graph features, wherein the pre-trained graph neural network model is obtained by training the graph neural network model with an enhanced data set, and the enhanced data set is obtained by performing data enhancement on the training data set based on the dihedral angle relationship between the residues and the main chain of the protein-protein interaction interface, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics; Among them, the PDB (Protein Data Bank) file of the complex molecule is a standard format for storing the structural information of three-dimensional biological macromolecules (such as proteins, nucleic acids, etc.). The PDB file contains data such as the atomic coordinates, residue information, chain information, and connection information of the molecule. The PDB file of the complex molecule can be downloaded from the PDB database or used to build a molecular model and export it to the PDB format when performing molecular modeling using software (such as PyMOL, Chimera, GROMACS, etc.).
[0027] When the graph neural network model (GNN) extracts graph features from the PDB file of a complex molecule, the molecule can be regarded as a graph, where atoms are nodes and chemical bonds are edges. GNN can naturally process this graph structure and capture the relationship and interaction between atoms in the molecule. It learns local features by aggregating information from neighboring nodes, and can also capture global features through a multi-layer network to obtain graph features.
[0028] Before the graph neural network model is applied, it needs to be trained first. After the training, a pre-trained graph neural network model is obtained, and then the PDB files of the complex molecules before and after the mutation are input into the pre-trained graph neural network model to obtain the graph characteristics. In order to fully consider the relative position relationship and geometric characteristics between the residues at the protein-protein interaction interface when training the graph neural network model, this embodiment performs data enhancement on the training data set based on the dihedral angle relationship between the residues at the protein-protein interaction interface and the main chain, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics when training the graph neural network model. Through data enhancement, the pre-trained graph neural network model can learn the graph characteristics that characterize the interactions between atoms and the characteristics of the protein-protein interface.
[0029] like Figure 2 As shown, the steps for training the graph neural network model include: Step S110: obtaining composite data; The PDB file (Protein Data Bank file) of the complex is obtained from a data set such as PDB-BIND, which stores and manages molecular structures and corresponding chemical information. Specifically, since the PDB-BIND database provides experimentally measured binding affinity data of protein-ligand complexes, the complex data recorded in the PDB-BIND database is used as the basic data for constructing the training data set.
[0030] Step S120: changing the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface of the complex molecule in the complex data to obtain first enhanced data; The side chain residues at the protein-protein interaction interface are usually involved in the formation of hydrogen bonds, hydrophobic interactions, ionic bonds and other interaction forces. Changes in the three-dimensional coordinates of the side chain residues may cause changes in the spatial configuration of the protein, thereby affecting its function and interaction ability.
[0031] The specific process of changing the three-dimensional coordinates of the side chain residues is as follows: for each complex molecule in the complex data, randomly select side chain residues and randomly sample side chain angles according to the Ramachandran distribution to perturb the structure of the complex molecule; use the von Mises probability density function as the kernel function, estimate the probability of the rotational configuration of the complex molecule according to the Bayesian rule, and obtain the probability density; finally, perform random sampling based on the probability density to generate complex molecules with different rotational configurations and obtain the first enhanced data.
[0032] Specifically, according to research, the configuration probability distribution of side chain residues can be regarded as Ramachandran distribution, and the density function can be estimated by kernel density estimation. The von Mises probability density function (a probability distribution used to describe periodic data, especially suitable for angle or direction data) is used as the kernel function. The expression of the von Mises probability density function is: , Where x is the angle, is the zero-order modified Bessel function with concentrated parameters is inversely proportional to the square width of the von Mises kernel. On this circle, The larger the value, the narrower the kernel width.
[0033] Each dihedral configuration will become the center of the kernel function, which means that there will be a distribution of kernel functions for each possible configuration. For each rotational configuration r of a given residue type, a probability density estimate is determined, that is, , and for each rotational configuration r, it obeys the Ramachandran distribution. By randomly selecting residues and randomly sampling their side chain angles according to the Ramachandran distribution, the original structure of the protein can be disturbed to simulate its dynamic changes in the biological environment. For example: randomly select one or more side chain residues from the protein structure, randomly sample the possible values of the side chain angle according to the Ramachandran distribution, generate a random number, determine the value of the side chain angle, update the three-dimensional coordinates of the side chain, and ensure that the sampled angle conforms to the selected probability distribution. By randomly sampling the side chain angles, it can be used to simulate the dynamic behavior of proteins in different environments.
[0034] Then, the probability of the rotational configuration is estimated by Bayes' rule: , that is, the probability density is obtained, and the specific expression is: , in, is the probability that the rotational configuration r is not affected by the main chain, Refers to the chain from the N atom of the nth amino acid residue to the n+1th amino acid residue. Dihedral angles between atoms. Refers to the number of amino acids in the main chain starting from the nth amino acid residue The dihedral angle between the atom and the N atom of the n+1th amino acid residue, is the von Mises probability density function, For the rotational configuration Other possible rotational configurations related to the side chain residues. This probability density reflects the relative likelihood of different configurations that the side chain residues may adopt under specific conditions.
[0035] Next, random sampling is performed on the probability density to obtain different selected configurations, generate complex molecules with different rotational configurations, and obtain the complex structure data after data enhancement of the basic data by the dihedral angle between the residues and the main chain of the protein-protein interaction interface, i.e., the first enhanced data. In this step, the basic data is enhanced by the dihedral angle relationship between the residues and the main chain, which can expand the basic data by 10,000 times.
[0036] Step S130: perturbing the coordinates of the complex molecules in the complex data to obtain second enhanced data; For protein-protein complexes, due to certain interactions between molecules, this interaction often causes changes in the three-dimensional structure of the protein, thereby affecting the length and angle of the internal covalent bonds. When amino acids in proteins mutate, changes in local conformations or interactions between atoms usually cause changes in the bond length and bond angle of residues located at the protein-protein interface. Therefore, the contraction and stretching of side chain residues and main chain chemical bonds and the bond angle characteristics can be used to enhance the basic data. Specifically, for residues located at the protein-protein interface, the bond angle is randomly adjusted within a preset angle range. When the covalent bond is a single bond, the bond length is randomly adjusted within the first preset range; when the covalent bond is a double bond, the bond length is randomly adjusted within the second preset range.
[0037] In this embodiment, for residues located at the protein-protein interface, when the covalent bond is a single bond, the bond length is randomly adjusted within the range of 0.01-0.05A; when the covalent bond is a double bond, the bond length is randomly adjusted within the range of 0.01-0.03A. For the bond angle, the range of 0-10 o Randomly adjust within the range to obtain enhanced data, namely the second enhanced data. Through the above steps, the basic data can be expanded 100 times.
[0038] In the process of data augmentation of the basic data, considering changes in bond length and bond angle can help the graph neural network model learn how to identify protein-protein interactions.
[0039] Step S140: combining the composite data, the first enhanced data, and the second enhanced data into a training data set; Step S150: Use the training data set to train the graph neural network model.
[0040] The composite data, the first enhanced data, and the second enhanced data are combined into a training data set, and the training data set is divided into a training set and a validation set in a ratio of 9:1, and the graph neural network model is trained to obtain a pre-trained graph neural network model. The training steps of the graph neural network model are conventional techniques in the art and will not be repeated here.
[0041] Step S200: inputting the graph features into the pre-trained binding affinity change model to obtain the binding affinity change results.
[0042] Deep learning models usually contain a large number of parameters and complex structures, such as multiple hidden layers and a large number of neurons. This makes the deep learning model have a strong expressive ability and can fit very complex functions. However, this strong expressive ability also makes it easier for the model to perform well on training data, but perform poorly on unseen data, resulting in overfitting problems. Therefore, the present invention does not directly use the graph neural network model to predict the change in affinity of the complex before and after mutation, but inputs the graph features into the pre-trained binding affinity change model to obtain the binding affinity change results, and predicts the mutation characteristics of the complex obtained by combining the binding affinity change model with the pre-trained graph neural network learning, which can improve the prediction accuracy.
[0043] In this embodiment, the binding affinity change model is an XGBoost model. The XGBoost model can capture the complex nonlinear relationship between graph features and performs well in processing data sets with complex relationships due to its excellent prediction performance and computational efficiency.
[0044] Among them, the process of training the binding affinity change model is: taking the PDB file of the complex molecule before and after mutation as the input of the pre-trained graph neural network model, and obtaining the spectral features of the complex molecule before and after mutation output by the graph neural network model; combining all the spectral features of the complex molecules before and after mutation into a training data set, and using the training data set to train the binding affinity change model to obtain the pre-trained binding affinity change model. The pre-trained binding affinity change model can output the binding affinity change results such as enhanced binding affinity, weakened binding affinity, and binding constant after mutation according to the spectral features output by the graph neural network model.
[0045] When training the graph neural network model and the binding affinity change model, two evaluation indicators, Pearson's correlation coefficient and root mean square error (RMSE), were used to evaluate the model prediction results. For the prediction result x and the benchmark true value y of n samples, the calculation formula of Pearson's correlation coefficient is: , in, and represent the means of x and y respectively.
[0046] The calculation formula of RMSE is: .
[0047] In summary, this embodiment designs a data enhancement method with physicochemical significance through the physicochemical properties of the complex molecules, namely: taking into account the dihedral angle between the residues and the main chain of the protein-protein interaction interface, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics of the side chain residues, by changing the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface and by perturbing the coordinates of the complex molecules to obtain data-enhanced complex molecule data, which can fully train the graph neural network model. After training, the pre-trained graph neural network model uses the PDB files of the complex molecules before and after mutation as input to efficiently extract the graph features of the complex molecules, and uses the XGBoost model to use the learned graph features to predict the affinity changes before and after the protein mutation, which can predict the effect of amino acid mutations on the affinity of the complex with high prediction accuracy.
[0048] like Figure 3 As shown, based on the above-mentioned method for predicting the effect of mutation on protein-protein binding affinity, an embodiment of the present invention discloses a device for predicting the effect of mutation on protein-protein binding affinity, comprising: A graph feature module 600 is used to input the PDB files of the complex molecules before and after the mutation into a pre-trained graph neural network model to obtain graph features, wherein the pre-trained graph neural network model is obtained by training the graph neural network model with an enhanced data set, wherein the enhanced data set is obtained by performing data enhancement on the training data set based on the dihedral angle relationship between the residues and the main chain of the protein-protein interaction interface, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics; The prediction module 610 is used to input the graph features into a pre-trained binding affinity change model to obtain a binding affinity change result.
[0049] In some embodiments, a training module is further included, wherein the training module includes a data acquisition unit, a data enhancement unit and a training unit; The data acquisition unit is used to acquire composite data; The data enhancement unit is used to change the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface of the complex molecules in the complex data to obtain first enhanced data; and to perturb the coordinates of the complex molecules in the complex data to obtain second enhanced data; The training unit is used to combine the composite data, the first enhanced data, and the second enhanced data into a training data set, and use the training data set to train the graph neural network model.
[0050] like Figure 4 As shown, an embodiment of the present invention discloses an electronic device, including a memory 401 storing executable program codes and a processor 402 coupled to the memory 401; The processor 402 calls the executable program code stored in the memory 401 to execute the method for predicting the effect of mutation on protein-protein binding affinity described in the above embodiments.
[0051] The embodiments of the present invention further disclose a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the method for predicting the effect of mutations on protein-protein binding affinity described in the above embodiments.
[0052] The purpose of the above embodiments is to exemplarily reproduce and deduce the technical solution of the present invention, and to fully describe the technical solution, purpose and effect of the present invention. Its purpose is to make the public understand the disclosed content of the present invention more thoroughly and comprehensively, and it does not limit the scope of protection of the present invention.
[0053] The above embodiments are not exhaustive enumerations of the present invention, and there may be multiple other implementations not listed. Any replacement and improvement made without violating the concept of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for predicting the effect of mutation on protein-protein binding affinity, characterized in that: include: Inputting the PDB files of the complex molecules before and after mutation into a pre-trained graph neural network model to obtain graph features, and inputting the graph features into a pre-trained binding affinity change model to obtain binding affinity change results; The pre-trained graph neural network model is obtained by training the graph neural network model with an enhanced data set, wherein the enhanced data set is obtained by performing data enhancement on the training data set based on the dihedral angle relationship between the residues at the protein-protein interaction interface and the main chain, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics.
2. The method for predicting the effect of mutation on protein-protein binding affinity according to claim 1, characterized in that: The steps of training the graph neural network model include: Get composite data; Changing the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface of the complex molecule in the complex data to obtain first enhanced data; perturbing the coordinates of the complex molecules in the complex data to obtain second enhanced data; combining the composite data, the first enhanced data, and the second enhanced data into a training data set; The graph neural network model is trained using the training data set.
3. The method for predicting the effect of mutation on protein-protein binding affinity according to claim 2, characterized in that: Changing the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface of the complex molecule in the complex data to obtain the first enhanced data comprises: For each complex molecule in the complex data, side chain residues are randomly selected and side chain angles are randomly sampled according to Ramachandran distribution to perturb the structure of the complex molecule; the von Mises probability density function is used as the kernel function, and the probability of the rotational configuration of the complex molecule is estimated according to the Bayesian rule to obtain the probability density; random sampling is performed according to the probability density to generate complex molecules with different rotational configurations to obtain the first enhanced data.
4. The method for predicting the effect of mutation on protein-protein binding affinity according to claim 3, characterized in that: The expression of the probability density is: , in, Rotational configuration The probability of not being affected by the main chain, The number of amino acids from the N atom of the nth amino acid residue to the n+1th amino acid residue in the main chain The dihedral angles between atoms, The nth amino acid residue in the main chain The dihedral angle between the atom and the N atom of the n+1th amino acid residue, is the von Mises probability density function, For the rotational configuration Other possible rotation configurations are related.
5. The method for predicting the effect of mutation on protein-protein binding affinity according to claim 2, characterized in that: The step of perturbing the coordinates of the complex molecules in the complex data to obtain second enhanced data comprises: For residues located at the protein-protein interface, the bond angle is randomly adjusted within the preset angle range. When the covalent bond is a single bond, the bond length is randomly adjusted within the first preset range; when the covalent bond is a double bond, the bond length is randomly adjusted within the second preset range.
6. The method for predicting the effect of mutation on protein-protein binding affinity according to claim 1, characterized in that: The binding affinity change model is an XGBoost model, and the steps of training the binding affinity change model include: The PDB files of the complex molecules before and after mutation are used as the input of the pre-trained graph neural network model to obtain the graph features of the complex molecules before and after mutation; All of the graph features are combined into a training set to train the binding affinity change model.
7. A device for predicting the effect of mutation on protein-protein binding affinity, characterized in that: include: A graph feature module, for inputting the PDB files of the complex molecules before and after mutation into a pre-trained graph neural network model to obtain graph features, wherein the pre-trained graph neural network model is obtained by training the graph neural network model with an enhanced data set, wherein the enhanced data set is obtained by performing data enhancement on the training data set based on the dihedral angle relationship between the residues and the main chain of the protein-protein interaction interface, the shrinkage and stretching of the chemical bonds between the residues and the main chain, and the bond angle characteristics; The prediction module is used to input the graph features into a pre-trained binding affinity change model to obtain a binding affinity change result.
8. The device for predicting the effect of mutation on protein-protein binding affinity according to claim 7, characterized in that: Also includes a training module, the training module includes a data acquisition unit, a data enhancement unit and a training unit; The data acquisition unit is used to acquire composite data; The data enhancement unit is used to change the three-dimensional coordinates of the side chain residues of the protein-protein interaction interface of the complex molecule in the complex data to obtain first enhanced data; perturbing the coordinates of the complex molecules in the complex data to obtain second enhanced data; The training unit is used to combine the composite data, the first enhanced data, and the second enhanced data into a training data set, and use the training data set to train the graph neural network model.
9. An electronic device, characterized in that: It comprises a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the method for predicting the effect of mutation on protein-protein binding affinity according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program enables a computer to execute the method for predicting the effect of a mutation on protein-protein binding affinity according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for training antibody-protein binding affinity prediction model
CN115206415A
Protein and ligand affinity prediction method based on graph neural network decoupling
CN116312758A
Protein ligand affinity prediction method and system for virtual screening, electronic equipment and storage medium
CN118782136A
Protein-protein binding affinity prediction method and device, equipment and medium
CN119601080A
Systems and methods for polymer side-chain conformation prediction
US20250014673A1