Method, device, equipment and medium for predicting protein-small molecule binding structure
By combining the TANKBind model with molecular dynamics simulations, the automated prediction of the binding structure of proteins and small molecule compounds has been achieved, solving the problems of inaccurate prediction and low efficiency in existing technologies, and improving the accuracy and efficiency of drug design.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN READLINE BIOTECH CO LTD
- Filing Date
- 2024-07-26
- Publication Date
- 2026-08-04
AI Technical Summary
Existing technologies struggle to accurately predict the structural stability of protein-small molecule compounds in drug design, and are computationally expensive, inefficient, and heavily influenced by human factors, leading to process uncertainty and resource waste.
By employing the TANKBind model combined with molecular dynamics simulations, deep learning is used to predict the interactions between proteins and small molecules, generate interaction matrices, and perform molecular dynamics simulations to automatically assess the stability of the binding structure.
It improves the accuracy and efficiency of predicting the structure of protein and small molecule complexes, reduces experimental costs and time, and provides more precise guidance for drug design, applicable to drug development and biochemical research.
Smart Images

Figure CN121415918B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of protein-small molecule binding prediction technology, and particularly to methods, apparatus, equipment and media for predicting protein-small molecule binding structures. Background Technology
[0002] First, drug design typically begins with a deep understanding of the disease. Scientists study the biological mechanisms of disease, including how pathogens grow, cell signaling pathways, and the structure and function of proteins. Through research in these areas, scientists can identify potential drug targets—biomolecules that play a crucial role in the disease process. Next, drug design requires finding compounds that can interact with these targets and alter their function. These compounds are often called drug candidates. Molecular docking is a commonly used technique for predicting how drugs interact with targets. It uses calculations and simulations to predict binding patterns and affinities between molecules, thus helping scientists screen for promising drug candidates.
[0003] Molecular docking typically involves the following steps: First, computational methods are used to generate the structure of candidate drugs. Then, through simulation and calculation, the conformational and energy changes of the candidate drug upon binding to the target are predicted. Finally, based on these predictions, the affinity and pharmacodynamic properties of the candidate drug are evaluated, and the design is further optimized. The development of molecular docking technology has made the drug design process more efficient and precise. It can help scientists quickly screen for potential drugs from a large number of compounds and provide theoretical guidance for drug development. However, despite the important role of molecular docking technology in drug design, it still faces some challenges, such as accuracy and computational costs, requiring continuous technological and methodological improvements.
[0004] In summary, how to achieve automated prediction and assessment of the stability of protein-small molecule compound binding structures is a technical problem that needs to be solved in this field. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for predicting the binding structure of proteins and small molecules, capable of automatically predicting and evaluating the stability of the binding structure of proteins and small molecule compounds. The specific solution is as follows:
[0006] In a first aspect, this application discloses a method for predicting the binding structure of proteins and small molecules, including:
[0007] Obtain the first and second modeling structures of the protein block and the small molecule compound to be tested, respectively;
[0008] The first modeling graphical structure and the second modeling graphical structure are input into the TANKBind model so that the TANKBind model outputs an interaction matrix that integrates the first modeling graphical structure and the second modeling graphical structure. The interaction embedding is updated according to the interaction matrix, and then the interaction embedding is predicted to output the predicted protein small molecule complex. The interaction matrix describes the interaction between the protein block and the small molecule compound to be tested.
[0009] Generate the protein structure file and the small molecule structure file of the protein complex, and their respective first and second topology files containing coordinate information and force field parameter information;
[0010] Molecular dynamics simulations are performed using a preset molecular dynamics simulation tool and a simulation file generated based on the first and second topology files to obtain a molecular dynamics simulation trajectory file.
[0011] The molecular dynamics simulation trajectory file, the first topology file, and the second topology file are analyzed to obtain the simulation analysis results of the protein-small molecule binding structure.
[0012] Optionally, obtaining the first and second modeled graphical structures of the protein block and the small molecule compound to be tested, respectively, includes:
[0013] The protein is divided into protein blocks according to the preset protein block radius;
[0014] The protein block and the small molecule compound to be tested are subjected to their respective graphic modeling processes to obtain a first modeling graphic structure and a second modeling graphic structure.
[0015] Optionally, the step of inputting the first modeling graphical structure and the second modeling graphical structure into the TANKBind model, so as to output an interaction matrix that integrates the first modeling graphical structure and the second modeling graphical structure through the TANKBind model, includes:
[0016] The first and second modeling graphic structures are input into the TANKBind model so that the first and second modeling graphic structures can be transformed by the trigonometric function module of the TANKBind model to obtain the distance and angle information between the protein block and the small molecule compound to be tested, and an interaction matrix is constructed based on the distance and angle information.
[0017] Optionally, the first topology file and the second topology file, which contain coordinate information and force field parameter information, respectively, corresponding to the protein structure file and the small molecule structure file for generating the protein small molecule complex, include:
[0018] Set the force field information and simulation parameters for the protein structure file and the small molecule structure file respectively to create corresponding parameter files;
[0019] Invoke the format conversion command and convert the protein structure file into a first coordinate file and a first topology file based on the parameter file;
[0020] Based on the parameter file, the small molecule charge is calculated on the small molecule structure file, and small molecule position parameter information in various formats is generated to obtain the converted second coordinate file and second topology file.
[0021] Optionally, before performing molecular dynamics simulation using a preset molecular dynamics simulation tool and a simulation file generated based on the first topology file and the second topology file to obtain a molecular dynamics simulation trajectory file, the method further includes:
[0022] The first coordinate file, the first topology file, the second coordinate file, the second topology file, and the solvent file are merged to obtain a merged file, and the boundary information and size information of the simulation box used to accommodate the protein small molecule complex for molecular dynamics simulation are set.
[0023] A corresponding simulation file is generated based on the merged file, the boundary information, and the size information.
[0024] Optionally, the step of performing molecular dynamics simulation using a preset molecular dynamics simulation tool and a simulation file generated based on the first topology file and the second topology file to obtain a molecular dynamics simulation trajectory file includes:
[0025] The simulation file is input into a preset molecular dynamics simulation tool, which uses the particle mesh Ewald method to process long-distance electrostatic interactions in the simulation file and truncates non-bonded interactions in the simulation file at a preset cutoff distance. By setting the transformation distance, the non-bonded interactions are controlled to smoothly return to zero within the transformation distance, and the lengths of all bonds involving hydrogen atoms are constrained to obtain a molecular dynamics simulation trajectory file.
[0026] Optionally, the analysis of the molecular dynamics simulation trajectory file, the first topology file, and the second topology file to obtain the simulation analysis results of the protein-small molecule binding structure includes:
[0027] Read the molecular dynamics simulation trajectory file, the first topology file, and the second topology file;
[0028] The difference between the particle coordinates and the particle reference coordinates at different time points in the molecular dynamics simulation trajectory file is calculated using the root mean square fluctuation calculation instruction to obtain the position fluctuation of the particle in the simulation trajectory.
[0029] By combining the free energy calculation instructions with the molecular dynamics simulation trajectory file, the first topology file, and the second topology file, binding free energy calculation and related energy term calculation are performed to obtain the calculation results characterizing the binding strength between the protein and the small molecule and the contribution information of each residue to the binding.
[0030] Secondly, this application discloses a protein-small molecule binding structure prediction device, comprising:
[0031] The modeling module is used to obtain the first and second modeling graphic structures of the protein block and the small molecule compound to be tested, respectively, after modeling.
[0032] The prediction module is used to input the first modeling graphical structure and the second modeling graphical structure into the TANKBind model, so that the TANKBind model outputs an interaction matrix that integrates the first modeling graphical structure and the second modeling graphical structure, updates the interaction embedding according to the interaction matrix, and then predicts the interaction embedding to output the predicted protein small molecule complex; wherein, the interaction matrix describes the interaction between the protein block and the small molecule compound to be tested;
[0033] The file generation module is used to generate a protein structure file and a first topology file and a second topology file containing coordinate information and force field parameter information, respectively, for the protein structure file and the small molecule structure file of the protein complex.
[0034] The dynamics simulation module is used to perform molecular dynamics simulations using preset molecular dynamics simulation tools and simulation files generated based on the first topology file and the second topology file, so as to obtain molecular dynamics simulation trajectory files;
[0035] The results analysis module is used to analyze the molecular dynamics simulation trajectory file, the first topology file, and the second topology file to obtain the simulation analysis results of the protein-small molecule binding structure.
[0036] Thirdly, this application discloses an electronic device, including:
[0037] Memory, used to store computer programs;
[0038] A processor is used to execute the computer program to implement the steps of the aforementioned disclosed method for predicting the binding structure of proteins and small molecules.
[0039] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed method for predicting the binding structure of proteins and small molecules.
[0040] As can be seen, this application discloses a method for predicting the binding structure of proteins and small molecules, comprising: obtaining a first modeled graphical structure and a second modeled graphical structure after modeling a protein block and a small molecule compound to be tested, respectively; inputting the first modeled graphical structure and the second modeled graphical structure into a TANKBind model, so that the TANKBind model outputs an interaction matrix that integrates the first modeled graphical structure and the second modeled graphical structure, and updating the interaction embedding according to the interaction matrix, and then predicting the interaction embedding to output the predicted protein-small molecule complex; wherein, the interaction matrix describes the interaction between the protein block and the small molecule compound to be tested; generating a first topology file and a second topology file containing coordinate information and force field parameter information for the protein structure file and the small molecule structure file of the protein-small molecule complex, respectively; performing molecular dynamics simulation using a preset molecular dynamics simulation tool and a simulation file generated according to the first topology file and the second topology file to obtain a molecular dynamics simulation trajectory file; and analyzing the molecular dynamics simulation trajectory file, the first topology file, and the second topology file to obtain the simulation analysis results of the protein-small molecule binding structure. Therefore, by integrating the modeling graphs of proteins and small molecules using the TANKBind model, the interactions between them can be described more comprehensively and accurately, thereby improving the accuracy of predicting the structure of protein-small molecule complexes. Generating topology files and performing molecular dynamics simulations allows for the simulation of the dynamic behavior of complexes at the atomic level, enabling in-depth exploration of structural changes and interaction mechanisms during protein-small molecule binding. More accurate binding structure predictions and a deeper understanding of binding mechanisms contribute to the design of more effective drug molecules. Small molecules can be optimized and modified based on simulation analysis results, improving the binding affinity and specificity of drugs to proteins. In the early stages of drug development, predicting and analyzing the binding structure of proteins and small molecules through computer simulations can reduce the number of experiments, lower experimental costs, and reduce time investment. It can quickly provide information on the potential structures and properties of protein-small molecule binding, accelerating research progress in fields such as drug development and biochemistry. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0042] Figure 1 This is a flowchart of a method for predicting the binding structure of proteins and small molecules disclosed in this application;
[0043] Figure 2 This is a flowchart of a method for predicting and simulating small protein molecule complexes disclosed in this application;
[0044] Figure 3 This is a schematic diagram of the structure of a protein-small molecule binding structure prediction device disclosed in this application.
[0045] Figure 4 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0046] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0047] First, drug design typically begins with a deep understanding of the disease. Scientists study the biological mechanisms of disease, including how pathogens grow, cell signaling pathways, and the structure and function of proteins. Through research in these areas, scientists can identify potential drug targets—biomolecules that play a crucial role in the disease process. Next, drug design requires finding compounds that can interact with these targets and alter their function. These compounds are often called drug candidates. Molecular docking is a commonly used technique for predicting how drugs interact with targets. It uses calculations and simulations to predict binding patterns and affinities between molecules, thus helping scientists screen for promising drug candidates.
[0048] Molecular docking typically involves the following steps: First, computational methods are used to generate the structure of the candidate drug. Then, through simulation and calculation, the conformational and energy changes of the candidate drug upon binding to the target are predicted. Finally, based on these predictions, the affinity and pharmacodynamic properties of the candidate drug are evaluated, and the design is further optimized.
[0049] The development of molecular docking technology has made drug design more efficient and precise. It helps scientists quickly screen for promising drugs from a large number of compounds and provides theoretical guidance for drug development. However, despite its important role in drug design, molecular docking technology still faces challenges such as accuracy and computational costs, requiring continuous technological and methodological improvements.
[0050] Currently, simulations based on hand-designed rules and assumptions neglect complex biophysical processes, resulting in insufficient accuracy. Relying on physics-based force field assessments of interaction energies is time-consuming when processing large-scale structural data. They are limited to predefined ligand binding pockets, ignoring global interactions. Furthermore, they lack flexibility, making it difficult to adapt to complex biological systems and diverse ligand libraries, thus affecting the identification and assessment of specific binding modes. In summary, the limitations of traditional methods may affect the accurate simulation and prediction of protein-ligand interactions.
[0051] Furthermore, relying solely on algorithmic scoring is insufficient to determine whether a small molecule can stably bind in the active pocket. Molecular docking algorithms assess binding affinity, but these are only theoretical indicators and cannot fully reflect the situation in biological systems. Biological systems are complex and influenced by multiple factors such as solvent effects, conformational changes, and dynamics. Therefore, even if a small molecule scores highly in simulations, it may not necessarily bind stably in the pocket in a real biological environment.
[0052] Finally, due to reliance on manual operation and data transfer, each step in the process is often affected by human factors and individual efficiency, leading to increased uncertainty in the overall process. Secondly, because the time required for each step cannot be accurately predicted, waiting times may become unpredictable and uncontrollable, resulting in wasted resources and reduced efficiency. Furthermore, non-automated processes are prone to errors and delays, especially in a large amount of repetitive work, where the risk of human negligence and errors is higher, thus affecting the smooth operation of the entire process.
[0053] Therefore, the present invention provides a scheme for predicting the binding structure of proteins and small molecules, which can realize automated prediction and evaluation of the stability of the binding structure of proteins and small molecule compounds.
[0054] Reference Figure 1 As shown in the figure, this invention discloses a method for predicting the binding structure of proteins and small molecules, including:
[0055] Step S11: Obtain the first and second modeling graphic structures of the protein block and the small molecule compound to be tested, respectively.
[0056] In this embodiment, the protein is divided into protein blocks according to a preset protein block radius; the protein blocks and the small molecule compound to be tested are subjected to their respective graphic modeling processes to obtain a first modeling graphic structure and a second modeling graphic structure. It is understood that, referring to... Figure 2 As shown, the entire protein is divided into blocks, each with a radius of 20 Å (Ångström). This division method is commonly used in protein structure research and analysis to study and model local structures or specific regions of the protein. Both the protein blocks and the drug compound to be tested (small molecule compound) are modeled as graphics to obtain a first modeling graphic structure and a second modeling graphic structure, which better represents their structure and interactions.
[0057] Step S12: Input the first modeling graphical structure and the second modeling graphical structure into the TANKBind model so that the TANKBind model outputs an interaction matrix that integrates the first modeling graphical structure and the second modeling graphical structure, and updates the interaction embedding according to the interaction matrix. Then, the interaction embedding is predicted to output the predicted protein small molecule complex; wherein, the interaction matrix describes the interaction between the protein block and the small molecule compound to be tested.
[0058] In this embodiment, the first and second modeling graphical structures are input into the TANKBind model. The TANKBind model uses a trigonometric function module to perform coordinate transformation on the first and second modeling graphical structures to obtain the distance and angle information between the protein block and the small molecule compound to be tested. An interaction matrix is then constructed based on this distance and angle information. It is understood that TANKBind is a novel deep learning model used to predict the binding structure and affinity of protein-small molecule ligands. This model has the following characteristics: Neural network architecture: It includes a Drug module, a Protein module, and an Interact module. The Drug and Protein modules are used to process information about the small molecule compound and the protein, respectively, while the Interact module is used to predict the interaction between them. Trigonometry-Aware feature extraction: To better capture the geometric information between the drug and the protein, the coordinates of the small molecule compound and the protein are converted into polar coordinates relative to a certain reference point, thereby better expressing the angle and distance relationships between them. This method is based on the constraints of a three-dimensional structure. Multi-scale information fusion: The TANKBind model incorporates feature extraction modules at multiple scales and fuses them using a specific fusion method to consider feature information at different scales. Loss function design: A maximum edge contrast affinity loss function is designed to drive the model to fully utilize affinity information and global 3D structural information, thereby better reflecting the interaction between the drug and protein. Therefore, referring to... Figure 2 As shown, the first and second modeling graph structures are input into the trained TANKBind model to perform multiple evolutions of the block-small molecule interaction matrix using the TANKBind model. In this process, trigonometric function modules and distance maps between protein blocks and small molecules are used to obtain additional inputs. These inputs include information such as distance and angle.
[0059] Specifically, by converting the coordinates of the drug and protein to polar coordinates relative to a reference point, the angular and distance relationships between them can be better expressed. This helps to more accurately describe the interactions between protein blocks and small molecules. For example, when predicting the binding structure of a protein-small molecule ligand, this method can capture the geometric information between the small molecule and the protein, thereby improving the accuracy of the prediction. Specifically, a trigonometric function module is used to obtain additional input based on the distance map between the protein block and the small molecule. These inputs include information such as distance and angle, which helps to more accurately describe the interaction between the block and the small molecule. Then, the interaction embedding is updated based on the evolved interaction matrix. This embedding includes information such as the binding affinity of the small molecule to the protein block and the block-ligand distance map. Then, this information is used for prediction, including the binding affinity of the small molecule to the protein block and the protein block-small molecule distance map.
[0060] Integrating multi-scale information is also a crucial aspect of this model. This involves introducing feature extraction modules at multiple scales and fusing them together using a specific fusion method to account for feature information at different scales. Furthermore, to optimize model performance, a maximum edge contrastive affinity loss function is designed. This function drives the model to fully utilize affinity information and global 3D structural information, better reflecting the interaction between the drug and protein. To ensure the binding advantage of protein blocks to small molecules, a contrastive loss function is employed. This loss function helps the model learn to distinguish the binding between real small molecules and camouflage, thereby improving the accuracy and reliability of predictions. Finally, based on the model's internal scoring ranking, the structure of the highest-scoring complex is extracted as the predicted protein-small molecule complex, and the protein structure and small molecule structure of the protein-small molecule complex are extracted separately.
[0061] Experiments on multiple datasets demonstrate that the TANKBind model exhibits high accuracy and robustness in predicting protein-small molecule ligand structures and affinities, outperforming existing methods. This showcases its practical value and superiority in fields such as drug discovery. Overall, the TANKBind model achieves more accurate predictions of block-compound interaction matrices and protein-small molecule interactions through techniques such as trigonometric function-aware feature extraction, multi-scale information fusion, and specific loss function design. These techniques contribute to a deeper understanding of drug-protein binding mechanisms, providing valuable information for drug design and development.
[0062] Step S13: Generate the protein structure file and the small molecule structure file of the protein complex, and the first topology file and the second topology file containing coordinate information and force field parameter information respectively.
[0063] In this embodiment, force field information and simulation parameters are set for the protein structure file and the small molecule structure file respectively to create corresponding parameter files; a format conversion instruction is called and the protein structure file is converted into a first coordinate file and a first topology file based on the parameter files; small molecule charge calculations are performed on the small molecule structure file based on the parameter files, and small molecule field parameter information in various formats is generated to obtain the converted second coordinate file and second topology file. It is understood that, referring to... Figure 2 As shown, the force field information for proteins and small molecules, as well as the parameters required for simulation, are pre-defined to create a parameter file. The pre-defined force field (a mathematical model describing intermolecular interactions) and various simulation parameters (such as temperature, pressure, and time step) are organized into a specific format and structure to form the parameter file. This parameter file contains various configuration information required by the Gromacs software (molecular dynamics simulation software) during subsequent simulations. The Gromacs pdb2gmx command is called, and the file generation parameters are automatically established based on the parameter file. Using the processed protein structure file as input, a first coordinate .gro file and a first topology file containing the force field parameters are generated corresponding to the protein structure file. Then, using the small molecule structure file as input, the program identifies and adds missing hydrogen atoms based on software such as OpenBabel. ORCA and Multiwfn are called to calculate the RESP charge of the small molecule. Sobtop is called to generate force field parameters for small molecules in various mainstream formats (such as Amber, Gromacs, CHARMM, etc.) to obtain a second coordinate .gro file and a second topology file containing the force field parameters corresponding to the small molecule structure file.
[0064] Step S14: Perform molecular dynamics simulation using a preset molecular dynamics simulation tool and a simulation file generated based on the first topology file and the second topology file to obtain a molecular dynamics simulation trajectory file.
[0065] In this embodiment, before performing molecular dynamics simulation using a preset molecular dynamics simulation tool and a simulation file generated based on the first topology file and the second topology file to obtain a molecular dynamics simulation trajectory file, the method further includes: merging the first coordinate file, the first topology file, the second coordinate file, the second topology file, and the solvent file to obtain a merged file; and setting the boundary information and size information of the simulation box used to accommodate the protein small molecule complex for molecular dynamics simulation; and generating a corresponding simulation file based on the merged file, the boundary information, and the size information. It is understood that, referring to... Figure 2As shown, before obtaining the molecular dynamics simulation trajectory file, it is necessary to first construct the simulation system to obtain the corresponding simulation file. The specified components are then processed separately, and the Gro and topology files of macromolecules, water, or small molecules are automatically merged. The specified components are the various components contained in the system used for molecular dynamics simulation, such as proteins, small molecules (drug candidate ligand molecules, etc.), solvents (e.g., water), ions, etc. In molecular dynamics simulations, these different components need to be processed and the simulation system constructed. Specifically, this involves merging the coordinate Gro and topology files of these components, setting the size of the simulation box and the distance between the structure and the box, performing solvation processing, and adding anions and cations to neutralize the system's charge, among other operations.
[0066] For example, it is necessary to merge the GRO and topology files of proteins, small molecules, and solvents (such as water molecules) according to certain rules and requirements to construct a complete simulation system. This allows for a more accurate simulation of the interactions and behaviors of these components under specific conditions. This enables the study of the binding between proteins and small molecules, the influence of solvents on the system, and the role of ions within the system. Different simulation systems and research objectives may involve different components and their combinations.
[0067] Specifically, the `editconf` command is automatically invoked to set the size of the simulation box and the distance between the structure and the box. In other words, the `editconf` command determines the boundaries and dimensions of the simulation box based on the set rules and parameters. This process may involve considering the geometry and spatial distribution of the molecular system to ensure that the box can reasonably accommodate the molecules under study and provide a suitable spatial range for subsequent simulations.
[0068] For example, some simulation software defines the origin of the simulation box and several edge vectors starting from the origin. These vectors determine the shape and size of the box. By adjusting these parameters, the volume and shape of the simulation box can be controlled to adapt it to different molecular systems and simulation requirements.
[0069] Furthermore, the size and shape of the simulation box can be influenced by factors such as the number of molecules, their size and shape, and the required simulation accuracy. In practice, some trial and error may be necessary to determine the most suitable simulation box size and structure, thereby ensuring the accuracy and efficiency of the simulation.
[0070] The simulation box constructed in this way can provide a limited spatial range for molecular dynamics simulations, allowing the study of molecular motion, interactions, and other behaviors within this space, while reducing the influence of boundary effects by setting periodic boundary conditions.
[0071] The system automatically invokes Gromacs's `solvate` command for solvation, automatically selecting and adjusting parameters for different water models, and automatically merging water molecules from the input file. It automatically determines the system's charge and adds an ion model to the neutral condition for the selected force field scheme by invoking Gromacs's `grompp` and `genion` commands.
[0072] Finally, based on the merged file, distance, boundary, and neutral conditions described above, the corresponding simulation file is generated.
[0073] Then, the simulation file is input into a preset molecular dynamics simulation tool, which uses the particle mesh Ewald method to process long-range electrostatic interactions in the simulation file and truncates non-bonded interactions at a preset cutoff distance. By setting a transition distance, the non-bonded interactions are smoothly reduced to zero within the transition distance, and the lengths of all bonds involving hydrogen atoms are constrained to obtain a molecular dynamics simulation trajectory file. It is understood that, referring to... Figure 2As shown, the simulation file described above is used as input for Gromacs molecular dynamics simulations. During the simulation, the Ewald particle mesh is used for long-distance electrostatic interactions. A cutoff distance is set to truncate non-bonded interactions. By default, non-bonded interactions are significantly truncated at the cutoff distance. By setting the switch distance using the switch distance setting, the interactions smoothly reach zero within the switch distance, and the lengths of all bonds involving hydrogen atoms are constrained. Gromacs is used for energy minimization to optimize the simulation system. The Velocity-rescale method is automatically invoked in Gromacs to simulate isothermal effects, and the Parrinello-Rahman method is used to simulate isobaric effects. The Parrinello-Rahman method is a technique used to achieve isobaric conditions in molecular dynamics simulations; that is, the simulation file is used as the material for Gromacs simulation. During the simulation, the Ewald particle mesh method is used to handle long-distance electrostatic interactions. This allows for better consideration of the mutual influence between long-distance charges, thus more realistically simulating the behavior of molecular systems. For example, it helps to accurately predict protein folding, the binding affinity of drug molecules to receptors, etc. A cutoff distance needs to be set; beyond this distance, non-bonded interactions will be significantly truncated. This improves computational efficiency while maintaining a certain level of accuracy. Calculating all interactions with complete precision is often computationally very time-consuming; by using reasonable truncation, a good balance can be found between accuracy and efficiency. Additionally, a switch distance is set using `Switch Distance`, allowing interactions to gradually decrease to zero within this distance. Furthermore, the lengths of all hydrogen-related bonds must be fixed. Gromacs is used to minimize the energy of the simulation system, performing energy optimization. Gromacs automatically uses the Velocity-rescale method to simulate the case of constant temperature and the Parrinello-Rahman method to simulate the case of constant pressure. Finally, the corresponding molecular dynamics simulation trajectory files obtained after performing the above simulations are obtained.
[0074] Step S15: Analyze the molecular dynamics simulation trajectory file, the first topology file, and the second topology file to obtain the simulation analysis results of the protein-small molecule binding structure.
[0075] In this embodiment, the molecular dynamics simulation trajectory file, the first topology file, and the second topology file are read. The difference between the particle coordinates at different time points in the molecular dynamics simulation trajectory file and the particle reference coordinates is calculated using root mean square deviation (RMSD) calculation instructions to obtain the positional fluctuations of the particles in the simulated trajectory. Binding free energy and related energy terms are calculated for the molecular dynamics simulation trajectory file, the first topology file, and the second topology file using combined free energy calculation instructions to obtain calculation results characterizing the binding strength between proteins and small molecules and the contribution of each residue to the binding. It is understood that the MDAnalysis toolkit automatically reads the molecular dynamics simulation trajectory file and all topology files, calculates the RMSD (Root Mean Square Deviation) using instructions, and calculates the solvent-accessible surface area using instructions. RMSD measures the degree of change in molecular structure over time. It is obtained by calculating the square root of the average of the squares of the deviations between the structure of each frame in the trajectory and the structure of the reference frame. In MDAnalysis, correlation functions can be used to calculate RMSD. Specifically, it compares the differences between the particle coordinates at different time points in the molecular dynamics simulation trajectory file and the particle coordinates of the reference structure. For example, use the following code to calculate RMSD: import MDAnalysis as mda
[0076] from MDAnalysis.analysis import rms
[0077] # Load topology and trajectory files
[0078] u = mda.Universe('topology_file.pdb', 'trajectory_file.xtc')
[0079] # Select protein atoms
[0080] protein = u.select_atoms('protein')
[0081] # Select a reference frame (usually the first frame)
[0082] ref = mda.Universe('topology_file.pdb')
[0083] ref_protein = ref.select_atoms('protein')
[0084] # Initialize RMSD analysis
[0085] R = rms.RMSD(protein, ref_protein, select='backbone')
[0086] R.run().
[0087] In the code above, the topology file and trajectory file are loaded first, then the protein atoms to be analyzed are selected, and a reference frame is specified. RMSD analysis is performed using the rms.RMSD function, where select='backbone' indicates that the main chain atoms of the protein are selected for calculation.
[0088] The RMSF (Root Mean Square Fluctuation) is obtained by first calculating the variance of all coordinates of the particle in the trajectory, then summing them, and finally taking the square root. RMSF measures the positional fluctuation of a particle in a trajectory. It is calculated by first calculating the variance of all coordinates of the particle in the trajectory, then summing them, and finally taking the square root. MDAnalysis has a corresponding class for calculating RMSF.
[0089] Furthermore, MMPB / GBSA (Molecular Mechanics Poisson-Boltzmann / Generalized Born Surface Area) calculations are performed using topology and trajectory files. Specifically, MMPB / GBSA is a method for calculating molecular binding free energy. It calculates the binding free energy by separating it into molecular mechanical energy (including van der Waals forces and electrostatic interactions) and solvation free energy. When performing MMPB / GBSA calculations, it is necessary to prepare a complex topology file, a trajectory file of the kinetic simulation results (such as an mdcrd file), and an input parameter file (such as mmpbsa.in). Based on the settings and parameters in the input file, the trajectory and topology files are analyzed to obtain the calculated binding free energy and related energy terms. These results help to understand the binding strength between ligands and acceptors, as well as the contribution of each residue to the binding.
[0090] The specific calculations involve several thermodynamic formulas and energy terms, including gas-phase energy (composed of electrostatic and van der Waals terms) and solvation free energy (including polar and nonpolar solvation free energies). The polar solvation free energy can be obtained using the generalized Born model or by solving the Poisson-Boltzmann equation, while the nonpolar solvation free energy is typically calculated using the solvent's accessible surface area. Interaction entropy may also be included in the calculations to more accurately describe the energy changes during the binding process. Therefore, by using molecular dynamics simulations to calculate parameters such as RMSD, binding free energy, and the number of hydrogen bonds, the binding of small molecules in the active pocket can be more accurately assessed. RMSD reflects the similarity between the simulated structure and the experimentally resolved structure, while binding free energy and the number of hydrogen bonds provide an assessment of binding stability and affinity. Considering these parameters comprehensively allows for a more precise determination of whether small molecules can effectively bind in the active pocket.
[0091] After the calculations are complete, tools such as matplotlib can be used to visualize the results for a more intuitive analysis and understanding of the data. This helps researchers gain a deeper understanding of the dynamic behavior of molecules, binding patterns, and the importance of each residue in the binding process. This method has wide applications in biochemistry, drug design, and other fields, used to evaluate ligand-receptor interactions and predict the binding effects of potential drug molecules.
[0092] The method of this invention is automated, enabling the immediate initiation of the next step upon completion of the previous one, without waiting for human intervention or manual data transfer. This not only improves overall efficiency but also minimizes resource waste caused by waiting time. This continuous and seamless workflow transition not only enhances work efficiency but also makes the workflow more predictable and stable. This means that large quantities of protein-small molecule complexes can be processed effectively without human intervention. This highly efficient automated process saves time and labor costs, improving work efficiency. Deep learning can improve prediction accuracy by learning from large amounts of data, automatically extracting features through large-scale data learning, more efficiently modeling protein-ligand interactions, and accelerating the docking process. Simultaneously, processing larger-scale data results in higher prediction accuracy, enabling more precise prediction of binding ability and improving docking accuracy and reliability. Therefore, deep learning molecular docking has become an important tool for drug design and biomedical research, supporting rapid and accurate drug candidate discovery. Molecular dynamics simulations can take into account the details of physical interactions. Therefore, combining the two can improve prediction accuracy, making it more consistent with reality. This invention is applicable to different types of proteins and small molecules because both deep learning and molecular dynamics simulations have a certain degree of versatility. Therefore, this method can be applied to various medical fields, such as drug design and efficacy evaluation.
[0093] As can be seen, this application discloses a method for predicting the binding structure of proteins and small molecules, comprising: obtaining a first modeled graphical structure and a second modeled graphical structure after modeling a protein block and a small molecule compound to be tested, respectively; inputting the first modeled graphical structure and the second modeled graphical structure into a TANKBind model, so that the TANKBind model outputs an interaction matrix that integrates the first modeled graphical structure and the second modeled graphical structure, and updating the interaction embedding according to the interaction matrix, and then predicting the interaction embedding to output the predicted protein-small molecule complex; wherein, the interaction matrix describes the interaction between the protein block and the small molecule compound to be tested; generating a first topology file and a second topology file containing coordinate information and force field parameter information for the protein structure file and the small molecule structure file of the protein-small molecule complex, respectively; performing molecular dynamics simulation using a preset molecular dynamics simulation tool and a simulation file generated according to the first topology file and the second topology file to obtain a molecular dynamics simulation trajectory file; and analyzing the molecular dynamics simulation trajectory file, the first topology file, and the second topology file to obtain the simulation analysis results of the protein-small molecule binding structure. Therefore, by integrating the modeling graphs of proteins and small molecules using the TANKBind model, the interactions between them can be described more comprehensively and accurately, thereby improving the accuracy of predicting the structure of protein-small molecule complexes. Generating topology files and performing molecular dynamics simulations allows for the simulation of the dynamic behavior of complexes at the atomic level, enabling in-depth exploration of structural changes and interaction mechanisms during protein-small molecule binding. More accurate binding structure predictions and a deeper understanding of binding mechanisms contribute to the design of more effective drug molecules. Small molecules can be optimized and modified based on simulation analysis results, improving the binding affinity and specificity of drugs to proteins. In the early stages of drug development, predicting and analyzing the binding structure of proteins and small molecules through computer simulations can reduce the number of experiments, lower experimental costs, and reduce time investment. It can quickly provide information on the potential structures and properties of protein-small molecule binding, accelerating research progress in fields such as drug development and biochemistry.
[0094] Reference Figure 3 As shown, the present invention also discloses a protein-small molecule binding structure prediction device, comprising:
[0095] Modeling module 11 is used to obtain the first and second modeling graphic structures of the protein block and the small molecule compound to be tested after modeling, respectively.
[0096] The prediction module 12 is used to input the first modeling graphical structure and the second modeling graphical structure into the TANKBind model, so that the TANKBind model outputs an interaction matrix that integrates the first modeling graphical structure and the second modeling graphical structure, updates the interaction embedding according to the interaction matrix, and then predicts the interaction embedding to output the predicted protein small molecule complex; wherein, the interaction matrix describes the interaction between the protein block and the small molecule compound to be tested.
[0097] The file generation module 13 is used to generate a first topology file and a second topology file containing coordinate information and force field parameter information, respectively, for the protein structure file and the small molecule structure file of the protein complex.
[0098] The dynamics simulation module 14 is used to perform molecular dynamics simulation using a preset molecular dynamics simulation tool and a simulation file generated based on the first topology file and the second topology file, so as to obtain a molecular dynamics simulation trajectory file.
[0099] The results analysis module 15 is used to analyze the molecular dynamics simulation trajectory file, the first topology file, and the second topology file to obtain the simulation analysis results of the protein-small molecule binding structure.
[0100] As can be seen, this application discloses obtaining a first modeled graphical structure and a second modeled graphical structure after modeling a protein block and a small molecule compound to be tested, respectively; inputting the first modeled graphical structure and the second modeled graphical structure into the TANKBind model, so that the TANKBind model outputs an interaction matrix that integrates the first modeled graphical structure and the second modeled graphical structure, and updating the interaction embedding according to the interaction matrix, and then predicting the interaction embedding to output the predicted protein-small molecule complex; wherein, the interaction matrix describes the interaction between the protein block and the small molecule compound to be tested; generating a first topology file and a second topology file containing coordinate information and force field parameter information for the protein structure file and the small molecule structure file of the protein-small molecule complex, respectively; performing molecular dynamics simulation using a preset molecular dynamics simulation tool and the simulation file generated according to the first topology file and the second topology file to obtain a molecular dynamics simulation trajectory file; analyzing the molecular dynamics simulation trajectory file, the first topology file, and the second topology file to obtain the simulation analysis results of the protein-small molecule binding structure. Therefore, by integrating the modeling graphs of proteins and small molecules using the TANKBind model, the interactions between them can be described more comprehensively and accurately, thereby improving the accuracy of predicting the structure of protein-small molecule complexes. Generating topology files and performing molecular dynamics simulations allows for the simulation of the dynamic behavior of complexes at the atomic level, enabling in-depth exploration of structural changes and interaction mechanisms during protein-small molecule binding. More accurate binding structure predictions and a deeper understanding of binding mechanisms contribute to the design of more effective drug molecules. Small molecules can be optimized and modified based on simulation analysis results, improving the binding affinity and specificity of drugs to proteins. In the early stages of drug development, predicting and analyzing the binding structure of proteins and small molecules through computer simulations can reduce the number of experiments, lower experimental costs, and reduce time investment. It can quickly provide information on the potential structures and properties of protein-small molecule binding, accelerating research progress in fields such as drug development and biochemistry.
[0101] Furthermore, embodiments of this application also disclose an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0102] Figure 4This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the protein-small molecule binding structure prediction method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0103] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0104] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0105] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0106] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. It can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the protein and small molecule binding structure prediction method disclosed in any of the foregoing embodiments executed by the electronic device 20, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.
[0107] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned method for predicting the binding structure of proteins and small molecules. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0108] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0109] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs (Compact Disc-Read Only Memory), or any other form of storage medium known in the art.
[0110] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0111] The solution provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for predicting the binding structure of proteins and small molecules, characterized in that, include: Obtain the first and second modeling structures of the protein block and the small molecule compound to be tested, respectively; The first modeling graphical structure and the second modeling graphical structure are input into the TANKBind model so that the TANKBind model outputs an interaction matrix that integrates the first modeling graphical structure and the second modeling graphical structure. The interaction embedding is updated according to the interaction matrix, and then the interaction embedding is predicted to output the predicted protein small molecule complex. The interaction matrix describes the interaction between the protein block and the small molecule compound to be tested. Generate the protein structure file and the small molecule structure file of the protein complex, and their respective first and second topology files containing coordinate information and force field parameter information; Using a preset molecular dynamics simulation tool, molecular dynamics simulations are performed on the simulation files generated based on the first topology file and the second topology file to obtain molecular dynamics simulation trajectory files; The molecular dynamics simulation trajectory file, the first topology file, and the second topology file are read; the difference between the particle coordinates and the particle reference coordinates at different time points in the molecular dynamics simulation trajectory file is calculated using the root mean square fluctuation calculation instruction to obtain the position fluctuation of the particle in the simulation trajectory; the binding free energy and related energy terms are calculated on the molecular dynamics simulation trajectory file, the first topology file, and the second topology file by combining the free energy calculation instruction to obtain the calculation results characterizing the binding strength between the protein and the small molecule and the contribution information of each residue to the binding. The first and second topology files, each containing coordinate information and force field parameter information, corresponding to the protein structure file and the small molecule structure file that generate the protein small molecule complex, include: The force field information and simulation parameters used by the protein structure file and the small molecule structure file are set respectively to create corresponding parameter files; the format conversion command is called and the protein structure file is converted into a first coordinate file and a first topology file based on the parameter file; the small molecule charge is calculated on the small molecule structure file based on the parameter file, and small molecule force field parameter information in various formats is generated to obtain the converted second coordinate file and second topology file; The step of using a preset molecular dynamics simulation tool to perform molecular dynamics simulations on simulation files generated based on the first and second topology files to obtain molecular dynamics simulation trajectory files includes: The simulation file is input into a preset molecular dynamics simulation tool, which uses the particle mesh Ewald method to process long-distance electrostatic interactions in the simulation file and truncates non-bonded interactions in the simulation file at a preset cutoff distance. By setting the transformation distance, the non-bonded interactions are controlled to smoothly return to zero within the transformation distance, and the lengths of all bonds involving hydrogen atoms are constrained to obtain a molecular dynamics simulation trajectory file.
2. The method for predicting the binding structure of proteins and small molecules according to claim 1, characterized in that, The acquisition of the first and second modeling graphic structures after modeling the protein block and the small molecule compound to be tested includes: The protein is divided into protein blocks according to the preset protein block radius; The protein block and the small molecule compound to be tested are subjected to their respective graphic modeling processes to obtain a first modeling graphic structure and a second modeling graphic structure.
3. The method for predicting the binding structure of proteins and small molecules according to claim 1, characterized in that, The step of inputting the first modeling graphical structure and the second modeling graphical structure into the TANKBind model, so as to output an interaction matrix that integrates the first modeling graphical structure and the second modeling graphical structure through the TANKBind model, includes: The first and second modeling graphic structures are input into the TANKBind model so that the first and second modeling graphic structures can be transformed by the trigonometric function module of the TANKBind model to obtain the distance and angle information between the protein block and the small molecule compound to be tested, and an interaction matrix is constructed based on the distance and angle information.
4. The method for predicting the binding structure of proteins and small molecules according to claim 1, characterized in that, Before performing molecular dynamics simulations on the simulation files generated based on the first and second topology files using a preset molecular dynamics simulation tool to obtain the molecular dynamics simulation trajectory files, the method further includes: The first coordinate file, the first topology file, the second coordinate file, the second topology file, and the solvent file are merged to obtain a merged file, and the boundary information and size information of the simulation box used to accommodate the protein small molecule complex for molecular dynamics simulation are set. A corresponding simulation file is generated based on the merged file, the boundary information, and the size information.
5. A device for predicting the binding structure of proteins and small molecules, characterized in that, include: The modeling module is used to obtain the first and second modeling graphic structures of the protein block and the small molecule compound to be tested, respectively, after modeling. The prediction module is used to input the first modeling graphical structure and the second modeling graphical structure into the TANKBind model, so that the TANKBind model outputs an interaction matrix that integrates the first modeling graphical structure and the second modeling graphical structure, updates the interaction embedding according to the interaction matrix, and then predicts the interaction embedding to output the predicted protein small molecule complex; wherein, the interaction matrix describes the interaction between the protein block and the small molecule compound to be tested; The file generation module is used to generate a protein structure file and a first topology file and a second topology file containing coordinate information and force field parameter information, respectively, for the protein structure file and the small molecule structure file of the protein complex. The dynamics simulation module is used to perform molecular dynamics simulations on simulation files generated based on the first topology file and the second topology file using preset molecular dynamics simulation tools, so as to obtain molecular dynamics simulation trajectory files; The results analysis module is used to read the molecular dynamics simulation trajectory file, the first topology file, and the second topology file; calculate the difference between the particle coordinates and the particle reference coordinates at different time points in the molecular dynamics simulation trajectory file using root mean square fluctuation calculation instructions to obtain the position fluctuation of the particle in the simulation trajectory; and perform binding free energy calculation and related energy term calculation on the molecular dynamics simulation trajectory file, the first topology file, and the second topology file using free energy calculation instructions to obtain calculation results characterizing the binding strength between the protein and the small molecule and the contribution information of each residue to the binding. The file generation module is specifically used to set the force field information and simulation parameters used by the protein structure file and the small molecule structure file respectively to create corresponding parameter files; call the format conversion instruction and convert the protein structure file into a first coordinate file and a first topology file based on the parameter file; perform small molecule charge calculation on the small molecule structure file based on the parameter file, and generate small molecule force field parameter information in multiple formats to obtain the converted second coordinate file and second topology file; The dynamics simulation module is specifically used to input the simulation file into a preset molecular dynamics simulation tool, so that the preset molecular dynamics simulation tool uses the particle grid Ewald method to process the long-distance electrostatic interactions in the simulation file, and truncates the non-bonded interactions in the simulation file at a preset cutoff distance. By setting the transformation distance, the non-bonded interactions are controlled to smoothly return to zero within the transformation distance, and the lengths of all bonds involving hydrogen atoms are constrained to obtain a molecular dynamics simulation trajectory file.
6. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the protein-small molecule binding structure prediction method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the protein-small molecule binding structure prediction method as described in any one of claims 1 to 4.