Method, device and terminal for determining three-dimensional structure of protein-ligand complex

By simultaneously processing protein and ligand data, generating and voxelizing the data, and then using a neural network for sorting, the problem of long time and low accuracy in determining the three-dimensional structure of protein-ligand complexes in existing technologies has been solved, achieving faster and more accurate structure determination.

CN116312763BActive Publication Date: 2025-11-28CHONGQING KANGZHOU ZHITONG PHARM TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310274035.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2025-11-28
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

The existing technology for determining the three-dimensional structure of protein-ligand complexes is time-consuming and has low accuracy, which cannot effectively guide drug development.

Method used

By simultaneously processing protein and ligand data, protein pocket data and ligand pose data are generated, matched to form a first three-dimensional structure, and voxelized into a grid. The grid is then input into a trained neural network for sorting, determining the optimal binding site and pose, and finally obtaining the three-dimensional structure of the protein-ligand complex.

Benefits of technology

It enables faster and more accurate determination of the three-dimensional structure of protein-ligand complexes, shortens processing time, and improves processing efficiency and accuracy. It is applicable to the processing of different databases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312763B_ABST
    Figure CN116312763B_ABST
Patent Text Reader

Abstract

The application discloses a kind of protein-ligand complex three-dimensional structure determination method, device and terminal, the method includes: simultaneously processing protein data and ligand data, generate protein pocket data and ligand pose data;Matching protein pocket data and ligand pose data, obtain first three-dimensional structure;Voxelization first three-dimensional structure, obtain several grids, including corresponding position protein-ligand binding information in grid;The grid is input into trained neural network, and the grid is sorted;And according to the first grid of sorting, obtain optimal binding site and optimal ligand pose, and combine to obtain protein-ligand complex three-dimensional structure.Through simultaneously extracting and processing protein data and ligand data, then in the form of voxelization is formed including protein-ligand binding information grid, again through neural network for grid sorting, the three-dimensional structure of protein-ligand complex can be more quickly and accurately determined by the application.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of three-dimensional structure design of drug molecules, and particularly relates to a method and device for determining the three-dimensional structure of a protein-ligand complex. BACKGROUND

[0002] In the field of structure design of drug molecules, a key step is to determine the three-dimensional structure of a protein-ligand complex, and thus one of the tasks in the development of new drugs is to determine the position on the surface of a protein where a ligand can bind. With the development of X-ray crystallography and nuclear magnetic resonance spectroscopy and other related technologies, the surface structure of a protein and a protein-ligand complex can be more clearly displayed to people, but even the highest efficiency of these two technologies cannot keep up with the pace of drug development, and cannot determine the optimal binding position on the protein, and thus cannot guide how to optimize the binding of the protein and the ligand.

[0003] In the prior art, in order to determine the optimal binding position between a protein and a ligand, a three-dimensional structure of a protein-ligand complex is generally simulated by a molecular docking method on a computer, and then a score is given to the three-dimensional structure according to a specific algorithm, and the three-dimensional structure with the highest score corresponds to the optimal binding position of the protein and the ligand. However, in actual application, the complete three-dimensional structure of the protein-ligand complex cannot be simulated, and the simulation process requires a large amount of time, thereby resulting in a long process time and low accuracy in determining the three-dimensional structure of the protein-ligand complex.

[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY

[0005] The present application provides a method and device for determining the three-dimensional structure of a protein-ligand complex, thereby solving the problems of long process time and low accuracy in determining the three-dimensional structure of a protein-ligand complex in the prior art.

[0006] The technical solutions adopted by the present application to solve the problems are as follows:

[0007] In a first aspect, the present application provides a method for determining the three-dimensional structure of a protein-ligand complex, which comprises:

[0008] processing protein data and ligand data simultaneously to generate protein pocket data and ligand pose data;

[0009] matching the protein pocket data and the ligand pose data to obtain a first three-dimensional structure;

[0010] voxelizing the first three-dimensional structure to obtain a plurality of grids, wherein the grids include protein-ligand binding information at corresponding positions;

[0011] inputting the grid into a trained neural network to rank the grid;

[0012] obtaining an optimal binding site and an optimal ligand pose according to the first ranked grid, and obtaining a protein-ligand complex three-dimensional structure by combining the optimal binding site and the optimal ligand pose.

[0013] In an embodiment, before inputting the grid into a trained neural network, the method comprises:

[0014] deleting the grid with zero protein-ligand binding information at the corresponding position.

[0015] In an embodiment, before processing the protein data and the ligand data simultaneously, the method comprises:

[0016] obtaining protein molecular structure data in a protein database, ranking the protein molecular structure data according to a selection standard characteristic value, and selecting the first ranked protein molecular structure data as the protein data;

[0017] obtaining molecular structure data of a ligand, adding polar hydrogen to the molecular structure data to obtain the ligand data.

[0018] In an embodiment, the selection standard characteristic value comprises:

[0019] resolution, which shows the resolution level of the electron density map;

[0020] completeness, which shows the percentage of the part that can be simulated in the total structure in the protein molecular structure data;

[0021] R value, which represents the difference between the diffraction pattern of the model obtained from the protein molecular structure data and the diffraction pattern obtained from the experiment;

[0022] R-free value, which represents the degree of overfitting.

[0023] In an embodiment, the processing of the protein data and the ligand data simultaneously to generate protein pocket data and ligand pose data comprises:

[0024] processing the protein data and the ligand data by using a blind docking method to obtain the ligand pose data, which reflects the state of the ligand in the protein-ligand binding process;

[0025] processing the protein data by using a pocket detection method to obtain the protein pocket data, which shows the binding site information on the surface of the protein;

[0026] The above two steps are processed simultaneously and in parallel.

[0027] In an embodiment, the voxelizing the first three-dimensional structure comprises:

[0028] dividing the first three-dimensional structure into cubes with an edge length of ;

[0029] assigning protein-ligand binding information on the first three-dimensional structure to the cubes at corresponding positions to form the grid.

[0030] In an embodiment, the training process of the neural network comprises:

[0031] obtaining experimental characterization data of protein-ligand complexes;

[0032] generating training three-dimensional structures using the experimental characterization data;

[0033] voxelizing the training three-dimensional structures to obtain a plurality of training grids, wherein the training grids include protein-ligand binding information at corresponding positions;

[0034] forming a training grid set by deleting training grids with zero protein-ligand binding information at corresponding positions, and iteratively training the neural network according to the training grid set, wherein the neural network updates parameters according to the free energy of the protein-ligand binding information in the training grid set.

[0035] In an embodiment, the sorting manner of the neural network for the grid comprises sorting the grid according to a sorting feature value, wherein the sorting feature value comprises:

[0036] a quantity feature value, which reflects the number of protein-ligand binding sites in a single grid;

[0037] a distance feature value, which reflects the distance between the center of a single grid and a protein-ligand binding site;

[0038] a pose feature value, which reflects the binding energy of a protein-ligand binding site in a single grid;

[0039] a pocket feature value, which reflects the biochemical properties of a protein pocket corresponding to a protein-ligand binding site in a single grid.

[0040] In a second aspect, an embodiment of the present application further provides a device for determining a three-dimensional structure of a protein-ligand complex, wherein the device comprises:

[0041] a data processing module, configured to simultaneously process protein data and ligand data to generate protein pocket data and ligand pose data;

[0042] a structure generating module configured to match the protein pocket data and the ligand pose data to obtain a first three-dimensional structure;

[0043] a voxelization module configured to voxelize the first three-dimensional structure to obtain a plurality of grids, wherein the grids include protein-ligand binding information at corresponding positions;

[0044] a calculation module configured to input the grids into a trained neural network to rank the grids;

[0045] a binding module configured to obtain an optimal binding site according to the first ranked grid, and bind a ligand at the optimal binding site to obtain a protein-ligand complex three-dimensional structure.

[0046] In a third aspect, an embodiment of the present application further provides a terminal, wherein the terminal comprises a computer readable storage medium and one or more processors; the computer readable storage medium stores one or more programs; the programs contain instructions for executing the method for determining a protein-ligand complex three-dimensional structure as described above; and the processors are configured to execute the programs.

[0047] The present application discloses a method, device and terminal for determining a protein-ligand complex three-dimensional structure. The method comprises simultaneously processing protein data and ligand data to generate protein pocket data and ligand pose data; matching the protein pocket data and the ligand pose data to obtain a first three-dimensional structure; voxelizing the first three-dimensional structure to obtain a plurality of grids, wherein the grids include protein-ligand binding information at corresponding positions; inputting the grids into a trained neural network to rank the grids; and obtaining an optimal binding site and an optimal ligand pose according to the first ranked grid, and combining the optimal binding site and the optimal ligand pose to obtain a protein-ligand complex three-dimensional structure. By simultaneously extracting and processing protein data and ligand data, and then forming grids containing protein-ligand binding information in a voxelized form, and ranking the grids through a neural network, the present application can more quickly and accurately determine a three-dimensional structure of a protein-ligand complex, has a wider application range and stronger robustness. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0049] Figure 1 is a flow chart of steps of the method for determining the three-dimensional structure of protein-ligand according to the embodiments of the present application.

[0050] Figure 2 is a flow chart of the method for determining the three-dimensional structure of protein-ligand according to the embodiments of the present application.

[0051] Figure 3 is a flow chart of the voxelization step of the method for determining the three-dimensional structure of protein-ligand according to the embodiments of the present application.

[0052] Figure 4 is a characterization diagram of the method for determining the three-dimensional structure of protein-ligand according to the embodiments of the present application and the two methods of Fpocket and P2rank in terms of pocket recognition accuracy.

[0053] Figure 5 is a characterization diagram of the method for determining the three-dimensional structure of protein-ligand according to the embodiments of the present application and the two methods of Fpocket and P2rank in terms of pose prediction accuracy.

[0054] Figure 6 is a schematic diagram of internal modules of the device for determining the three-dimensional structure of protein-ligand according to the embodiments of the present application.

[0055] Figure 7 is a principle block diagram of the terminal according to the embodiments of the present application. DETAILED DESCRIPTION

[0056] The present application discloses a method and device for determining the three-dimensional structure of protein-ligand complex, and a terminal. In order to make the purpose, technical scheme and effect of the present application more clear and explicit, the present application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0057] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an" and "the" used herein also include the plural forms. It should be further understood that the use of the term "comprising" in the specification of the present application means that the features, integers, steps, operations, elements and / or components described exist, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be intermediate elements. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any single unit and all combinations of the associated listed items.

[0058] As will be appreciated by one skilled in the art, the disclosure employs, unless otherwise indicated, a number of technical terms and scientific terms that have specialized meanings. Such terms are used consistent with their accepted meanings in the art.

[0059] By applying molecular docking method on computer, the binding form of receptor (protein) and ligand is simulated, and genetic algorithm, Monte Carlo simulation algorithm, simulated annealing algorithm and distance geometry method are used to sort all docking forms, so as to find the optimized protein-ligand complex three-dimensional structure. However, this way cannot simulate the complete three-dimensional structure of protein-ligand complex, and the simulation process takes too long time, resulting in long time and low accuracy in determining the three-dimensional structure of protein-ligand complex.

[0060] In the prior art, the protein surface is generally searched by blind docking method to find the binding site and binding mode corresponding to the specific ligand, so as to shorten the time and improve the accuracy. Blind docking selects the binding site on the protein surface randomly and simulates the binding, so as to more completely simulate the binding site of protein and ligand and determine the binding relationship of protein and ligand. Blind docking includes two schemes, one is traditional blind docking method, that is, the entire surface of target protein is searched by molecular binding software, and the position is randomly selected as the binding site for simulation; the other is sequential blind docking method, that is, the pocket position on the protein surface capable of binding with ligand is first determined by pocket detection tool, and the searched pocket is selected as the alternative binding site to simulate the binding form of protein-ligand.

[0061] However, both of the two schemes in blind docking can only take PDB id as input data and output original data, so blind docking is very dependent on protein database, and the output data need to be processed by other molecular binding method. When the database is too large, the processing time required by blind docking will sharply rise, and the work efficiency will be greatly affected. At this time, the user needs to manually select and optimize the protein structure saved in the database, and manually select the protein structure for subsequent cross-docking or reverse docking, which undoubtedly consumes a lot of time and is prone to errors in the processing process.

[0062] In view of the above defects of the prior art, the present application provides a consensus blind docking (Consensus Blind Dock, CoBDock) framework as a method for determining the three-dimensional structure of protein-ligand complex.

[0063] Unlike the direct identification of possible binding sites in the prior art, the CoBDock provided by the present application simultaneously processes protein data and ligand data through various docking software and pocket detection algorithms to obtain protein pocket data and ligand pose data, and matches the protein pocket data and the ligand pose data to obtain a first three-dimensional structure, at this time the first three-dimensional structure includes the state of the ligand being combined with the protein surface in different poses. Then the first three-dimensional structure is voxelized into a plurality of grids, wherein each grid includes the binding information of the protein and the ligand at the corresponding position of the first three-dimensional structure. All the grids after voxelization of the first three-dimensional structure are input into the trained neural network, so as to sort all the grids, and the corresponding position included in the first sorted grid is the optimal binding site. On the optimal binding site, the ligand can be combined to determine and obtain the three-dimensional structure of the protein-ligand complex.

[0064] Specifically, CoBDock first needs to obtain protein data and ligand data. For a specific protein database, because the number of protein entries is large, it often takes a lot of time to select the correct PDB id to obtain the corresponding protein. The present application first determines a protein database (such as the UniProtKB database), and then sorts the protein structure files in the protein database according to a selection standard characteristic value, and selects the protein structure file with the highest sorting to obtain the protein data. Specifically, the selection standard characteristic value includes:

[0065] 1. Resolution, the resolution shows the resolution level of the electron density map, and the smaller the value of the resolution, the higher the resolution. Specifically, the resolution should be less than 3.5.

[0066] 2. Completeness, the completeness shows the percentage of the part that can be simulated in the total structure of the protein structure.

[0067] 3. R value, the R value represents the difference between the diffraction pattern of the protein molecular structure data obtained model and the diffraction pattern obtained by experiment, that is, the protein structure corresponding to the simulation data identical with the experimental data, and the R value is 0. Specifically, the R value is about 0.20.

[0068] 4. R-free value, the R-free value represents the degree of overfitting, that is, the lower the degree of overfitting, the closer the R-free value to the R value. Specifically, 90% of the experimental data is used to fit and construct the protein structure, and then the obtained protein structure model is used to simulate the remaining 10% of the experimental data, and compared with the actual experimental data to obtain the R-free value. The R-free value is generally higher than the R value, and is about 0.26.

[0069] For obtaining ligand data, first, the molecular structure data of the ligand is obtained, which is in the format of SMILES, pdb, mol, mol2 or sdf, etc. which can be used for virtual screening, reverse docking and cross docking. Then, polar hydrogen is added to the molecular structure data to obtain the ligand data.

[0070] After obtaining the protein data and the ligand data, as shown in Figure 2 The protein-ligand three-dimensional structure determination method is shown in the flowchart. The protein pocket data of the protein surface is generated by the pocket detection method of the pocket detection tool, and the ligand pose data corresponding to the protein surface pocket is generated by the blind docking method of the docking program. Then, the protein pocket data and the ligand pose data are matched to obtain the first three-dimensional structure including the ligand combined with the protein surface in different ways. That is, the protein pocket data and the ligand pose data are obtained in parallel rather than in sequence, and the first three-dimensional structure is generated in real time by matching, thereby simplifying the overall process and shortening the time.

[0071] Specifically, the protein data and the ligand data are processed by the blind docking method, and the state of the ligand and the protein is scored by various molecular docking software included in the blind docking method to obtain ligand pose data, which reflects the state of the ligand in the protein-ligand binding process. Alternatively, the blind docking method includes AutoDock Vina software, PLANTS software, Z-Dock software and Galaxydock3 software, so that different databases can be handled to ensure the same processing effect. Among them, the false discovery rate of AutoDock Vina software is 81%, which is the most concerned docking software; on the Astex Diverse Set (ADS) database, the accuracy of PLANTS software for flexible protein side chains is as high as 87%; Z-Dock software has faster speed and higher efficiency, and the best performance can reach 85.71%; and Galaxydock3 software provides ligand freedom, so it is better than AutoDock Vina in this respect. By integrating various molecular docking software, the blind docking method can be run independently, has high accuracy, and does not consume excessive computing resources, which can save processing time.

[0072] Specifically, the protein data is processed by a pocket detection method, and a pocket detection tool of the pocket detection method is used to detect a region of a protein surface capable of binding a ligand to form the protein pocket data, which shows the binding site information of the protein surface. Optionally, the pocket detection method includes a P2rank algorithm and an Fpocket algorithm, so that good processing results can be achieved for different databases. The P2rank algorithm is based on local chemical neighborhood ligandization prediction to detect the binding site of the protein surface. The Fpocket algorithm is the most popular pocket recognition method, which realizes fast and accurate recognition of the protein surface pocket based on Voronoi tessellation and Alpha sphere. By combining multiple algorithms, the pocket detection method can detect the protein surface pocket (binding site) while realizing blind docking, thereby saving the processing time of the method for determining the three-dimensional structure of the protein-ligand complex.

[0073] After the first three-dimensional structure is generated by real-time matching, the first three-dimensional structure is voxelized to obtain a plurality of grids, wherein each grid contains protein-ligand binding information at the corresponding position of the first three-dimensional structure. Figure 3 As shown in the process diagram of voxelizing the first three-dimensional structure corresponding to protein 1A3E by using Pymol software. The grid of the first three-dimensional structure is a cube with an edge length of angstrom level. Optionally, the edge length of the grid is that is, the first three-dimensional structure is divided into cubes with an edge length of , wherein each cube includes structural information in a volume range of 1x10 3 cubic angstroms on the first three-dimensional structure as the protein-ligand binding information. The protein-ligand binding information includes the binding state information of the pocket on the protein and the ligand pose, so as to reflect the binding state of the pocket and the ligand at different positions of the first three-dimensional structure.

[0074] According to different positions of the grid on the first three-dimensional structure, the binding state information contained in the grid includes three kinds of protein structure only, ligand structure only, and protein structure and ligand structure at the same time. Among them, the grid with protein structure and ligand structure at the same time means that the docking of protein and ligand can be realized at the corresponding position of the first three-dimensional structure, that is, the possible binding site on the protein surface can be represented by several nearest grids, and by sorting the protein-ligand binding information in the grid, the optimal binding site can be obtained; the protein-ligand binding information in the grid with protein structure only and ligand structure only is zero, which can be considered as being not conducive to the binding of protein and ligand at the corresponding position of the first three-dimensional structure, so such grid can be deleted preferentially, thereby reducing the workload of subsequent sorting process, shortening the processing time, and improving the processing efficiency.

[0075] After screening the grid and deleting the grid with zero protein-ligand binding information at the corresponding position, the grid is input into the trained neural network, the grid is sorted, and the ranking of the corresponding binding site and binding pose in the grid is determined, wherein the protein-ligand binding information corresponding to the grid ranked first is confirmed as the most advantageous binding state in the actual environment.

[0076] Specifically, before sorting the grid, the neural network needs to be trained, and the training method of the neural network is as follows:

[0077] Obtain experimental characterization data of protein-ligand complex;

[0078] Generate a training three-dimensional structure using the experimental characterization data;

[0079] Voxelize the training three-dimensional structure to obtain a plurality of training grids, wherein the training grid includes protein-ligand binding information at the corresponding position;

[0080] Delete the training grid with zero protein-ligand binding information at the corresponding position to form a training grid set, and iteratively train the neural network according to the training grid set, wherein the neural network updates parameters according to the free energy of the protein-ligand binding information in the training grid set.

[0081] Specifically, the protein-ligand binding information reflects the binding state of the protein surface and the ligand in the experimental environment, i.e. the stability of the protein and the ligand after binding in different pockets and different postures, and the smaller the free energy of the protein-ligand binding information is, the more stable the binding of the protein and the ligand is, and the corresponding binding site and binding posture are in a more advantageous state. Therefore, the training grid with the lowest free energy of the protein-ligand binding information can be taken as a positive sample, and the training grid with the highest free energy of the protein-ligand binding information can be taken as a negative sample, and the parameters in the neural network are updated to achieve the effect of iterative training.

[0082] After the neural network is trained, the grid is input into the trained neural network to complete the sorting. Specifically, the neural network sorts the grid by using a sorting feature value to score the grid, and the higher the score of the grid is, the higher the sorting of the grid is, and the grid with the highest score is sorted first, wherein the sorting feature value includes:

[0083] The number feature value reflects the number of protein-ligand binding sites in a single grid, and the larger the number feature value is, the higher the possibility of binding of the protein and the ligand in the grid is, and the corresponding position is more likely to be a binding site in the real environment, and the score of the grid is higher;

[0084] The distance feature value reflects the distance between the center of a single grid and the protein-ligand binding site, and the smaller the distance feature value is, the closer the protein-ligand binding information in the grid is to the actual binding site, and the score of the grid is higher;

[0085] The posture feature value reflects the binding energy of the protein-ligand binding site in a single grid, and the smaller the posture feature value is, the lower the binding energy of the protein and the ligand in the grid is, and the more stable the corresponding binding site is, and the score of the grid is higher;

[0086] The pocket feature value reflects the biochemical performance of the protein pocket corresponding to the protein-ligand binding site in a single grid, and the larger the pocket feature value is, the better the druggability and solvent accessibility of the protein pocket corresponding to the grid are, and the more superior the biochemical performance of the protein pocket is, and the score of the grid is higher.

[0087] The importance ratios of the quantity feature values, the distance feature values, the pose feature values and the pocket feature values in the sorting are determined by training the neural network, the grid scores are determined in combination with the sorting feature values, and the protein-ligand binding information in the grid ranked first is obtained, so that the optimal binding site and the optimal ligand pose are obtained, and the protein and the ligand are docked at the optimal binding site in the optimal ligand pose to obtain the required protein-ligand complex three-dimensional structure. Alternatively, since a complete protein-ligand binding site cannot be included in a single grid, the grid ranked first and the grids within a certain range around the grid ranked first can be selected to comprehensively obtain the protein-ligand binding information, so that the optimal binding site and the optimal ligand pose are obtained.

[0088] In specific implementation, as shown in the method for determining the protein-ligand complex three-dimensional structure provided by the present application comprises the following steps: Figure 1 In specific implementation, as shown in the method for determining the protein-ligand complex three-dimensional structure provided by the present application comprises the following steps:

[0089] In specific implementation, as shown in the method for determining the protein-ligand complex three-dimensional structure provided by the present application comprises the following steps:

[0090] In one embodiment, before step S100, the method further comprises:

[0091] In specific implementation, as shown in the method for determining the protein-ligand complex three-dimensional structure provided by the present application comprises the following steps:

[0092] In specific implementation, as shown in the method for determining the protein-ligand complex three-dimensional structure provided by the present application comprises the following steps:

[0093] In specific implementation, the selection standard feature values include resolution, completeness, R value and R-free value, so that correct protein molecular data can be obtained from different protein databases more quickly.

[0094] In specific implementation, the molecular structure data of the ligand is data in SMILES format, pdb format, mol format, mol2 format or sdf format, so that the data can be used for virtual screening, reverse docking and cross-docking.

[0095] In one embodiment, step S100 specifically comprises:

[0096] In specific implementation, as shown in the method for determining the protein-ligand complex three-dimensional structure provided by the present application comprises the following steps:

[0097] Step S120, processing the protein data by a pocket detection method to obtain protein pocket data, which shows the binding site information of the protein surface.

[0098] In implementation, step S110 and step S120 are performed simultaneously and in parallel, so as to simplify the overall process and shorten the processing time. In the blind docking method, AutoDock Vina software, PLANTS software, Z-Dock software and Galaxydock3 software are comprehensively applied, so as to cope with different databases and ensure the same processing effect. In the pocket detection method, P2rank algorithm and Fpocket algorithm are comprehensively applied. Both algorithms can be applied alone, but the comprehensive application of the two algorithms can obtain more accurate processing effect.

[0099] In implementation, as shown in Figure 1 After step S100, the method for determining the three-dimensional structure of the protein-ligand complex further includes:

[0100] Step S200, matching the protein pocket data and the ligand pose data to obtain a first three-dimensional structure.

[0101] In an embodiment, the protein pocket data and the ligand pose data are matched by using a plurality of molecular docking software to predict the binding state of the protein and the ligand and simulate the formation of the first three-dimensional structure, so as to include all possible binding states of all protein pockets and all ligand poses in the first three-dimensional structure. Specifically, the protein pocket data and the ligand pose data are matched by using AutoDockVina software, PLANTS software, Z-Dock software, Galaxydock3 software, P2rank algorithm software and Fpocket algorithm software, and the first three-dimensional structure is simulated according to the prediction result.

[0102] In implementation, as shown in Figure 1 After step S200, the method for determining the three-dimensional structure of the protein-ligand complex further includes:

[0103] Step S300, voxelizing the first three-dimensional structure to obtain a plurality of grids, and the grids include protein-ligand binding information of corresponding positions.

[0104] In an embodiment, voxelizing the first three-dimensional structure includes dividing the first three-dimensional structure into a plurality of cubes with the same edge length and angstrom level, and assigning the protein-ligand binding information on the first three-dimensional structure to the cubes at corresponding positions to form the grids. Optionally, the edge length of the grid is Each of the grids includes 1x103 cubic angstroms The structural information in the volume range is the protein-ligand binding information. Specifically, the protein-ligand binding information includes the number of protein-ligand binding sites, the distance between protein-ligand binding sites, the protein-ligand binding energy, and protein pocket information.

[0105] In specific implementation, as shown in Figure 1 After step S300, the method for determining the three-dimensional structure of the protein-ligand complex according to the present application further includes:

[0106] In step S400, the grid is input into the trained neural network to rank the grid.

[0107] In specific implementation, the grid with zero protein-ligand binding information at the corresponding position is deleted, and the remaining grids form a test grid set. The test grid set is input into the neural network, and the remaining grids are ranked.

[0108] In an embodiment, the training process of the neural network is as follows:

[0109] In step S401, experimental characterization data of the protein-ligand complex is obtained.

[0110] In step S402, a training three-dimensional structure is generated using the experimental characterization data.

[0111] In step S403, the training three-dimensional structure is voxelized to obtain a plurality of training grids, and the training grids include protein-ligand binding information at the corresponding position.

[0112] In step S404, the training grid with zero protein-ligand binding information at the corresponding position is deleted to form a training grid set. The neural network is iteratively trained according to the training grid set, and the parameters of the neural network are updated according to the free energy of the protein-ligand binding information in the training grid set.

[0113] In specific implementation, the training grid with the lowest free energy of the protein-ligand binding information is taken as a positive sample, and the training grid with the highest free energy of the protein-ligand binding information is taken as a negative sample. The parameters in the neural network are updated. Alternatively, the protein-ligand binding information is ranked from low to high free energy, where the protein-ligand binding information with the highest free energy is marked as 0, and the protein-ligand binding information with the lowest free energy is marked as 1. A score in the range of 0-1 is assigned according to the change of the free energy of the protein-ligand binding information between the lowest free energy and the highest free energy, so as to obtain a continuous sequence, thereby better training the neural network.

[0114] In an embodiment, the neural network ranks the grids according to ranking feature values, which include:

[0115] a quantity feature value representing the number of protein-ligand binding sites in a single grid;

[0116] a distance feature value representing the distance between the center of a single grid and a protein-ligand binding site;

[0117] a pose feature value representing the binding energy of a protein-ligand binding site in a single grid; and

[0118] a pocket feature value representing the biochemical properties of a protein pocket corresponding to a protein-ligand binding site in a single grid.

[0119] In implementation, the importance proportions of the quantity feature value, the distance feature value, the pose feature value and the pocket feature value in ranking are determined through the training process of the neural network, so as to rank the grids by using the quantity feature value, the distance feature value, the pose feature value and the pocket feature value in the actual ranking process.

[0120] In implementation, the quantity feature value of a single grid is evaluated by using AutoDock Vina software, PLANTS software, Z-Dock software and Galaxydock3 software. The greater the quantity feature value is, the higher the possibility of protein-ligand binding in the grid is, and the more likely the corresponding position is a binding site in the real environment. The grid has a higher score. When multiple software determines that the quantity feature value of a grid is the maximum, it means that the protein surface region corresponding to the grid is the most likely binding site. In an embodiment, the quantity feature value of protein 1GS4 under the MTi standard is evaluated by using Z-Dock software and PLANTS software, and the grid with a greater quantity feature value is used to determine the binding site on the surface of protein 1GS4.

[0121] In specific implementation, the AutoDock Vina software, the PLANTS software, the Z-Dock software and the Galaxydock3 software are used to evaluate the distance characteristic value of a single grid, and the smaller the distance characteristic value is, the closer the protein-ligand binding information in the grid to the actual binding site, and the higher the score of the grid. It should be noted that when the pocket on the surface of the protein is large, only the number of binding sites in a single grid is used to determine the binding site, and the detailed structure of the sub-pocket in the large pocket can be lost, so that the optimal binding site is incorrectly positioned. By selecting the grid with a smaller distance between the center of the grid and the binding site, it can be ensured that the grid and the surrounding grid are positioned in the area of the more optimal binding site on the surface of the protein. Alternatively, when there are multiple binding sites, the average distance between the center of the grid and the binding sites is calculated as the distance characteristic value. In an embodiment, the AutoDock Vina software, the PLANTS software and the Galaxydock3 software are used to evaluate the distance characteristic value of the protein 1FM9 under the MTi standard, and the binding site on the surface of the protein 1FM9 is determined by the grid positioning with a smaller distance characteristic value in the case of fewer binding sites.

[0122] By combining the number characteristic value and the distance characteristic value, a grid sorting result with higher robustness can be provided.

[0123] In practice, AutoDock Vina, PLANTS, Z-Dock, and Galaxydock3 software are used to evaluate the attitude characteristic values ​​of individual grids. These attitude characteristic values ​​include protein structure, ligand structure, and molecular binding structure. A smaller attitude characteristic value indicates a lower binding energy between the protein and ligand in the grid, a more stable binding site, and a higher grid score. The attitude characteristic values ​​determine the binding effectiveness of ligands with different attitudes at a given binding site. In one embodiment, AutoDock Vina, PLANTS, and Galaxydock3 software are used to evaluate the attitude characteristic values ​​of protein 2YDO and rosiglitazone under the MTi standard. Regions with lower surface scores for protein 2YDO are selected to determine the binding site of 2YDO and the rosiglitazone ligand, as well as the corresponding ligand attitude. It should be noted that different software programs use different scoring methods for protein-ligand binding energy. Among them, AutoDock Vina, PLANTS, and Galaxydock3 software give evaluation results that are inversely proportional to binding energy. The higher the evaluation score, the lower the binding energy, the smaller the final attitude feature value, and the more stable the corresponding binding site. Z-Dock software gives evaluation results that are directly proportional to binding energy. The higher the evaluation score, the higher the binding energy, the larger the final attitude feature value, and the less stable the corresponding binding site.

[0124] In practice, the P2rank and Fpocket algorithms are used to evaluate the pocket feature values ​​of individual grids. A larger pocket feature value indicates better biochemical performance of the protein pocket corresponding to the grid, resulting in a higher grid score. The biochemical performance of a protein includes drug-like properties and solvent accessibility. Drug-like properties represent the likelihood that the protein pocket will exhibit pharmaceutical properties; the more easily the protein pocket undergoes structural changes, the higher its drug-like properties and the better its biochemical performance. Solvent accessibility represents the surface area (SASA) of the protein surface that can be contacted by a solvent; better solvent accessibility indicates better biochemical performance. In one embodiment, the P2rank and Fpocket algorithms are used to evaluate the pocket feature values ​​of protein 3MXF under the MTi standard, and the protein pocket corresponding to the grid with the highest score is selected as the more favorable pocket location for binding on the protein 3MXF surface.

[0125] In specific implementation, such as Figure 1 As shown, after step S500, the method for determining the three-dimensional structure of the protein-ligand complex of the present invention further includes:

[0126] Step S500, obtaining the optimal binding site and the optimal ligand pose according to the grid ranked first, and obtaining the protein-ligand complex three-dimensional structure by combining the optimal binding site and the optimal ligand pose.

[0127] In an embodiment, the grid ranked first and the grids within a certain range are selected to obtain the protein-ligand binding information, thereby obtaining the optimal binding site and the optimal ligand pose.

[0128] In an embodiment, we predict the three-dimensional structure of the protein 1FM6 complex with the ligand rosiglitazone in the MTi standard by the CoBDock framework corresponding to the method for determining the protein-ligand complex three-dimensional structure according to the present application. First, the protein data of the protein 1FM6 is obtained from the UniProtKB database, and the ligand data of the ligand rosiglitazone is obtained, and the protein pocket data and the ligand pose data are obtained by processing. The protein pocket data of the protein 1FM6 is matched with the ligand pose data of the rosiglitazone, and a first three-dimensional structure is generated. The first three-dimensional structure is voxelized to obtain the first three-dimensional structure composed of 480 grids, and after deleting the grids not containing protein-ligand binding information, the remaining grids are input into the trained neural network for ranking. In the case that the actual binding site is only ranked sixth in Fpocket and P2rank, CoBDock combines the scores given by Z-Dock software and AutoDock Vina software to ensure that the actual binding site is ranked first, thereby accurately identifying the structure of the actual binding site in all binding sites. In another embodiment, CoBDock accurately predicts the binding site and binding pose of the protein 2YEK matching with the ligand rosiglitazone (rosiglitazone) in the MTi standard by GalaxyDock3 software, Z-Dock software, AutoDock Vina software and PLANTS software. Therefore, by combining a plurality of different molecular docking software, CoBDock can comprehensively evaluate and accurately predict the three-dimensional structure of the protein-ligand complex.

[0129] In a specific implementation, in order to detect the final effect of the method for determining the three-dimensional structure of the protein-ligand complex, we verify the evaluation effect of the CoBDock framework corresponding to the method for determining the three-dimensional structure of the protein-ligand complex on the pocket recognition accuracy and the atomic mean distance of the existing P2rank algorithm software, the Fpocket algorithm software, the CB-Dock framework, and the CB-Dock2 framework and other molecular docking software based on the Astex Diverse Set (ADS) dataset, the MTi AutoDock Set (MTi) dataset, the COACH dataset, the DUDE dataset, and the PDBBind dataset. The standard of the pocket recognition accuracy is that the distance between the predicted binding site and the actual binding site is less than 5 A. That is, the accurate recognition; the atomic mean distance is the root-mean-square deviation (RMSD) of the predicted main chain atom position relative to the actual position.

[0130] As shown in Table 1, the performance of the CoBDock framework based on five different datasets. Since only the accurate determination of the binding site can determine the ideal binding mode in the field of molecular docking technology, the more accurate the prediction of the binding site is, the more accurate the prediction of the three-dimensional structure of the protein-ligand complex can be. Among them, the average distance refers to the average distance between the center of the prediction result and the center of the actual result (unit: A). The median distance is the median distance between the center of the prediction result and the center of the actual result (unit: A). The accuracy refers to the proportion of the case that the center of the prediction result is within the center of the actual result range to all cases.

[0131] Table 1

[0132]

[0133] From Table 1, it can be seen that CoBDock is overall superior to CB-Dock, and has better performance than CB-Dock based on all five data sets. Meanwhile, CoBDock is also very competitive relative to Fpocket and P2rank. Although for a particular data set, Fpocket and P2rank can achieve better prediction results due to individual optimization (such as P2rank relative to the COACH data set), they cannot maintain the corresponding effect when facing other data sets, while CoBDock can maintain good prediction results when facing different data sets, and has a wider range of application. This means that when considering multiple data sets, whether in the prediction of binding sites, pocket recognition or the prediction of binding poses, CoBDock embodies better prediction results and is more robust.

[0134] In the process of predicting the final binding pose, since there are various binding poses between the protein and the ligand, even if the binding site is determined, the actual binding pose needs to be determined from the various binding poses. For the prediction effectiveness of the binding pose, RMSD is generally used to judge. In the case where the RMSD value between the predicted binding pose and the actual binding pose is less than 2.0 A, the binding pose is judged to be effective, otherwise it is determined to be invalid.

[0135] Table 2

[0136] Benchmark Fpocket P2rank CB-0 CB-1 CB-2 CoB ADS 1 ]] 0.388 0.588 0.388 0.376 0.106 0.671 MTi 1 ]]> 0.593 0.593 0.37 0.259 0.148 0.778 DUDE 2 ]] 0.294 0.52 0.01 0.402 0 0.608 PDBBind 3 ]]> 0.485 0.441 0.044 0.002 0.011 0.517 Mean 0.44 0.536 0.203 0.26 0.066 0.644

[0137] As shown in Table 2, the judgment effectiveness of various different pocket recognition algorithms based on different data sets for the binding pose, wherein in the determined binding site, the ligand is re-docked to the same binding site using the predicted binding pose of each algorithm by using the PLANTS software, and then the RMSD value between the predicted binding pose of each algorithm and the actual binding pose is calculated. The case where the RMSD value is less than 2.0 A is taken as the effective value, and the value in the table is the proportion of the number of effective values predicted by each algorithm to the total number of predictions. Specifically, CB-0 is the original CB-Dock, CB-1 is CB-Dock-2 (Structure-based blind docking), and CB-2 is CB-Dock-2 (Template-based blind docking).

[0138] From Table 2, it can be seen that no matter which data set is based on, the prediction effectiveness of the binding pose given by CoBDock is higher than that of other algorithms, which embodies more excellent prediction results.

[0139] As shown in Table 2, the judgment effectiveness of various different pocket recognition algorithms based on different data sets for the binding pose, wherein in the determined binding site, the ligand is re-docked to the same binding site using the predicted binding pose of each algorithm by using the PLANTS software, and then the RMSD value between the predicted binding pose of each algorithm and the actual binding pose is calculated. The case where the RMSD value is less than 2.0 A is taken as the effective value, and the value in the table is the proportion of the number of effective values predicted by each algorithm to the total number of predictions. Specifically, CB-0 is the original CB-Dock, CB-1 is CB-Dock-2 (Structure-based blind docking), and CB-2 is CB-Dock-2 (Template-based blind docking). Figure 4 ​​The diagram illustrates the accuracy variations of CoBDock, Fpocket, and P2rank in pocket recognition. This accuracy refers to the distance between the predicted binding site and the actual binding site, based on the PDBBind dataset. The proportion of quantities within a certain range to the total quantity. Figure 4 (a) shows the changes in accuracy of CoBDock, Fpocket, and P2rank as the number of atoms in the protein changes; Figure 4 (b) represents the variation with the selected volume range (unit: cubic angstroms). The accuracy changes of CoBDock, Fpocket, and P2rank. Figure 4 As can be seen, CoBDock's prediction accuracy is less affected by the number of atoms or the volume range, while P2rank's accuracy drops significantly in the range of 6000-14000 atoms and 200000-500000 volumes. Fpocket's accuracy drops significantly when the volume range is above 500000 cubic angstroms. Therefore, CoBDock exhibits better overall prediction stability, and within the PDBBind dataset, CoBDock performs better than Fpocket or P2rank in pocket recognition.

[0140] like Figure 5 The diagram illustrates the accuracy variations of CoBDock, Fpocket, and P2rank in pose prediction. This accuracy refers to the RMSD value between the predicted combined pose and the actual combined pose, based on the PDBBind dataset. The proportion of quantities within a certain range to the total quantity. Figure 5 (a) shows the changes in accuracy of CoBDock, Fpocket, and P2rank as the number of atoms in the protein changes; Figure 5 (b) represents the variation with the selected volume range (unit: cubic angstroms). The accuracy changes of CoBDock, Fpocket, and P2rank. Figure 5 As can be seen, CoBDock's prediction accuracy is less affected by the number of atoms or the volume range, while P2rank's accuracy drops significantly in the range of 4000-14000 atoms and 200000-500000 volumes. Therefore, CoBDock exhibits better overall prediction stability, and within the PDBBind dataset, CoBDock performs better than Fpocket or P2rank in pose prediction.

[0141] In summary, CoBDock can give excellent prediction results of the three-dimensional structure of protein-ligand complex on the basis of different data sets by combining various molecular docking software and algorithms, and can give more accurate prediction structure for proteins with different numbers of atoms or structures. Compared with existing molecular docking software or algorithms, the method for determining the three-dimensional structure of the protein-ligand complex can more accurately predict the three-dimensional structure of the protein-ligand complex, has a wider application range, and has stronger robustness.

[0142] Based on the above embodiments, the application further provides a device for determining the three-dimensional structure of a protein-ligand complex, as shown in Figure 6 The device comprises:

[0143] A data processing module 100 is configured to process protein data and ligand data simultaneously to generate protein pocket data and ligand pose data.

[0144] A structure generation module 200 is configured to match the protein pocket data and the ligand pose data to obtain a first three-dimensional structure.

[0145] A voxelization module 300 is configured to voxelize the first three-dimensional structure to obtain a plurality of grids, wherein the grids include protein-ligand binding information at corresponding positions.

[0146] A calculation module 400 is configured to input the grids into a trained neural network to sort the grids.

[0147] A combination module 500 is configured to obtain an optimal binding site according to the first sorted grid, combine a ligand at the optimal binding site, and obtain a three-dimensional structure of a protein-ligand complex, thereby determining the three-dimensional structure of the protein-ligand complex.

[0148] Based on the above embodiments, the application further provides a terminal, and a principle block diagram thereof can be as shown in Figure 7 The terminal comprises a processor, a computer readable storage medium, a network interface and a display screen connected through a system bus. The processor of the terminal is configured to provide computing and control capabilities. The computer readable storage medium of the terminal comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the terminal is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement the method for determining the three-dimensional structure of a protein-ligand complex. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.

[0149] Those skilled in the art can understand that, Figure 7The principle block diagram shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied. The specific terminal can include more or less components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0150] In an implementation, the computer readable storage medium of the terminal stores one or more programs, and the one or more processors are configured to execute the one or more programs include instructions for performing the method for determining the three-dimensional structure of the protein-ligand complex.

[0151] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, it can include the processes of the above-mentioned embodiments of the method. Any reference to memory, storage, database or other medium used in each embodiment of the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0152] In summary, the application discloses a method and device for determining the three-dimensional structure of a protein-ligand complex, and a terminal, the method comprising: simultaneously processing protein data and ligand data to generate protein pocket data and ligand pose data; matching the protein pocket data and the ligand pose data to obtain a first three-dimensional structure; voxelizing the first three-dimensional structure to obtain a plurality of grids, the grids including protein-ligand binding information at corresponding positions; inputting the grids into a trained neural network to sort the grids; and obtaining an optimal binding site and an optimal ligand pose according to the first grid in the sorting, and combining the optimal binding site and the optimal ligand pose to obtain the three-dimensional structure of the protein-ligand complex. By simultaneously extracting and processing protein data and ligand data, then forming grids containing protein-ligand binding information in the form of voxelization, and then sorting the grids through a neural network, the application can more quickly and accurately determine the three-dimensional structure of a protein-ligand complex, has a wider application range, and has stronger robustness.

[0153] It should be understood that the application is not limited to the examples described above, and that all modifications and variations that can be made by a person of ordinary skill in the art based on the above description should be within the scope of protection of the appended claims of the application.

Claims

1. A method for determining a three-dimensional structure of a protein-ligand complex, characterized by, The method comprises: simultaneously processing protein data and ligand data to generate protein pocket data and ligand pose data; matching the protein pocket data and the ligand pose data to obtain a first three-dimensional structure; voxelizing the first three-dimensional structure to obtain a plurality of grids, wherein the grids include protein-ligand binding information at corresponding positions; inputting the grids into a trained neural network to sort the grids; obtaining an optimal binding site and an optimal ligand pose according to the grid ranked first, and combining the optimal binding site and the optimal ligand pose to obtain a protein-ligand complex three-dimensional structure.

2. The method of determining the three-dimensional structure of a protein-ligand complex according to claim 1, wherein Before inputting the grids into the trained neural network, it comprises: deleting the grids with zero protein-ligand binding information at corresponding positions.

3. The method of determining the three-dimensional structure of a protein-ligand complex according to claim 1, wherein Before simultaneously processing the protein data and the ligand data, it comprises: obtaining protein molecular structure data in a protein database, sorting the protein molecular structure data according to selected standard eigenvalues, and selecting the protein molecular structure data ranked first as the protein data; obtaining molecular structure data of a ligand, adding polar hydrogen to the molecular structure data to obtain the ligand data.

4. The method of determining a three-dimensional structure of a protein-ligand complex according to claim 3, wherein The selected standard eigenvalues comprise: resolution, which shows the resolution level of an electron density map; completeness, which shows the percentage of the part that can be simulated in the protein molecular structure data in the total structure; R value, which represents the difference between the diffraction pattern of the model obtained from the protein molecular structure data and the diffraction pattern obtained from the experiment; R-free value, which represents the degree of overfitting.

5. The method of determining a three-dimensional structure of a protein-ligand complex according to claim 1, wherein The simultaneous processing of the protein data and the ligand data to generate the protein pocket data and the ligand pose data comprises: processing the protein data and the ligand data by using a blind docking method to obtain the ligand pose data, wherein the ligand pose data reflects the state of the ligand in the protein-ligand binding process; processing the protein data by using a pocket detection method to obtain the protein pocket data, wherein the protein pocket data shows the binding site information on the surface of the protein; the above two steps are simultaneously and parallelly processed.

6. The method of determining a three-dimensional structure of a protein-ligand complex according to claim 1, wherein The voxelization of the first three-dimensional structure comprises: said first three-dimensional structure is divided into cubes with edge length of 2. allocating the protein-ligand binding information on the first three-dimensional structure to the corresponding positions of the cubes to form the grids.

7. The method of determining a three-dimensional structure of a protein-ligand complex according to claim 1, wherein The training process of the neural network comprises: obtaining experimental characterization data of a protein-ligand complex; generating training three-dimensional structures by using the experimental characterization data; voxelizing the training three-dimensional structures to obtain a plurality of training grids, wherein the training grids include protein-ligand binding information at corresponding positions; deleting the training grids with zero protein-ligand binding information at corresponding positions to form a training grid set, and iteratively training the neural network according to the training grid set, wherein the neural network updates parameters according to the free energy of the protein-ligand binding information in the training grid set.

8. The method of determining a three-dimensional structure of a protein-ligand complex according to claim 1, wherein The sorting manner of the neural network for the grids comprises sorting the grids according to sorting eigenvalues, wherein the sorting eigenvalues comprise: a quantity characteristic value representing the number of protein-ligand binding sites in a single grid; a distance characteristic value representing the distance between the center of a single grid and a protein-ligand binding site; a pose characteristic value representing the binding energy of a protein-ligand binding site in a single grid; a pocket characteristic value representing the biochemical properties of a protein pocket corresponding to a protein-ligand binding site in a single grid.

9. An apparatus for determining a three-dimensional structure of a protein-ligand complex, characterized by The device comprises: a data processing module for simultaneously processing protein data and ligand data to generate protein pocket data and ligand pose data; a structure generation module for matching the protein pocket data and the ligand pose data to obtain a first three-dimensional structure; a voxelization module for voxelizing the first three-dimensional structure to obtain a plurality of grids, the grids including protein-ligand binding information at corresponding positions; a calculation module for inputting the grids into a trained neural network to sort the grids; a combination module for obtaining an optimal binding site and an optimal ligand pose according to the first sorted grid, and combining the optimal binding site and the optimal ligand pose to obtain a protein-ligand complex three-dimensional structure.

10. A terminal, characterized by comprising: The terminal comprises a computer readable storage medium and one or more processors; the computer readable storage medium stores one or more programs; the programs contain instructions for executing the method for determining the protein-ligand complex three-dimensional structure according to any one of claims 1-8; and the processor is configured to execute the programs.

Citation Information

Patent Citations

  • Virtual drug screening method and device, computing equipment and storage medium

    CN111462833A

  • System and method for prediction of protein-ligand bioactivity using point-cloud machine learning

    US11256995B1