Drug molecule screening method and device, computer equipment, readable storage medium and program product
By predicting the binding affinity of drug molecules output by predicting the model and combining with three-dimensional structural evaluation, the problem of inefficient drug molecular screening in the existing technology is solved, fast and accurate molecular screening is achieved, and the drug development process is shortened.
Patent Information
- Application Number
- CN202510648114.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The prior art has low processing efficiency for large compound libraries in drug molecular screening, large data processing volume, high screening difficulty and time-consuming, resulting in extended drug development process.
By obtaining the molecular characteristics of the molecule to be screened and the molecular characteristics of the target protein molecule, input it to the prediction model to output the first binding affinity, screen the candidate molecule, and determine the second binding affinity based on the three-dimensional structure of the candidate molecule, and determine the target molecule with both.
It achieves rapid and accurate screening of target molecules from the molecules to be screened, shortens the drug research and development process, improves the efficiency and accuracy of molecular screening, and reduces time costs.
Smart Images

Figure CN120164548A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of biomedical technologies, and particularly to a method, device, computer device, computer-readable storage medium, and computer program product for screening drug molecules. Background Art
[0002] Research shows that it takes more than 10 years for a drug to be developed and successfully launched. Therefore, accelerating the drug development process will facilitate the new drug R & D process. In the entire drug development process, the discovery and optimization of lead compounds is a very important link.
[0003] In related technologies, virtual screening can be used to reduce the number of compounds screened in experiments, effectively improving the discovery efficiency of lead compounds. Usually, local search algorithms that combine genetic algorithms and local optimization algorithms can be used for virtual screening. However, when dealing with large compound libraries, this method has a huge amount of data processing, high screening difficulty, large getting-started difficulty, requires a large amount of time, and is inefficient. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for screening drug molecules that can improve efficiency.
[0005] In a first aspect, the present application provides a method for screening drug molecules, including:
[0006] Obtain a molecule to be screened, and input the molecular characteristics of the molecule to be screened and the molecular characteristics of a target protein molecule into a prediction model, and the prediction model outputs a first binding affinity between the molecule to be screened and the target protein molecule, where the prediction model is trained based on the correspondence between molecular characteristic samples and binding affinity labels;
[0007] Determine candidate molecules from the molecules to be screened whose first binding affinity is greater than a preset threshold;
[0008] According to the three-dimensional structure of the candidate molecules, determine the binding energy between the candidate molecules and the target protein molecule to obtain a second binding affinity of the candidate molecules;
[0009] Determine target molecules from the candidate molecules according to the first binding affinity and the second binding affinity.
[0010] In one embodiment, the determining target molecules from the candidate molecules according to the first binding affinity and the second binding affinity includes:
[0011] Determine a target score of the candidate molecule according to the first binding affinity and the second binding affinity;
[0012] Determine an initial molecule from the candidate molecules according to the target score;
[0013] Cluster the initial molecules to obtain a plurality of molecule groups, and determine screened molecules corresponding to each molecule group according to the target score to obtain target molecules.
[0014] In one embodiment, the determining the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule includes:
[0015] Determine a target protein region according to information on a target binding pocket corresponding to the target protein molecule;
[0016] Within the target protein region, determine energy terms in multiple dimensions between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule;
[0017] Determine the binding energy between the candidate molecule and the target protein molecule according to the energy terms in multiple dimensions.
[0018] In one embodiment, the method for determining the target binding pocket includes:
[0019] Obtain structure information of the target protein molecule;
[0020] Input the structure information into a preset pocket prediction model, and output a candidate binding pocket corresponding to the target protein molecule and a corresponding pocket score through the preset pocket prediction model;
[0021] Determine the target binding pocket from the candidate binding pockets according to the pocket score.
[0022] In one embodiment, the obtaining the structure information of the target protein molecule includes:
[0023] Determine a missing structure in the target protein molecule according to the structure information of the target protein molecule;
[0024] Complete the missing structure to obtain an adjusted target protein molecule;
[0025] Obtain the structure information of the adjusted target protein molecule.
[0026] In one embodiment, the molecular characteristics of the molecule to be screened include the molecular characteristic map of the molecule to be screened; the molecular characteristics of the target protein molecule include the molecular characteristic map of the target protein molecule; the determination method of the molecular characteristics of the molecule to be screened and the molecular characteristics of the target protein molecule includes:
[0027] Generate the molecular characteristic map corresponding to the molecule to be screened according to the molecular structure of the molecule to be screened;
[0028] Input the structural sequence information of the target protein molecule into a preset protein language model, and construct the molecular characteristic map corresponding to the target protein molecule through the sequence characteristics output by the preset protein language model.
[0029] In a second aspect, the present application also provides a drug molecule screening device, including:
[0030] An acquisition module, configured to acquire a molecule to be screened, and input the molecular characteristics of the molecule to be screened and the molecular characteristics of the target protein molecule into a prediction model, and output the first binding affinity between the molecule to be screened and the target protein molecule through the prediction model, where the prediction model is trained based on the correspondence between molecular characteristic samples and binding affinity labels;
[0031] A first determination module, configured to determine, from the molecules to be screened, candidate molecules whose first binding affinity is greater than a preset threshold;
[0032] A second determination module, configured to determine the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule, and obtain the second binding affinity of the candidate molecule;
[0033] A third determination module, configured to determine a target molecule from the candidate molecules according to the first binding affinity and the second binding affinity.
[0034] In one embodiment, the third determination module is further configured to:
[0035] Determine the target score of the candidate molecule according to the first binding affinity and the second binding affinity;
[0036] Determine an initial molecule from the candidate molecules according to the target score;
[0037] Cluster the initial molecules to obtain multiple molecule groups, and determine the screened molecules corresponding to each molecule group according to the target score to obtain the target molecule.
[0038] In one embodiment, the second determination module is further configured to:
[0039] Determine the target protein region according to the information of the target binding pocket corresponding to the target protein molecule;
[0040] Within the target protein region, determine the energy terms of multiple dimensions between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule;
[0041] Determine the binding energy between the candidate molecule and the target protein molecule according to the energy terms of multiple dimensions.
[0042] In one embodiment, the device further includes a determination module for the target binding pocket, and the determination module for the target binding pocket is used for:
[0043] Obtain the structural information of the target protein molecule;
[0044] Input the structural information into a preset pocket prediction model, and output the candidate binding pocket corresponding to the target protein molecule and the corresponding pocket score through the preset pocket prediction model;
[0045] Determine the target binding pocket from the candidate binding pockets according to the pocket score.
[0046] In one embodiment, the determination module for the target binding pocket is further used for:
[0047] Determine the missing structure in the target protein molecule according to the structural information of the target protein molecule;
[0048] Complete the missing structure to obtain an adjusted target protein molecule;
[0049] Obtain the structural information of the adjusted target protein molecule.
[0050] In one embodiment, the molecular characteristics of the molecule to be screened include the molecular feature map of the molecule to be screened; the molecular characteristics of the target protein molecule include the molecular feature map of the target protein molecule; the device further includes a determination module for the molecular characteristics of the molecule to be screened and the molecular characteristics of the target protein molecule, which is used for:
[0051] Generate the molecular feature map corresponding to the molecule to be screened according to the molecular structure of the molecule to be screened;
[0052] Input the structural sequence information of the target protein molecule into a preset protein language model, and construct the molecular feature map corresponding to the target protein molecule through the sequence features output by the preset protein language model.
[0053] In a third aspect, an embodiment of the present disclosure further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of the embodiments of the present disclosure are implemented.
[0054] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the method according to any one of the embodiments of the present disclosure are implemented.
[0055] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of the embodiments of the present disclosure are implemented.
[0056] In the above drug molecule screening method, device, computer device, computer-readable storage medium, and computer program product, when performing drug molecule screening, first, a prediction model is used to output the first binding affinity between the molecule to be screened and the target protein molecule. Candidate molecules are screened from the molecules to be screened according to the first binding affinity, and based on the three-dimensional structure of the candidate molecules, the binding energy between the candidate molecules and the target protein molecule is determined to obtain the second binding affinity. The target molecule is determined from the candidate molecules according to the first binding affinity and the second binding affinity, so that the target molecule can be quickly and accurately screened from the molecules to be screened, effectively shortening the drug R & D process. Through the prediction model, rapid and accurate molecular high-throughput screening can be achieved, and some molecules with relatively low binding possibilities can be quickly screened out. Subsequently, for the preliminarily screened candidate molecules, a secondary evaluation of the binding affinity is performed according to the three-dimensional structure of the molecules, which increases the interpretability of the screening results. At the same time, it is not necessary to perform a secondary evaluation on all molecules to be screened, effectively reducing the time cost and improving the molecular screening efficiency and accuracy. Determining the target molecule based on the first binding affinity and the second binding affinity takes into account the initial evaluation result of the prediction model and the calculated secondary evaluation result, and the target molecule is screened from multiple dimensions, further improving the accuracy of the screening result, realizing a drug molecule screening process with high efficiency, high precision, and high interpretability, and being applicable to more application scenarios. Description of the Drawings
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0058] Figure 1 It is a schematic flow chart of a drug molecule screening method in an embodiment;
[0059] Figure 2 It is a schematic flow chart of a drug molecule screening method in another embodiment;
[0060] Figure 3 It is a schematic flow chart of a method for determining a target binding pocket in an embodiment;
[0061] Figure 4 It is a schematic flow chart of a drug molecule screening method in another embodiment;
[0062] Figure 5 It is a structural block diagram of a drug molecule screening device in an embodiment;
[0063] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0064] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0065] In one embodiment, as Figure 1 shown, a drug molecule screening method is provided. In this embodiment, it is exemplified that the method is applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0066] Step S110, obtain a molecule to be screened, and input the molecular characteristics of the molecule to be screened and the molecular characteristics of a target protein molecule into a prediction model, and the prediction model outputs a first binding affinity between the molecule to be screened and the target protein molecule, where the prediction model is trained based on the correspondence between molecular characteristic samples and binding affinity labels;
[0067] Exemplarily, a molecule to be screened is obtained, and the molecule to be screened can be obtained from a preset molecular library. In one example, the preset molecular library may include, but is not limited to, a structurally diverse small molecule library, a bioactive small molecule library, a natural product small molecule library, a traditional Chinese medicine small molecule library, etc. The molecule to be screened can be all or part of the molecules in the preset molecular library, and can be specifically determined according to the actual application scenario.
[0068] Optionally, when drug research and development is required, different databases can be processed to obtain a preset molecular library, and the molecules to be screened in the preset molecular library can be screened. In one example, by cleaning the data in databases such as ZINC and PubChem, compounds with diverse structures and different properties can be obtained to form a small molecule library with diverse structures; by cleaning the data in databases such as DrugBank, compounds that have been marketed or are in the clinical trial stage can be obtained to form a bioactive small molecule library; by cleaning the data in databases such as Natural Products Atlas, natural product compounds derived from bacteria, fungi, etc. can be obtained to form a natural product small molecule library; by cleaning the data in databases such as TCMSP, traditional Chinese medicine ingredient compounds can be obtained to form a traditional Chinese medicine small molecule library.
[0069] In some possible implementation manners, during the database construction process, taking the small molecule library with diverse structures as an example, molecular structure files are obtained from databases such as ZINC, PubChem, and ChEMBL, and then data cleaning work is performed on the data, including deleting invalid molecules, molecular standardization, deleting duplicate molecules, re-numbering according to new rules, and molecular property calculation operations. Specifically, first, RDKit is used to read the molecular structure files, and the files that cannot be read are regarded as invalid files and deleted. Subsequently, the molecules are converted into standard molecular graphs (rdmol) in a unique representation manner and deduplication operations are performed based on them to ensure that there are no duplicate molecules in the dataset and they are numbered uniformly according to the new standard. Subsequently, five properties specified by the rule of five for drug-likeness are calculated for each molecule and stored in the attributes of the molecule. The rule of five for drug-likeness (Lipinski's rule) is a rule for screening drug-like molecules. Usually, compounds that conform to Lipinski's rule tend to have better pharmacokinetic properties and higher bioavailability during the metabolic process in the body. The rule of five for drug-likeness includes molecular weight, number of hydrogen bond acceptors, number of hydrogen bond donors, number of rotatable bonds, and lipophilicity-hydrophilicity partition coefficient.
[0070] Exemplarily, after obtaining the molecule to be screened, the molecular features of the molecule to be screened and the molecular features of the target protein molecule are input into the prediction model, where the target protein molecule can be determined according to the actual application scenario. In one example, the target protein molecule can be determined according to various factors such as the product required in the actual application scenario, the properties of the drug to be developed, etc., and the target protein molecule can include one or more. In one example, the protein structure can be obtained according to the protein PDB ID or the protein file; in some possible implementation manners, according to the provided 4-digit protein PDB ID, the corresponding protein three-dimensional structure is automatically downloaded from the PDB database and stored for later use. If an incorrect ID is provided, it is prompted to provide the correct ID or provide the protein file by oneself; if a protein three-dimensional structure file (.pdb) is provided, the file is stored at a specified location for later use.
[0071] Optionally, the prediction model can be trained based on a preset artificial intelligence algorithm, and the prediction model is trained according to the correspondence between the molecular feature samples and the binding affinity labels, where the molecular feature samples can be determined according to the representation manner of the molecules in the actual application scenario, and the representation manner can include but is not limited to molecular descriptors, molecular graphs, molecular structures, sequences, etc., and can be specifically determined according to the actual application scenario.
[0072] In one example, the prediction model can be implemented based on PSICHIC. PSICHIC models the protein-ligand (i.e., the molecule to be screened) interaction through three physicochemical graph convolutional layers. Specifically, two independent GNNs are used to respectively transmit the messages inside the protein residues and ligand atoms, and at the same time, physicochemical constraints are constructed. The ligand atoms are grouped into functional groups, and the protein residues are grouped into regions to limit the range of interactions and avoid overfitting. When processing the intermolecular interaction, the ligand functional groups are aggregated into "ligand spheres", and the cross-attention mechanism is used to calculate the interaction strength between the ligand spheres and the protein regions. After the calculation is completed, the ligand spheres are depolymerized into updated ligand atoms, and the protein regions are depolymerized into updated residues. After passing through three layers of PGC, PSICHIC generates an interaction fingerprint, which contains the features of the protein and the ligand, as well as the importance scores of them in the interaction. Based on this interaction fingerprint, the binding affinity between the ligand molecule and the protein is predicted to obtain the first binding affinity.
[0073] Exemplarily, the prediction model obtains the predicted first binding affinity according to the molecular features of the molecule to be screened and the molecular features of the target protein molecule. In one example, the greater the first binding affinity, the higher the binding possibility between the molecule to be screened and the target protein molecule.
[0074] Step S120, determine the candidate molecules with the first binding affinity greater than the preset threshold from the molecules to be screened;
[0075] Exemplarily, after determining the first binding affinity of the molecule to be screened, candidate molecules are determined from the molecules to be screened according to the first binding affinity. In some examples, a preset threshold is set according to the actual application scenario, and candidate molecules with a first binding affinity greater than the preset threshold are determined, molecules with a higher binding possibility are obtained, and molecules with a lower binding possibility are screened out.
[0076] In one example, the target number corresponding to the candidate molecules can be set according to the actual application scenario, and the preset threshold is determined according to the target number and the first binding affinity. For example, the target number can be set to 10,000. The molecules to be screened are sorted in descending order according to the first binding affinity, and the preset threshold is determined according to the first binding affinities of the molecules to be screened ranked at the 10,000th and 10,001st positions, so as to screen out 10,000 candidate molecules.
[0077] Step S130: Determine the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule, and obtain the second binding affinity of the candidate molecule;
[0078] Exemplarily, a secondary evaluation of the binding affinity of the candidate molecule is performed. Optionally, according to the three-dimensional structure of the candidate molecule, the binding energy between the candidate molecule and the target protein molecule is determined, and the second binding affinity of the candidate molecule is obtained.
[0079] Step S140: Determine the target molecule from the candidate molecules according to the first binding affinity and the second binding affinity.
[0080] Exemplarily, the target molecule is determined according to the first binding affinity and the second binding affinity. In some possible implementation manners, the first binding affinity and the second binding affinity can be weighted and summed to obtain a new affinity score, and candidate molecules with a higher binding possibility are screened out as the target molecules according to the affinity score; it is also possible to determine a new affinity score according to the first binding affinity and the second binding affinity, cluster the candidate molecules, and screen the clustered candidate molecules according to the new affinity score to determine the target molecule.
[0081] In the embodiments of the present disclosure, when performing drug molecule screening, first, a prediction model is used to output the first binding affinity between the molecule to be screened and the target protein molecule. Candidate molecules are screened from the molecules to be screened according to the first binding affinity. Then, according to the three-dimensional structure of the candidate molecules, the binding energy between the candidate molecules and the target protein molecule is determined to obtain the second binding affinity. According to the first binding affinity and the second binding affinity, the target molecule is determined from the candidate molecules, so that the target molecule can be quickly and accurately screened from the molecules to be screened, effectively shortening the drug R & D process. Through the prediction model, rapid and accurate molecular high-throughput screening can be achieved, and some molecules with relatively low binding possibilities can be quickly screened out. Subsequently, for the candidate molecules preliminarily screened, according to the three-dimensional structure of the molecules, a secondary evaluation of the binding affinity is carried out, which increases the interpretability of the screening results. At the same time, it is not necessary to perform a secondary evaluation on all molecules to be screened, effectively reducing the time cost and improving the molecular screening efficiency and accuracy. Determining the target molecule based on the first binding affinity and the second binding affinity takes into account the initial evaluation result of the prediction model and the calculated secondary evaluation result, and the target molecule is screened from multiple dimensions, further improving the accuracy of the screening result, realizing a drug molecule screening process with high efficiency, high precision and high interpretability, and being applicable to more application scenarios.
[0082] In one embodiment, as Figure 2 shown, the determining the target molecule from the candidate molecules according to the first binding affinity and the second binding affinity includes:
[0083] Step S141, determining the target score of the candidate molecule according to the first binding affinity and the second binding affinity;
[0084] Step S142, determining the initial molecule from the candidate molecules according to the target score;
[0085] Step S143, clustering the initial molecules to obtain multiple molecule groups, and determining the screened molecules corresponding to each molecule group according to the target score to obtain the target molecule.
[0086] Exemplarily, according to the first binding affinity and the second binding affinity, the target score of the candidate molecule is determined. In one example, the weighted sum of the first binding affinity and the second binding affinity can be used as the target score. In one example, the first binding affinity is proportional to the binding possibility, the second binding affinity is proportional to the binding possibility, and the target score , where is the first binding affinity, is the second binding affinity, and the larger the value, the stronger the binding ability of the molecule to the protein.
[0087] Optionally, the higher the target score, the higher the likelihood that the molecule binds to the protein molecule. Based on the target score, initial molecules with a higher binding likelihood can be screened out from the candidate molecules for subsequent processing.
[0088] Exemplarily, clustering processing is performed on the initial molecules to obtain multiple molecule groups. In one example, hierarchical clustering can be performed using the Tanimoto coefficient as an index to obtain X molecule groups. Based on the clustering result and the target score, the screened molecules corresponding to each molecule group are determined to obtain the target molecules. In some possible implementation manners, the top three molecules with the highest target scores in each molecule group can be output as the screened molecules. If the number of molecules in the molecule group is less than three, all the molecules in the group are output.
[0089] In one possible implementation manner, the ECFP4 fingerprint of each molecule can be calculated. Based on the local chemical environment of the atoms, the complexity of the chemical structure around the atoms is iteratively increased to form a fingerprint vector of a fixed length, which is the ECFP4 fingerprint of the molecule. In one example, the length of the fingerprint vector can be 1024. The corresponding Tanimoto coefficient is determined according to the fingerprint vector of the candidate molecule , , where a is the number of 1s in vector S, b is the number of 1s in vector T, and c is the number of 1s at the same positions in both vectors.
[0090] In one example, after the target molecules are determined, the complex structure of the target molecules and the protein is output.
[0091] In the embodiments of the present disclosure, when determining the target score according to the first binding affinity and the second binding affinity, the target score of the candidate molecule is determined according to the first binding affinity and the second binding affinity, and the initial molecules are preliminarily screened out. The screened initial molecules are clustered to obtain multiple molecule groups, and the screened molecules are determined from each molecule group to obtain the target molecules. Thus, when determining the target molecules, the structural diversity among the molecules and the binding likelihood between the molecules and the protein molecule can be taken into account, effectively improving the accuracy of molecule screening and shortening the drug R & D cycle.
[0092] In one embodiment, the determining the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule includes:
[0093] Determining the target protein region according to the information of the target binding pocket corresponding to the target protein molecule;
[0094] Within the target protein region, determining multiple-dimensional energy terms between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule;
[0095] Determine the binding energy between the candidate molecule and the target protein molecule according to the energy terms of the multiple dimensions.
[0096] Exemplarily, when determining the binding energy, according to the information of the target binding pocket corresponding to the target protein molecule, determine the target protein region, where the target protein region may include a docking box, and the configuration of the docking box may include the central position of the docking box, the size information of the docking box, etc.
[0097] Optionally, within the range of the target protein region, calculate the energy terms of multiple dimensions between the candidate molecule and the target protein molecule. Among them, the energy terms of multiple dimensions may include, but are not limited to, van der Waals force, hydrogen bond, hydrophobic interaction, torsional entropy penalty, electrostatic interaction, desolvation effect, etc., and can be specifically determined according to the actual application scenario. According to the energy terms of multiple dimensions, determine the binding energy between the candidate molecule and the target protein molecule. In one example, different energy terms of different dimensions correspond to different weights, and according to the weighted sum of the energy terms of multiple dimensions, determine the binding energy between the candidate molecule and the target protein molecule. In one example, the greater the binding energy, the smaller the binding possibility.
[0098] In some possible implementation manners, when determining the binding energy, perform a conformational search on the candidate molecule, generate 3 conformations for each molecule and save them respectively for subsequent molecular docking; process the ligand molecule (i.e., the candidate molecule) into a pdbqt format file for subsequent molecular docking; docking box preparation: based on the information of the target binding pocket of the target protein molecule, generate a docking box file, and the configured docking box is as follows: center_x, center_y, center_z are the central positions of the docking box, and size_x, size_y, size_z are the sizes of the docking box; perform molecular docking and calculate the binding energy between the candidate molecule and the target protein molecule. In one example, the binding energy can be calculated by AutoDock Vina.
[0099] In the embodiments of the present disclosure, when calculating the binding energy between the candidate molecule and the target protein molecule, according to the information of the target binding pocket corresponding to the target protein molecule, determine the energy terms of multiple dimensions within the target protein region, and obtain the binding energy. Calculate the binding energy from the energy terms of multiple dimensions, which improves the interpretability of the screening results, and the results are accurate, ensuring the reliability and accuracy of the screening process.
[0100] In one embodiment, as Figure 3 shown, the determination method of the target binding pocket includes:
[0101] Step S310, obtain the structural information of the target protein molecule;
[0102] Step S320: Input the structural information into a preset pocket prediction model, and output the candidate binding pockets corresponding to the target protein molecule and the corresponding pocket scores through the preset pocket prediction model.
[0103] Step S330: Determine the target binding pocket from the candidate binding pockets according to the pocket scores.
[0104] Optionally, obtain the structural information of the target protein molecule, input the structural information into a preset pocket prediction model, and output the candidate binding pockets and the corresponding pocket scores through the preset pocket prediction model.
[0105] In some examples, when the preset pocket prediction model performs binding pocket prediction, it extracts atomic features according to the received structural information of the protein molecule, represents each atom in the protein as a feature vector, including information such as atom type, charge, and hydrophobicity; when encoding spatial information, it divides the three-dimensional space of the protein into a network and maps the atomic features into the grid, thereby converting the protein structure into a format suitable for processing by a deep learning model; updates and learns the processed features through a convolutional neural network, and finally inputs them into a fully connected layer to predict the binding site. In one example, the model outputs a probability map indicating the probability that each position is a binding site. By setting a threshold, the high-probability region is determined as the binding site, that is, the candidate binding pocket, and the candidate binding pocket is scored to obtain the corresponding pocket score.
[0106] Optionally, according to the pocket scores, select the pocket with a higher score as the binding pocket. Among them, the higher the pocket score, the higher the probability that the candidate binding pocket is a binding site. In one example, the structural information of the protein can include the three-dimensional structure file of the protein.
[0107] In the embodiments of the present disclosure, the candidate binding pockets are determined according to the structural information of the target protein molecule, and the target binding pocket is determined according to the scores, so that the target binding pocket can be quickly and accurately determined according to the structural information of the target protein molecule, and then used to determine the target protein region in the subsequent binding affinity evaluation process, effectively improving the accuracy and reliability of the binding affinity determination and being applicable to more application scenarios.
[0108] In one embodiment, the obtaining the structural information of the target protein molecule includes:
[0109] Determine the missing structure in the target protein molecule according to the structural information of the target protein molecule;
[0110] Complete the missing structure to obtain an adjusted target protein molecule;
[0111] Obtain the structural information of the adjusted target protein molecule.
[0112] Exemplarily, before determining the binding pocket, the protein structure is first optimized. According to the structural information of the target protein molecule, the missing structure in the target protein molecule is determined, where the missing structure may include, but is not limited to, protein residues, atoms, and protons. The missing structure is complemented to obtain the adjusted target protein molecule.
[0113] In one example, in addition to complementing the missing structure, other chains of the protein except the first chain are deleted; non-standard residues are replaced with standard residues; heterologous substances other than the protein in the structure file are deleted to obtain the adjusted target protein molecule.
[0114] In the embodiments of the present disclosure, when determining the structural information of the target protein, the missing structure is determined according to the structural information of the target protein and complemented to obtain the adjusted target protein molecule, so that the structure of the target protein molecule can be optimized and improved, making the structure of the target protein molecule in the subsequent calculation process more accurate and reliable, and ensuring the accuracy and reliability of the screening results.
[0115] In one embodiment, the molecular characteristics of the molecule to be screened include the molecular feature map of the molecule to be screened; the molecular characteristics of the target protein molecule include the molecular feature map of the target protein molecule; the determination methods of the molecular characteristics of the molecule to be screened and the molecular characteristics of the target protein molecule include:
[0116] Generate the molecular feature map corresponding to the molecule to be screened according to the molecular structure of the molecule to be screened;
[0117] Input the structural sequence information of the target protein molecule into a preset protein language model, and construct the molecular feature map corresponding to the target protein molecule through the sequence features output by the preset protein language model.
[0118] Exemplarily, the molecular characteristics of the molecule to be screened include the molecular feature map of the molecule to be screened. In one example, according to the molecular structure of the molecule to be screened, the molecular feature map corresponding to the molecule to be screened is generated, where the molecular graph can be generated from the SMILES string of the molecule by RDKit, and the calculated atomic features include, but are not limited to, atomic type and physicochemical properties (such as atomic degree, hybridization state, etc.). The SMILES string of the molecule can be obtained by converting according to the information of the molecule to be screened. For example, the molecule is converted into a SMILES formula that meets the input standard of the model molecule by using the RDKit toolkit and stored in a CSV file.
[0119] Optionally, the molecular characteristics of the target protein molecule include a molecular characteristics map. In some possible implementations, first obtain the structural sequence information of the target protein molecule. In one example, read the three-dimensional structure file of the protein, and then use the Biopython package to convert the protein into a FASTA sequence (i.e., structural sequence information) for input to the protein part of the model, and store it in the form of a CSV file. The preset protein language model may include, but is not limited to, the ESM2 protein language model. The preset protein language model predicts the contact map between protein residues based on the structural sequence information and embeds the residues into the map to obtain the molecular characteristics map of the target protein molecule. During the calculation process, the characteristics of each residue include, but are not limited to, residue type and physicochemical properties (such as pK value, hydrophobicity, etc.).
[0120] In the embodiments of the present disclosure, a molecular characteristics map corresponding to the molecule to be screened is generated according to the molecular structure of the molecule to be screened, and the molecular characteristics map of the target protein molecule is determined by using the structural sequence information of the target protein molecule and the preset protein language model, so that a characteristics map that can characterize the characteristics of the molecule to be screened and the characteristics of the target protein molecule can be obtained quickly and accurately. When using these characteristics maps to predict the binding affinity, the accuracy and reliability of the prediction results are ensured, and it is applicable to more application scenarios.
[0121] Figure 4 A drug molecule screening method shown according to an exemplary embodiment is referred to Figure 4 as shown, construct a data set, including molecular libraries such as a structurally diverse small molecule library, bioactive small molecules, natural product small molecule library, traditional Chinese medicine small molecule library, etc.; obtain the target protein molecule structure according to the provided protein PDB ID or protein file; predict the candidate binding pockets according to the protein structure and select the target binding pocket with the best score; perform a first binding affinity evaluation based on the prediction model to obtain the first binding affinity, and screen out candidate molecules; calculate the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule to obtain the second binding affinity; obtain the target score of the candidate molecule according to the first binding affinity and the second binding affinity, and perform molecular clustering to screen out the target molecule.
[0122] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0123] Based on the same inventive concept, an embodiment of the present application also provides a drug molecule screening device for implementing the above-mentioned drug molecule screening method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the drug molecule screening device provided below can refer to the limitations on the drug molecule screening method in the above text, and will not be repeated here.
[0124] In an exemplary embodiment, as Figure 5 shown, a drug molecule screening device 500 is provided, including:
[0125] An acquisition module 510, configured to acquire a molecule to be screened, and input the molecular characteristics of the molecule to be screened and the molecular characteristics of a target protein molecule into a prediction model, and output a first binding affinity between the molecule to be screened and the target protein molecule through the prediction model, where the prediction model is trained based on the correspondence between molecular characteristic samples and binding affinity labels;
[0126] A first determination module 520, configured to determine, from the molecules to be screened, candidate molecules whose first binding affinity is greater than a preset threshold;
[0127] A second determination module 530, configured to determine the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule, and obtain a second binding affinity of the candidate molecule;
[0128] A third determination module 540, configured to determine a target molecule from the candidate molecules according to the first binding affinity and the second binding affinity.
[0129] In one of the embodiments, the third determination module is further configured to:
[0130] Determine a target score of the candidate molecule according to the first binding affinity and the second binding affinity;
[0131] Determine an initial molecule from the candidate molecules according to the target score;
[0132] Cluster the initial molecules to obtain multiple molecule groups, and determine the screened molecules corresponding to each molecule group according to the target score to obtain the target molecules.
[0133] In one embodiment, the second determination module is further configured to:
[0134] Determine a target protein region according to the information of the target binding pocket corresponding to the target protein molecule;
[0135] Within the target protein region, determine the energy terms of multiple dimensions between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule;
[0136] Determine the binding energy between the candidate molecule and the target protein molecule according to the energy terms of multiple dimensions.
[0137] In one embodiment, the device further includes a determination module for the target binding pocket, and the determination module for the target binding pocket is configured to:
[0138] Obtain the structural information of the target protein molecule;
[0139] Input the structural information into a preset pocket prediction model, and output the candidate binding pocket corresponding to the target protein molecule and the corresponding pocket score through the preset pocket prediction model;
[0140] Determine the target binding pocket from the candidate binding pockets according to the pocket score.
[0141] In one embodiment, the determination module for the target binding pocket is further configured to:
[0142] Determine the missing structure in the target protein molecule according to the structural information of the target protein molecule;
[0143] Complete the missing structure to obtain an adjusted target protein molecule;
[0144] Obtain the structural information of the adjusted target protein molecule.
[0145] In one embodiment, the molecular characteristics of the molecule to be screened include the molecular feature map of the molecule to be screened; the molecular characteristics of the target protein molecule include the molecular feature map of the target protein molecule; the device further includes a determination module for the molecular characteristics of the molecule to be screened and the molecular characteristics of the target protein molecule, and is used for:
[0146] Generate a molecular feature map corresponding to the molecule to be screened according to the molecular structure of the molecule to be screened;
[0147] Input the structural sequence information of the target protein molecule into a preset protein language model, and construct a molecular feature map corresponding to the target protein molecule based on the sequence features output by the preset protein language model.
[0148] Each module in the above drug molecule screening device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0149] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data involved in the method described in this embodiment such as molecular features. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a drug molecule screening method.
[0150] Those skilled in the art can understand that Figure 6 the structure shown in
[0151] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0152] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0153] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0155] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0156] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0157] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for screening drug molecules, characterized in that: The method comprises: Acquire a molecule to be screened, and input the molecular features of the molecule to be screened and the molecular features of the target protein molecule into a prediction model, and output the first binding affinity between the molecule to be screened and the target protein molecule through the prediction model, wherein the prediction model is trained based on the correspondence between the molecular feature sample and the binding affinity label; From the molecules to be screened, determining candidate molecules whose first binding affinity is greater than a preset threshold; Determining the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule to obtain a second binding affinity of the candidate molecule; A target molecule is determined from the candidate molecules according to the first binding affinity and the second binding affinity.
2. The method according to claim 1, characterized in that Determining the target molecule from the candidate molecules according to the first binding affinity and the second binding affinity comprises: determining a target score for the candidate molecule based on the first binding affinity and the second binding affinity; determining an initial molecule from the candidate molecules according to the target score; The initial molecules are clustered to obtain a plurality of molecule groups, and the screened molecules corresponding to each of the molecule groups are determined according to the target scores to obtain target molecules.
3. The method according to claim 1, characterized in that Determining the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule includes: Determining the target protein region according to the information of the target binding pocket corresponding to the target protein molecule; In the target protein region, determining energy terms of multiple dimensions of the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule; The binding energy between the candidate molecule and the target protein molecule is determined according to the energy items of the multiple dimensions.
4. The method according to claim 3, characterized in that The method for determining the target binding pocket includes: Obtaining structural information of the target protein molecule; Inputting the structural information into a preset pocket prediction model, and outputting the candidate binding pocket corresponding to the target protein molecule and the corresponding pocket score through the preset pocket prediction model; According to the pocket score, a target binding pocket is determined from the candidate binding pockets.
5. The method according to claim 4, characterized in that The step of obtaining the structural information of the target protein molecule comprises: Determining the missing structure in the target protein molecule according to the structural information of the target protein molecule; Completing the missing structure to obtain an adjusted target protein molecule; The structural information of the adjusted target protein molecule is obtained.
6. The method according to claim 1, characterized in that The molecular characteristics of the molecule to be screened include the molecular characteristic graph of the molecule to be screened; the molecular characteristics of the target protein molecule include the molecular characteristic graph of the target protein molecule; The method for determining the molecular characteristics of the molecule to be screened and the molecular characteristics of the target protein molecule includes: Generating a molecular characteristic map corresponding to the molecule to be screened according to the molecular structure of the molecule to be screened; The structural sequence information of the target protein molecule is input into a preset protein language model, and a molecular feature map corresponding to the target protein molecule is constructed through the sequence features output by the preset protein language model.
7. A drug molecule screening device, characterized in that: The device comprises: An acquisition module, used for acquiring a molecule to be screened, and inputting the molecular features of the molecule to be screened and the molecular features of the target protein molecule into a prediction model, and outputting the first binding affinity between the molecule to be screened and the target protein molecule through the prediction model, wherein the prediction model is trained based on the correspondence between the molecular feature sample and the binding affinity label; A first determination module is used to determine, from the molecules to be screened, candidate molecules whose first binding affinity is greater than a preset threshold; A second determination module is used to determine the binding energy between the candidate molecule and the target protein molecule according to the three-dimensional structure of the candidate molecule to obtain a second binding affinity of the candidate molecule; The third determination module is used to determine the target molecule from the candidate molecules according to the first binding affinity and the second binding affinity.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method and device for model training, drug screening and affinity prediction
CN114333986A
Protein binding pocket prediction device and method
CN115620801A
Screening method and device of protein inhibitor, electronic equipment and medium
CN116486941A
Method for optimizing compound structure based on protein binding pocket
CN117037946A
Multi-level virtual screening method constructed based on mutant protein and application of multi-level virtual screening method
CN117854589A