Method and device for generating drug target affinity prediction model, computer device and storage medium

By generating covalent and non-covalent maps of drug molecules and protein atoms and revising the initial prediction model, the problem of insufficient accuracy in drug target affinity prediction methods is solved, and more efficient drug screening is achieved.

CN116525028BActive Publication Date: 2026-04-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2023-05-04
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing methods for predicting drug target affinity suffer from limitations in learning information, poor prediction accuracy, and unsuitability for large-scale virtual screening during drug development.

Method used

By acquiring the training dataset, covalent and non-covalent maps of drug molecules and protein atoms are generated. The prediction results are determined using the initial prediction model, and the initial prediction model is corrected based on the difference between the labeled data and the prediction results until the corrected prediction model is obtained.

Benefits of technology

It improves the accuracy of predicting affinity between drugs and targets, thereby increasing the efficiency and reliability of the drug screening process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116525028B_ABST
    Figure CN116525028B_ABST
Patent Text Reader

Abstract

This disclosure proposes a method, apparatus, computer device, and storage medium for generating a drug target affinity prediction model, relating to the field of computer technology. The method includes: after acquiring a training dataset, firstly generating covalent and non-covalent maps of the drug molecule and protein atoms based on their binding conformations; then inputting these maps into an initial prediction model to determine the prediction results between the drug molecule and protein atoms; and finally, revising the initial prediction model based on the differences between the labeled data and the prediction results, until a revised prediction model is obtained. Thus, by training the prediction model using the covalent and non-covalent maps between the drug and the target, the accuracy and reliability of the prediction model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, computer device, and storage medium for generating a drug target affinity prediction model. Background Technology

[0002] In the process of drug development, thousands of compounds are usually tested in order to find safe and effective drugs, which consumes a lot of manpower, resources and time. Therefore, predicting the affinity between drugs and targets has become a crucial step in the drug screening process.

[0003] However, as the demand for virtual drug screening continues to increase and our understanding of drug-target binding conformations deepens, existing affinity prediction methods have been found to have problems such as limited learning information, poor prediction accuracy, and unsuitability for large-scale virtual screening. Therefore, there is an urgent need to find a better method for predicting drug target affinity. Summary of the Invention

[0004] This disclosure aims to at least partially address one of the technical problems in the related art.

[0005] According to a first aspect of this disclosure, a method for generating a drug target affinity prediction model is proposed, comprising:

[0006] Obtain a training dataset, wherein each training dataset includes the binding conformation of a drug molecule and a protein atom and the annotation data between the drug molecule and the protein atom, the annotation data including a first affinity;

[0007] Based on the binding conformation, a covalent diagram and a non-covalent diagram of the drug molecule and the protein atoms are generated;

[0008] The covalent and non-covalent diagrams are input into the initial prediction model to determine the prediction results between the drug molecule and the protein atom, including the second affinity.

[0009] Based on the difference between the labeled data and the prediction results, the initial prediction model is corrected until a corrected prediction model is obtained.

[0010] A second aspect of this disclosure provides an apparatus for generating a drug target affinity prediction model, comprising:

[0011] The first acquisition module is used to acquire training datasets, wherein each training dataset includes the binding conformation of drug molecules and protein atoms and the annotation data between the drug molecules and the protein atoms, and the annotation data includes a first affinity.

[0012] The first generation module is used to generate a covalent diagram and a non-covalent diagram of the drug molecule and the protein atoms according to the binding conformation.

[0013] The first determining module is used to input the covalent diagram and the non-covalent diagram into the initial prediction model to determine the prediction result between the drug molecule and the protein atom, wherein the prediction result includes a second affinity.

[0014] The second acquisition module is used to correct the initial prediction model based on the difference between the labeled data and the prediction results until the corrected prediction model is obtained.

[0015] A third aspect of this disclosure provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements a method for generating a drug target affinity prediction model as proposed in a first aspect of this disclosure.

[0016] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for generating a drug target affinity prediction model as proposed in a first aspect of this disclosure.

[0017] A fifth aspect of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements a method for generating a drug target affinity prediction model as proposed in a first aspect of this disclosure.

[0018] The method, apparatus, computer equipment, and storage medium for generating drug target affinity prediction models disclosed herein have the following beneficial effects:

[0019] In this embodiment, after obtaining the training dataset, covalent and non-covalent maps of the drug molecule and protein atoms are first generated based on their binding conformations. These maps are then input into an initial prediction model to determine the prediction results between the drug molecule and protein atoms. The initial prediction model is then corrected based on the differences between the labeled data and the prediction results, until a corrected prediction model is obtained. This improves the model's accuracy in predicting the affinity between the drug and the target.

[0020] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0021] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which:

[0022] Figure 1 This is a flowchart illustrating a method for generating a drug target affinity prediction model according to an embodiment of this disclosure.

[0023] Figure 2 This is a flowchart illustrating a method for generating a drug target affinity prediction model according to another embodiment of this disclosure.

[0024] Figure 3 This is a flowchart illustrating a method for determining drug target affinity according to an embodiment of the present disclosure.

[0025] Figure 4 This is a schematic diagram of the structure of a device for generating a drug target affinity prediction model according to an embodiment of the present disclosure;

[0026] Figure 5 A schematic diagram of a device for determining drug target affinity provided in an embodiment of this disclosure;

[0027] Figure 6 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0028] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.

[0029] The method for generating the drug target affinity prediction model in this embodiment can be executed by the device for generating the drug target affinity prediction model in this embodiment. This device can be configured in a computer device, and this disclosure does not limit its use. The computer device can be any device with computing capabilities, such as a mobile phone, tablet computer, personal computer, personal digital assistant, or other hardware device with various operating systems, touchscreens, and / or displays. This embodiment uses the example of the drug target affinity prediction model generation device being configured in a drug target affinity prediction model generation system.

[0030] The following description, with reference to the accompanying drawings, outlines a method, apparatus, computer device, and storage medium for generating drug target affinity prediction models according to embodiments of this disclosure.

[0031] Figure 1 This is a flowchart illustrating a method for generating a drug target affinity prediction model according to an embodiment of this disclosure.

[0032] like Figure 1 As shown, the method for generating the drug target affinity prediction model may include:

[0033] Step 101: Obtain training datasets, wherein each training dataset includes the binding conformation of drug molecules and protein atoms and the annotation data between drug molecules and protein atoms, and the annotation data includes the first affinity.

[0034] In this context, protein atoms refer to the protein atoms at the target site that are located within a certain range near the drug molecule after the drug binds to the target. Binding conformation refers to the spatial arrangement of the drug molecule and protein atoms when they bind. Affinity refers to the strength of the binding between the drug molecule and the protein atoms; a higher affinity indicates a stronger binding, which increases the likelihood of drug activity at the target site.

[0035] It should be noted that the training dataset can be any dataset containing the binding conformations of drug molecules and protein atoms, such as the Protein Data Bank bind (PDBbind) dataset, etc., and this disclosure does not limit it.

[0036] Step 102: Based on the binding conformation, generate covalent and non-covalent diagrams of the drug molecule and protein atoms.

[0037] In some possible implementations, the system for generating drug target affinity prediction models first uses drug molecules and protein atoms as atoms in covalent and non-covalent diagrams. Then, the covalent bonds between protein atoms and drug molecules are used as edges of the covalent diagram. Based on the binding conformation of drug molecules and protein atoms, every two atoms within a certain distance are connected to form edges of the non-covalent diagram, thus generating both covalent and non-covalent diagrams.

[0038] A covalent bond is formed when two or more atoms share their outer electrons, ideally reaching an electron saturation state, thus creating a relatively stable and robust chemical structure. In other words, when a covalent bond exists between a protein atom and a drug molecule, it can be determined that the binding strength between them is relatively high. In this disclosure, the covalent bond information between protein atoms and drug molecules can be obtained from a dataset.

[0039] It should be noted that distance refers to the distance between protein atoms and drug molecules. The system for generating drug target affinity prediction models can determine whether the distance between atoms in the non-covalent diagram is within a set distance threshold. This threshold can be fixed or can vary according to the properties of the target or drug; this disclosure does not limit it in this way.

[0040] For example, when the distance threshold set in the generation system of the drug target affinity prediction model is a, in the non-covalent graph, the distance between atom A and atom B is b, and a > b, then an edge in the non-covalent graph can be formed between atom A and atom B; the distance between atom A and atom D is d, and a < d, so there is no edge between atom A and atom D in the non-covalent graph.

[0041] Step 103: Input the covalent graph and the non-covalent graph into the initial prediction model to determine the prediction result between the drug molecule and the protein atom. The prediction result includes the second affinity.

[0042] Among them, the initial prediction model can be a network model with any structure, and the present disclosure does not limit this.

[0043] In some possible implementation forms, the labeled data in the training dataset may also include the first interaction information between the drug molecule and the protein atom. Therefore, the generation system of the drug target affinity prediction model can first use the first sub-network in the initial prediction model to process the covalent graph to obtain the second affinity, and then use the second sub-network in the initial prediction model to process the non-covalent graph and the atomic vectors synchronized by the first sub-network to obtain the second interaction information between the drug molecule and the protein atom.

[0044] Among them, the interaction information may include at least one of hydrogen bonds, hydrophobic interactions, salt bridges, metal ions, halogen bonds, molecular access system keys MACCS, pharmacophore molecular fingerprints, and energy distributions, etc. The first sub-network and the second sub-network are two branch networks in the initial prediction model, and they can process the covalent graph and the non-covalent graph respectively.

[0045] It should be noted that the first interaction information can be calculated by the Open Drug Discovery Toolkit (Oddt) and is used to assist in the training of the affinity prediction model.

[0046] In some possible implementation forms, the first sub-network of the initial prediction model can update the atomic vectors based on the edge vectors connected to the atoms in the covalent graph to obtain the second affinity. At the same time, the atomic vectors before and after the update in the covalent graph are synchronized to the second sub-network. Then, the second sub-network can update the edge vectors based on the atomic vectors synchronized by the first sub-network, thereby obtaining the second interaction information between the drug molecule and the protein atom, realizing the flow of information between the covalent graph and the non-covalent graph, and improving the accuracy of the model's prediction of the affinity between the drug and the target.

[0047] Step 104: Based on the difference between the labeled data and the prediction result, correct the initial prediction model until the corrected prediction model is obtained.

[0048] In some possible implementations, the system for generating a drug target affinity prediction model can determine a correction gradient based on the second difference between the second affinity and the first affinity in the prediction results, and then perform reverse correction on the initial prediction model based on the correction gradient until the final prediction model is obtained.

[0049] In some possible implementations, the prediction results may also include second interaction information. In this case, the system generating the drug target affinity prediction model can first modify the second sub-network based on the first difference between the first and second interaction information in the prediction results. Since the acquisition of the second interaction information depends on the atomic vectors provided by the first sub-network, the system generating the drug target affinity prediction model can modify the first sub-network based on the first difference and the second difference between the first and second affinity.

[0050] The generation system for the drug target affinity prediction model can determine the correction gradient corresponding to the first sub-network based on the same weights and by combining the first difference and the second difference; or it can determine the correction gradient corresponding to the first sub-network based on different weights and by combining the first difference and the second difference. This disclosure does not limit this.

[0051] It should be noted that this disclosure primarily predicts the affinity between drug molecules and protein molecules based on covalent graphs, while predicting the interaction information between drug molecules and protein atoms based on non-covalent graphs and atomic vectors within the covalent graphs. Therefore, the influence of interaction information needs to be considered when modifying the first sub-network. In other words, non-covalent graphs and the interaction information between drug molecules and protein atoms assist in the training process of the first sub-network, thereby making the obtained first sub-network (affinity prediction network) more accurate and reliable.

[0052] In this embodiment, after acquiring the training dataset, the system generating the drug target affinity prediction model first generates covalent and non-covalent maps of the drug molecule and protein atoms based on their binding conformations. These maps are then input into the initial prediction model to determine the prediction results between the drug molecule and protein atoms. Finally, the initial prediction model is corrected based on the differences between the labeled data and the prediction results, until a corrected prediction model is obtained. Thus, by training the prediction model using the covalent and non-covalent maps between the drug and the target, the accuracy and reliability of the prediction model are improved.

[0053] Figure 2 This is a flowchart illustrating a method for generating a drug target affinity prediction model, as provided in another embodiment of this disclosure.

[0054] like Figure 2As shown, the method for generating the drug target affinity prediction model may include:

[0055] Step 201: Obtain training datasets, wherein each training dataset includes the binding conformation of drug molecules and protein atoms and the annotation data between drug molecules and protein atoms, and the annotation data includes the first affinity.

[0056] Step 202: Based on the binding conformation, generate covalent and non-covalent diagrams of the drug molecule and protein atoms.

[0057] The implementation of steps 201-202 above can be found in the detailed description of any of the above embodiments, and will not be repeated here.

[0058] Step 203: Process the covalent graph using the first sub-network in the initial prediction model to obtain the second affinity.

[0059] In some possible implementations, when the system generating the drug target affinity prediction model processes the covalent graph using the first sub-network in the initial prediction model, it first determines the edge connected to each atom based on the relationship between atoms and connecting edges in the covalent graph. Then, it updates the atom vectors using the vectors of the connecting edges to obtain the updated atom vectors. Finally, it determines the second affinity based on the vectors of the connecting edges in the covalent graph and the updated atom vectors.

[0060] It should be noted that during the processing of the covalent diagram, the atom vectors can be updated multiple times. The number of updates can be preset in the generation system of the drug target affinity prediction model. Alternatively, the number of updates can be determined during training. For example, if the change between two consecutive updates of the atom vector is within a preset range, then the accuracy of the atom vector can be considered to have met the requirements, and updates can be stopped. This disclosure does not limit this.

[0061] In some possible implementations, an atom may have covalent bonds with multiple atoms, meaning an atom may have multiple connecting edges. Therefore, when updating the vector of this atom, the weights of the connecting edges need to be considered, and the weights of different connecting edges may be the same or different. For example, the generation system of the drug target affinity prediction model can determine the weight of each connecting edge based on the correlation between each connecting edge and the atom (for example, the connecting edge with a larger covalent bond value corresponds to a larger weight). Alternatively, the generation system of the drug target affinity prediction model can also determine that each connecting edge has the same weight; this disclosure does not limit this. Furthermore, when updating the vector of an atom, the weight of the original vector of the atom should be greater than the weight of the edge vector.

[0062] For example, an atom in a covalent graph may have covalent bonds with three other atoms, meaning it has three connecting edges: edge x, edge y, and edge z. If the original weights of the atom's vectors are 0.6, edge x's vector weight is 0.2, edge y's vector weight is 0.1, and edge z's vector weight is 0.1, then the updated vector of the atom can be calculated as: original atom vector * 0.6 + edge x vector * 0.2 + edge y vector * 0.1 + edge z vector * 0.1. By updating the atom vectors in the covalent graph, the accuracy of the model's predicted affinity is further improved.

[0063] Step 204: Use the second sub-network in the initial prediction model to process the atomic vectors synchronized with the non-covalent graph and the first sub-network to obtain the second interaction information between drug molecules and protein atoms.

[0064] In some possible implementations, when the system generating the drug target affinity prediction model processes the atomic vectors synchronized with the non-covalent graph and the first sub-network in the initial prediction model using the second sub-network, it first determines the atoms connected to the edges based on the relationship between atoms and connecting edges in the non-covalent graph. Then, it updates the edge vectors in the non-covalent graph using the atomic vectors synchronized with the first sub-network to obtain the updated edge vectors. Finally, it determines the second interaction information based on the relationship between atoms and connecting edges in the non-covalent graph, the atomic vectors synchronized with the first sub-network, and the updated edge vectors.

[0065] It should be noted that the update of the edge vectors in the non-covalent graph of the drug target affinity prediction model generation system can be performed after the atomic vectors in the covalent graph have stopped updating, or the edge vectors can be updated based on the atomic vectors synchronized with the first sub-network after each update of the atomic vectors. This disclosure does not limit this.

[0066] In some possible implementations, each edge in the non-covalent graph connects two atoms. Therefore, when updating the edge vector, it is necessary to consider the weight of each atom with respect to the edge, as well as the original weight of the edge vector. The weights of different atoms may be the same or different. For example, the generation system of the drug target affinity prediction model can determine the degree of influence of an atom on the edge based on the number of protons contained in each atom, and thus determine the weight of the atom; or, it can determine the weight of an atom on each edge based on the number of edges connected to each atom, and so on. Alternatively, the generation system of the drug target affinity prediction model can also determine that the weights corresponding to each atom are the same, and this disclosure does not limit this.

[0067] For example, consider two atoms, M and N, connected by an edge. The original weight of the edge vector was 0.5. After the edge vector is updated, the weight of atom M becomes 0.2, and the weight of atom N becomes 0.3. Therefore, the updated edge vector can be calculated as: (original edge vector * 0.5) + (updated atom M vector * 0.2) + (updated atom N vector * 0.3). By updating the edge vectors of non-covalent graphs based on the updated atom vectors in the covalent graph, information exchange between the covalent and non-covalent graphs is achieved. This further assists in training the affinity prediction model, thereby improving its accuracy.

[0068] Step 205: Based on the first difference between the first interaction information and the second interaction information, the second sub-network is corrected.

[0069] Step 206: Based on the first difference and the second difference between the first affinity and the second affinity, the first sub-network is modified.

[0070] The specific implementation of steps 205-206 above can be found in the detailed description of any of the above embodiments, and will not be repeated here.

[0071] In this embodiment, the system for generating the drug target affinity prediction model acquires a training dataset and obtains the binding conformation and annotation data of drug molecules and protein atoms. First, based on the binding conformation of drug molecules and protein atoms, it generates covalent and non-covalent maps of the drug molecules and protein atoms. Then, it processes the covalent and non-covalent maps using an initial prediction model to obtain prediction results. Finally, based on a first difference and a second difference, it corrects the first and second sub-networks of the initial prediction model, respectively. Thus, by training the prediction model using the covalent and non-covalent maps between the drug and the target, the accuracy and reliability of the prediction model are improved.

[0072] Figure 3 This is a flowchart illustrating a method for determining drug target affinity according to an embodiment of this disclosure.

[0073] like Figure 3 As shown, the method for determining drug target affinity may include:

[0074] Step 301: Obtain the binding conformation of the candidate drug molecule and protein atoms.

[0075] Among them, candidate drug molecules refer to active compound molecules that are intended to undergo systematic preclinical trials and enter clinical research.

[0076] Step 302: Based on the binding conformation, generate a covalent diagram of the candidate drug molecule and protein atoms.

[0077] Step 303: Input the covalent diagram into a preset prediction model to obtain the affinity between the candidate drug molecule and protein atoms. The preset prediction model is based on... Figure 1 , Figure 2 It was generated through the steps in the process.

[0078] The specific implementation of steps 301-303 above can be found in any embodiment of this disclosure. Further details are omitted here.

[0079] In this embodiment, the system for determining drug target affinity first obtains the binding conformation of the candidate drug molecule and protein atoms. Then, based on this binding conformation, it generates a covalent map of the candidate drug molecule and protein atoms. The covalent map is then input into a preset prediction model to obtain the affinity between the candidate drug molecule and protein atoms. Therefore, by utilizing a prediction model trained based on both covalent and non-covalent maps, the accuracy of the determined affinity is improved, thus enhancing the accuracy and reliability of candidate drug molecule screening.

[0080] To achieve the above embodiments, this disclosure also proposes an apparatus for generating a drug target affinity prediction model.

[0081] Figure 4 This is a schematic diagram of the structure of the device for generating a drug target affinity prediction model provided in an embodiment of this disclosure.

[0082] like Figure 4 As shown, the device 400 for generating the drug target affinity prediction model includes:

[0083] The first acquisition module 401 is used to acquire training datasets, wherein each training dataset includes the binding conformation of drug molecules and protein atoms and the annotation data between drug molecules and protein atoms, and the annotation data includes the first affinity.

[0084] The first generation module 402 is used to generate covalent and non-covalent diagrams of drug molecules and protein atoms based on the binding conformation.

[0085] The first determining module 403 is used to input covalent and non-covalent diagrams into the initial prediction model to determine the prediction results between drug molecules and protein atoms, including the second affinity.

[0086] The second acquisition module 404 is used to correct the initial prediction model based on the difference between the labeled data and the prediction results until the corrected prediction model is obtained.

[0087] Optionally, the first determining module 403 further includes:

[0088] The third acquisition module 405 is used to process the covalent graph using the first sub-network in the initial prediction model to obtain the second affinity.

[0089] The fourth acquisition module 406 is used to process the non-covalent graph and the atomic vectors synchronized with the first sub-network using the second sub-network in the initial prediction model to obtain the second interaction information between drug molecules and protein atoms. The second sending module 506 is used to send the second video stream to the server.

[0090] Optionally, the third acquisition module 405 mentioned above further includes:

[0091] The fifth acquisition module 407 is used to update the vector of the atom based on the relationship between the atom and the connecting edge in the covalent graph, using the vector of the connecting edge, and to obtain the updated atom vector.

[0092] The second determining module 408 is used to determine the second affinity based on the vectors of the connecting edges in the covalent graph and the updated atom vectors.

[0093] Optionally, the fourth acquisition module 406 mentioned above further includes:

[0094] The sixth acquisition module 409 is used to update the edge vectors in the non-covalent graph based on the relationship between atoms and connecting edges in the non-covalent graph and using the atomic vectors synchronized by the first sub-network, and to obtain the updated edge vectors.

[0095] The third determining module 410 is used to determine the second interaction information based on the relationship between atoms and connecting edges in the non-covalent graph, the atomic vector synchronized in the first sub-network, and the updated edge vector.

[0096] Optionally, the second acquisition module 404 described above further includes:

[0097] The first correction module 411 is used to correct the second sub-network based on the first difference between the first interaction information and the second interaction information.

[0098] The second correction module 412 is used to correct the first sub-network based on the first difference and the second difference between the first affinity and the second affinity.

[0099] In this embodiment, after acquiring the training dataset, the drug target affinity prediction model generation system first generates covalent and non-covalent maps of the drug molecule and protein atoms based on their binding conformations. These maps are then input into the initial prediction model to determine the prediction results between the drug molecule and protein atoms. Finally, based on the differences between the labeled data and the prediction results, the initial prediction model is corrected until a corrected prediction model is obtained. Thus, by training the prediction model using the covalent and non-covalent maps between the drug and the target, the accuracy and reliability of the prediction model are improved.

[0100] Figure 5 This is a schematic diagram of a device for determining drug target affinity provided in an embodiment of the present disclosure.

[0101] like Figure 5 As shown, the apparatus 500 for determining drug target affinity includes:

[0102] The seventh acquisition module 501 is used to acquire the binding conformation of candidate drug molecules and protein atoms;

[0103] The second generation module 502 is used to generate a covalent diagram of the candidate drug molecule and protein atoms based on the binding conformation.

[0104] The eighth acquisition module 503 is used to input the covalent diagram into a preset prediction model in order to obtain the affinity between the candidate drug molecule and the protein atom.

[0105] In this embodiment, the system for generating the drug target affinity prediction model acquires a training dataset and obtains the binding conformation and annotation data of drug molecules and protein atoms. First, based on the binding conformation of drug molecules and protein atoms, it generates covalent and non-covalent maps of the drug molecules and protein atoms. Then, it processes the covalent and non-covalent maps using an initial prediction model to obtain prediction results. Finally, based on a first difference and a second difference, it corrects the first and second sub-networks of the initial prediction model, respectively. Thus, by training the prediction model using the covalent and non-covalent maps between the drug and the target, the accuracy and reliability of the prediction model are improved.

[0106] To implement the above embodiments, this disclosure also proposes a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method for generating a drug target affinity prediction model as proposed in the foregoing embodiments of this disclosure.

[0107] To implement the above embodiments, this disclosure also proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a method for generating a drug target affinity prediction model as proposed in the foregoing embodiments of this disclosure.

[0108] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program that, when executed by a processor, implements the charging method proposed in the foregoing embodiments of this disclosure.

[0109] Figure 6 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Figure 6 The computer device 12 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0110] like Figure 6 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).

[0111] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0112] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0113] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 5 Not shown; usually referred to as a "hard drive"). Although Figure 6 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.

[0114] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this disclosure.

[0115] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with computer device 12, and / or with any device that enables computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0116] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the foregoing embodiments.

[0117] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0118] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0119] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.

[0120] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0121] It should be understood that various parts of this disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0122] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiments.

[0123] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0124] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A method for generating a drug target affinity prediction model, comprising: Obtain a training dataset, wherein each training dataset includes the binding conformation of a drug molecule and a protein atom and the annotation data between the drug molecule and the protein atom, the annotation data including a first affinity; Based on the binding conformation, a covalent graph and a non-covalent graph composed of the drug molecule and the protein atom are generated, wherein the process includes: using the drug molecule and the protein atom as atoms in the covalent graph and the non-covalent graph, using the covalent bond between the protein atom and the drug molecule as the edge of the covalent graph, and connecting every two atoms whose distance is within a distance threshold to form the edge of the non-covalent graph, thereby generating the covalent graph and the non-covalent graph. The covalent and non-covalent diagrams are input into the initial prediction model to determine the prediction results between the drug molecule and the protein atom, and the prediction results include the second affinity. Based on the difference between the labeled data and the prediction results, the initial prediction model is corrected until a corrected prediction model is obtained. The labeled data also includes first interaction information between the drug molecule and the protein atom. The step of inputting the covalent and non-covalent graphs into the initial prediction model to determine the prediction results between the drug molecule and the protein atom includes: The covalent graph is processed using the first sub-network in the initial prediction model to obtain the second affinity; The second sub-network in the initial prediction model is used to process the non-covalent graph and the atomic vector synchronized with the first sub-network to obtain the second interaction information between the drug molecule and the protein atom; The step of correcting the initial prediction model based on the difference between the labeled data and the prediction results includes: Based on the first difference between the first interaction information and the second interaction information, the second sub-network is corrected; The first sub-network is modified based on the first difference and the second difference between the first affinity and the second affinity.

2. The method as described in claim 1, wherein, The step of processing the covalent graph using the first sub-network in the initial prediction model to obtain the second affinity includes: Based on the relationship between atoms and connecting edges in the covalent graph, the vectors of the atoms are updated using the vectors of the connecting edges to obtain the updated atom vectors. The second affinity is determined based on the vectors of the connecting edges in the covalent graph and the updated atom vectors.

3. The method as described in claim 1, wherein, The step of processing the non-covalent graph and the atomic vectors synchronized with the first sub-network using the second sub-network in the initial prediction model to obtain the second interaction information between the drug molecule and the protein atom includes: Based on the relationship between atoms and connecting edges in the non-covalent graph, the edge vectors in the non-covalent graph are updated using the atom vectors synchronized by the first sub-network to obtain the updated edge vectors. The second interaction information is determined based on the relationship between atoms and connecting edges in the non-covalent graph, the atom vector synchronized in the first sub-network, and the updated edge vector.

4. The method as described in any one of claims 2-3, wherein, The interactive information includes at least one of the following: hydrogen bonds, hydrophobic interactions, salt bridges, metal ions, halogen bonds, molecular access system bonds (MACCS), pharmacophore molecular fingerprints, and energy distribution.

5. A method for determining drug target affinity, comprising: Obtain the binding conformation of candidate drug molecules to protein atoms; Based on the binding conformation, a covalent diagram of the candidate drug molecule and the protein atoms is generated; The covalent diagram is input into a preset prediction model to obtain the affinity between the candidate drug molecule and the protein atom, wherein the preset prediction model is generated based on the method described in any one of claims 1-4.

6. A device for generating a drug target affinity prediction model, comprising: The first acquisition module is used to acquire training datasets, wherein each training dataset includes the binding conformation of drug molecules and protein atoms and the annotation data between the drug molecules and the protein atoms, and the annotation data includes a first affinity. The first generation module is used to generate a covalent diagram and a non-covalent diagram composed of the drug molecule and the protein atom according to the binding conformation. The drug molecule and the protein atom are used as atoms in the covalent diagram and the non-covalent diagram, the covalent bond between the protein atom and the drug molecule is used as the edge of the covalent diagram, and according to the binding conformation of the drug molecule and the protein atom, every two atoms within a distance threshold are connected to form the edge of the non-covalent diagram to generate the covalent diagram and the non-covalent diagram. The first determining module is used to input the covalent diagram and the non-covalent diagram into the initial prediction model to determine the prediction result between the drug molecule and the protein atom, wherein the prediction result includes a second affinity. The second acquisition module is used to correct the initial prediction model based on the difference between the labeled data and the prediction result until the corrected prediction model is obtained. The labeled data also includes first interaction information between the drug molecule and the protein atom, wherein the first determining module further includes: The third acquisition module is used to process the covalent graph using the first sub-network in the initial prediction model to acquire the second affinity; The fourth acquisition module is used to process the non-covalent graph and the atomic vector synchronized with the first sub-network using the second sub-network in the initial prediction model to obtain the second interaction information between the drug molecule and the protein atom. The second acquisition module also includes: The first correction module is used to correct the second sub-network based on the first difference between the first interaction information and the second interaction information. The second correction module is used to correct the first sub-network based on the first difference and the second difference between the first affinity and the second affinity.

7. The apparatus of claim 6, wherein the third acquisition module further comprises: The fifth acquisition module is used to update the vector of the atom based on the relationship between the atom and the connecting edge in the covalent graph, using the vector of the connecting edge, and to obtain the updated atom vector; The second determining module is used to determine the second affinity based on the vectors of the connecting edges in the covalent graph and the updated atom vectors.

8. The apparatus of claim 6, wherein the fourth acquisition module further comprises: The sixth acquisition module is used to update the edge vectors in the non-covalent graph based on the relationship between atoms and connecting edges in the non-covalent graph and using the atom vectors synchronized by the first sub-network, so as to obtain the updated edge vectors. The third determining module is used to determine the second interaction information based on the relationship between atoms and connecting edges in the non-covalent graph, the atom vector synchronized by the first sub-network, and the updated edge vector.

9. The apparatus according to any one of claims 6-8, wherein, The interactive information includes at least one of the following: hydrogen bonds, hydrophobic interactions, salt bridges, metal ions, halogen bonds, molecular access system bonds (MACCS), pharmacophore molecular fingerprints, and energy distribution.

10. An apparatus for determining drug target affinity, comprising: The seventh acquisition module is used to acquire the binding conformation of candidate drug molecules with protein atoms; The second generation module is used to generate a covalent diagram of the candidate drug molecule and the protein atoms based on the binding conformation. The eighth acquisition module is used to input the covalent diagram into a preset prediction model to obtain the affinity between the candidate drug molecule and the protein atom, wherein the preset prediction model is generated based on the method described in any one of claims 1-4.

11. A computer device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method for generating a drug target affinity prediction model as described in any one of claims 1-4.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for generating a drug target affinity prediction model as described in any one of claims 1-4.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method for generating a drug target affinity prediction model as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Method and device for predicting affinity between protein and ligand molecule

    CN115148279A