Method, system, device and medium for predicting the affinity between biomacromolecules and ligands
By combining virtual screening and pilot compound optimization in the affinity prediction of biological macromolecules and ligands, fine-tuning using pre-trained models and known affinity ligands, the problem of task interdependence in existing methods is solved, achieving more accurate affinity prediction and higher experimental success rates.
Patent Information
- Application Number
- CN202510157970.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-13
AI Technical Summary
Existing machine learning methods ignore interdependence and complementarity between tasks in virtual screening and pilot compound optimization, affecting the accuracy of the affinity prediction of biological macromolecules and ligands.
A method for predicting affinity between biological macromolecules and ligands is provided. Candidate ligands are screened based on a pre-trained affinity prediction model, and the model is fine-tuned using known affinity ligands, combining virtual screening and pilot compound optimization to establish a global molecular pocket-ligand activity relationship.
It improves the accuracy of binding affinity prediction between biological macromolecules and ligands, improves the success rate of experiments, and reduces R&D costs.
Smart Images

Figure CN119649895B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of artificial intelligence and drug discovery, and particularly relates to a method, system, device and medium for predicting the affinity between a biomacromolecule and a ligand. Background Art
[0002] With significant breakthroughs in the field of protein structure prediction, models represented by AlphaFold and RoseTTAFold can accurately predict the three-dimensional structures of almost all proteins. Recently, AlphaFold3 has shown remarkable progress in predicting the structures of protein-ligand complexes, greatly accelerating the drug discovery process by providing new insights into the interactions between disease-related proteins and ligands. In the drug discovery pipeline, the next challenge after structure prediction is to predict protein-ligand binding affinity. In particular, protein-ligand binding affinity reflects the strength of the interaction between a potential ligand and a disease-related protein, and is a key factor determining drug efficacy and target selectivity.
[0003] In early drug discovery, protein-ligand binding affinity is mainly applied to two scenarios in early drug research and development: virtual screening and lead compound optimization. Virtual screening aims to discover active compounds that can bind to target human proteins from large-scale compound libraries; while lead compound optimization is dedicated to optimizing the structures of these active ligands to improve their binding affinity and drug-like properties and enhance their drug efficacy. In recent years, machine learning methods have made significant progress in both of these fields, greatly improving the computational efficiency while maintaining comparable performance to traditional computational chemistry methods.
[0004] However, existing machine learning methods generally study virtual screening and lead compound optimization separately, ignoring the interdependence and complementarity between the virtual screening and lead compound optimization tasks, thereby affecting the accuracy of predicting the affinity between biomacromolecules and ligands.
[0005] Therefore, the existing technology still needs to be improved. Summary of the Invention
[0006] The technical problem to be solved by the present application is to provide a method, system, device and medium for predicting the affinity between a biomacromolecule and a ligand in view of the deficiencies of the existing technology.
[0007] To solve the above technical problem, the first aspect of the present application provides a method for predicting the affinity between a biomacromolecule and a ligand, wherein the method for predicting the affinity between the biomacromolecule and the ligand specifically includes:
[0008] Based on a pre-trained affinity prediction model, screen a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library;
[0009] Based on the molecular pocket and a number of known affinity ligands corresponding to the molecular pocket, fine-tune the pre-trained affinity prediction model to obtain a fine-tuned affinity prediction model;
[0010] Predict the first affinity between the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model, and select a preset number of target ligands from the number of candidate ligands based on the first affinity.
[0011] The method for predicting the affinity between a biomacromolecule and a ligand, wherein screening a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library based on the pre-trained affinity prediction model specifically includes:
[0012] Obtain a first pocket embedding vector of the molecular pocket of the biomacromolecule and first ligand embedding vectors of each preset ligand in the preset ligand library based on the pre-trained affinity prediction model;
[0013] Predict the second affinity between the molecular pocket and each preset ligand based on the first pocket embedding vector and the first ligand embedding vectors, and select a number of candidate ligands from the preset ligand library based on the second affinity.
[0014] The method for predicting the affinity between a biomacromolecule and a ligand, wherein obtaining a first pocket embedding vector of the molecular pocket of the biomacromolecule based on the pre-trained affinity prediction model specifically includes:
[0015] Construct a molecular pocket-ligand knowledge graph for the molecular pocket of the biomacromolecule;
[0016] Convert the pocket nodes in the molecular pocket-ligand knowledge graph into second pocket embedding vectors and convert the ligand nodes in the molecular pocket-ligand knowledge graph into second ligand embedding vectors through the pre-trained affinity prediction model to obtain an embedded knowledge graph;
[0017] Input the embedded knowledge graph into a preset graph neural network, and output a first pocket embedding vector of the molecular pocket of the biomacromolecule through the graph neural network.
[0018] The method for predicting the affinity between a biomacromolecule and a ligand, wherein constructing a molecular pocket-ligand knowledge graph for the molecular pocket of the biomacromolecule specifically includes:
[0019] Obtain similar molecular pockets and known ligands of the molecular pocket corresponding to the biomacromolecule;
[0020] Use the molecular pocket and similar molecular pockets as pocket nodes and known ligands as ligand nodes;
[0021] Edges are constructed between the pocket nodes and ligand nodes based on the structural similarity between the pocket nodes and the affinity between the ligand nodes and the pocket nodes to obtain a molecular pocket-ligand knowledge graph.
[0022] The method for predicting the affinity between a biological macromolecule and a ligand, wherein the fine-tuning of the pre-trained affinity prediction model based on the molecular pocket and a number of known-affinity ligands corresponding to the molecular pocket to obtain a fine-tuned affinity prediction model specifically includes:
[0023] Encoding the molecular pocket into a third pocket embedding vector through the pre-trained affinity prediction model;
[0024] Encoding each known-affinity ligand into a third ligand embedding vector through the pre-trained affinity prediction model;
[0025] Predicting the third affinity between the molecular pocket and each known-affinity ligand based on the third pocket embedding vector and the third ligand embedding vector;
[0026] Constructing a first ranking loss term based on the third affinity between the molecular pocket and each known-affinity ligand and the first experimental affinity between the molecular pocket and each known-affinity ligand;
[0027] Fine-tuning the pre-trained affinity prediction model based on the first ranking loss term to obtain a fine-tuned affinity prediction model.
[0028] The method for predicting the affinity between a biological macromolecule and a ligand, wherein the prediction of the first affinity between the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model specifically includes:
[0029] Encoding the molecular pocket into a fourth pocket embedding vector through the fine-tuned affinity prediction model;
[0030] Encoding each candidate ligand into a fourth ligand embedding vector through the fine-tuned affinity prediction model;
[0031] Predicting the first affinity between the molecular pocket and each known-affinity ligand based on the fourth pocket embedding vector and the fourth ligand embedding vector.
[0032] The method for predicting the affinity between a biological macromolecule and a ligand, wherein the affinity prediction model includes a pocket encoder for encoding a molecular pocket into an embedding vector and a ligand encoder for encoding a ligand into an embedding vector.
[0033] The method for predicting the affinity between a biological macromolecule and a ligand, wherein the pre-training process of the pre-trained affinity prediction model specifically includes:
[0034] Obtain a training data set, where the training data set includes a number of experimental measurement groups, and each experimental measurement group includes a training molecular pocket, a number of training ligands, and a second experimental affinity between each training ligand and the training molecular pocket;
[0035] Encode the training molecular pocket in each experimental measurement group into a training pocket embedding vector through the affinity prediction model to be pre-trained, and encode each training ligand in each experimental measurement group into a training ligand embedding vector;
[0036] Based on the training pocket embedding vector and the training ligand embedding vector, determine the fourth affinity between the training molecular pocket in each experimental measurement group and each training ligand in each experimental measurement group;
[0037] Construct a loss function term based on the fourth affinity and the second experimental affinity, and train the affinity prediction model to be pre-trained based on the loss function term to obtain a pre-trained affinity prediction model.
[0038] The method for predicting the affinity between a biomacromolecule and a ligand, wherein the obtaining of the training data set specifically includes:
[0039] Collect known affinity data of biomacromolecules and ligands determined by experiments;
[0040] Divide the obtained known affinity data into a number of initial experimental measurement groups based on the same biomacromolecule and experimental measurement method, and match a molecular pocket for each initial experimental measurement group;
[0041] Use the molecular pocket matched for each initial experimental measurement group as the training molecular pocket, the ligand in each initial experimental measurement group as the training ligand, and the experimentally determined affinity in the initial experimental measurement group as the second experimental affinity to obtain the experimental measurement group corresponding to each initial experimental measurement group;
[0042] Use the data set composed of the experimental measurement groups as the training data set.
[0043] The method for predicting the affinity between a biomacromolecule and a ligand, wherein the matching of a molecular pocket for each initial experimental measurement group specifically includes:
[0044] Select a biomacromolecule-ligand complex structure for the biomacromolecule in each initial experimental measurement group in the known biomacromolecule structure database based on biomacromolecule similarity;
[0045] Calculate the ligand similarity between the crystal ligand bound to the biomacromolecule-ligand complex structure selected for each initial experimental measurement group and the ligand in each initial experimental measurement group respectively;
[0046] Match a target crystal ligand for each initial experimental measurement group based on ligand similarity, and use the molecular pocket corresponding to the target crystal ligand matched by each initial experimental measurement group as the molecular pocket corresponding to each initial experimental measurement group.
[0047] The method for predicting the affinity between a biomacromolecule and a ligand, wherein constructing the loss function term based on the fourth affinity and the second experimental affinity specifically includes:
[0048] Construct a second ranking loss term for each experimental measurement value group based on the fourth affinity and the second experimental affinity between the training molecular pocket in each experimental measurement group and the training ligand in each experimental measurement group;
[0049] Construct a contrast loss term based on the fourth affinity between the training molecular pocket in each experimental measurement group and the training ligands in each experimental measurement group;
[0050] Calculate the loss function term based on the second ranking loss term of each experimental measurement value group and the contrast loss term.
[0051] The second aspect of the present application provides a system for predicting the affinity between a biomacromolecule and a ligand, wherein the system for predicting the affinity between a biomacromolecule and a ligand specifically includes:
[0052] A screening module for screening a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library based on a pre-trained affinity prediction model;
[0053] A fine-tuning module for fine-tuning the pre-trained affinity prediction model based on the molecular pocket and a number of known-affinity ligands corresponding to the molecular pocket to obtain a fine-tuned affinity prediction model;
[0054] A lead compound optimization module for predicting the first affinity between the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model, and selecting a preset number of target ligands from a number of candidate ligands based on the first affinity.
[0055] The system for predicting the affinity between a biomacromolecule and a ligand, wherein the affinity prediction model includes a pocket encoder for encoding a molecular pocket into an embedding vector and a ligand encoder for encoding a ligand into an embedding vector.
[0056] The third aspect of the present application provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps in any one of the above methods for predicting the affinity between a biomacromolecule and a ligand, and / or,
[0057] The computer-readable storage medium stores the affinity prediction system of the biomacromolecule and the ligand as described above.
[0058] The fourth aspect of the present application provides a terminal device, which includes: a processor and a memory;
[0059] A computer-readable program executable by the processor is stored on the memory;
[0060] When the processor executes the computer-readable program, the steps in any of the above-described affinity prediction methods of the biomacromolecule and the ligand are implemented.
[0061] Beneficial effects: Compared with the prior art, the present application provides an affinity prediction method, system, device and medium for biomacromolecules and ligands. The method includes screening a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library based on a pre-trained affinity prediction model; fine-tuning the pre-trained affinity prediction model based on the molecular pocket and a number of known affinity ligands corresponding to the molecular pocket to obtain a fine-tuned affinity prediction model; predicting the first affinity between the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model, and selecting a preset number of target ligands from a number of candidate ligands based on the first affinity. The present application combines virtual screening and lead compound optimization through a pre-trained affinity prediction model, establishes a global molecular pocket-ligand activity relationship through virtual screening, and at the same time uses lead compound optimization data to deeply study the influence of subtle structural differences on affinity, so as to more accurately predict the binding affinity between biomacromolecules and ligands, improve the experimental success rate, and reduce the R & D cost. Description of the Drawings
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0063] Figure 1 It is a flowchart of the affinity prediction method of the biomacromolecule and the ligand provided by the embodiment of the present application.
[0064] Figure 2 It is a principle flowchart of the screening process of the candidate ligands.
[0065] Figure 3 It is a principle flowchart of the pre-training process of the affinity prediction model.
[0066] Figure 4It is a principle flowchart of the fine-tuning process of the affinity prediction model and the prediction process of predicting the affinity by fine-tuning the affinity prediction model.
[0067] Figure 5 It is a principle block diagram of the affinity prediction system for biomacromolecules and ligands provided by the embodiments of the present application.
[0068] Figure 6 It is a principle block diagram of the terminal device provided by the embodiments of the present application. Detailed implementation manners
[0069] The embodiments of the present application provide a method, a system, a device and a medium for predicting the affinity between a biomacromolecule and a ligand. To make the objectives, technical solutions and effects of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0070] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" to another element, it can be directly connected to other elements, or there may also be intermediate elements. In addition, the "connection" used herein may include wireless connection. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0071] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0072] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0073] The following further describes the content of the application by describing the embodiments with reference to the accompanying drawings.
[0074] This embodiment provides a method for predicting the affinity between a biomacromolecule and a ligand, as follows Figure 1 The method includes:
[0075] S10. Based on a pre-trained affinity prediction model, screen a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library.
[0076] Specifically, a biomacromolecule refers to a macromolecule such as protein, nucleic acid, polysaccharide, etc. present in the cells of an organism, and a ligand is a small molecule used to bind to a biomacromolecule, such as some small molecule components of traditional Chinese medicine, antibiotics, etc. A molecular pocket is a partial structure of a biomacromolecule, and the molecular pocket can bind to a ligand. That is to say, the molecular pocket is the binding position in the biomacromolecule that can bind to the ligand. Compared with the whole biomacromolecule, the molecular pocket can provide more detailed binding site information, so as to more accurately predict the affinity between the biomacromolecule and the ligand.
[0077] The affinity prediction model is a pre-trained deep learning network model. Through the pre-trained affinity prediction model, virtual screening can be carried out, that is, screen a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library. Among them, when performing virtual screening through the pre-trained affinity prediction model, the embedded vectors of the molecular pocket of the biomacromolecule and the preset ligands in the preset ligand library can be extracted respectively through the pre-trained affinity prediction model, and then the affinity can be predicted by aligning the two embedded vectors, and finally candidate ligands are screened based on the predicted affinity.
[0078] Exemplarily, the step of screening a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library based on the pre-trained affinity prediction model specifically includes:
[0079] S11. Based on the pre-trained affinity prediction model, obtain the first pocket embedded vector of the molecular pocket of the biomacromolecule and the first ligand embedded vectors of the respective preset ligands in the preset ligand library;
[0080] S12. Predict the second affinity between the molecular pocket and each preset ligand based on the first pocket embedded vector and the first ligand embedded vectors, and select a number of candidate ligands from the preset ligand library based on the second affinity.
[0081] Specifically, in step S11, the first pocket embedding vector and the first ligand embedding vector are vector representations in the same vector space. That is to say, the vector dimensions of the first pocket embedding vector and the first ligand embedding vector are the same. For example, the vector dimensions of both the first pocket embedding vector and the first ligand embedding vector are 128 dimensions, etc. The first pocket embedding vector can be obtained by directly mapping the three-dimensional structure and polar surface of the molecular pocket to the vector space through the affinity prediction model, or the embedding vector corresponding to the molecular pocket can be extracted first through the affinity prediction model, and then the prior knowledge that structurally similar molecular pockets may share similar ligands can be used to optimize the embedding vector corresponding to the molecular pocket extracted through the affinity prediction model.
[0082] In a specific implementation manner, obtaining the first pocket embedding vector of the molecular pocket of the biological macromolecule based on the pre-trained affinity prediction model specifically includes:
[0083] Constructing a molecular pocket-ligand knowledge graph for the molecular pocket of the biological macromolecule;
[0084] Converting the pocket nodes in the molecular pocket-ligand knowledge graph into second pocket embedding vectors and converting the ligand nodes in the molecular pocket-ligand knowledge graph into second ligand embedding vectors through the pre-trained affinity prediction model to obtain an embedded knowledge graph;
[0085] Inputting the embedded knowledge graph into a preset graph neural network, and outputting the first pocket embedding vector of the molecular pocket of the biological macromolecule through the graph neural network.
[0086] Specifically, the molecular pocket-ligand knowledge graph is used as the prior knowledge of the molecular pocket. It is constructed by the molecular pocket, similar molecular pockets with similar structures to the molecular pocket, and several known ligands. Through the molecular pocket-ligand knowledge graph, the relationship between the molecular pocket and the similar molecular pockets and known ligands can be obtained. That is to say, the affinity (strength of interaction) relationship between the same molecular pocket and the similar molecular pockets and known ligands in the molecular pocket-ligand knowledge graph is used to provide prior knowledge for the molecular pocket, and based on this prior knowledge, the pocket nodes extracted by the pocket encoder in the affinity prediction model are converted into second pocket embedding vectors for optimization, so as to achieve rapid and accurate screening of a large number of candidate ligands.
[0087] Exemplarily, constructing the molecular pocket-ligand knowledge graph for the molecular pocket of the biological macromolecule specifically includes:
[0088] Obtain similar molecular pockets and known ligands corresponding to the biological macromolecule, and use the molecular pocket and the similar molecular pockets as pocket nodes and the known ligands as ligand nodes;
[0089] Based on the structural similarity between the pocket nodes and the affinity between the ligand nodes and the pocket nodes, construct edges between the pocket nodes and the ligand nodes to obtain a molecular pocket-ligand knowledge graph.
[0090] Specifically, a similar molecular pocket is a molecular pocket with a structure similar to that of the molecular pocket, and a known ligand is a ligand known to be able to bind to the molecular pocket. For example, if the molecular pocket is the binding pocket of the human sodium ion channel protein, the similar molecular pocket can be the binding pocket of an ion channel protein with a similar structure, and the known ligand can be a known sodium ion channel blocker (such as lidocaine, mexiletine, etc.).
[0091] After obtaining the similar molecular pockets and known ligands corresponding to the molecular pocket of the biological macromolecule, use the molecular pocket and the similar molecular pockets as pocket nodes and the known ligands as ligand nodes, and then construct edges between the nodes to form a molecular pocket-ligand knowledge graph. Among them, the edges can include edges from pocket node to pocket node for reflecting the structural similarity between molecular pockets, and can also include edges from ligand node to pocket node for reflecting the affinity between the ligand node and the pocket node. Therefore, when constructing edges between nodes, the structural similarity between pocket nodes and the affinity between the ligand nodes and the pocket nodes can be used as the basis. Among them, the structural similarity between pocket nodes can be determined by calculating the similarity score between pocket nodes (such as sequence similarity), and the edge from ligand node to pocket node can be established based on the experimentally measured affinity data. Specifically, a structural similarity threshold between pocket nodes and an affinity threshold between the ligand nodes and the pocket nodes can be preset in advance, and then edges are established between pocket nodes with a structural similarity greater than the structural similarity threshold, and edges are established between ligand nodes and pocket nodes with an affinity greater than the affinity threshold between the ligand node and the pocket node.
[0092] Illustrate with an example: Suppose the pocket nodes include the human sodium ion channel and the rat sodium ion channel, and the ligand node includes lidocaine. The experimentally measured affinity between lidocaine and the sodium ion channel is greater than the affinity threshold, and the sequence similarity of the binding pockets between the human sodium ion channel and the rat sodium ion channel is higher than the structural similarity threshold. Then, edges are established between lidocaine and the human sodium ion channel, and between the human sodium ion channel and the rat sodium ion channel pairwise.
[0093] Furthermore, as Figure 2As shown, after constructing a molecular pocket-ligand knowledge graph for the molecular pocket of the biomacromolecule, in order to obtain the embedding vector of the molecular pocket, the pocket nodes and ligand nodes in the pocket-ligand knowledge graph can both be converted into embedding vectors. That is to say, the nodes in the pocket-ligand knowledge graph are converted into embedding vectors to obtain an embedded knowledge graph. Specifically, when converting the nodes in the pocket-ligand knowledge graph into embedding vectors, a pre-trained affinity prediction model can be used. Among them, the affinity prediction model includes a pocket encoder (such as Uni-Mol) for encoding the molecular pocket into an embedding vector and a ligand encoder (such as Uni-Mol, SmilesBERT) for encoding the ligand into an embedding vector. The pocket nodes are encoded by the pocket encoder to convert the pocket nodes into second pocket embedding vectors, and the ligand nodes are encoded by the ligand encoder to convert the ligand nodes into second ligand embedding vectors. In addition, after obtaining the embedded knowledge graph, the embedded knowledge graph can be input into a preset graph neural network. The preset graph neural network is a heterogeneous graph neural network (such as GraphRec). The information in the embedded knowledge graph is integrated through the graph neural network to output the first pocket embedding vector corresponding to the molecular pocket. In this way, the prior knowledge can be combined to optimize the embedding vector corresponding to the molecular pocket, so that the first pocket embedding vector corresponding to the molecular pocket carries more accurate knowledge information.
[0094] Further, in step S22, after obtaining the first pocket embedding vector of the molecular pocket and the first ligand embedding vectors of the preset ligands in the preset ligand library, the second affinity between the molecular pocket and each preset ligand can be predicted based on the first pocket embedding vector and the first ligand embedding vectors, and the second affinity between the molecular pocket and the preset ligand can be determined by means of vector alignment. For example, the vector alignment can adopt the cosine similarity method, that is, the cosine similarity between the first pocket embedding vector and each first ligand embedding vector can be calculated, and the calculated cosine similarity can be used as the second affinity between the molecular pocket and the preset ligand. Moreover, if the positions of the first ligand embedding vector and the first pocket embedding vector are closer, the cosine similarity between the first pocket embedding vector and the first ligand embedding vector is greater, which indicates that the non-covalent interaction between the molecular pocket and the preset ligand is also higher, and thus the second affinity between the molecular pocket and the preset ligand is higher. On the contrary, if the positions of the first ligand embedding vector and the first pocket embedding vector are farther, the cosine similarity between the first pocket embedding vector and the first ligand embedding vector is smaller, which indicates that the non-covalent interaction between the molecular pocket and the preset ligand is also lower, and thus the second affinity between the molecular pocket and the preset ligand is lower. In addition, after obtaining the second affinity between the molecular pocket and each preset ligand, several candidate ligands can be selected for the molecular pocket from the preset ligand library in the order of the second affinity from high to low. This can quickly search for high-affinity active ligands for the molecular pocket, and use N active ligands with high similarity as candidate ligands. Its search speed exceeds that of traditional methods (such as Vina Docking) by six orders of magnitude.
[0095] In the embodiment of the present application, a molecular pocket-ligand knowledge graph is constructed for the molecular pocket, and the pocket nodes and ligand nodes in the molecular pocket-ligand knowledge graph are both converted into embedding vectors in the same vector space through a pre-trained affinity prediction model, and then the information of the obtained embedded knowledge graph is learned through a graph neural network to obtain the first pocket embedding vector of the molecular pocket. In this way, the prior knowledge that structurally similar protein pockets may share similar ligands can be fully utilized, and the screening accuracy is improved. At the same time, after obtaining the first pocket embedding vector and the first ligand embedding vectors of the preset ligands in the preset ligand library, the affinity is directly predicted by vector alignment, and several active ligands are selected as candidate ligands for the molecular pocket from the preset ligand library in the order of the affinity from high to low. In this way, the screening efficiency of candidate ligands is greatly improved.
[0096] The above is the description of the virtual screening process in the embodiment of the present application. In the virtual screening process, a pre-trained affinity prediction model is used. Next, the pre-training process of the affinity prediction model is specifically described by using a training data set to train the pre-trained affinity prediction model. Among them, the pre-trained affinity prediction model has the same model structure as the affinity prediction model, as Figure 3As shown, the pre-trained affinity prediction model includes a pocket encoder (such as Uni-Mol, etc.) and a ligand encoder (such as Uni-Mol, SmilesBERT, etc.). The difference between the two lies in the model parameters. The initial model parameters are used for the affinity prediction model to be pre-trained, and the model parameters after training with the training dataset are used for the pre-trained affinity prediction model.
[0097] Exemplarily, the pre-training process of the pre-trained affinity prediction model specifically includes:
[0098] H10. Obtain a training dataset;
[0099] H20. Encode the training molecular pockets in each experimental measurement group into training pocket embedding vectors through the affinity prediction model to be pre-trained, and encode each training ligand in each experimental measurement group into training ligand embedding vectors;
[0100] H30. Based on the training pocket embedding vectors and the training ligand embedding vectors, determine the fourth affinity between the training molecular pockets in each experimental measurement group and each training ligand in each experimental measurement group;
[0101] H40. Construct a loss function term based on the fourth affinity and the second experimental affinity, and train the affinity prediction model to be pre-trained based on the loss function term to obtain a pre-trained affinity prediction model.
[0102] Specifically, in step H10, the training dataset is used to train the affinity prediction model to be pre-trained to obtain a pre-trained affinity prediction model. Among them, the training dataset includes several experimental measurement groups, each experimental measurement group includes a training molecular pocket, several training ligands, and the second experimental affinity between each training ligand and the training molecular pocket. The second experimental affinity is obtained by experimental means, and the experimental means corresponding to the training ligands in each experimental measurement group are the same. In addition, for any two experimental measurement groups among several experimental measurement groups, the training molecular pockets in the two experimental measurement groups can be the same or different. For example, several experimental measurement groups include experimental measurement group A, experimental measurement group B, and experimental measurement group C. The training molecular pockets in the three experimental groups can be the same or different, but the experimental means corresponding to each training ligand within each experimental measurement group are the same.
[0103] Based on this, in one implementation, the obtaining of the training dataset specifically includes:
[0104] Obtain the known affinity data of the experimentally measured biomacromolecules and ligands;
[0105] Taking a biological macromolecule, a ligand that measures the affinity between the biological macromolecule and the ligand using the same experimental measurement method, and the known affinity data obtained from the measurement as an initial experimental measurement group to obtain a number of initial experimental measurement groups, and matching the molecular pocket of the biological macromolecule for each of the initial experimental measurement groups;
[0106] Taking the molecular pocket matched for each of the initial experimental measurement groups as the training molecular pocket, the ligand in each of the initial experimental measurement groups as the training ligand, and the experimentally determined affinity in the initial experimental measurement group as the second experimental affinity, to obtain the experimental measurement group corresponding to each of the initial experimental measurement groups;
[0107] Taking the data set composed of the experimental measurement groups as the training data set.
[0108] Specifically, the known affinity data can be collected from relevant databases (such as ChEMBL, BindingDB), relevant literature databases, and wet experiments. The known affinity data includes the biological macromolecule, the ligand, the experimental measurement method for the affinity between the biological macromolecule and the ligand, and the experimental measurement affinity, and the experimental measurement affinity is obtained by experimentally measuring the affinity between the biological macromolecule and the ligand using this experimental measurement method.
[0109] After collecting the known affinity data, the known affinity data is divided based on the same biological macromolecule and experimental measurement method, so that the known affinity data obtained with the same biological macromolecule and using the same experimental measurement method is divided into one initial experimental measurement group. For example, among the collected known affinity data, there are known affinity data for the affinities of 10 different ligands to the human sodium ion channel protein measured using the same experimental measurement method. Then, the known affinity data corresponding to these 10 ligands is classified into one initial experimental measurement group, and this initial experimental measurement group is about the human sodium ion channel protein.
[0110] Furthermore, after dividing to obtain a number of initial experimental measurement groups, match the molecular pocket for each initial experimental measurement group. Among them, the molecular pocket is the binding position where the ligand can bind to the biological macromolecule. By matching the molecular pocket for the biological macromolecule, more refined binding point information can be provided. Specifically, when matching the molecular pocket for each initial experimental measurement group, the molecular pocket can be matched for each initial experimental measurement group in a preset biological macromolecule database, or the molecular pocket can be matched for each initial experimental measurement group through human-computer interaction, or the molecular pocket can be matched for the initial experimental measurement group through a large language model, etc.
[0111] In a specific implementation manner, the matching the molecular pocket for each of the initial experimental measurement groups specifically includes:
[0112] Select a biomolecule-ligand complex structure for each biomolecule in the initial experimental measurement group in a known biomolecule structure database based on the similarity of biomolecules;
[0113] Calculate the ligand similarity between the crystal ligand bound to the selected biomolecule-ligand complex structure for each initial experimental measurement group and the ligand in each initial experimental measurement group;
[0114] Match a target crystal ligand for each initial experimental measurement group based on the ligand similarity, and use the molecular pocket corresponding to the target crystal ligand matched by each initial experimental measurement group as the molecular pocket corresponding to each initial experimental measurement group.
[0115] Specifically, the known biomolecule structure database is an existing biomolecule structure database (such as PDB, etc.). The known biomolecule structure database includes several biomolecule-ligand complex structures. That is to say, the known biomolecule structures in the known biomolecule structure database include biomolecules and crystal ligands. Therefore, when matching molecular pockets for the initial experimental measurement group, several biomolecule-ligand complex structures can be selected in the known biomolecule structure database based on biomolecule similarity (such as sequence similarity exceeding 40%, etc.), and then a target crystal ligand can be selected for the biomolecule from the candidate biomolecule-ligand complex structures based on ligand similarity (such as ligand similarity being greater than a preset ligand similarity threshold, etc.). Finally, the molecular pocket corresponding to the target crystal ligand is used as the molecular pocket corresponding to the initial experimental measurement group. Among them, biomolecule similarity refers to the sequence similarity between the biomolecule in the initial experimental measurement group and the biomolecule in the known biomolecule structure, and ligand similarity refers to the similarity between the ligand in the initial experimental measurement group and the crystal ligand in the known biomolecule structure (such as Tanimoto similarity, etc.).
[0116] For example: Suppose the initial experimental measurement group includes the known affinity data of the human sodium ion channel protein and 10 different ligands, and the known biomolecule structure database is PDB. First, calculate the sequence similarity between the human sodium ion channel protein and the biomolecules in each PDB structure, and select the PDB structures with sequence similarity exceeding 40% as candidate PDB structures; then, calculate the ligand similarity (such as Tanimoto similarity) between the crystal ligands in the candidate PDB structures and the 10 ligands in the initial experimental measurement group respectively. When the ligand similarity with any one of the 10 ligands exceeds a certain threshold, the molecular pocket of this crystal ligand is used as the molecular pocket of this initial experimental measurement group. In addition, it should be noted that when multiple molecular pockets are matched, the molecular pocket of the crystal ligand with the highest ligand similarity is used as the molecular pocket of this initial experimental measurement group.
[0117] Further, when matching a molecular pocket for each initial experimental measurement group, the molecular pocket matched by each initial experimental measurement group is used as the training molecular pocket corresponding to each initial experimental measurement group, the ligand in each initial experimental measurement group is used as the training ligand corresponding to each initial experimental measurement group, and the experimentally determined affinity between the ligand and the biological macromolecule in each initial experimental measurement group is used as the second experimental affinity corresponding to each initial experimental measurement group; then, the set composed of the training molecular pocket, the training ligand, and the second experimental affinity corresponding to each initial experimental measurement group is used as the experimental measurement group corresponding to each initial experimental measurement group to obtain a number of experimental measurement groups; finally, the data set composed of all experimental measurement groups is used as the training data set.
[0118] In the embodiment of the present application, by constructing experimental measurement groups based on biological macromolecules and experimental measurement methods, the comparability of affinity data within the experimental measurement groups is ensured, and batch effects caused by different experimental methods are avoided (for example, the pH value affects the affinity result between the ligand and the human sodium ion channel protein, and it is unreasonable to directly compare the data under different pH values). At the same time, in the present application, by collecting known affinity data from relevant databases (such as ChEMBL, BindingDB), relevant literature databases, and wet experiments, a large-scale training data set of experimental measurement groups is constructed, covering more than 400,000 ligands and 40,000 protein pockets, ensuring the richness and diversity of the training data set.
[0119] Further, in step H20, after obtaining the training data set, the affinity prediction model to be pre-trained is trained based on the training data set. Specifically, each experimental measurement group in the training data set is input into the affinity prediction model to be pre-trained, and each training molecular pocket in the experimental measurement group is encoded into a training pocket embedding vector by the pocket encoder in the affinity prediction model to be pre-trained, and each training ligand in the experimental measurement group is encoded into a training ligand embedding vector by the ligand encoder in the affinity prediction model to be pre-trained.
[0120] In addition, it should be noted that when training the affinity prediction model to be pre-trained based on the training data set, the training data set can be divided into several training batches, each training batch includes a preset number of experimental measurement value groups, and then each training batch is used as the training data for a training iteration process, and the training process of each training batch is the same. Among them, the number of experimental measurement groups included in each training batch can be determined according to actual needs, and no specific limitation is made here.
[0121] Further, in step H30, after obtaining the training pocket embedding vector and the training ligand embedding vector, the fourth affinity between the training molecular pocket and the training ligand can be predicted by means of vector alignment. For example, the fourth affinity between the training molecular pocket and the training ligand can be predicted by calculating the cosine similarity between the training pocket embedding vector and the training ligand embedding vector, etc.
[0122] Further, in step H40, after obtaining the fourth affinity between the training molecular pocket and the training ligand, a loss function term can be constructed based on the fourth affinity and the second experimental affinity. The loss function term can include a permutation loss term between the training molecular pocket and the training ligand in the experimental measurement group, or can include a contrast loss term between experimental measurement groups, or can also include both the permutation loss term between the training molecular pocket and the training ligand in the experimental measurement group and the contrast loss term between experimental measurement groups, etc.
[0123] In a specific implementation manner, constructing the loss function term based on the fourth affinity and the second experimental affinity specifically includes:
[0124] Constructing a second ranking loss term for each experimental measurement value group based on the fourth affinity and the second experimental affinity between the training molecular pocket and the training ligand in each experimental measurement group;
[0125] Constructing a contrast loss term based on the fourth affinity between the training molecular pocket in each experimental measurement group and the training ligands in each experimental measurement group;
[0126] Calculating the loss function term based on the second ranking loss term for each experimental measurement value group and the contrast loss term.
[0127] Specifically, the second ranking loss term is used for list ranking learning, that is, ranking learning is performed on the training molecular pocket and the training ligand within each experimental measurement group, so that the predicted affinity ranking by the model is as consistent as possible with the experimentally measured affinity ranking. Among them, the determination process of the second ranking loss term can be to first use a ranking method to perform affinity ranking based on the fourth affinity between the training molecular pocket and the training ligand in the experimental measurement group to form a predicted affinity sequence, and perform ranking based on the second experimental affinity between the training molecular pocket and the training ligand in the experimental measurement group to form an experimental affinity sequence, and then calculate the second ranking loss term based on the predicted affinity sequence and the experimental affinity sequence. Among them, the existing ranking methods can be used for the ranking method, for example, the Plackett-Luce model, etc.
[0128] The contrastive loss term is used for contrastive learning across the entire training dataset, enabling the pre-trained affinity prediction model to quickly distinguish active ligands from inactive ligands. Among them, when determining the contrastive loss term, the training molecule pocket-training ligand pairs with the second experimental affinity are used as positive samples (such as the binding pocket of a certain sodium channel blocker and the human sodium channel protein), and the training molecule pocket-training ligand pairs without the second experimental affinity are used as negative samples. This can maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs, enabling the pre-trained affinity prediction model to distinguish active ligands from inactive ligands. Among them, the active ligands are the top N ligands sorted by affinity from large to small as active ligands.
[0129] Furthermore, after obtaining the second ranking loss term and the contrastive loss term, the second ranking loss term and the contrastive loss term can be weighted to determine the loss function term, enabling the loss function term to combine list ranking learning and contrastive learning. List ranking learning enables the pre-trained affinity prediction model to accurately predict the affinity ranking of ligands within the same experimental measurement group (such as the activity ranking of 10 sodium channel blockers), while contrastive learning enables the pre-trained affinity prediction model to distinguish active and inactive ligands (such as distinguishing sodium channel blockers from non-blockers). This makes the pre-trained affinity prediction model suitable not only for virtual screening scenarios but also for lead compound optimization scenarios, and can utilize the inherent correlation between virtual screening scenarios and lead compound optimization scenarios to improve the performance in virtual screening scenarios and lead compound optimization scenarios. Of course, the pre-trained affinity prediction model is also applicable to the affinity prediction scenario, that is, by encoding the embedding vectors of the molecule pocket and the ligand through the pre-trained affinity prediction model, and then directly predicting the affinity between the molecule pocket and the ligand through vector alignment.
[0130] S20. Fine-tune the pre-trained affinity prediction model based on the molecule pocket and several known affinity ligands corresponding to the molecule pocket to obtain a fine-tuned affinity prediction model.
[0131] Specifically, after applying the pre-trained affinity prediction model to the virtual screening scenario and obtaining several candidate ligands, several active ligands are obtained. Further, the pre-trained affinity prediction model and several active ligands can be used for lead compound optimization. Among them, to ensure the prediction performance of the pre-trained affinity prediction model in the new scenario, such as Figure 4As shown, the pre-trained affinity prediction model can be fine-tuned based on known experimental data first, and then the fine-tuned affinity prediction model can be used for lead compound optimization. When fine-tuning the pre-trained affinity prediction model, the pre-trained affinity prediction model is controlled to perform ranking learning. By learning the relative affinity ranking relationship between ligands, the fine-tuned affinity prediction model can have better prediction performance on the molecular pocket.
[0132] Exemplarily, the method for fine-tuning the pre-trained affinity prediction model based on the molecular pocket and a plurality of known affinity ligands corresponding to the molecular pocket to obtain a fine-tuned affinity prediction model specifically includes:
[0133] Encoding the molecular pocket into a third pocket embedding vector through a pocket encoder in the pre-trained affinity prediction model;
[0134] Encoding each known affinity ligand into a third ligand embedding vector through a ligand encoder in the pre-trained affinity prediction model;
[0135] Predicting the third affinity between the molecular pocket and each known affinity ligand based on the third pocket embedding vector and the third ligand embedding vector;
[0136] Constructing a first ranking loss term based on the third affinity between the molecular pocket and each known affinity ligand and the first experimental affinity between the molecular pocket and each known affinity ligand;
[0137] Fine-tuning the pre-trained affinity prediction model based on the first ranking loss term to obtain a fine-tuned affinity prediction model.
[0138] Specifically, a known affinity ligand is a ligand that has been experimentally tested with a molecular pocket. That is, experimental measurements of the affinity between a number of ligands and the molecular pocket are pre - performed to obtain the affinities between the ligands and the molecular pocket. These ligands are used as known affinity ligands. Among them, the ligands for lead compound optimization include a number of candidate ligands obtained from virtual screening. That is, in addition, the known affinity ligands are the candidate ligands among the candidate ligands screened in step S10. That is, after screening a number of candidate ligands, the candidate ligands can be divided into two groups: an experimental ligand group and an optimization ligand group. Then, the known affinity between the molecular pocket and each candidate ligand in the experimental ligand group is determined through experiments, and the candidate ligands in the experimental ligand group are used as known affinity ligands for fine - tuning the affinity prediction model. In the embodiment of the present application, affinity data of a small number of candidate ligands among the candidate ligands obtained through experiments are used to train a pre - trained affinity model to provide knowledge information on the relative affinity ranking relationship between the ligands of this structural class of compounds, so that the fine - tuned affinity prediction model can learn the relative affinity ranking relationship between the ligands of this structural class of compounds rather than directly predicting the absolute affinity value, thereby enabling the fine - tuned affinity prediction model to accurately predict the affinity of this structural analog.
[0139] In addition, the third pocket embedding vector is obtained by encoding through a pocket encoder, and the third ligand embedding vector is obtained by encoding through a ligand encoder. The third affinity is obtained by using a vector alignment method (such as calculating the cosine similarity between the third pocket embedding vector and the third ligand embedding vector). These processes are the same as the prediction processes of the above - mentioned embedding vectors and affinities, and will not be elaborated here.
[0140] After obtaining the third affinity between the molecular pocket and each known affinity ligand, a sorting method can be used to determine a predicted affinity sequence based on the third affinity, and an experimental affinity sequence based on the first experimental affinity. Finally, a first sorting loss term is determined based on the predicted affinity sequence and the experimental affinity sequence, and the affinity prediction model is fine - tuned based on the first sorting loss term so that the fine - tuned affinity prediction model can have better prediction performance on the molecular pocket. Among them, the process of determining the predicted affinity sequence and the experimental affinity sequence using the sampling sorting method is the same as the sorting method used in determining the second sorting loss term above. For specific details, reference can be made to the description of the second sorting loss term, and it will not be elaborated here.
[0141] For example: Suppose the binding affinity data of lidocaine and its 10 structural analogs have been experimentally measured. Then, during the fine-tuning process, the similarity of the embedding vectors of these 10 ligands with the sodium ion channel protein pocket will be calculated as the third affinity. Then, the third affinity will be compared with the first experimental affinity in the experimental data, and the first ranking loss term (such as the Plackett-Luce model) will be calculated. Finally, the parameters of the pocket encoder and ligand encoder in the affinity prediction model will be adjusted according to the first ranking loss term to fine-tune the pre-trained affinity prediction model to obtain a fine-tuned affinity prediction model.
[0142] S30. Predict the first affinity of the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model, and select a preset number of target ligands from several candidate ligands based on the first affinity.
[0143] Specifically, the first affinity is determined based on the fine-tuned affinity prediction model, and its determination process is the same as the above-mentioned affinity determination process, that is, first encode the molecular pocket into a fourth pocket embedding vector through the fine-tuned affinity prediction model; encode each candidate ligand into a fourth ligand embedding vector through the fine-tuned affinity prediction model; finally, predict the first affinity of the molecular pocket and each ligand with known affinity based on the fourth pocket embedding vector and the fourth ligand embedding vector, and select a preset number of target ligands from several candidate ligands based on the first affinity. Among them, the preset number of target ligands is the first preset number of candidate ligands in descending order of the first affinity.
[0144] In the embodiment of the present application, the pre-trained affinity prediction model is fine-tuned by using a small amount of experimental data, and then the fine-tuned affinity prediction model is used to select target ligands for the molecular pocket, so that the structure-affinity relationship information of known compounds can be fully utilized to improve the prediction accuracy of the affinity between the candidate ligand and the molecular pocket. Moreover, with the generation of new experimental data, the fine-tuned affinity prediction model can be fine-tuned again with the new experimental data, so that the affinity prediction model can be continuously updated and optimized. Thus, the affinity of newly designed compounds can be predicted through the fine-tuned affinity prediction model, so as to timely adjust and optimize the method. This can not only significantly shorten the lead compound optimization cycle, but also reduce the number of compounds that need to be synthesized and tested, improve the experimental success rate, and thus reduce the R & D cost.
[0145] In summary, this embodiment provides a method for predicting the affinity between a biomacromolecule and a ligand. The method includes screening a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library based on a pre-trained affinity prediction model; fine-tuning the pre-trained affinity prediction model based on the molecular pocket and a number of known-affinity ligands corresponding to the molecular pocket to obtain a fine-tuned affinity prediction model; predicting the first affinity between the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model, and selecting a preset number of target ligands from the number of candidate ligands based on the first affinity. This application uses a pre-trained affinity prediction model that combines permutation learning and contrastive learning to jointly study virtual screening and lead compound optimization. By virtual screening, a global molecular pocket-ligand activity relationship is established, and at the same time, lead compound optimization data is used to deeply study the influence of subtle structural differences on affinity, so as to more accurately predict the binding affinity between a biomacromolecule and a ligand, improve the experimental success rate, and reduce the R & D cost.
[0146] Based on the above method for predicting the affinity between a biomacromolecule and a ligand, this embodiment provides a system for predicting the affinity between a biomacromolecule and a ligand, as Figure 5 shown. The system for predicting the affinity between a biomacromolecule and a ligand specifically includes:
[0147] A screening module 100, configured to screen a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library based on a pre-trained affinity prediction model;
[0148] A fine-tuning module 200, configured to fine-tune the pre-trained affinity prediction model based on the molecular pocket and a number of known-affinity ligands corresponding to the molecular pocket to obtain a fine-tuned affinity prediction model;
[0149] A lead compound optimization module 300, configured to predict the first affinity between the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model, and select a preset number of target ligands from the number of candidate ligands based on the first affinity.
[0150] Based on the above method for predicting the affinity between a biomacromolecule and a ligand, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for predicting the affinity between a biomacromolecule and a ligand as described in the above embodiment.
[0151] Based on the above method for predicting the affinity between a biomacromolecule and a ligand, this application also provides a terminal device, as Figure 6As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communications interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communications interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communications interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiments.
[0152] In addition, when the logical instructions in the above-mentioned memory 22 can be implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.
[0153] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the methods in the above embodiments.
[0154] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc can also be a transient storage medium.
[0155] In addition, the specific processes of loading and executing multiple instructions by the instruction processor in the above storage medium and the terminal device have been described in detail in the above method, and will not be repeated here one by one.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for predicting the affinity between a biomacromolecule and a ligand, characterized in that: The method for predicting the affinity between a biomacromolecule and a ligand specifically includes: Based on the pre-trained affinity prediction model, a plurality of candidate ligands are screened for the molecular pocket of the biomacromolecule in a preset ligand library, wherein the loss function used in the pre-training process of the pre-trained affinity prediction model includes a permutation loss term and a contrast loss term; Based on the molecular pocket and several known affinity ligands corresponding to the molecular pocket, fine-tuning the pre-trained affinity prediction model to obtain a fine-tuned affinity prediction model, wherein the fine-tuned affinity prediction model is used for lead compound optimization; The first affinity between the molecular pocket and each candidate ligand is predicted based on the fine-tuned affinity prediction model, and a preset number of target ligands are selected from a number of candidate ligands based on the first affinity.
2. The method for predicting affinity between a biomacromolecule and a ligand according to claim 1, characterized in that: The method of screening a number of candidate ligands for the molecular pocket of the biomacromolecule in the preset ligand library based on the pre-trained affinity prediction model specifically includes: Acquire a first pocket embedding vector of a molecular pocket of a biomacromolecule and a first ligand embedding vector of each preset ligand in the preset ligand library based on a pre-trained affinity prediction model; The second affinity between the molecular pocket and each preset ligand is predicted based on the first pocket embedding vector and the first ligand embedding vector, and a number of candidate ligands are selected from the preset ligand library based on the second affinity.
3. The method for predicting affinity between a biomacromolecule and a ligand according to claim 2, characterized in that: The method of obtaining the first pocket embedding vector of the molecular pocket of the biomacromolecule based on the pre-trained affinity prediction model specifically includes: Constructing a molecular pocket-ligand knowledge graph for the molecular pocket of the biomacromolecule; The pocket nodes in the molecular pocket-ligand knowledge graph are converted into second pocket embedding vectors through the pre-trained affinity prediction model, and the ligand nodes in the molecular pocket-ligand knowledge graph are converted into second ligand embedding vectors to obtain an embedded knowledge graph; The embedded knowledge graph is input into a preset graph neural network, and the first pocket embedding vector of the molecular pocket of the biological macromolecule is output through the graph neural network.
4. The method for predicting affinity between a biomacromolecule and a ligand according to claim 3, characterized in that: Constructing a molecular pocket-ligand knowledge graph for the molecular pocket of the biomacromolecule specifically includes: Obtaining similar molecular pockets and known ligands of the molecular pocket corresponding to the biomacromolecule; The molecular pocket and similar molecular pockets are used as pocket nodes, and the known ligand is used as a ligand node; Based on the structural similarity between the pocket nodes and the affinity between the ligand node and the pocket node, edges are constructed between the pocket node and the ligand node to obtain a molecular pocket-ligand knowledge graph.
5. The method for predicting affinity between a biomacromolecule and a ligand according to claim 1, characterized in that: The step of fine-tuning the pre-trained affinity prediction model based on the molecular pocket and a plurality of known affinity ligands corresponding to the molecular pocket to obtain a fine-tuned affinity prediction model specifically includes: encoding the molecular pocket into a third pocket embedding vector by the pre-trained affinity prediction model; encoding each known affinity ligand into a third ligand embedding vector by the pre-trained affinity prediction model; Predicting the third affinity of the molecular pocket to each known affinity ligand based on the third pocket embedding vector and the third ligand embedding vector; constructing a first ranking loss term based on the third affinity of the molecular pocket to each known affinity ligand and the first experimental affinity of the molecular pocket to each known affinity ligand; The pre-trained affinity prediction model is fine-tuned based on the first ranking loss term to obtain a fine-tuned affinity prediction model.
6. The method for predicting affinity between a biomacromolecule and a ligand according to claim 1, characterized in that: The predicting the first affinity between the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model specifically includes: encoding the molecular pocket into a fourth pocket embedding vector by the fine-tuned affinity prediction model; encoding each candidate ligand as a fourth ligand embedding vector by the fine-tuned affinity prediction model; The first affinity of the molecular pocket to each ligand with known affinity is predicted based on the fourth pocket embedding vector and the fourth ligand embedding vector.
7. The method for predicting affinity between a biomacromolecule and a ligand according to claim 1, characterized in that: The affinity prediction model includes a pocket encoder for encoding a molecular pocket into an embedding vector and a ligand encoder for encoding a ligand into an embedding vector.
8. The method for predicting the affinity between a biomacromolecule and a ligand according to any one of claims 1 to 7, characterized in that: The pre-training process of the pre-trained affinity prediction model specifically includes: Acquire a training data set, wherein the training data set includes a plurality of experimental measurement groups, each of which includes a training molecular pocket, a plurality of training ligands, and a second experimental affinity of each training ligand to the training molecular pocket; encoding the training molecule pocket in each experimental measurement group into a training pocket embedding vector through the affinity prediction model to be pre-trained, and encoding each training ligand in each experimental measurement group into a training ligand embedding vector; Determining a fourth affinity between the training molecule pocket in each experimental measurement group and each training ligand in each experimental measurement group based on the training pocket embedding vector and the training ligand embedding vector; A loss function term is constructed based on the fourth affinity and the second experimental affinity, and the affinity prediction model to be pre-trained is trained based on the loss function term to obtain a pre-trained affinity prediction model.
9. The method for predicting affinity between a biomacromolecule and a ligand according to claim 8, characterized in that: The obtaining of the training data set specifically includes: Collect experimentally determined known affinity data between biomacromolecules and ligands; Dividing the acquired known affinity data into a number of initial experimental measurement groups based on the same biomacromolecule and experimental measurement method, and matching a molecular pocket for each of the initial experimental measurement groups; Using the molecular pocket matched by each of the initial experimental measurement groups as a training molecular pocket, the ligand in each of the initial experimental measurement groups as a training ligand, and the experimentally determined affinity in the initial experimental measurement group as a second experimental affinity, to obtain an experimental measurement group corresponding to each of the initial experimental measurement groups; The data set formed by the experimental measurement group is used as a training data set.
10. The method for predicting affinity between a biomacromolecule and a ligand according to claim 9, characterized in that: The matching of molecular pockets for each of the initial experimental measurement groups specifically includes: Selecting a biomacromolecule-ligand complex structure for each biomacromolecule in the initial experimental measurement group based on biomacromolecule similarity in a known biomacromolecule structure database; respectively calculating the ligand similarity between the crystal ligand bound to the biomacromolecule-ligand complex structure selected in each of the initial experimental measurement groups and the ligand in each of the initial experimental measurement groups; The target crystal ligands are matched for each initial experimental measurement group based on ligand similarity, and the molecular pocket corresponding to the target crystal ligand matched by each initial experimental measurement group is used as the molecular pocket corresponding to each initial experimental measurement group.
11. The method for predicting affinity between a biomacromolecule and a ligand according to claim 8, characterized in that: The step of constructing a loss function term based on the fourth affinity and the second experimental affinity specifically includes: constructing a second ranking loss term for each experimental measurement value group based on a fourth affinity of the training molecule pocket in each experimental measurement group to the training ligand in each experimental measurement group and a second experimental affinity; constructing a comparative loss term based on the fourth affinity between the training molecule pocket in each experimental measurement group and each training ligand in each experimental measurement group; The loss function term is calculated based on the second ranking loss term of each experimental measurement value group and the contrast loss term.
12. A system for predicting affinity between a biomacromolecule and a ligand, characterized in that: The affinity prediction system for biomacromolecules and ligands specifically includes: A screening module, for screening a number of candidate ligands for the molecular pocket of the biomacromolecule in a preset ligand library based on a pre-trained affinity prediction model, wherein the loss function used in the pre-training process of the pre-trained affinity prediction model includes a permutation loss term and a comparison loss term; A fine-tuning module, for fine-tuning the pre-trained affinity prediction model based on the molecular pocket and a number of known affinity ligands corresponding to the molecular pocket to obtain a fine-tuned affinity prediction model, wherein the fine-tuned affinity prediction model is used for lead compound optimization; The lead compound optimization module is used to predict the first affinity between the molecular pocket and each candidate ligand based on the fine-tuned affinity prediction model, and select a preset number of target ligands from a number of candidate ligands based on the first affinity.
13. The system for predicting affinity between biomacromolecules and ligands according to claim 12, characterized in that: The affinity prediction model includes a pocket encoder for encoding a molecular pocket into an embedding vector and a ligand encoder for encoding a ligand into an embedding vector.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for predicting the affinity between a biomacromolecule and a ligand as described in any one of claims 1-11.
15. A terminal device, characterized in that: include: Processor and memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the steps of the method for predicting the affinity between a biomacromolecule and a ligand as described in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Drug screening method and device, electronic equipment and storage medium
CN114121180A
Combination conformation prediction method and device, model training method and device and storage medium
CN118506855A