Methods and apparatus for model training, drug screening, and affinity prediction
By constructing topological maps and using graph neural networks to process feature vectors, the problems of low accuracy and poor interpretability of predicting small and medium-sized compounds and proteins in the prior art are solved, and a more efficient and accurate drug screening process is achieved.
Patent Information
- Application Number
- CN202111039673.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-06
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-09-06
AI Technical Summary
In the prior art, when predicting the affinity between small molecule compounds and proteins, there are problems of low accuracy, poor interpretability and poor repeatability, resulting in high cost and low efficiency of drug screening.
By constructing a topological map, the features of small-molecule compounds and protein access regions are quantified, and a trained machine learning model, especially graph neural networks, are processed to predict affinity.
The efficiency, interpretability, repeatability, accuracy and accuracy of affinity prediction between small molecule compounds and proteins is improved, thereby reducing the cost of drug screening-related work and improving the efficiency of drug screening.
Smart Images

Figure CN114333986B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a method and apparatus for model training, drug screening, and affinity prediction. More specifically, the present application relates to a method and apparatus for predicting the affinity between a small molecule compound and a protein, a drug screening method and apparatus, and a method and apparatus for training a machine learning model. Background Art
[0002] As is well known, the research and development of new drugs is very long, complex, and depends on many factors. At the same time, the development of new drugs is also a very expensive process. It is estimated that for each new drug approved, pharmaceutical companies spend an average of $2.6 billion on research and development, mainly because most candidate drugs fail.
[0003] Machine Learning (ML) improves the discovery and decision-making of specified problems through rich and high-quality data. Machine learning is applied in all stages of drug discovery: target validation, identification of biomarkers, and analysis of digital pathology data in clinical trials. In currently commonly used computational methods, it is necessary to consider the three-dimensional (3D) structure of the protein-ligand complex, perform molecular docking, and then evaluate the binding activity through a scoring function. However, the docking poses generated by molecular docking and the scoring function for estimating the protein-drug ligand binding affinity are not accurate enough, resulting in a high false positive rate. In addition, the main challenges in current machine learning also lie in the lack of interpretability of the results and poor reproducibility. Summary of the Invention
[0004] Embodiments of the present application provide a method and apparatus for model training, drug screening, and affinity prediction, so as to improve the efficiency, interpretability, reproducibility, accuracy, and precision of predicting the affinity between a small molecule compound and a protein, thereby reducing the cost of drug screening-related work and improving the efficiency of drug screening.
[0005] In a first aspect, an embodiment of the present application provides a method for predicting the affinity between a small molecule compound and a protein, which includes: determining an access region based on the three-dimensional conformation of a complex formed by the small molecule compound to be analyzed and the protein; constructing a topological graph G based on the characteristics of atoms and chemical bonds within the access region; determining a feature vector based on the topological graph G; and processing the feature vector using a trained machine learning model to obtain the affinity between the compound and the protein.
[0006] In some embodiments, the three-dimensional conformation is obtained through the following steps:
[0007] Generating candidate conformations using access software based on the information of the small molecule compound and the protein; and
[0008] Using a conformational evaluation model, the three-dimensional conformation is determined based on the candidate conformation, and the conformational evaluation model is trained using proteins and small molecule compounds known to have interactions.
[0009] In some embodiments, the conformational evaluation model is obtained through the following steps:
[0010] Based on the information of the small molecule compound and the protein with co-crystal data, a plurality of first conformation samples are generated using the access software;
[0011] Based on the deviation of the first conformation samples from the co-crystal structure, the plurality of first conformation samples are classified into positive samples and negative samples;
[0012] Using the first conformation samples as a training set, a preliminary conformational evaluation model is trained;
[0013] Based on the information of the small molecule compound and the protein that do not have co-crystal data but have known activity data, a plurality of predicted conformation samples are generated using the access software;
[0014] Using the preliminary conformational evaluation model, the plurality of predicted conformation samples are evaluated to select a second conformation sample including positive samples and negative samples;
[0015] Using the first conformation samples and the second conformation samples, the preliminary conformational evaluation model is optimized to obtain the conformational evaluation model.
[0016] In some embodiments, the access region is determined based on the atoms of the small molecule compound and the pocket atoms on the protein, and the distance between the pocket atoms and the atoms of the small molecule compound is less than a predetermined distance threshold.
[0017] In some embodiments, the distance threshold is 1 to 100 angstroms, and optionally, the distance threshold is 1 to 10 angstroms.
[0018] In some embodiments, the feature vector includes atomic features, bond features, and corner and edge features of the topological graph.
[0019] The atomic features include at least one of the following features: atomic type, number of neighbors, number of free electrons, chiral type of the atom, valence of the atom, hybridization type of the atom, whether the atom has a predetermined property, whether the atom is included in a 3- to 8-membered ring, charge distribution of the atom, whether the atom belongs to a protein or a compound, amino acid type to which the atom belongs, distance between the atom and each neighbor, and number of hydrogen atoms connected to the atom, and
[0020] The key features include at least one of the following features: the number of bonds formed by an atom with other atoms, the type of bond, the distance between the bonding atoms, whether the two atoms connected by the bond are in the same ring, hydrogen bonds, π-π stacking, π-ion, hydrophobicity, salt bridges, and X-bonds.
[0021] In some embodiments, the machine learning model is provided with an attention readout layer for determining the contribution weights of each atom in the access region to the affinity.
[0022] In some embodiments, the machine learning model includes a graph neural network provided with at least one of the following: at least one convolutional layer, at least one feedforward neural network, at least one attention layer, at least one information bottleneck unit.
[0023] In some embodiments, the machine learning model sequentially includes: an attention layer, a zero-th graph convolutional neural network layer, a first graph convolutional neural network layer, a linear transformation layer, a second graph convolutional neural network layer, a third graph convolutional neural network layer, an attention readout layer, and a feedforward neural network. Among them, the attention layer performs a dimensionality increase transformation on its input matrix, the first graph convolutional neural network layer performs a dimensionality reduction transformation on its input matrix, the linear transformation layer does not change the dimension of its input matrix, and the second graph convolutional neural network layer performs a dimensionality increase transformation on its input matrix.
[0024] In a second aspect, an embodiment of the present application provides a drug screening method, which includes:
[0025] Based on the structural formula of the candidate compound and the amino acid sequence of the protein, determine the three-dimensional conformation of the candidate compound and the protein complex, where the protein is related to a predetermined disease;
[0026] According to the method described above, predict the affinity between the candidate compound and the protein. That the affinity is higher than a predetermined threshold is an indication that the candidate compound can treat the predetermined disease.
[0027] In some embodiments, the candidate compound is obtained by modifying a starting compound.
[0028] In some embodiments, according to the method described above, determine the affinity between the starting compound and the protein and determine the contribution weights of each atom in the starting compound to the affinity; and based on the contribution weights of each atom in the starting compound to the affinity, determine the candidate sites for modification.
[0029] In a third aspect, an embodiment of the present application provides a method for training a machine learning model, where the machine learning model is used to predict the affinity between a small molecule compound and a protein. The method includes: obtaining three-dimensional conformations of a plurality of complexes formed by a small molecule compound with a known affinity and a protein; determining an access region based on the three-dimensional conformations of the complexes; constructing a topological graph G based on the characteristics of atoms and chemical bonds within the access region; determining a feature vector based on the topological graph G; and using the known affinity as a label to train the machine learning model with the feature vector to obtain a trained machine learning model.
[0030] In a fourth aspect, an embodiment of the present application provides a device for predicting the affinity between a small molecule compound and a protein, which includes: an access region determination unit for determining an access region based on the three-dimensional conformation of a complex formed by the small molecule compound to be analyzed and the protein; a feature vector determination unit for constructing a topological graph G based on the characteristics of atoms and chemical bonds within the access region and determining a feature vector based on the topological graph G; and a prediction unit for processing the feature vector with a trained machine learning model to obtain the affinity between the compound and the protein.
[0031] In a fifth aspect, an embodiment of the present application provides a drug screening device, which includes: a three-dimensional conformation determination unit for determining the three-dimensional conformation of a complex of the candidate compound and the protein based on the structural formula of the candidate compound and the amino acid sequence of the protein, where the protein is related to a predetermined disease; and a prediction unit for predicting the affinity between the candidate compound and the protein according to the method described above, where an affinity higher than a predetermined threshold indicates that the candidate compound can treat the predetermined disease.
[0032] In a sixth aspect, an embodiment of the present application provides a device for training a machine learning model, where the machine learning model is used to predict the affinity between a small molecule compound and a protein. The device includes: an acquisition unit for obtaining three-dimensional conformations of a plurality of complexes formed by a small molecule compound with a known affinity and a protein; an access region determination unit for determining an access region based on the three-dimensional conformations of the complexes; a feature vector determination unit for constructing a topological graph G based on the characteristics of atoms and chemical bonds within the access region and determining a feature vector based on the topological graph G; and a training unit for using the known affinity as a label to train the machine learning model with the feature vector to obtain a trained machine learning model.
[0033] In a seventh aspect, an embodiment of the present application provides a computing device, which includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program to implement the method described in any one of the first to third aspects above.
[0034] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, the storage medium includes computer instructions, when the instructions are executed by a computer, the computer is enabled to implement the method described in any one of the first to third aspects.
[0035] The method and device for predicting the affinity between a small molecule compound and a protein, the drug screening method and device, and the method and device for training a machine learning model provided by the embodiments of the present application, through three-dimensional structure analysis of multiple small molecule compounds with known affinity and proteins, after molecular docking, perform binding graph analysis on the formed docking region, and after obtaining the corner features of the graph and the relevant features of each atom and the formed bonds, train the machine learning model to obtain the ability to predict the affinity between a small molecule compound and a protein. Since more relevant atomic and bond features of compounds and proteins are obtained after molecular docking, the training accuracy of the machine learning model can be further improved. Therefore, when using the accurately trained machine learning model to perform related prediction work on the affinity between a small molecule compound and a protein, etc., the efficiency, interpretability, repeatability, accuracy, and precision of predicting the affinity between a small molecule compound and a protein can be improved, thereby reducing the cost of drug screening-related work and improving the efficiency of drug screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0037] Figure 1 It is a schematic diagram of a system architecture related to an embodiment of the present application;
[0038] Figure 2 It is a schematic flowchart of a method for predicting the affinity between a small molecule compound and a protein related to an embodiment of the present application;
[0039] Figure 3 It is a schematic diagram of a three-dimensional conformation of a complex formed by a small molecule compound and a protein related to another embodiment of the present application;
[0040] Figure 4 It is a schematic diagram of extracting feature vectors from a topological graph related to another embodiment of the present application;
[0041] Figure 5A framework diagram for predicting affinity from a feature matrix provided by another embodiment of the present application;
[0042] Figure 6 A framework diagram for predicting affinity from a feature matrix provided by another embodiment of the present application;
[0043] Figure 7 A schematic diagram of a method for screening drugs provided by another embodiment of the present application;
[0044] Figure 8 A schematic structural diagram of a device for predicting the affinity between a small molecule compound and a protein according to an embodiment of the present application;
[0045] Figure 9 A drug screening device according to an embodiment of the present application is shown;
[0046] Figure 10 A device for training a machine learning model according to an embodiment of the present application is shown;
[0047] Figure 11 A block diagram of a computing device related to an embodiment of the present application;
[0048] Figure 12 The affinity prediction results of C25H26N8O3 (Schembl20951758) and tyrosine kinase are shown;
[0049] Figure 13 The affinity prediction results of 2-amino-4-methoxybenzoic acid and anthranilate phosphoribosyltransferase are shown;
[0050] Figure 14 The affinity prediction results of ADP and ribonuclease A are shown;
[0051] Figure 15 A schematic flow diagram of obtaining a three-dimensional conformation according to an embodiment of the present method is shown;
[0052] Figure 16 A schematic flow diagram of obtaining conformation evaluation according to an embodiment of the present invention is shown; and
[0053] Figure 17 A schematic structural diagram of the attention mechanism according to an embodiment of the present invention is shown. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0055] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean that B is determined only according to A, and B can also be determined according to A and / or other information.
[0056] In the description of the present application, unless otherwise specified, "a plurality of" means two or more than two.
[0057] In addition, for the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily mean different.
[0058] For the convenience of understanding the embodiments of the present application, the following is a brief introduction to the relevant concepts involved in the embodiments of the present application:
[0059] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0060] Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0061] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0062] A neural network (NN), in the fields of machine learning and cognitive science, is a mathematical model or computational model that mimics the structure and function of a biological neural network (the central nervous system of an animal, especially the brain) and is used to estimate or approximate a function. A neural network consists of a large number of artificial neurons connected for computation. In most cases, a neural network can change its internal structure based on external information and is an adaptive system. A neural network is usually optimized through a learning method based on mathematical statistics, so it is also a practical application of mathematical statistics methods. Through standard mathematical methods of statistics, we can obtain a large number of local structure spaces that can be expressed by functions. Like other machine learning methods, neural networks have been used to solve various problems, such as machine vision and speech recognition. These problems are very difficult to solve by traditional rule-based programming.
[0063] The attention mechanism in this article refers to a vector used to represent the importance weights of each feature. To predict or infer a target element (such as a node in a topological graph), an attention vector can be used to estimate the degree of association between the target element and other elements, and the sum obtained by multiplying the values of these elements by the attention vector for weighting is used as an approximation of the target element.
[0064] The term "small molecule compound" as used herein refers to a compound molecule with a molecular weight not exceeding 1000 daltons, such as not exceeding 900 daltons, not exceeding 800, not exceeding 700, not exceeding 600, or not exceeding 500 daltons, which includes organic small molecules and inorganic small molecules. Currently, most drugs are small molecule drugs, and the basic building blocks of biological macromolecules such as proteins and nucleic acids (such as amino acids, ribonucleotides, deoxyribonucleotides) are also small molecules. Generally speaking, small molecule drugs exert their functions by interacting with proteins inside cells, especially by inhibiting or activating target proteins of certain diseases to achieve therapeutic effects. Due to the relatively small molecular weight of small molecule compounds, they can diffuse into cells relatively quickly in the human body and reach the action targets.
[0065] The term "access region" as used herein refers to the position where a small molecule compound interacts with a protein, which includes the small molecule compound and the protein pocket. The protein pocket refers to the amino acid part demarcated by drawing a circle with the atoms of the small molecule as the center and a predetermined radius, such as 1 - 100 angstroms.
[0066] The term "affinity" as used herein is the force characterizing the interaction strength between two or more substances, which can be quantitatively represented by pIC50. That is, pIC50 is a numerical index of affinity, and the larger this value, the stronger the affinity.
[0067] How to find small molecule compounds that can be used to treat specific diseases from a vast number of compounds and how to modify existing compounds to further improve the interaction between the compounds and disease targets have always been the main tasks in new drug research and development by major pharmaceutical companies. Generally speaking, drug screening usually relies on the artificial experience of drug experts and is improved through continuous trial and error and verification. For example, a large number of new compound structures are designed for synthesis and biological activity testing, which is extremely labor-intensive, material-intensive, and financially intensive.
[0068] The greatest advantage of AI technology is that it can, through the process of self-learning in a short period of time, digest a large amount of learning data and achieve the goal of teaching itself without a teacher.
[0069] Based on this, embodiments of the present application utilize AI technology to predict the affinity by constructing the three-dimensional conformations of small molecule compounds and proteins, and further selecting the access regions from the three-dimensional conformations for prediction. For the access regions, since the atoms and bonds in this region can form the vertices (V) and edges (E) in the topological graph, therefore, the topological graph of the access region can be quantified by using the vector G=(V, E), and other properties of the relevant atoms (vertices) and bonds (edges) can be further extracted, such as atom types, chemical bond types and other features. Based on a large amount of known affinity data, the training of a machine learning model (also referred to as "affinity prediction model", "target prediction model" or "SBDD-Poses" in this article) can be completed. Specifically, since the force between the compound and the protein is usually non-covalent binding, therefore, the features outside the access region are mostly futile for the training of the prediction model and have little practical significance for improving the accuracy. Therefore, according to some embodiments, the solution of the embodiments of the present application can improve the training efficiency and training accuracy. The trained prediction model can also quickly and accurately predict the binding sites between small molecule compounds and proteins, and the prediction cost is low. That is, embodiments of the present application utilize AI technology to assist in predicting the affinity between compounds and proteins, thereby reducing the overhead of manpower and material resources, improving the efficiency of subsequent drug screening, and reducing the cost of drug screening.
[0070] The application scenarios of the present application include, but are not limited to, fields such as medical treatment, biology, scientific research, etc., for example, for drug production, drug research and development, etc., and the entire recognition process does not require human intervention, and the recognition cost is low.
[0071] In some embodiments, the system architecture of the embodiments of the present application is as Figure 1 shown.
[0072] Figure 1 FIG. is a schematic diagram of a system architecture involved in an embodiment of the present application, including a user device 101, a data acquisition device 102, a training device 103, an execution device 104, a database 105, and a content library 106.
[0073] Among them, the data acquisition device 102 is used to read the training data from the content library 106 and store the read training data in the database 105. The training data involved in the embodiments of the present application includes the amino acid sequence of the protein or its crystal structure, the structural formula of the compound, and / or the co-crystal structure of the compound-protein, the topological graph features of the access region, etc.
[0074] In some embodiments, the training device 103 trains the prediction model based on the training data maintained in the database 105, so that the trained target prediction model can accurately predict the affinity between small molecule compounds and proteins. The target prediction model obtained by the training device 103 can be applied to different systems or devices.
[0075] In the appendix Figure 1 The execution device 104 is configured with an I / O interface 107 to interact with external devices for data. For example, it receives relevant information of proteins and small molecule compounds to be predicted sent by the user device 101 through the I / O interface, such as complex information, access area topology map information, etc. The computing module 109 in the execution device 104 processes the input information using a trained machine learning model, outputs the affinity between the small molecule compound and the protein, and sends the corresponding result to the user device 101 through the I / O interface.
[0076] Among them, the user device 101 may include a mobile phone, a tablet computer, a laptop computer, a personal digital assistant, a mobile internet device (MID), or other terminal devices with a browser installation function.
[0077] The execution device 104 may be a server.
[0078] Exemplarily, the server may be a computing device such as a rack server, a blade server, a tower server, or a cabinet server. The server may be an independent test server or a test server cluster composed of multiple test servers.
[0079] In this embodiment, the execution device 104 is connected to the user device 101 through a network. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, or a call network.
[0080] It should be noted that Figure 1 This is only a schematic diagram of a system architecture provided by the embodiments of the present application. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. In some embodiments, the above data acquisition device 102, the user device 101, the training device 103, and the execution device 104 may be the same device. The above database 105 may be distributed on one server or on multiple servers, and the above content library 106 may be distributed on one server or on multiple servers.
[0081] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These several embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0082] First, in combination with Figure 2 a method for predicting the affinity between a small molecule compound and a protein in the embodiments of the present application will be described in detail.
[0083] Figure 2 is a schematic flowchart of a method for predicting the affinity between a small molecule compound and a protein provided by an embodiment of the present application. As Figure 2 shown, the method includes:
[0084] S101: Determine an access region based on the three-dimensional conformation of the complex, where the complex is formed by the small molecule compound to be analyzed and the protein;
[0085] According to the embodiments of the present application, the three-dimensional conformation of the complex can be obtained through the co-crystal structure of the small molecule compound (sometimes directly referred to as "compound" in the present application) and the protein, that is, by co-crystallizing a solid-phase structure containing both the compound and the protein through a co-crystallization reaction of the small molecule compound and the protein in solution. After obtaining the co-crystal product, the three-dimensional conformation of the complex can be obtained by performing three-dimensional structure analysis on the co-crystal product, such as X-ray crystal diffraction analysis, electron microscopy three-dimensional reconstruction technology, and nuclear magnetic resonance technology. In addition, crystal data of related proteins or co-crystal products can also be obtained through publicly available databases, such as The Cambridge structural Database (CSD), The Protein Data Bank (PDB), The Inorganic Crystal Structure Database (ICSD), and the powder crystal database of the International Center for Diffraction Data (JCPDS-ICDD).
[0086] In addition, after determining the relevant information of the protein and the compound (such as amino acid sequence, structural formula, or partial crystal data), three-dimensional structure reconstruction can be performed through a variety of software, such as Figure 3As shown, by using molecular docking software (Molecular Docking), based on the amino acid sequence of HSA (human serum albumin) and the structural formula of the compound hippuric acid, the three-dimensional conformational structure of the HSA-hippuric acid complex can be obtained using molecular docking software (refer to doi:10.1371 / journal.pone.0071422.g007). Further, the corresponding access region can be selected in this three-dimensional conformational structure. Those skilled in the art can use a variety of known molecular docking software to obtain the three-dimensional structure of the complex, such as including but not limited to AutoDock, AutoDock Vina, LeDock, rDock, UCSF DOCK, LigandFit, GLIDE, GOLD, MOE Dock, and Surflex-Dock, etc. Generally, using molecular access software can generate a variety of access structures, and one or several optimal structures can be selected for subsequent analysis according to the performance such as the affinity preliminarily predicted by the corresponding software.
[0087] In addition, referring to Figure 15 and Figure 16 , in some embodiments, the three-dimensional conformation adopted can also be obtained through the following steps:
[0088] S510: Based on the information of the small molecule compound and the protein, use access software to generate candidate conformations; and
[0089] S520: Use a conformation evaluation model to determine the three-dimensional conformation based on the candidate conformations, and the conformation evaluation model is trained using known proteins and small molecule compounds that interact with each other.
[0090] Since the conformation evaluation model is trained using known proteins and small molecule compounds that interact with each other, this conformation evaluation model can effectively predict whether the conformation is close to the true three-dimensional structure of the protein and the small molecule compound. According to the embodiments of the present application, the information of the small molecule compound and the protein that can be used here includes any known information related to function, such as including but not limited to structural formula, amino acid sequence, atomic composition, protein three-dimensional structure, chiral molecule information, etc. In addition, the "known proteins and small molecule compounds that interact with each other" that can be used for training here refers to any proteins and small molecules that may have the possibility of forming a complex, such as those known to interact or bind with each other through chemical or biological experiments (such as proven by yeast two-hybrid, immunoprecipitation experiments, etc.), or those known to be able to form a co-crystal.
[0091] According to an embodiment of the present application, for a pair of a protein and a small molecule compound, multiple candidate conformations can be generated by using the access software, so that multiple three-dimensional conformations can be obtained. Thus, the trained model has higher robustness and generalization ability.
[0092] In some embodiments, the conformation evaluation model is obtained through the following steps:
[0093] S610: Based on the information of the small molecule compound and the protein with co-crystal data, use the access software to generate multiple first conformation samples. Since a pair of small molecule compounds and proteins with co-crystal data means that these members interact in a biological environment, this data can be effectively used to train the conformation evaluation model.
[0094] S620: Classify the multiple first conformation samples into positive samples and negative samples based on the deviation between the first conformation samples and the co-crystal structure. According to an embodiment of the present application, RMSD (root mean square deviation) can be used to characterize the deviation between the first conformation sample and the co-crystal structure. If the deviation is within a preset threshold, it can be considered that the predicted conformation sample is close to the co-crystal structure, and thus it can be considered a positive sample; otherwise, it can be considered a negative sample. The threshold used here can be no more than 5 angstroms, such as 4 angstroms, 3 angstroms, or 2 angstroms. Thus, according to an embodiment of the present application, on the one hand, the number of training samples is increased. For example, from ten thousand co-crystal structures, hundreds of thousands of first conformation samples can be derived (for example, for each co-crystal structure, a certain number of positive samples and negative samples can be selected, such as 10 to 20 positive samples and 10 to 20 negative samples). On the other hand, the training efficiency and evaluation accuracy of the conformation evaluation model can be improved by jointly training with positive and negative samples.
[0095] S630: Use the first conformation samples as a training set to train a preliminary conformation evaluation model. According to an embodiment of the present application, by using the first conformation samples as a training set, positive and negative samples can be used as labels to train a machine learning model to obtain a model that can output a conformation evaluation result. According to an embodiment of the present application, the machine learning model used here can be a neural network, such as a graph neural network. This machine learning model can output both classification results and evaluation quantification results.
[0096] S640: Based on the information of the small molecule compound and the protein that do not have co-crystal data but have known activity data, use the access software to generate multiple predicted conformation samples.
[0097] S650: Using the preliminary conformation evaluation model, evaluate the multiple predicted conformation samples to select a second conformation sample including positive samples and negative samples.
[0098] S660: Using the first conformation sample and the second conformation sample, optimize the preliminary conformation evaluation model to obtain the conformation evaluation model.
[0099] By using the information of the small molecule compounds and the proteins that do not have co-crystal data but have known activity data, the scale of the training set can be further expanded. Since having activity data is similar to having co-crystal data, it can indicate that the small molecule compound and the protein will form a stable complex structure. Therefore, these pairs of small molecule compounds and proteins can be effectively used for conformation evaluation. In fact, there is a vast amount of activity data of compounds and proteins currently, which can effectively improve the efficiency of model training. According to an embodiment of the present application, multiple predicted conformation samples can be obtained by using access software. Next, through the preliminarily constructed preliminary conformation evaluation model, the predicted conformation samples are evaluated, and at least one positive sample and at least one negative sample can be obtained respectively. Since such compounds and proteins do not have co-crystal data and cannot distinguish positive and negative samples by means such as RMSD (root mean square deviation), one or several conformation samples with the highest ranking in the output result of the preliminary conformation evaluation model can be selected as positive samples, and one or several conformation samples with the lowest ranking can be selected as negative samples. Thus, the training set for optimizing the subsequent conformation evaluation model can be further expanded. According to some embodiments of the present application, for combinations of tens of thousands or even hundreds of thousands of small molecule compounds and proteins with activity data, one or several positive and negative samples can be obtained respectively.
[0100] Reference Figure 3 , after obtaining the three-dimensional conformation structure of the complex, the access region can be determined. According to some embodiments, the access region is determined based on the atoms of the small molecule compound and the pocket atoms on the protein, and the distance between the pocket atoms and the atoms of the small molecule compound is less than a predetermined distance threshold. That is, the atoms on the protein that are no more than the predetermined threshold away from the compound molecule can be selected as pocket atoms. The predetermined threshold can be about 1 to 100 angstroms, such as about 1 to 90 angstroms, about 1 to 80 angstroms, about 1 to 70 angstroms, about 1 to 60 angstroms, about 1 to 50 angstroms, about 1 to 40 angstroms, about 1 to 30 angstroms, about 1 to 20 angstroms or about 1 to 10 angstroms. It should be noted that the above ranges cover all the values involved within the ranges. Additionally, unless otherwise specified, the term "about" as used herein means a 10% fluctuation up and down. The protein atoms selected in this way constitute the protein pocket, and the protein pocket and the atoms of the compound molecule form the access region.
[0101] S102: Based on the characteristics of atoms and chemical bonds within the access area, construct a topological graph G;
[0102] After determining the access area, a topological graph G can be constructed by modeling the atoms and bonds within the access area. Compounds can be modeled through graphs, where each vertex represents an atom or a chemical group, and the edges represent chemical bonds.
[0103] In some embodiments, when selecting atoms within the access area to construct the topological graph G, hydrogen atoms may not be considered. Since in organic substances such as organic small molecules and proteins, hydrogen atoms are abundantly present, these hydrogen atoms will cause a large amount of background data on the topological graph, and usually these hydrogen atoms do not contribute much to the affinity between the compound and the protein. Thus, by removing hydrogen atoms, the waste of computing resources can be reduced, and the training efficiency, prediction efficiency, precision, accuracy, etc. of machine learning can be improved.
[0104] S103: Based on the topological graph G, determine the feature vector;
[0105] Reference Figure 4 , according to the embodiments of the present application, after obtaining the topological graph, the feature vector can be determined from the topological graph. According to the embodiments of the present application, the feature vector adopted here can include the corner and edge features of the topological graph, and can also include the features of the atoms involved in the topological graph, as well as the features of related bonds such as chemical bonds. The related features can be summarized as a multi-dimensional vector matrix, thereby realizing the quantitative characterization of the access area.
[0106] Regarding the corner and edge features of the topological graph, the adjacency matrix and degree matrix can be used for characterization. Among them, the degree matrix is a diagonal matrix, and the elements on the diagonal are the degrees of each vertex. The degree of a vertex represents the number of edges associated with that vertex. The adjacency matrix represents whether there is a relationship between vertices. For a given topological graph, those skilled in the art can determine the adjacency matrix and degree matrix features manually, or can calculate them through some publicly available software, such as RDKit (https: / / www.rdkit.org / ).
[0107] As Figure 4As shown, the circles represent nodes. The smaller nodes with only element names are nodes of small molecules, and the larger ones are atoms of amino acid residues. The generated topological graph will only retain atom pairs (small molecule atoms and pocket atoms) with interactions within a certain distance, and each node will have characteristic values of a fixed dimension. According to an embodiment of the present application, for the convenience of description, [N, M] is used herein to represent the data input to the machine learning model, which means that M features are respectively set for each of the N nodes (atoms) (that is, except for the node number, the parameter feature is M-dimensional). Thus, an N×M matrix is obtained. Those skilled in the art can understand that during the processing of the machine learning model, the numbers of N and M will change with operations such as dimension elevation and dimension reduction.
[0108] Regarding atomic features, the atomic features that can be obtained include at least one selected from the following: atomic type, number of neighbors, number of free electrons, chiral type of the atom, valence of the atom, hybridization type of the atom, whether the atom has a predetermined attribute, whether the atom is included in a 3- to 8-membered ring, charge distribution of the atom, whether the atom belongs to a protein or a compound, amino acid type to which the atom belongs, distance between the atom and each neighbor, and number of hydrogen atoms connected to the atom, etc. Regarding bond features, the bond features include at least one of the following features: number of bonds formed by the atom with other atoms, type of the bond, distance between the bonding atoms, whether the two atoms connected by the bond are in the same ring, hydrogen bond, π-π stacking, π-ion, hydrophobicity, salt bridge, and X-bond. In some embodiments of the present invention, for the above atomic features and bond features, one-hot encoding can be used for characterization.
[0109] S104: Process the feature vector using the trained machine learning model to obtain the affinity between the compound and the protein.
[0110] In the embodiment of the present application, the execution subject of this step is a device with a trained machine learning model, such as an affinity prediction device. The prediction device can be a computing device, or a part of the computing device, such as a processor in the computing device. Exemplarily, the above measurement device can be Figure 1 the computing module in. Wherein Figure 1 the computing module in can be understood as a computing device, or a processor in the computing device, etc.
[0111] Referring to Figure 4 , after determining the edge angle features, atomic features, and bond features, these features can be integrated to obtain a multi-dimensional feature matrix. Further, the feature matrix is input to the machine learning model for analysis, so as to obtain data representing the affinity, such as pIC50 value.
[0112] The prediction model in the embodiments of the present application is a graph neural network model. The specific type of the prediction model in the embodiments of the present application is not limited, as long as it is a deep neural network model that can predict the affinity between a compound and a protein.
[0113] Reference Figure 5 , in a possible implementation, the prediction model in the embodiments of the present application is a graph neural network (GNN). Optionally, before inputting into the GNN, the feature matrix is pre-processed by an attention mechanism, so as to improve the interpretability of the output result.
[0114] According to the embodiments of the present application, the GNN that can be adopted is not particularly limited, and may include but is not limited to at least one selected from graph convolutional neural network (GCN), recurrent neural network (GRN), and graph attention network (GAT).
[0115] According to the embodiments of the present application, the machine learning model that can be adopted may include a graph neural network provided with at least one of the following: at least one convolutional layer, at least one feedforward neural network, at least one attention layer, and at least one information bottleneck unit. Those skilled in the art can understand that each information processing layer itself can also nest multiple neural networks, such as a conventional feedforward neural (FFN).
[0116] Reference Figure 17 Taking... as an example to illustrate the attention mechanism, the input x at the bottom layer 1 , x 2 , x 3 …, x T x 1 can respectively represent the feature matrix of a certain node. First, they are preliminarily embedded through an embedding layer (optional) to obtain a 1 , a 2 , a 3 …, a T ; then, they are respectively multiplied by three matrices W Q , W K and W V to obtain q i , k i , v i , i ∈ (1, 2, 3…T). Figure 17 Shows how the output b 1 corresponding to the input x 1 is obtained. That is: using q 1 to calculate the vector dot product with k 1 , k 2 , k 3 …, k T to obtain α 1,1 , α 1,2 , α1,3 …, α 1,T ; Add α 1,1 , α 1,1 , α 1,3 …, α 1,T Input into the softmax layer to obtain attention weight values all between 0 and 1: Multiply the obtained in the previous step respectively with the corresponding v 1 , v 2 , v 3 …, v T and then sum them up, thus obtaining the output b 1 corresponding to the input x 1 . Similarly, the output b 2 corresponding to the input x 2 is also obtained according to a similar process, except that at this time, the vector dot product is calculated using the q 2 corresponding to b 1 respectively with k 1 , k 2 , k 3 …, k T . The same processing process is also applied to other input nodes, and they can share the parameters W q 、W K and W V of these networks. These matrices also need to be optimized and learned during the training process of the machine learning model.
[0117] In addition, according to some embodiments of the present application, referring to Figure 6 , the machine learning model that can be adopted can be provided with an attention readout layer for determining the contribution weights of each atom in the access area to the affinity. Thus, the final output result will show the contributions of each atom to the affinity value, and it also makes the prediction model interpretable, so that it can determine which atoms have the greatest impact on the affinity. Further, it can provide important reference data for subsequent modification of compounds to improve compound performance.
[0118] Readout refers to aggregating the features of all nodes (such as each atom) updated through each layer into a vector representation representing the entire graph. According to some embodiments of the present application, the attention readout layer can obtain the contribution weights of each atom to the final output result (such as affinity) by adopting the following operations:
[0119] First, sum the values of the same feature dimension of each node in the input matrix H: [N, M'] of the attention readout layer to obtain the matrix For example, for the input matrix H
[0120] The first column represents the atom numbers, and the other columns represent the updated feature values of the M' dimension of each atomic node after multiple layers of processing. After performing the summation process, the result matrix H_sum is
[0121] [1 8 19 17 19].
[0122] Next, transpose the obtained matrix H_sum to get the matrix That is
[0123] Next, perform a dot product of the input matrix H and the matrix H_sum^T to obtain the matrix [N 1,]. Further, process the matrix [N 1,] through a normalization exponential function, such as the softmax() function, to obtain the weight of each of the N nodes for the output result (within the range of 0 to 1).
[0124] Thus, the above operation can be expressed as softmax(HxH_sum^T).
[0125] Reference Figure 6 , the machine learning model may include a graph neural network that sets at least one of the following: at least one convolutional layer, at least one feedforward neural network, at least one attention layer, at least one information bottleneck unit. By adopting the information bottleneck unit, the robustness of the machine learning model can be further improved.
[0126] Specifically, the machine learning model sequentially includes: an attention layer, a zero-th graph convolutional neural network layer (GCN-0), a first graph convolutional neural network layer (GCN-1), a linear transformation layer, a second graph convolutional neural network layer (GCN-2), a third graph convolutional neural network layer (GCN-3), an attention readout layer, and a feedforward neural network.
[0127] To improve the robustness of the prediction model, in some embodiments, an information bottleneck structure is introduced. Thus, during the training process, the model can be made to select relatively key features for calculation. In other words, in the prediction model, first convert the high-dimensional Embedding into a low-dimensional Embedding, and then output it as a high-dimensional Embedding. Thereby, the robustness performance of the model can be significantly improved. For example, according to some embodiments, the attention layer performs a dimensionality increase transformation on its input matrix, the first graph convolutional neural network layer performs a dimensionality reduction transformation on its input matrix, the linear transformation layer does not change the dimension of its input matrix, and the second graph convolutional neural network layer performs a dimensionality increase transformation on its input matrix. Thus, GCN-0, GCN-1, and GCN-2 together constitute an information bottleneck structure, which can improve the robustness of the prediction model.
[0128] In addition, according to some embodiments of the present application, in the model, a residual connection processing method is also adopted. In other words, a non-linear transformation function is used to describe the input and output of a network, that is, the input is X and the output is F(x). F usually includes operations such as convolution and activation. An input can be added to the output of the function, that is, the linear superposition of F(x) and X is used as the actual output or the input of the next layer. The X for linear superposition here can be the input of this layer or the input of other layers. In addition, normalization processing can also be performed after summation, such as Batch Normalization. For example, the output result of the second graph convolutional neural network layer and the output result of the first graph convolutional neural network layer are added and batch-normalized and then used as the input matrix of the third graph convolutional neural network layer. The output result of the third graph convolutional neural network layer and the output result of the zero-th graph convolutional input layer are added and batch-normalized and then used as the input matrix of the attention readout layer. Thereby, the accuracy and precision of the prediction model can be further improved. During the training process, the model can be more easily backpropagated to the previous layers during training, improving the efficiency of model training.
[0129] According to some embodiments of the present application, after the above processing, affinity prediction is performed through a feed-forward neural network (FFN), and then the predicted affinity value is output. The finally output 1D value is the pIC50 (affinity magnitude). The model then compares with the Label with known affinity parameters, uses a Loss function such as MSE (Mean Square Error), and then updates the model parameters through backpropagation.
[0130] Regarding the above-mentioned multiple graph convolutional neural network (GCN) layers, it should be noted that those skilled in the art can further nest more neural networks in the corresponding layers. In each GCN, the following can be used as the propagation rule of the convolutional layer:
[0131]
[0132] Among them, represents the adjacency matrix A of the topological graph G plus the identity matrix IN representing self-connection,
[0133] represents the degree matrix of the topological graph G, that is,
[0134] H (l) represents the activation unit matrix of the l-th layer (including the 0-th layer, that is, the input layer),
[0135] W (l) represents the convolutional kernel parameter matrix of the l-th layer.
[0136] Thus, in some embodiments of the present application, the performance of the machine learning model in predicting the affinity between proteins and ligands is further improved based on the three-dimensional conformational data of the complex. Specifically, based on the atom-level graph neural network (GNN) architecture, an interactive key topology graph constructed based on three-dimensional data according to atom types, amino acid types, and their additional features is used to train a deep learning model. The method of Readout Attention is further adopted to explain the reasons for the model prediction and the interpretability of the interaction between small molecules and pockets.
[0137] In addition, based on the weight output of each node, the importance of atoms can be visualized in visualization software. Visualization is very useful for drug experts and can greatly facilitate subsequent drug screening and compound modification.
[0138] In addition, the modeling method based on information bottleneck can improve the robustness of the model when the conformational data of the molecular access region is input.
[0139] The prediction models of the prior art, for example, in 3D CNN, use the 3D Grid method as input features. The disadvantage is that there is a lot of redundant information (noise) in the blank areas (where there are no atoms). The method of the embodiment of the present invention can represent the atomic structure in the form of a graph and select the best access region, reducing a lot of computational complexity and enabling the model to be trained better.
[0140] By adopting the above-mentioned machine learning model, the efficiency, interpretability, repeatability, accuracy, and precision of predicting the affinity between compounds and proteins can be improved, thereby reducing the cost of drug screening-related work and improving the efficiency of drug screening. Many existing affinity prediction models are usually trained only based on co-crystal data, resulting in low prediction performance. In addition, the lack of interpretability is a common problem in current deep learning. Therefore, the existing methods do not provide any insightful explanations for how to predict the affinity between proteins and drug ligands, that is, which specific features lead to the results of model inference. This important defect greatly hinders the popularization and application of the model in the actual process. The technical solution of the present application effectively solves the two major defects of low precision and lack of interpretability in the prior art, predicts the interaction between proteins and drug ligands, and obtains better generalization and prediction accuracy.
[0141] Next, the technical effects of the embodiments of the present application will be further introduced in combination with specific experiments.
[0142] Comparison between Example 1 and Other Prediction Models
[0143] The inventors conducted experiments on the prediction model (SBDD-Poses) of the embodiments of this application and other known models on the Pdbbind dataset (the general eutectic dataset used to train the model accuracy is the PDBbind v2019 refined dataset, and the test set is the PDBbind v2016 core set. Among them, the 2016 Core set is the commonly used gold test set in the industry recently because this test set has been calibrated with high precision manually and the data targets are relatively scattered, so the industry and academia often use this test set to verify). The SBDD-Poses model obtained the optimal performance of 0.82, which is significantly higher than other models. The results are as follows:
[0144]
[0145] *N represents the number of data points, and T represents the number of targets
[0146] In addition, the inventors constructed 2 test sets, which are test sets composed of 3400 data points including 46 GPCR + Kinase + Protease targets, named the docking_test dataset. The comparison results with other models are as follows:
[0147]
[0148]
[0149] Thus, it can be seen that the prediction model of the embodiments of this application can outperform other models in predicting the affinity between proteins and ligands.
[0150] Example 2 Interpretability demonstration
[0151] The inventors predicted the affinities between the following compounds and proteins respectively according to the prediction model of the embodiments of this application, and showed the weights of each atom on the affinity, Figures 12 - 14 The corresponding visualization results are shown respectively. For easy understanding, the corresponding proteins and compounds are summarized as follows:
[0152] Figure 12 The prediction result of the affinity between C25H26N8O3 (Schembl20951758) and tyrosine kinase is shown;
[0153] Figure 13 The prediction result of the affinity between 2-amino-4-methoxybenzoic acid and anthranilate phosphoribosyltransferase is shown;
[0154] Figure 14 The prediction result of the affinity between ADP and ribonuclease A is shown.
[0155] Thus, it can be clearly seen from the figure the weights of each atom for the affinity, and subsequently, key modifications or protections can be carried out on the atoms with large weights.
[0156] The method for predicting the binding affinity between a protein and a compound based on the structural information of the protein and the compound has been described above. Next, the application of this method, namely the drug screening method, will be described. In another aspect of the present invention, the present invention proposes a drug screening method, referring to Figure 7 , according to some embodiments, the method includes:
[0157] S201: Based on the structural formula of the candidate compound and the amino acid sequence of the protein, determine the three-dimensional conformation of the complex of the candidate compound and the protein, where the protein is related to a predetermined disease.
[0158] The method for obtaining the three-dimensional structure based on the sequence or structure of the amino acids of the compound and the protein has been described in detail above, and will not be elaborated here.
[0159] It should be noted that generally, the occurrence of diseases is related to the abnormality of cell signal pathways. Therefore, various enzymes, cytokines, etc. related to signal pathways often become the key targets for drug screening. Such proteins are also called drug screening targets, that is, the binding sites of drugs in the body, including biological macromolecules such as gene sites, receptors, enzymes, ion channels, nucleic acids, etc. So far, the total number of therapeutic drug targets discovered is about 500, among which receptors, especially G-protein coupled receptors (GPCRs), account for the vast majority, and there are also targets for enzymes, antibacterial, antiviral, and antiparasitic drugs. Rational drug design can design drug molecules based on the potential drug action targets, such as enzymes, receptors, ion channels, nucleic acids, etc., or the chemical structure characteristics of their endogenous ligands and natural substrates revealed in life science research, in order to discover new drugs that selectively act on the targets.
[0160] In addition, it should be noted that the source of the candidate compound is not particularly limited, and it can be obtained by any method. According to the embodiments of the present application, the candidate compound is obtained by modifying a starting compound. Specifically, the starting compound can be modified through the following steps:
[0161] First, determine the affinity of the starting compound for the protein and determine the contribution weights of each atom in the starting compound to the affinity; and
[0162] Next, based on the contribution weights of each atom in the starting compound to the affinity, determine the candidate sites for modification.
[0163] Since the machine learning model described in the present invention can analyze the contribution weight of each atom to the affinity, it is possible to focus on selecting atoms with high contribution weights for modification, such as replacing the type of atoms, such as replacing carbon with oxygen or nitrogen, etc., or using biological electronic isosteres for replacement. No further details will be given here.
[0164] The starting compound mentioned here can be any compound that may interact with the protein, for example, a known drug, a lead compound, a prospect compound, etc. In particular, a known drug whose affinity needs to be improved.
[0165] S202: According to the method described in the first aspect above, predict the affinity of the candidate compound to the protein, wherein the affinity being higher than a predetermined threshold is an indication that the candidate compound can treat the predetermined disease.
[0166] The threshold used here can be obtained after parallel treatment with a control compound. Those skilled in the art can also detect the affinity of known compounds and proteins through biological experiments to use it as the reference threshold.
[0167] As mentioned above, in some embodiments of the present application, by adopting the above-mentioned machine learning model, the efficiency, explainability, repeatability, accuracy and precision of predicting the affinity of compounds with proteins can be improved, so as to reduce the cost of drug screening related work and improve the efficiency of drug screening. Specifically, the method of the embodiment of the present application refers to finding out the small molecules that are most likely to become drugs (as lead compounds or lead compounds) from a large number of small molecule drug libraries through model prediction sorting. This process is a very important link in the pharmaceutical industry. Usually, general pharmaceutical companies will screen out small molecules that may be tens of thousands, and then conduct wet experiments to verify whether they are active. This process is very resource-saving. The method of the embodiment of the present invention can effectively improve the proportion of active molecules screened out, which will save pharmaceutical companies hundreds of millions of dollars in research and development funds.
[0168] The above describes a method for predicting the binding affinity between a protein and a compound based on the structural information of the protein and the compound. The following describes the method for training a machine learning model. In a third aspect of the present invention, the present invention proposes a method for training a machine learning model, wherein the machine learning model is used to predict the affinity between a small molecule compound and a protein, and the method comprises:
[0169] Obtain the three-dimensional conformations of multiple complexes formed by small molecule compounds and proteins with known affinities; determine the access regions based on the three-dimensional conformations of the complexes; construct a topological graph G based on the characteristics of the atoms and chemical bonds within the access regions; determine feature vectors based on the topological graph G; use the known affinities as labels to train the machine learning model with the feature vectors to obtain a trained machine learning model.
[0170] In the first aspect mentioned above, the construction of three-dimensional conformations, topological graph analysis, graph neural networks, etc. have been described in detail and will not be elaborated here.
[0171] It should be noted that the "known affinity" mentioned here can be the affinity reported in the literature through biological experiments or the affinity predicted based on existing software.
[0172] In some embodiments, during the training of the machine learning model, the known affinity can be used as a label, compare the value obtained by machine learning with the label, use MSE (Mean Square Error) as the Loss function, and update the parameters of the model through Back Propagation to obtain the finally trained machine learning model. According to some embodiments, since only the access regions are selected for feature analysis and some irrelevant atoms, such as hydrogen atoms, are removed, the efficiency of machine learning can be greatly improved. By using the above machine learning model, the efficiency, interpretability, repeatability, accuracy, and precision of predicting the affinity between compounds and proteins can be improved, thereby reducing the cost of drug screening-related work and improving the efficiency of drug screening.
[0173] According to some embodiments of the present application, the present application also proposes a processing method for input invariance. That is, a training method in which the output result remains unchanged by rotating or translating the input data can further improve the reliability of the model.
[0174] In addition, according to some embodiments of the present application, as described above, the present application also proposes a training method for a conformation evaluation model, which includes the following steps:
[0175] S610: Based on the information of the small molecule compound and the protein with co-crystal data, use the access software to generate multiple first conformation samples.
[0176] S620: Classify the multiple first conformation samples into positive samples and negative samples based on the deviation between the first conformation samples and the co-crystal structure.
[0177] S630: Use the first conformation samples as the training set to train a preliminary conformation evaluation model.
[0178] S640: Based on the information of the small molecule compound and the protein without eutectic data but with known activity data, use the access software to generate multiple predicted conformation samples.
[0179] S650: Use the preliminary conformation evaluation model to evaluate the multiple predicted conformation samples, so as to select a second conformation sample including positive samples and negative samples.
[0180] S660: Use the first conformation sample and the second conformation sample to optimize the preliminary conformation evaluation model, so as to obtain the conformation evaluation model.
[0181] The training of the conformation evaluation model has been described in detail above and will not be elaborated here.
[0182] Next, in combination with Figures 8 - 10 , the device embodiments of the present application will be described in detail.
[0183] Figure 8 Fig. shows a schematic structural diagram of a device for predicting the affinity between a small molecule compound and a protein according to an embodiment of the present application. The device can be a computing device or a component of a computing device (e.g., an integrated circuit, a chip, etc.), and is used to execute the method for predicting the affinity between a compound and a protein. The device includes:
[0184] An access area determination unit 210, configured to determine an access area based on the three-dimensional conformation of a complex formed by the small molecule compound to be analyzed and the protein;
[0185] A feature vector determination unit 220, configured to construct a topological graph G based on the characteristics of atoms and chemical bonds within the access area, and determine a feature vector based on the topological graph G;
[0186] A prediction unit 230, configured to process the feature vector using a trained machine learning model to obtain the affinity between the compound and the protein.
[0187] Figure 9 Fig. shows a drug screening device according to an embodiment of the present application. The device can be a computing device or a component of a computing device (e.g., an integrated circuit, a chip, etc.), and is used to execute the above drug screening method. The device includes:
[0188] A three-dimensional conformation determination unit 310, configured to determine the three-dimensional conformation of a complex of the candidate compound and the protein based on the structural formula of the candidate compound and the amino acid sequence of the protein, where the protein is related to a predetermined disease;
[0189] A prediction unit 320, for the method described in the first aspect, predicts the affinity between the candidate compound and the protein, wherein an affinity higher than a predetermined threshold is an indication that the candidate compound can treat the predetermined disease.
[0190] Figure 10 Disclosed is a device for training a machine learning model according to an embodiment of the present application. The machine learning model is used to predict the affinity between a small molecule compound and a protein. The device can be a computing device or a component of a computing device (e.g., an integrated circuit, a chip, etc.), and is used to execute the method for training the machine learning model described above. The device includes:
[0191] An acquisition unit 410, configured to acquire three-dimensional conformations of a plurality of complexes, where the complexes are formed by small molecule compounds with known affinities and proteins;
[0192] An access area determination unit 420, configured to determine an access area based on the three-dimensional conformations of the complexes;
[0193] A feature vector determination unit 430, configured to construct a topological graph G based on the features of atoms and chemical bonds within the access area, and determine a feature vector based on the topological graph G; and
[0194] A training unit 440, configured to use the known affinity as a label and train the machine learning model using the feature vector to obtain a trained machine learning model.
[0195] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, details are not described herein again.
[0196] In the foregoing, the device according to the embodiments of the present application has been described from the perspective of functional modules. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware of the processor and / or instructions in software. The steps of the method disclosed in combination with the embodiments of the present application can be directly implemented by the execution of the hardware processor, or implemented by the combination of the hardware and software modules in the processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0197] Figure 11 Is a block diagram of a computing device according to an embodiment of the present application. The device can beFigure 1 The server shown is used to execute the method described in the above embodiments. For specific details, refer to the description in the above method embodiments.
[0198] Figure 11 The computing device 200 shown includes a memory 201, a processor 202, and a communication interface 203. The memory 201, the processor 202, and the communication interface 203 are communicatively connected to each other. For example, the memory 201, the processor 202, and the communication interface 203 can be communicatively connected by means of a network connection. Alternatively, the above computing device 200 may further include a bus 204. The memory 201, the processor 202, and the communication interface 203 are communicatively connected to each other through the bus 204. Figure 14 The computing device 200 is such that the memory 201, the processor 202, and the communication interface 203 are communicatively connected to each other through the bus 204.
[0199] The memory 201 can be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 201 can store a program. When the program stored in the memory 201 is executed by the processor 202, the processor 202 and the communication interface 203 are used to execute the above method.
[0200] The processor 202 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits.
[0201] The processor 202 can also be an integrated circuit chip with the ability to process signals. In the implementation process, the method of this application can be completed by the integrated logic circuit in the hardware of the processor 202 or the instructions in the form of software. The above-mentioned processor 202 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 201, and the processor 202 reads the information in the memory 201 and combines its hardware to complete the method of the embodiment of this application.
[0202] The communication interface 203 uses a transceiver module such as, but not limited to, a transceiver to implement the communication between the computing device 200 and other devices or communication networks. For example, a data set can be obtained through the communication interface 203.
[0203] When the above-mentioned computing device 200 includes a bus 204, the bus 204 can include a path for transmitting information between various components of the computing device 200 (for example, the memory 201, the processor 202, the communication interface 203).
[0204] According to the present application, there is also provided a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method of the above method embodiment. Or rather, the embodiment of the present application also provides a computer program product including instructions. When the instructions are executed by the computer, the computer executes the method of the above method embodiment.
[0205] According to the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of the above method embodiment.
[0206] In other words, when implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0207] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0208] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.
[0209] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0210] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In addition, each method embodiment and each device embodiment can also refer to each other, and the same or corresponding content in different embodiments can be cited from each other without elaboration.
Claims
1. A method for predicting the affinity between a small molecule compound and a protein, characterized in that, comprising: determining an access region based on the three-dimensional conformation of a complex formed by the small molecule compound to be analyzed and the protein; constructing a topological graph G based on the characteristics of atoms and chemical bonds within the access region ; determining feature vectors based on the topological graph G; processing the feature vectors using a trained machine learning model to obtain the affinity between the compound and the protein; wherein the three-dimensional conformation is obtained through the following steps: generating candidate conformations using docking software based on the information of the small molecule compound and the protein; and determining the three-dimensional conformation based on the candidate conformations using a conformation evaluation model, and the conformation evaluation model is trained using proteins and small molecule compounds known to have interactions; the conformation evaluation model is obtained through the following steps: generating a plurality of first conformation samples using the docking software based on the information of the small molecule compound and the protein with co-crystal data; classifying the plurality of first conformation samples into positive samples and negative samples based on the deviation between the first conformation samples and the co-crystal structures corresponding to the first conformation samples; using the first conformation samples as a training set to train a preliminary conformation evaluation model; generating a plurality of predicted conformation samples using the docking software based on the information of the small molecule compound and the protein without co-crystal data but with known activity data; evaluating the plurality of predicted conformation samples using the preliminary conformation evaluation model to select a second conformation sample including positive samples and negative samples; optimizing the preliminary conformation evaluation model using the first conformation samples and the second conformation samples to obtain the conformation evaluation model.
2. The method according to claim 1, characterized in that, the access region is determined based on the atoms of the small molecule compound and the pocket atoms on the protein, and the distance between the pocket atoms and the atoms of the small molecule compound is less than a predetermined distance threshold.
3. The method according to claim 1, characterized in that, the feature vectors include atomic features, bond features, and corner and edge features of the topological graph, the atomic features include at least one of the following features: atom type, number of neighbors, number of free electrons, chiral type of the atom, valence of the atom, hybridization type of the atom, whether the atom has a predetermined property, whether the atom is included in a 3- to 8-membered ring, charge distribution of the atom, whether the atom belongs to the protein or the compound, amino acid type to which the atom belongs, distance between the atom and each neighbor, and number of hydrogen atoms connected to the atom, and the bond features include at least one of the following features: number of chemical bonds between the atom and other atoms, type of the bond, distance between the bonding atoms, whether the two atoms connected by the bond are in the same ring, hydrogen bond, π-π stacking, π-ion, hydrophobicity, salt bridge, and X-bond.
4. The method according to claim 1, characterized in that, The machine learning model is provided with an attention readout layer for determining the contribution weights of each atom in the access region to the affinity.
5. The method according to claim 1, wherein, the machine learning model includes a graph neural network provided with at least one of the following: at least one convolutional layer, at least one feedforward neural network, at least one attention layer, at least one information bottleneck unit.
6. The method according to any one of claims 1 to 5, wherein, the machine learning model sequentially includes: an attention layer, a zero-th graph convolutional neural network layer, a first graph convolutional neural network layer, a linear transformation layer, a second graph convolutional neural network layer, a third graph convolutional neural network layer, an attention readout layer, a feedforward neural network, wherein, the attention layer performs a dimensionality increase transformation on its input matrix, the first graph convolutional neural network layer performs a dimensionality reduction transformation on its input matrix, the linear transformation layer does not change the dimension of its input matrix, the second graph convolutional neural network layer performs a dimensionality increase transformation on its input matrix.
7. A drug screening method, wherein, it includes: Based on the structural formula of the candidate compound and the amino acid sequence of the protein, determining the three-dimensional conformation of the candidate compound and the protein complex, the protein being related to a predetermined disease; According to the method according to any one of claims 1 to 6, predicting the affinity between the candidate compound and the protein, the affinity being higher than a predetermined threshold being an indication that the candidate compound can treat the predetermined disease.
8. The drug screening method according to claim 7, wherein, the candidate compound is obtained by modifying a starting compound.
9. The drug screening method according to claim 8, wherein, it includes: Determining the affinity between the starting compound and the protein and determining the contribution weights of each atom in the starting compound to the affinity; and Based on the contribution weights of each atom in the starting compound to the affinity, determining the candidate sites for modification.
10. A method for training a machine learning model, wherein, the machine learning model is used to predict the affinity between a small molecule compound and a protein, and the method includes: Obtaining the three-dimensional conformations of multiple complexes formed by small molecule compounds with known affinities and proteins; Determining an access region based on the three-dimensional conformations of the complexes; Based on the atoms and the characteristics of chemical bonds within the access region, constructing a topological graph G; Based on the topological graph G, determining feature vectors; Using the known affinity as a label, training the machine learning model with the feature vectors to obtain a trained machine learning model; wherein, the three-dimensional conformation is obtained through the following steps: Based on the information of the small molecule compound and the protein, generating candidate conformations using access software; and Using a conformation evaluation model, determining the three-dimensional conformation based on the candidate conformations, the conformation evaluation model being trained using proteins and small molecule compounds known to have interactions; The conformation evaluation model is obtained through the following steps: Based on the information of the small molecule compound and the protein with eutectic data, use the access software to generate a plurality of first conformation samples; Based on the deviation between the first conformation sample and the eutectic structure corresponding to the first conformation sample, classify the plurality of first conformation samples into positive samples and negative samples; Use the first conformation sample as a training set to train a preliminary conformation evaluation model; Based on the information of the small molecule compound and the protein that do not have eutectic data but have known activity data, use the access software to generate a plurality of predicted conformation samples; Use the preliminary conformation evaluation model to evaluate the plurality of predicted conformation samples in order to select a second conformation sample including positive samples and negative samples; Use the first conformation sample and the second conformation sample to optimize the preliminary conformation evaluation model in order to obtain the conformation evaluation model.
11. A device for predicting the affinity between a small molecule compound and a protein, characterized in that, it includes: An access area determination unit for determining an access area based on the three-dimensional conformation of a complex formed by the small molecule compound to be analyzed and the protein; A feature vector determination unit for constructing a topological graph G based on the characteristics of atoms and chemical bonds within the access area, and determining a feature vector based on the topological graph G; A prediction unit for using a trained machine learning model to process the feature vector to obtain the affinity between the compound and the protein; wherein, the three-dimensional conformation is obtained through the following steps: Based on the information of the small molecule compound and the protein, use access software to generate candidate conformations; and Use a conformation evaluation model to determine the three-dimensional conformation based on the candidate conformations, and the conformation evaluation model is trained using proteins and small molecule compounds known to have interactions; The conformation evaluation model is obtained through the following steps: Based on the information of the small molecule compound and the protein with eutectic data, use the access software to generate a plurality of first conformation samples; Based on the deviation between the first conformation sample and the eutectic structure corresponding to the first conformation sample, classify the plurality of first conformation samples into positive samples and negative samples; Use the first conformation sample as a training set to train a preliminary conformation evaluation model; Based on the information of the small molecule compound and the protein that do not have eutectic data but have known activity data, use the access software to generate a plurality of predicted conformation samples; Use the preliminary conformation evaluation model to evaluate the plurality of predicted conformation samples in order to select a second conformation sample including positive samples and negative samples; Use the first conformation sample and the second conformation sample to optimize the preliminary conformation evaluation model in order to obtain the conformation evaluation model.
12. A drug screening device, characterized in that, it includes: A three-dimensional conformation determination unit for determining the three-dimensional conformation of a complex of the candidate compound and the protein based on the structural formula of the candidate compound and the amino acid sequence of the protein, and the protein is related to a predetermined disease; A prediction unit for predicting the affinity between the candidate compound and the protein according to the method described in any one of claims 1 to 6, wherein the fact that the affinity is higher than a predetermined threshold is an indication that the candidate compound can treat the predetermined disease.
13. An apparatus for training a machine learning model, characterized in that the machine learning model is used to predict the affinity between a small molecule compound and a protein, and the apparatus includes: an acquisition unit for acquiring the three-dimensional conformations of a plurality of complexes formed by small molecule compounds with known affinities and proteins; an access region determination unit for determining the access region based on the three-dimensional conformations of the complexes; a feature vector determination unit for constructing a topological graph G based on the characteristics of atoms and chemical bonds within the access region, and determining a feature vector based on the topological graph G; and a training unit for training the machine learning model using the known affinity as a label and the feature vector, so as to obtain a trained machine learning model; wherein the three-dimensional conformation is obtained through the following steps: generating candidate conformations using access software based on the information of the small molecule compound and the protein; and determining the three-dimensional conformation based on the candidate conformations using a conformation evaluation model, and the conformation evaluation model is trained using proteins and small molecule compounds known to have interactions; the conformation evaluation model is obtained through the following steps: generating a plurality of first conformation samples using the access software based on the information of the small molecule compound and the protein with co-crystal data; classifying the plurality of first conformation samples into positive samples and negative samples based on the deviation between the first conformation samples and the co-crystal structures corresponding to the first conformation samples; training a preliminary conformation evaluation model using the first conformation samples as a training set; generating a plurality of predicted conformation samples using the access software based on the information of the small molecule compound and the protein without co-crystal data but with known activity data; evaluating the plurality of predicted conformation samples using the preliminary conformation evaluation model to select a second conformation sample including positive samples and negative samples; optimizing the preliminary conformation evaluation model using the first conformation samples and the second conformation samples to obtain the conformation evaluation model.
14. A computing device, characterized in that it includes: a processor and a memory; the memory for storing a computer program; the processor for executing the computer program to implement the method described in any one of claims 1 to 10.
15. A computer-readable storage medium, characterized in that the storage medium includes computer instructions, and when the instructions are executed by a computer, the computer is caused to implement the method described in any one of claims 1 to 10.
Citation Information
Patent Citations
Human body behavior recognition method and system based on human body skeleton
CN111950485A
Method and apparatus for training predictive model for determining molecular binding force
CN113241126A
Computational method for classifying and predicting ligand docking conformations
US20180341754A1