Method, equipment and medium for obtaining protein small molecule evaluation model
By using a protein small molecule evaluation model and distance distribution prediction and binding affinity prediction models, atomic characterization and physical energy functions are generated, solving the problems of accuracy and multi-objective prediction in the prediction of protein-small molecule binding affinity in existing technologies, and realizing accurate assessment of affinity before binding.
Patent Information
- Application Number
- CN202511271673.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing molecular docking and molecular structure prediction methods have several drawbacks when predicting the binding affinity between proteins and small molecules. These include the inability to accurately predict binding affinity, difficulty in handling unfamiliar chemical structures, excessive task specificity, inability to perform multi-target prediction simultaneously, and predictions that do not conform to physical laws.
A small protein molecule evaluation model is adopted, which combines a distance distribution prediction model and a binding affinity prediction model. By generating atomic characterization, distance distribution and physical energy function, a data-driven approach is used to capture interatomic interactions and predict binding affinity.
It enables precise assessment of the binding affinity of small proteins before they actually bind to each other, taking into account both physical laws and the accuracy of multi-objective prediction.
Smart Images

Figure CN121122384A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics technology, and in particular to a technique for obtaining small protein molecule evaluation models. Background Technology
[0002] In the field of structure-based drug and protein design, molecular docking and molecular structure prediction are widely used computational methods, both aiming to predict the binding mode of ligands to proteins and estimate the binding affinity of protein-ligand complexes. Force fields (scoring functions) are the core component of these computational methods, and their technical approaches can be broadly categorized into two main types. Force fields (scoring functions) are primarily responsible for rapidly evaluating and ranking the conformations generated during docking and simulation, and for predicting the binding conformation and affinity of molecules in protein pockets by applying scoring functions, thus guiding the design of molecules or proteins.
[0003] The first type is traditional docking tools that use physical force field functions and calculate physical interaction energy. They have the advantage of predicting structures that conform to physical and chemical laws. However, this scoring method relies on simplified physical equations and linear energy superposition with fixed energy parameters, making it difficult to accurately model weak interactions and predict binding affinity.
[0004] The second category is scoring functions for structure prediction based on machine learning. These functions use machine learning models to learn structures from structure-affinity data and predict binding free energies. While this can improve computational efficiency and the accuracy of structure prediction, the models rely too heavily on a limited number of structures and cannot handle chemical structures that have not been learned. They are also too task-specific, making it difficult to simultaneously perform multiple objectives such as affinity prediction and conformation screening. The predicted structures do not conform to physical laws, resulting in phenomena such as severe collisions between small molecules and proteins that do not conform to physical laws. Furthermore, the scoring functions have low interpretability and cannot correctly assess the rationality of the structure. In many real-world project scenarios, the predicted structure must strictly conform to and limit the distance between specific small molecules and protein atoms, and the scoring functions of existing methods cannot assess whether the predicted structure conforms to the specific distance. Summary of the Invention
[0005] One object of this application is to provide a method, apparatus, and medium for obtaining small molecule protein evaluation models.
[0006] According to one aspect of this application, a method for obtaining a small protein molecule evaluation model is provided, the method comprising:
[0007] The protein and small molecule information used as training data are input into the protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model.
[0008] The protein small molecule evaluation model generates corresponding atomic representations based on the protein information and the small molecule information. The atomic representations are then input into the distance distribution prediction model to obtain the first distance distribution of atom pairs with first short-range interactions output by the distance distribution prediction model. The atomic representations include the distances between atoms.
[0009] The first distance distribution is input into the binding affinity prediction model, and the binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function to obtain the affinity evaluation information output by the binding affinity prediction model based on the first energy term;
[0010] The protein small molecule evaluation model is fitted based on the distance label corresponding to the training data and the affinity label to obtain a trained protein small molecule evaluation model. The trained protein small molecule evaluation model is used to determine the corresponding protein small molecule evaluation information according to the target affinity evaluation information output by the affinity prediction model after inputting the target protein information and target small molecule information as the data to be predicted, and output the protein small molecule evaluation information.
[0011] According to another aspect of this application, a method for evaluating small protein molecules is provided, the method comprising:
[0012] The protein and small molecule information, which are the data to be predicted, are input into the protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model.
[0013] The protein small molecule evaluation model generates corresponding atomic representations based on the protein information and the small molecule information. The atomic representations are then input into the distance distribution prediction model to obtain the first distance distribution of atom pairs with first short-range interactions output by the distance distribution prediction model. The atomic representations include the distances between atoms.
[0014] The first distance distribution is input into the binding affinity prediction model, and the binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function to obtain the affinity evaluation information output by the binding affinity prediction model based on the first energy term;
[0015] The protein small molecule evaluation model determines the corresponding protein small molecule evaluation information based on the affinity evaluation information and outputs the protein small molecule evaluation information.
[0016] According to one aspect of this application, a computer device for small molecule protein evaluation is provided, comprising a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described above.
[0017] According to one aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0018] According to one aspect of this application, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of any of the methods described above.
[0019] According to one aspect of this application, a first apparatus for obtaining a small protein molecule evaluation model is provided, the first apparatus comprising:
[0020] A module is used to input protein and small molecule information, which are used as training data, into a protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model.
[0021] The first and second modules are used to generate corresponding atomic representations based on the protein information and the small molecule information through the protein small molecule evaluation model, input the atomic representations into the distance distribution prediction model, and obtain the first distance distribution of atomic pairs with first short-range interactions output by the distance distribution prediction model, wherein the atomic representations include the distances between atoms;
[0022] The first and third modules are used to input the first distance distribution into the binding affinity prediction model, generate a corresponding first energy term based on a preset physical energy function through the binding affinity prediction model, and obtain the affinity evaluation information output by the binding affinity prediction model based on the first energy term;
[0023] The first four modules are used to fit the protein small molecule evaluation model based on the distance label corresponding to the training data and the affinity label to obtain a trained protein small molecule evaluation model. The trained protein small molecule evaluation model is used to determine the corresponding protein small molecule evaluation information based on the target affinity evaluation information output by the affinity prediction model after inputting the target protein information and target small molecule information as the data to be predicted, and output the protein small molecule evaluation information.
[0024] According to another aspect of this application, a second device for small molecule protein evaluation is provided, the second device comprising:
[0025] The second module is used to input protein information and small molecule information, which are the data to be predicted, into the protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model.
[0026] The second module is used to generate corresponding atomic representations based on the protein information and the small molecule information through the protein small molecule evaluation model, input the atomic representations into the distance distribution prediction model, and obtain the first distance distribution of atomic pairs with first short-range interactions output by the distance distribution prediction model, wherein the atomic representations include the distances between atoms;
[0027] The second and third modules are used to input the first distance distribution into the binding affinity prediction model, generate a corresponding first energy term based on a preset physical energy function through the binding affinity prediction model, and obtain the affinity evaluation information output by the binding affinity prediction model based on the first energy term;
[0028] The second and fourth modules are used to determine the corresponding small protein molecule evaluation information based on the affinity evaluation information through the small protein molecule evaluation model, and output the small protein molecule evaluation information.
[0029] Compared with existing technologies, this application inputs protein and small molecule information, used as training data, into a protein-small molecule evaluation model. The protein-small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model. The protein-small molecule evaluation model generates corresponding atomic representations based on the protein and small molecule information. These atomic representations are then input into the distance distribution prediction model to obtain a first distance distribution of atom pairs exhibiting first short-range interactions, as output by the distance distribution prediction model. The atomic representations include the distances between atoms. This first distance distribution is then input into the binding affinity prediction model, which generates a corresponding first energy term based on a preset physical energy function. The binding affinity prediction model is then used to obtain the first energy term. The output affinity assessment information; by fitting the protein small molecule assessment model based on the distance label and binding affinity label corresponding to the training data, a trained protein small molecule assessment model is obtained. The trained protein small molecule assessment model is used to determine the corresponding protein small molecule assessment information according to the target affinity assessment information output by the affinity prediction model after inputting the target protein information and target small molecule information as the data to be predicted, and outputs the protein small molecule assessment information. Thus, by proposing a protein small molecule assessment model based on a combination of statistical physics and machine learning, the interaction between atoms is captured in a data-driven manner to predict the binding affinity of protein small molecules, realizing the assessment of the binding affinity of protein small molecules before the actual binding of proteins and small molecules. Attached Figure Description
[0030] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0031] Figure 1 This diagram illustrates a method for obtaining a small protein molecule evaluation model according to an embodiment of the present application.
[0032] Figure 2 This diagram illustrates a method for small molecule protein evaluation according to one embodiment of the present application.
[0033] Figure 3 This diagram illustrates a structural diagram of a first device for constructing a small protein molecule evaluation model according to an embodiment of this application;
[0034] Figure 4 This diagram illustrates a structural diagram of a second device for small molecule protein evaluation according to an embodiment of this application;
[0035] Figure 5 Exemplary systems that can be used to implement the various embodiments described in this application are shown.
[0036] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation
[0037] The present application will now be described in further detail with reference to the accompanying drawings.
[0038] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a central processing unit (CPU)), input / output interfaces, network interfaces, and memory.
[0039] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.
[0040] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0041] The devices referred to in this application include, but are not limited to, user equipment, network equipment, or devices composed of user equipment and network equipment integrated through a network. The user equipment includes, but is not limited to, any mobile electronic product capable of human-computer interaction (e.g., via a touchpad), such as smartphones and tablets. These mobile electronic products can use any operating system, such as Android or iOS. The network equipment includes an electronic device capable of automatically performing numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. The network equipment includes, but is not limited to, computers, network hosts, single network servers, multiple network server clusters, or clouds composed of multiple servers. Here, a cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network, etc. Preferably, the device can also be a program running on the user equipment, network device, or a device formed by integrating user equipment and network device, network device, touch terminal, or network device and touch terminal through a network.
[0042] Of course, those skilled in the art should understand that the above-described devices are merely examples, and other existing or future devices that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0043] In the description of this application, "multiple" means two or more, unless otherwise expressly and specifically defined.
[0044] Figure 1This document illustrates a flowchart of a method for obtaining a protein small molecule evaluation model according to an embodiment of this application. The method includes steps S11, S12, S13, and S14. In step S11, protein information and small molecule information, used as training data, are input into the protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model. In step S12, the protein small molecule evaluation model generates corresponding atomic representations based on the protein information and the small molecule information, and inputs these atomic representations into the distance distribution prediction model to obtain a first distance distribution of atom pairs exhibiting a first short-range interaction, as output by the distance distribution prediction model. The atomic representations include the distances between atoms. In step S13, the first distance distribution is input into the binding affinity prediction model, and... The binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function, and obtains the affinity assessment information output by the binding affinity prediction model based on the first energy term; in step S14, the protein small molecule assessment model is fitted based on the distance label and binding affinity label corresponding to the training data to obtain a trained protein small molecule assessment model, wherein the trained protein small molecule assessment model is used to determine the corresponding protein small molecule assessment information according to the target affinity assessment information output by the affinity prediction model after inputting the target protein information and target small molecule information as the data to be predicted, and outputs the protein small molecule assessment information.
[0045] In step S11, the protein information and small molecule information used as training data are input into the protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model.
[0046] In some embodiments, protein information includes, but is not limited to, the atom type, chemical bond type, and atomic coordinates of each atom in the protein. Small molecule information includes, but is not limited to, the atom type, chemical bond type, and atomic coordinates of each atom in the small molecule. This example embodiment does not impose any special limitations on this. A small molecule refers to a ligand that can bind to a protein to form a small protein molecule; that is, a small protein molecule is a protein-ligand complex formed after a protein and a small molecule bind. Compared to a protein, a small molecule is an organic or inorganic compound with a very small molecular weight (typically less than 1000 Daltons) and a relatively simple structure. In some embodiments, during the model training phase, the protein information and small molecule information used as training data are input into a small protein molecule evaluation model. This small protein molecule evaluation model includes two sub-models: a distance distribution prediction model and a binding affinity prediction model.
[0047] In step S12, the protein small molecule evaluation model generates corresponding atomic representations based on the protein information and the small molecule information, and inputs the atomic representations into the distance distribution prediction model to obtain the first distance distribution of atomic pairs with first short-range interactions output by the distance distribution prediction model, wherein the atomic representations include the distances between atoms.
[0048] In some embodiments, the protein small molecule evaluation model first uses GatedGAT (a graph neural network model that combines two technologies: Graph Attention Network (GAT) and Gating Mechanism) and Interaction Network (IN, a graph neural network architecture specifically designed for reasoning about the physical interactions between objects) to generate corresponding atomic representations based on protein and small molecule information. Atomic representation refers to the precise description of the chemical and structural characteristics of each atom in a protein or small molecule using mathematical vectors (a set of numbers). The purpose of this representation is to convert the complex and abstract chemical information of atoms (e.g., element type, chemical environment, bonding situation, etc.) into a numerical form that a computer can understand and process.
[0049] In some embodiments, the protein small molecule evaluation model inputs the distances between atoms in the protein and the small molecule from the atomic characterization into a distance distribution prediction model. The distance distribution prediction model uses a probability distribution to infer the probability distribution of the most likely distances between atom pairs with a first short-range interaction (i.e., the first distance distribution). Here, an atom pair refers to two atoms consisting of one atom in the protein and one atom in the small molecule. The first short-range interaction refers to a non-bonded interaction occurring between the two atoms of an atom pair at a very short distance (usually within 4 Å), depending on electron cloud overlap or precise spatial coordination. The first short-range interaction includes, but is not limited to, van der Waals interactions and hydrogen bonding interactions, etc., which are not specifically limited in this example embodiment. In some embodiments, the first distance distribution is mainly based on a mixture Gaussian density function, with parameters including mean and variance.
[0050] In step S13, the first distance distribution is input into the binding affinity prediction model, and the binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function to obtain the affinity assessment information output by the binding affinity prediction model based on the first energy term.
[0051] In some embodiments, the protein small molecule evaluation model inputs the first distance distribution output by the distance distribution prediction model into the affinity prediction model. The affinity prediction model converts the first distance distribution into actual physical energies (i.e., first energy terms) based on a preset physical energy function, that is, converting the distance between atomic pairs into first energy terms. Each type of short-range interaction corresponds to one first energy term; for example, van der Waals interactions correspond to one first energy term, and hydrogen bonding interactions correspond to one first energy term. Therefore, the affinity prediction model generates two first energy terms. In some embodiments, the affinity prediction model obtains the corresponding binding affinity based on each first energy term. For example, it adds the first energy terms together, divides the result by a preset coefficient, and uses the calculated value as the binding affinity. In some embodiments, binding affinity can be directly used as the corresponding affinity assessment information, or the corresponding affinity assessment information can be further determined based on binding affinity. For example, binding affinity can be converted into a score or a piece of text information. Here, affinity assessment information refers to the assessment information about the binding affinity between a protein and a small molecule, that is, the assessment information about the binding affinity between a protein and a small molecule obtained before the protein and the small molecule actually bind (affinity assessment information). The specific form of affinity assessment information can be a specific numerical value (e.g., a score) used to characterize the strength of the binding affinity, or it can be a piece of text information.
[0052] In step S14, the protein small molecule evaluation model is fitted based on the distance label and the affinity label corresponding to the training data to obtain a trained protein small molecule evaluation model. The trained protein small molecule evaluation model is used to determine the corresponding protein small molecule evaluation information according to the target affinity evaluation information output by the affinity prediction model after inputting the target protein information and target small molecule information as the data to be predicted, and output the protein small molecule evaluation information.
[0053] In some embodiments, the protein small molecule evaluation model is trained by fitting model parameters (including fitting the parameters of each model included in the protein small molecule evaluation model) based on the distance labels and binding affinity labels corresponding to the training data, thereby obtaining a trained protein small molecule evaluation model. Here, the distance label refers to the true distance between the first short-range interacting atom pairs in the bound protein small molecule, and the binding affinity label refers to the true binding affinity of the bound protein small molecule. In some embodiments, model parameter fitting can be performed by applying a mean squared error loss function to the distance labels, or by applying an MDN (Mixture Density Network) loss function to the binding affinity labels.
[0054] In some embodiments, during the model inference phase, the target protein information and target small molecule information, which are the data to be predicted, are input into a trained protein small molecule evaluation model. The binding affinity prediction model in the protein small molecule evaluation model will output the corresponding target affinity evaluation information. Then, the protein small molecule evaluation model will determine the corresponding protein small molecule evaluation information based on the target affinity evaluation information. The protein small molecule evaluation information refers to the evaluation information that predicts the binding of proteins and small molecules. The specific form of the protein small molecule evaluation information can be a specific numerical value (e.g., a score), which is used to characterize the degree of goodness or badness of the expected binding. Alternatively, it can be a piece of text information. For example, the target affinity evaluation information can be directly used as the corresponding protein small molecule evaluation information. Alternatively, the target affinity evaluation information can be converted into the corresponding protein small molecule evaluation information through text conversion, numerical conversion, preset formula conversion, etc. This example embodiment does not make any special limitations on this.
[0055] This application inputs protein and small molecule information, used as training data, into a protein-small molecule evaluation model. The protein-small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model. The protein-small molecule evaluation model generates corresponding atomic representations based on the protein and small molecule information. These atomic representations are then input into the distance distribution prediction model to obtain a first distance distribution of atom pairs exhibiting first short-range interactions, as output by the distance distribution prediction model. The atomic representations include the distances between atoms. This first distance distribution is then input into the binding affinity prediction model, which generates a corresponding first energy term based on a preset physical energy function. The binding affinity prediction model then outputs the affinity based on the first energy term. The protein small molecule evaluation model is obtained by fitting the distance label and binding affinity label corresponding to the training data to the protein small molecule evaluation model. The trained protein small molecule evaluation model is used to determine the corresponding protein small molecule evaluation information based on the target affinity evaluation information output by the affinity prediction model after inputting the target protein information and target small molecule information as the data to be predicted, and outputs the protein small molecule evaluation information. Thus, by proposing a protein small molecule evaluation model based on a combination of statistical physics and machine learning, the model captures the interaction between atoms in a data-driven manner, predicts the binding affinity of protein small molecules, and realizes the evaluation of the binding affinity of protein small molecules before the actual binding of proteins and small molecules.
[0056] In some embodiments, the protein information includes information about amino acids in the protein whose nearest distance to the small molecule is less than or equal to a preset threshold. In some embodiments, the protein information includes, but is not limited to, the atom type, chemical bond type, and atomic coordinates of amino acids in the protein whose distance to the small molecule is less than or equal to a preset threshold (e.g., 5 Å) (these amino acids are considered as binding pockets of the protein). This example embodiment does not impose any special limitations on this.
[0057] In some embodiments, the distance distribution prediction model includes a hybrid density network, and the combination affinity prediction model includes a neural network. In some embodiments, the distance distribution prediction model can be a hybrid density network, a neural network architecture specifically designed to handle regression problems with inherent uncertainty and multi-valued mappings. The output of a hybrid density network is not a single predicted value, but rather the parameters of a complete conditional probability distribution. In some embodiments, the combination affinity prediction model can be a neural network, also known as an artificial neural network (ANN), composed of interconnected nodes or artificial neurons arranged in a hierarchical structure, transmitting and processing data through weighted connections.
[0058] In some embodiments, the first distance distribution includes a first specific distance distribution corresponding to van der Waals interactions and a second specific distance distribution corresponding to hydrogen bond interactions. In some embodiments, the first short-range interaction includes van der Waals interactions and hydrogen bond interactions, the first specific distance distribution refers to the probability distribution of the most likely distances between atom pairs with van der Waals interactions, and the second specific distance distribution refers to the probability distribution of the most likely distances between atom pairs with hydrogen bond interactions.
[0059] In some embodiments, the method further includes: obtaining the distance between atomic pairs exhibiting a second interaction based on the atomic characterization using the protein small molecule evaluation model, wherein the second interaction includes hydrophobic interactions and / or metallic interactions; generating a corresponding second energy term based on the preset physical energy function according to the distance between the atomic pairs; wherein obtaining the affinity assessment information output by the binding affinity prediction model based on the first energy term includes: obtaining the affinity assessment information output by the binding affinity prediction model based on the first energy term and the second energy term. In some embodiments, the protein small molecule evaluation model obtains the distance between atomic pairs exhibiting a second interaction based on the distance between atoms in the protein and the distance between atoms in the small molecule in the atomic characterization, without needing to output a corresponding distance distribution through a distance distribution prediction model, wherein an atomic pair refers to two atoms composed of one atom in the protein and one atom in the small molecule, and the second interaction is an interaction other than the first short-range interaction, i.e., the second interaction is distinct from the first short-range interaction. In some embodiments, the protein small molecule evaluation model directly converts the distance between atomic pairs with second interactions into actual physical energies (i.e., the second energy term) based on a preset physical energy function. As mentioned earlier, the binding affinity prediction model converts the first distance distribution of atomic pairs with first short-range interactions into corresponding first energy terms based on a preset physical energy function. Each interaction corresponds to one energy term; for example, van der Waals interactions correspond to one first energy term, hydrogen bonding interactions to one first energy term, hydrophobic interactions to one second energy term, and metal interactions to one second energy term. In some embodiments, the binding affinity prediction model obtains the corresponding binding affinity based on each first energy term and each second energy term. For example, the result of adding each first energy term and each second energy term is divided by a preset coefficient, and the calculated value is used as the binding affinity. The method of determining affinity evaluation information based on binding affinity has been detailed above and will not be repeated here.
[0060] In some embodiments, the method further includes: obtaining prediction weights corresponding to each interaction based on the atomic characterization of the atomic pairs corresponding to various interactions using the protein small molecule evaluation model; wherein, obtaining the affinity evaluation information output by the binding affinity prediction model based on the first energy term and the second energy term includes: obtaining the affinity evaluation information output by the binding affinity prediction model based on the first energy term, the second energy term, and the prediction weights. In some embodiments, the protein small molecule evaluation model also predicts the prediction weights corresponding to each interaction by passing the atomic characterization through multiple linear networks to allow the atomic pairs to predict under different atomic environments, i.e., the prediction weights corresponding to van der Waals interactions, hydrogen bonding interactions, hydrophobic interactions, and metal interactions. In some embodiments, the protein small molecule evaluation model also obtains the prediction weights corresponding to each interaction by passing the atomic characterizations of the two atoms (one atom in the protein and one atom in the small molecule) corresponding to each interaction through multiple linear networks. In some embodiments, after obtaining the first or second energy term corresponding to each interaction using the affinity prediction model, the corresponding binding affinity is obtained based on the predicted weight of the interaction corresponding to each first or second energy term. For example, each first energy term is weighted and summed with its corresponding predicted weight, and each second energy term is weighted and summed with its corresponding predicted weight. The result is then divided by a preset coefficient (e.g., a penalty term for small molecule torsion bonds), and the calculated value is taken as the binding affinity. The method for determining affinity assessment information based on binding affinity has been described in detail above and will not be repeated here.
[0061] In some embodiments, the protein small molecule evaluation model further includes a small molecule rationality module; wherein, the method further includes: obtaining a distance constraint matrix of the small molecule based on the small molecule information through the small molecule rationality module, wherein the lower triangle of the distance constraint matrix contains the minimum allowed distance between each pair of atoms in the small molecule; generating corresponding small molecule rationality evaluation information based on the distance constraint matrix and the actual distance between all atoms in the small molecule through the small molecule rationality module; wherein, the trained protein small molecule evaluation model is used to determine the corresponding protein small molecule evaluation information based on the target affinity evaluation information output by the affinity prediction model and the target small molecule rationality evaluation information output by the small molecule rationality module after inputting target protein information and target small molecule information as data to be predicted, and outputting the protein small molecule evaluation information.
[0062] In some embodiments, the protein small molecule evaluation model further includes a small molecule rationality module. The small molecule rationality module is used to output corresponding small molecule rationality evaluation information. The small molecule rationality evaluation information refers to the evaluation information on the rationality of the small molecule structure. This evaluation information is used to assess whether there will be any phenomena that do not conform to the laws of physics or whether there will be collisions after the protein and small molecule bind together before the protein and small molecule actually bind together. The specific form of the small molecule rationality evaluation information can be a specific numerical value (e.g., a score), which is used to characterize the degree of rationality of the small molecule structure, or it can be a piece of text information. In some embodiments, the small molecule rationality module obtains a distance constraint matrix for the small molecule based on the small molecule information. For example, the distance constraint matrix is generated using the GetMoleculeBoundsMatrix function in RDKit (an open-source Python cheminformatics library). This distance constraint matrix defines the lower and upper limits (i.e., the allowed distance range) of the distance between each pair of atoms in the small molecule. The lower triangle of the distance constraint matrix contains the minimum allowed distance between each pair of atoms in the small molecule, and the upper triangle contains the maximum allowed distance between each pair of atoms in the small molecule. The distance constraints are based on the basic chemical structure of the small molecule (such as bond lengths and bond angles) and represent the geometric rules that the small molecule must follow. In some embodiments, the small molecule rationality module can generate corresponding small molecule rationality assessment information based on the distance constraint matrix and the actual distances between all atoms in the small molecule. For example, the minimum allowed distance between each pair of atoms in the small molecule, as contained in the distance constraint matrix, is compared with the actual distances between all atoms in the small molecule. The corresponding small molecule rationality assessment information is determined based on whether the actual distance is greater than or equal to the minimum distance and / or the difference between the actual distance and the minimum distance. In some embodiments, the small molecule rationality assessment information is in the form of a fractional value. First, the actual distance between all atoms in the small molecule is checked. If the actual distance between a pair of atoms is less than the minimum juice allowed by the "distance constraint matrix", it is considered a "violation". Based on the severity of all these "violations" (the larger the gap, the greater the penalty), a total penalty value is finally calculated and used as the small molecule rationality assessment information.
[0063] In some embodiments, during the model inference phase, the target protein information and target small molecule information, which are the data to be predicted, are input into a trained protein small molecule evaluation model. The binding affinity prediction model in the protein small molecule evaluation model outputs the corresponding target affinity evaluation information, and the small molecule rationality module outputs the corresponding target small molecule rationality evaluation information. Then, the protein small molecule evaluation model determines the corresponding protein small molecule evaluation information based on the target affinity evaluation information and the target small molecule rationality evaluation information. For example, if both the target affinity evaluation information and the target small molecule rationality evaluation information are in numerical form, the numerical value obtained by adding the target affinity evaluation information and the target small molecule rationality evaluation information can be used as the corresponding protein small molecule evaluation information. This application evaluates the structure of protein small molecules to assess whether there will be physically inconsistent phenomena or collisions after the protein and small molecule actually bind.
[0064] In some embodiments, generating corresponding small molecule rationality assessment information by the small molecule rationality module based on the distance constraint matrix and the actual distances between all atoms in the small molecule includes: the small molecule rationality module generating corresponding small molecule rationality assessment information by comparing the distance constraint matrix and a distance matrix generated based on the actual distances between all atoms in the small molecule. In some embodiments, the small molecule rationality module can compare the distance constraint matrix and the distance matrix generated based on the actual distances between all atoms in the small molecule, obtain a difference matrix between the two matrices by comparing the two matrices, and then determine the corresponding small molecule rationality assessment information based on the difference matrix.
[0065] In some embodiments, the protein small molecule evaluation model further includes a distance constraint module; wherein, the method further includes: generating corresponding atomic distance constraint evaluation information by the distance constraint module based on pre-set distance constraint conditions for specific atom pairs and the actual distance of the specific atom pairs in the protein small molecule; wherein, the trained protein small molecule evaluation model is used to determine the corresponding protein small molecule evaluation information based on the target affinity evaluation information output by the affinity prediction model, the target small molecule rationality evaluation information output by the small molecule rationality module, and the target atomic distance constraint evaluation information output by the distance constraint module after inputting target protein information and target small molecule information as data to be predicted, and outputting the protein small molecule evaluation information.
[0066] In some embodiments, the protein small molecule evaluation model further includes a distance constraint module. This module outputs corresponding atomic distance constraint evaluation information. The atomic distance constraint evaluation information refers to the evaluation information regarding whether specific atomic pairs of the protein small molecule (each specific atomic pair includes a specific atom in the small molecule and a specific atom in the protein) meet pre-defined distance constraint conditions. This evaluation information is used to evaluate whether specific atomic pairs of the protein small molecule meet pre-defined distance constraint conditions before the protein actually binds to the small molecule. The distance constraint conditions include one or more specific atomic pairs, specifying the distance conditions between each specific atomic pair in the one or more specific atomic pairs. Each specific atom pair includes a specific atom in a small molecule and a specific atom in a protein. The distance constraint specifies the distance conditions between the two specific atoms in each atom pair, and the distance conditions between the specific atoms in the small molecule and the specific atoms in the protein. The distance condition can mean that the actual distance between the two atoms needs to be within a preset distance range, or that the actual distance between the two atoms needs to satisfy a specific magnitude relationship relative to a preset distance threshold, or, in addition to satisfying a specific magnitude relationship, that the difference between the actual distance between the two atoms and the preset distance threshold needs to be less than or equal to the preset threshold. This example embodiment does not specifically limit this. In some embodiments, the specific form of the atomic distance constraint evaluation information can be a specific numerical value (e.g., a score) used to characterize the rationality of the small molecule structure, or it can be a piece of text information. In some embodiments, the distance constraint module generates corresponding atomic distance constraint evaluation information based on pre-set distance constraints for specific atom pairs (specific atoms in a small molecule and specific atoms in a protein) and the actual distance between the specific atom pairs in the protein molecule (i.e., the actual distance between a specific atom in the protein and a specific atom in the small molecule). For example, the corresponding atomic distance constraint evaluation information is determined based on whether the actual distance satisfies the distance constraint. Alternatively, if the actual distance does not satisfy the distance constraint, the corresponding atomic distance constraint evaluation information is determined based on the degree of difference (e.g., difference) between the actual distance and the distance constraint. In some embodiments, the atomic distance constraint evaluation information is in the form of a fractional value. For each user-specified specific atom pair in the protein molecule (each specific atom pair includes a specific atom in the small molecule and a specific atom in the protein), the actual distance between the two atoms in the specific atom pair is compared with the pre-set distance constraint for that specific atom pair. The comparison is sufficient to determine the penalty value for that specific atom pair. Then, the sum of the penalty values for all user-specified specific atom pairs is calculated and used as the atomic distance constraint evaluation information.
[0067] In some embodiments, during the model inference phase, the target protein information and target small molecule information, which are to be predicted, are input into a trained protein small molecule evaluation model. The binding affinity prediction model within the protein small molecule evaluation model outputs corresponding target affinity evaluation information, the small molecule rationality module outputs corresponding target small molecule rationality evaluation information, and the distance constraint module outputs corresponding target atom distance constraint evaluation information. Then, the protein small molecule evaluation model determines the corresponding protein small molecule evaluation information based on the target affinity evaluation information, the target small molecule rationality evaluation information, and the target atom distance constraint evaluation information. For example, if the target affinity evaluation information, the target small molecule rationality evaluation information, and the target atom distance constraint evaluation information are all in numerical form, the numerical value obtained by adding the three can be used as the corresponding protein small molecule evaluation information. This application, by evaluating the structure of the protein small molecule, enables the evaluation of whether specific atom pairs in the protein small molecule conform to pre-set distance constraints before the actual binding of the protein and the small molecule.
[0068] Figure 2 The diagram illustrates a method for evaluating small protein molecules according to an embodiment of this application, the method comprising steps S21, S22, S23, and S24. In step S21, protein information and small molecule information, which are the data to be predicted, are input into a protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model; in step S22, the protein small molecule evaluation model generates corresponding atomic representations based on the protein information and the small molecule information, and inputs the atomic representations into the distance distribution prediction model to obtain a first distance distribution of atomic pairs with first short-range interactions output by the distance distribution prediction model, wherein the atomic representations include the distance between atoms; in step S23, the first distance distribution is input into the binding affinity prediction model, and the binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function to obtain the affinity evaluation information output by the binding affinity prediction model based on the first energy term; in step S24, the protein small molecule evaluation model determines the corresponding protein small molecule evaluation information based on the affinity evaluation information and outputs the protein small molecule evaluation information.
[0069] In step S21, the protein information and small molecule information, which are the data to be predicted, are input into the protein small molecule evaluation model. The protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model. The specific method has been detailed above and will not be repeated here.
[0070] In step S22, the protein small molecule evaluation model generates corresponding atomic representations based on the protein information and the small molecule information. These atomic representations are then input into the distance distribution prediction model to obtain the first distance distribution of atom pairs exhibiting first short-range interactions, as output by the distance distribution prediction model. The atomic representations include the distances between atoms. The specific method has been detailed above and will not be repeated here.
[0071] In step S23, the first distance distribution is input into the binding affinity prediction model. The binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function, and the affinity assessment information output by the binding affinity prediction model based on the first energy term is obtained. The specific method has been described in detail above and will not be repeated here.
[0072] In step S24, the corresponding protein small molecule evaluation information is determined based on the affinity evaluation information using the protein small molecule evaluation model, and the protein small molecule evaluation information is output. The specific method has been described in detail above and will not be repeated here.
[0073] In some embodiments, the protein small molecule evaluation model further includes a small molecule rationality module; wherein the method further includes: obtaining a distance constraint matrix of the small molecule based on the small molecule information through the small molecule rationality module, wherein the lower triangle of the distance constraint matrix contains the minimum allowed distance between each pair of atoms in the small molecule; generating corresponding small molecule rationality evaluation information based on the distance constraint matrix and the actual distance between all atoms in the small molecule through the small molecule rationality module; wherein determining the corresponding protein small molecule evaluation information based on the affinity evaluation information through the protein small molecule evaluation model includes: determining the corresponding protein small molecule evaluation information based on the affinity evaluation information and the small molecule rationality evaluation information through the protein small molecule evaluation model. The specific method has been detailed above and will not be repeated here.
[0074] In some embodiments, the protein small molecule evaluation model further includes a distance constraint module; wherein, the method further includes: generating corresponding atomic distance constraint evaluation information by the distance constraint module based on pre-set distance constraint conditions for specific atom pairs and the actual distances of the specific atom pairs in the protein small molecule; wherein, determining the corresponding protein small molecule evaluation information by the protein small molecule evaluation model based on the affinity evaluation information and the small molecule rationality evaluation information includes: determining the corresponding protein small molecule evaluation information by the protein small molecule evaluation model based on the affinity evaluation information, the small molecule rationality evaluation information, and the atomic distance constraint evaluation information. The specific method has been detailed above and will not be repeated here.
[0075] Figure 3 The diagram illustrates a first device structure for obtaining a protein small molecule evaluation model according to an embodiment of this application. The first device includes a first module 11, a second module 12, a third module 13, and a fourth module 14. The first module 11 is used to input protein information and small molecule information, which are used as training data, into the protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model. The second module 12 is used to generate corresponding atomic representations based on the protein information and the small molecule information through the protein small molecule evaluation model, and input the atomic representations into the distance distribution prediction model to obtain a first distance distribution of atom pairs with a first short-range interaction, as output by the distance distribution prediction model, wherein the atomic representation includes the distance between atoms. The third module 13 is used to input the first distance distribution into the binding affinity prediction model, and... The binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function, and obtains the affinity assessment information output by the binding affinity prediction model based on the first energy term; Module 14 is used to fit the protein small molecule assessment model based on the distance label and binding affinity label corresponding to the training data to obtain a trained protein small molecule assessment model, wherein the trained protein small molecule assessment model is used to determine the corresponding protein small molecule assessment information based on the target affinity assessment information output by the affinity prediction model after inputting the target protein information and target small molecule information as the data to be predicted, and output the protein small molecule assessment information.
[0076] Module 11 is used to input protein and small molecule information, which serve as training data, into a protein-small molecule evaluation model. This model includes a distance distribution prediction model and a binding affinity prediction model. The related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0077] Module 12 is used to generate corresponding atomic representations based on the protein information and the small molecule information using the protein small molecule evaluation model, input the atomic representations into the distance distribution prediction model, and obtain the first distance distribution of atom pairs with first short-range interactions output by the distance distribution prediction model. The atomic representations include the distances between atoms. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0078] Module 13 is used to input the first distance distribution into the binding affinity prediction model, generate a corresponding first energy term based on a preset physical energy function through the binding affinity prediction model, and obtain the affinity assessment information output by the binding affinity prediction model based on the first energy term. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0079] Module 14 is used to fit the protein small molecule evaluation model based on the distance label and affinity label corresponding to the training data to obtain a trained protein small molecule evaluation model. The trained protein small molecule evaluation model is used to determine the corresponding protein small molecule evaluation information based on the target affinity evaluation information output by the affinity prediction model after inputting target protein information and target small molecule information as input data to be predicted, and then outputs the protein small molecule evaluation information. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0080] In some embodiments, the protein information includes amino acid information within the protein whose nearest distance to a small molecule is less than or equal to a preset threshold. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0081] In some embodiments, the distance distribution prediction model includes a hybrid density network, and the combined affinity prediction model includes a neural network. Here, the relevant operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0082] In some embodiments, the first distance distribution includes a first specific distance distribution corresponding to van der Waals interactions and a second specific distance distribution corresponding to hydrogen bond interactions. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0083] In some embodiments, the first device is further configured to: obtain the distance between atomic pairs exhibiting a second interaction based on the atomic characterization using the protein small molecule evaluation model, wherein the second interaction includes hydrophobic interactions and / or metallic interactions; generate a corresponding second energy term based on the preset physical energy function according to the distance between the atomic pairs; wherein obtaining the affinity assessment information output by the binding affinity prediction model based on the first energy term includes: obtaining the affinity assessment information output by the binding affinity prediction model based on the first energy term and the second energy term. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0084] In some embodiments, the first device is further configured to: obtain prediction weights corresponding to each of the various interactions based on the atomic characterization of the atomic pairs corresponding to the various interactions using the protein small molecule evaluation model; wherein, obtaining the affinity evaluation information output by the binding affinity prediction model based on the first energy term and the second energy term includes: obtaining the affinity evaluation information output by the binding affinity prediction model based on the first energy term, the second energy term, and the prediction weights. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0085] In some embodiments, obtaining the predicted weights corresponding to various interactions through the protein small molecule evaluation model includes: obtaining the predicted weights corresponding to each interaction based on the atomic characterization of the atomic pairs corresponding to each interaction through the protein small molecule evaluation model. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0086] In some embodiments, the protein small molecule evaluation model further includes a small molecule rationality module; wherein, the first device is further configured to: obtain a distance constraint matrix of the small molecule based on the small molecule information through the small molecule rationality module, wherein the lower triangle of the distance constraint matrix contains the minimum allowed distance between each pair of atoms in the small molecule; generate corresponding small molecule rationality evaluation information based on the distance constraint matrix and the actual distance between all atoms in the small molecule through the small molecule rationality module; wherein, the trained protein small molecule evaluation model is configured to, after inputting target protein information and target small molecule information as data to be predicted, determine the corresponding protein small molecule evaluation information based on the target affinity evaluation information output by the affinity prediction model and the target small molecule rationality evaluation information output by the small molecule rationality module, and output the protein small molecule evaluation information. Here, related operations and Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0087] In some embodiments, the step of generating corresponding small molecule rationality evaluation information by the small molecule rationality module based on the distance constraint matrix and the actual distances between all atoms in the small molecule includes: generating corresponding small molecule rationality evaluation information by comparing the distance constraint matrix and a distance matrix generated based on the actual distances between all atoms in the small molecule through the small molecule rationality module. Here, the related operations are similar to... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0088] In some embodiments, the protein small molecule evaluation model further includes a distance constraint module; wherein, the first device is further configured to: generate corresponding atomic distance constraint evaluation information by means of the distance constraint module based on pre-set distance constraint conditions for specific atom pairs and the actual distance of the specific atom pairs in the protein small molecule; wherein, the trained protein small molecule evaluation model is configured to, after inputting target protein information and target small molecule information as data to be predicted, determine the corresponding protein small molecule evaluation information based on the target affinity evaluation information output by the affinity prediction model, the target small molecule rationality evaluation information output by the small molecule rationality module, and the target atomic distance constraint evaluation information output by the distance constraint module, and output the protein small molecule evaluation information. Here, related operations and Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0089] Figure 4The diagram shows a second device for small molecule protein evaluation according to an embodiment of this application. The second device includes a second-order module 21, a second-order module 22, a second-order module 23, and a second-order module 24. Module 21 is used to input protein information and small molecule information, which are the data to be predicted, into a protein-small molecule evaluation model, wherein the protein-small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model; Module 22 is used to generate corresponding atomic representations based on the protein information and small molecule information through the protein-small molecule evaluation model, input the atomic representations into the distance distribution prediction model, and obtain a first distance distribution of atomic pairs with a first short-range interaction output by the distance distribution prediction model, wherein the atomic representations include the distance between atoms; Module 23 is used to input the first distance distribution into the binding affinity prediction model, generate a corresponding first energy term based on a preset physical energy function through the binding affinity prediction model, and obtain the affinity evaluation information output by the binding affinity prediction model based on the first energy term; Module 24 is used to determine the corresponding protein-small molecule evaluation information based on the affinity evaluation information through the protein-small molecule evaluation model, and output the protein-small molecule evaluation information.
[0090] Module 21 is used to input protein and small molecule information, which are the data to be predicted, into a protein-small molecule evaluation model. This model includes a distance distribution prediction model and a binding affinity prediction model. The related operations are... Figure 2 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0091] Module 22 is used to generate corresponding atomic representations based on the protein information and the small molecule information using the protein small molecule evaluation model, input the atomic representations into the distance distribution prediction model, and obtain the first distance distribution of atom pairs with first short-range interactions output by the distance distribution prediction model, wherein the atomic representations include the distances between atoms. Here, the related operations are... Figure 2 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0092] Module 23 is used to input the first distance distribution into the binding affinity prediction model, generate a corresponding first energy term based on a preset physical energy function through the binding affinity prediction model, and obtain the affinity assessment information output by the binding affinity prediction model based on the first energy term. Here, the related operations are... Figure 2 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0093] Module 24 is used to determine the corresponding protein small molecule evaluation information based on the affinity evaluation information using the protein small molecule evaluation model, and output the protein small molecule evaluation information. Here, the related operations are... Figure 2 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0094] In some embodiments, the protein small molecule evaluation model further includes a small molecule rationality module; wherein, the second device is further configured to: obtain a distance constraint matrix of the small molecule based on the small molecule information through the small molecule rationality module, wherein the lower triangle of the distance constraint matrix contains the minimum allowed distance between each pair of atoms in the small molecule; generate corresponding small molecule rationality evaluation information based on the distance constraint matrix and the actual distance between all atoms in the small molecule through the small molecule rationality module; wherein, determining the corresponding protein small molecule evaluation information based on the affinity evaluation information through the protein small molecule evaluation model includes: determining the corresponding protein small molecule evaluation information based on the affinity evaluation information and the small molecule rationality evaluation information through the protein small molecule evaluation model. Here, related operations are similar to... Figure 2 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0095] In some embodiments, the protein small molecule evaluation model further includes a distance constraint module; wherein, the second device is further configured to: generate corresponding atomic distance constraint evaluation information by means of the distance constraint module based on pre-set distance constraint conditions for specific atom pairs and the actual distance of the specific atom pairs in the protein small molecule; wherein, determining the corresponding protein small molecule evaluation information by means of the protein small molecule evaluation model based on the affinity evaluation information and the small molecule rationality evaluation information includes: determining the corresponding protein small molecule evaluation information by means of the protein small molecule evaluation model based on the affinity evaluation information, the small molecule rationality evaluation information and the atomic distance constraint evaluation information.
[0096] Figure 5 Exemplary systems that can be used to implement the various embodiments described in this application are shown; such as Figure 5 As shown in some embodiments, system 300 can function as any of the devices described in each of the embodiments. In some embodiments, system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., one or more processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the module and thus perform the actions described in this application.
[0097] In one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the processors 305 and / or any suitable device or component communicating with the system control module 310.
[0098] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.
[0099] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. In one embodiment, system memory 315 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, system memory 315 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).
[0100] In one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage device 320 and (one or more) communication interfaces 325.
[0101] For example, NVM / storage device 320 may be used to store data and / or instructions. NVM / storage device 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).
[0102] NVM / storage device 320 may include storage resources that are physically part of a device on which system 300 is mounted, or that can be accessed by the device without necessarily being part of it. For example, NVM / storage device 320 may be accessed via a network through one or more communication interfaces 325.
[0103] One or more communication interfaces 325 may provide the system 300 with an interface to communicate over one or more networks and / or with any other suitable device. The system 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.
[0104] In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 (e.g., memory controller module 330). In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 to form a system-in-package (SiP). In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die. In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die to form a system-on-a-chip (SoC).
[0105] In various embodiments, system 300 may be, but is not limited to, a server, workstation, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0106] In addition to the methods and devices described in the above embodiments, this application also provides a computer-readable storage medium storing computer code that, when executed, performs the method described in any of the preceding embodiments.
[0107] This application also provides a computer program product that, when executed by a computer device, performs the method described in any of the preceding claims.
[0108] This application also provides a computer device, the computer device comprising:
[0109] One or more processors;
[0110] Memory, used to store one or more computer programs;
[0111] When the one or more computer programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in any of the preceding methods.
[0112] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0113] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0114] Communication media include media through which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media can include guided transmission media (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and wireless (unguided transmission) media capable of propagating energy waves, such as sound, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data can be embodied as modulated data signals in, for example, wireless media (such as carrier waves or similar mechanisms embodied as part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are altered or set in a manner that encodes information in the signal. Modulation can be analog, digital, or a hybrid modulation technique.
[0115] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memories such as random access memory (RAM, DRAM, SRAM); and non-volatile memories such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or hereafter developed capable of storing computer-readable information / data for use by a computer system.
[0116] Herein, one embodiment of this application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing embodiments of this application.
[0117] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.
Claims
1. A method for obtaining a small protein molecule evaluation model, wherein, The method includes: The protein and small molecule information used as training data are input into the protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model. The protein small molecule evaluation model generates corresponding atomic representations based on the protein information and the small molecule information. The atomic representations are then input into the distance distribution prediction model to obtain the first distance distribution of atom pairs with first short-range interactions output by the distance distribution prediction model. The atomic representations include the distances between atoms. The first distance distribution is input into the binding affinity prediction model, and the binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function to obtain the affinity evaluation information output by the binding affinity prediction model based on the first energy term; The protein small molecule evaluation model is fitted based on the distance label corresponding to the training data and the affinity label to obtain a trained protein small molecule evaluation model. The trained protein small molecule evaluation model is used to determine the corresponding protein small molecule evaluation information according to the target affinity evaluation information output by the affinity prediction model after inputting the target protein information and target small molecule information as the data to be predicted, and output the protein small molecule evaluation information.
2. The method according to claim 1, wherein, The protein information includes information on amino acids in the protein whose nearest distance to small molecules is less than or equal to a preset threshold.
3. The method according to claim 1, wherein, The distance distribution prediction model includes a hybrid density network, and the binding affinity prediction model includes a neural network.
4. The method according to claim 1, wherein, The first distance distribution includes a first specific distance distribution corresponding to van der Waals interactions and a second specific distance distribution corresponding to hydrogen bond interactions.
5. The method according to claim 4, wherein, The method further includes: The distance between atom pairs with a second interaction is obtained by using the protein small molecule evaluation model based on the atomic characterization, wherein the second interaction includes hydrophobic interaction and / or metal interaction; A second energy term is generated based on the distance between the atomic pairs and the preset physical energy function. Wherein, obtaining the affinity assessment information output by the binding affinity prediction model based on the first energy term includes: Obtain the affinity assessment information output by the binding affinity prediction model based on the first energy term and the second energy term.
6. The method according to claim 5, wherein, The method further includes: The predicted weights for each interaction are obtained by using the protein small molecule evaluation model based on the atomic characterization of the atomic pairs corresponding to various interactions. Wherein, obtaining the affinity assessment information output by the binding affinity prediction model based on the first energy term and the second energy term includes: Obtain the affinity assessment information output by the binding affinity prediction model based on the first energy term, the second energy term, and the prediction weights.
7. The method according to claim 1, wherein, The protein small molecule evaluation model also includes a small molecule rationality module; The method further includes: The small molecule rationality module obtains the distance constraint matrix of the small molecule based on the small molecule information, wherein the lower triangle of the distance constraint matrix contains the minimum allowed distance between each pair of atoms in the small molecule; The small molecule rationality module generates corresponding small molecule rationality assessment information based on the distance constraint matrix and the actual distances between all atoms in the small molecule. The trained small molecule protein evaluation model is used to determine the corresponding small molecule protein evaluation information based on the target affinity evaluation information output by the affinity prediction model and the target small molecule rationality evaluation information output by the small molecule rationality module after inputting the target protein information and target small molecule information as the data to be predicted, and outputs the small molecule protein evaluation information.
8. The method according to claim 7, wherein, The small molecule rationality module generates corresponding small molecule rationality assessment information based on the distance constraint matrix and the actual distances between all atoms in the small molecule, including: The small molecule rationality module generates corresponding small molecule rationality assessment information by comparing the distance constraint matrix with a distance matrix generated based on the actual distances between all atoms in the small molecule.
9. The method according to claim 8, wherein, The protein small molecule evaluation model also includes a distance constraint module; The method further includes: The distance constraint module generates corresponding atomic distance constraint evaluation information based on the pre-set distance constraint conditions for specific atomic pairs and the actual distance of the specific atomic pairs in the protein molecule. The trained small molecule protein evaluation model is used to determine the corresponding small molecule protein evaluation information based on the target affinity evaluation information output by the affinity prediction model, the target small molecule rationality evaluation information output by the small molecule rationality module, and the target atom distance constraint evaluation information output by the distance constraint module after inputting the target protein information and target small molecule information as the data to be predicted, and outputs the small molecule protein evaluation information.
10. A method for evaluating small protein molecules, wherein, The method includes: The protein and small molecule information, which are the data to be predicted, are input into the protein small molecule evaluation model, wherein the protein small molecule evaluation model includes a distance distribution prediction model and a binding affinity prediction model. The protein small molecule evaluation model generates corresponding atomic representations based on the protein information and the small molecule information. The atomic representations are then input into the distance distribution prediction model to obtain the first distance distribution of atom pairs with first short-range interactions output by the distance distribution prediction model. The atomic representations include the distances between atoms. The first distance distribution is input into the binding affinity prediction model, and the binding affinity prediction model generates a corresponding first energy term based on a preset physical energy function to obtain the affinity evaluation information output by the binding affinity prediction model based on the first energy term; The protein small molecule evaluation model determines the corresponding protein small molecule evaluation information based on the affinity evaluation information and outputs the protein small molecule evaluation information.
11. The method according to claim 10, wherein, The protein small molecule evaluation model also includes a small molecule rationality module; The method further includes: The small molecule rationality module obtains the distance constraint matrix of the small molecule based on the small molecule information, wherein the lower triangle of the distance constraint matrix contains the minimum allowed distance between each pair of atoms in the small molecule; The small molecule rationality module generates corresponding small molecule rationality assessment information based on the distance constraint matrix and the actual distances between all atoms in the small molecule. The step of determining the corresponding small protein molecule evaluation information based on the affinity evaluation information using the protein small molecule evaluation model includes: Based on the affinity assessment information and the small molecule rationality assessment information, the corresponding small molecule assessment information is determined using the protein small molecule assessment model.
12. The method according to claim 11, wherein, The protein small molecule evaluation model also includes a distance constraint module; The method further includes: The distance constraint module generates corresponding atomic distance constraint evaluation information based on the pre-set distance constraint conditions for specific atomic pairs and the actual distance of the specific atomic pairs in the protein molecule. Wherein, determining the corresponding protein small molecule evaluation information based on the affinity evaluation information and the small molecule rationality evaluation information through the protein small molecule evaluation model includes: The protein small molecule evaluation model determines the corresponding protein small molecule evaluation information based on the affinity evaluation information, the small molecule rationality evaluation information, and the atomic distance constraint evaluation information.
13. A computer device for evaluating small protein molecules, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 12.
14. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 12.
15. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Predictive scoring function for estimating binding affinity
CN101542284A
Construction method and system of protein-ligand docking scoring model based on graph Transform
CN119360952A
Method and device for predicting protein-ligand binding affinity
CN119785873A
Computing affinity for protein-protein interaction
US20240029820A1