Method and device for constructing a covalent small molecule drug design model based on artificial intelligence
Through the reinforcement learning framework based on artificial intelligence and MMGBSA calculation, the problems of low design efficiency and low binding capacity prediction quality in the existing technology are solved, and efficient automated design processes and high-quality covalent small molecule generation are achieved.
Patent Information
- Application Number
- CN202410441133.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-04-12
AI Technical Summary
The existing calculation methods are difficult to apply to the design of covalent small molecule drugs, mainly due to the dependence on the scale of the molecular library and the need for manual inspection and design, resulting in low design efficiency and low binding capacity prediction quality.
Adopting a reinforcement learning framework based on artificial intelligence, a generative model is trained through reverse constraints to generate high-quality covalent small molecules, and the quality of binding capacity prediction is improved by introducing MMGBSA calculations to realize an automated design process.
Get rid of the dependence on commercial molecular libraries, improve design efficiency and binding capacity prediction quality, significantly reduce artificial workload, and realize the specific molecular library of target proteins.
Smart Images

Figure CN118230852B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer-aided drug design, and in particular to a method and device for constructing a covalent small molecule drug design model based on artificial intelligence. Background Art
[0002] Small molecule drugs form complexes with specific target proteins related to diseases, causing their conformations to change and lose their biological activity, thereby achieving the effect of treating diseases. According to the way they bind to target proteins, small molecule drugs can be divided into two categories: non-covalent and covalent. Non-covalent small molecule drugs rely on non-bonded interactions such as hydrophobic effects, hydrogen bonds, and electrostatic effects to reversibly bind to target proteins and can dissociate freely. Covalent small molecules, on the basis of non-bonded interactions, further form covalent bonds with certain chemically reactive amino acid residues on the target protein, which can maintain the binding state for a long time without dissociation, thereby achieving a more lasting therapeutic effect. Covalent small molecule drugs have received increasing attention in recent years.
[0003] In the field of drug development, the use of computational methods to assist molecular design has become an important means to reduce R&D costs, shorten R&D cycles, and improve industry competitiveness. The current mainstream computational methods usually focus on molecular docking methods, searching for molecules with good binding ability to target proteins in existing compound libraries. However, this method is not suitable for the design of covalent small molecule drugs due to two reasons. First, the results of the molecular docking algorithm are very dependent on the scale of the small molecule compound library used, and a large-scale compound library is usually required to ensure the success rate of the algorithm. According to literature reports, the screening success rate is 24% for molecular docking in a compound library with a scale of 138 million; the screening success rate is increased to 39% for molecular docking in a compound library with a scale of 858 million. However, the scale of existing covalent small molecule libraries is significantly insufficient. For example, the Ukrainian Enamine company can provide a compound library containing more than 48 billion small molecules, but can only provide a covalent compound library containing about 15,000 small molecules. This scale obviously cannot meet the needs of molecular docking algorithms; second, the existing computational design model requires manual inspection of the molecular docking results first, and then re-do the molecular docking after designing new molecules based on medicinal chemistry experience, and repeatedly iterate these two steps. This not only relies on manual experience and is time-consuming and labor-intensive, but also limits the application scale of the computational method. Summary of the invention
[0004] The present application provides a method and device for constructing a covalent small molecule drug design model based on artificial intelligence, generates a large number of high-quality covalent small molecules through a reinforcement learning framework, and improves the prediction quality of the binding ability of covalent small molecules by introducing MMGBSA calculation, thereby breaking through the limitations of existing calculation methods that are difficult to apply to covalent small molecule design and promoting the development of covalent small molecule drugs.
[0005] In a first aspect, the present application provides a method for constructing a covalent small molecule drug design model based on artificial intelligence, the method comprising:
[0006] Obtaining a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound, wherein the reactive molecular fragment set includes at least one reactive molecular fragment that can react with a covalent reaction site in the target protein to be bound;
[0007] Using the reaction molecule fragment set to reversely constrain the first small molecule drug generation model to obtain a second small molecule drug generation model, wherein the first small molecule drug generation model is pre-trained to generate a correct small molecule structure expression, and the second small molecule drug generation model is trained to generate a small molecule structure expression constrained by the reaction molecule fragment;
[0008] Formulate drugability conditions according to drugability rules, and train a second small molecule drug generation model based on the reverse constraints of the drugability conditions to obtain a third small molecule drug generation model, wherein the third small molecule drug generation model is trained to generate small molecule structural expressions that meet the drugability conditions;
[0009] The comprehensive scoring platform is based on the comprehensive scoring information of the tertiary small molecule structure expression generated by the third small molecule drug generation model, and the third small molecule drug generation model is trained by reverse constraints to obtain a covalent small molecule drug generation model, wherein the comprehensive scoring platform includes a control layer that generates an optimized molecule file based on the tertiary small molecule structure expression and an execution layer that calculates the comprehensive scoring information based on the optimized molecule file.
[0010] In a second aspect, the present application embodiment provides a covalent small molecule drug design method based on artificial intelligence, comprising the following steps:
[0011] Sampling at least one covalent small molecule drug generated by an artificial intelligence-based covalent small molecule drug design model.
[0012] In a third aspect, the present application embodiment provides a device for constructing a covalent small molecule drug design model based on artificial intelligence, comprising:
[0013] An acquisition module, used for acquiring a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound, wherein the reactive molecular fragment set is provided with at least one reactive molecular fragment that can react with a covalent reaction site in the target protein to be bound;
[0014] A first training module is used to train a first small molecule drug generation model using the reverse constraints of the reaction molecule fragment set to obtain a second small molecule drug generation model, wherein the first small molecule drug generation model is pre-trained to generate a correct small molecule structure expression, and the second small molecule drug generation model is trained to generate a small molecule structure expression constrained by the reaction molecule fragment;
[0015] A second training module is used to formulate drugability conditions according to drugability rules, and train the second small molecule drug generation model based on the reverse constraints of the drugability conditions to obtain a third small molecule drug generation model, wherein the third small molecule drug generation model is trained to generate a small molecule structure expression that meets the drugability conditions;
[0016] The scoring constraint module reversely constrains the comprehensive scoring information of the tertiary small molecule structure expression generated by the third small molecule drug generation model based on the comprehensive scoring platform to train the third small molecule drug generation model to obtain a covalent small molecule drug generation model, wherein the comprehensive scoring platform includes a control layer for generating optimized molecule files based on the tertiary small molecule structure expression and an execution layer for calculating the comprehensive scoring information based on the optimized molecule file.
[0017] In a fourth aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute a method for constructing a covalent small molecule drug design model based on artificial intelligence or a covalent small molecule drug design method based on artificial intelligence.
[0018] In a fifth aspect, an embodiment of the present application provides a readable storage medium, in which a computer program is stored. The computer program includes a program code for controlling a process to execute a process. The process includes a method for constructing a covalent small molecule drug design model based on artificial intelligence or a covalent small molecule drug design method based on artificial intelligence.
[0019] The main contributions and innovations of the present invention are as follows:
[0020] This solution uses a reinforcement learning framework to train generative models, thereby continuously generating high-quality covalent small molecules and getting rid of dependence on commercial molecular libraries; this solution combines molecular docking in the field of computational chemistry and MMGBSA calculations under the framework of reinforcement learning by establishing a comprehensive scoring platform, thus realizing a closed loop of molecule generation and molecular evaluation, and being able to provide docking conformations and binding abilities to target proteins while continuously generating new molecules, thus avoiding time-consuming and labor-intensive post-analysis steps, improving design efficiency, and improving the quality of prediction of the binding abilities of covalent small molecules; all processes in this solution are fully automated, significantly reducing manual workload, and adopting a step-by-step strategy to train the generative model, with specific reaction fragments, drugability conditions, and good binding abilities as training targets, thus avoiding the difficulty of simultaneously training multiple targets and improving training efficiency; under the guidance of the comprehensive scoring platform, the generated molecules in this solution are gradually enriched in high binding ability regions, thus realizing a specific molecular library for the target protein and avoiding the need to search the complete chemical space.
[0021] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0023] Figure 1 It is a flow chart of a method for constructing a covalent small molecule drug design model based on artificial intelligence according to an embodiment of the present application;
[0024] Figure 2 It is a structural block diagram of a device for constructing a covalent small molecule drug design model based on artificial intelligence according to an embodiment of the present application;
[0025] Figure 3 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0027] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0028] In order to facilitate the understanding of this solution, the conventional technical terms in this field that appear in this solution are explained here:
[0029] Small molecule structure expression: An expression generated in SMILES, used to describe the three-dimensional chemical structure of a molecule. SMILES (Simplified molecular input line entry system), simplified molecular linear input specification, is a specification that explicitly describes the structure of a molecule using an ASCII string. SMILES was developed by Arthur Weininger and David Weininger in the late 1980s and has been modified and extended by others, especially Daylight Chemical Information Systems Inc.
[0030] Residue: In the protein sequence, the amino and carboxyl groups between amino acids are dehydrated to form bonds. Since some of their groups participate in the formation of peptide bonds, the remaining structural parts are called amino acid residues.
[0031] MMGBSA calculation: Molecular Mechanics Generalized Born Surface Area calculation. The principle is to calculate the free energy of the target protein and small molecule in the two states of binding and dissociation, and to obtain the binding free energy by difference. The internal energy of the molecule is calculated using the molecular force field, while the solvation energy is calculated based on the generalized Born equation and the solvent accessible surface.
[0032] Embodiment 1
[0033] The embodiment of the present application provides a method for constructing a covalent small molecule drug design model based on artificial intelligence, which trains the artificial intelligence model through a reinforcement learning framework so that the covalent small molecule drug design model generates a large number of high-quality covalent small molecules, and introduces MMGBSA calculation in the training process of the covalent small molecule drug design model to improve the prediction quality of the covalent small molecule drug design model for the binding ability of covalent small molecules. Specifically, reference Figure 1 , the method comprising:
[0034] Obtaining a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound, wherein the reactive molecular fragment set includes at least one reactive molecular fragment that can react with a covalent reaction site in the target protein to be bound;
[0035] The first small molecule drug generation model is trained by reverse constraint of the reaction molecule fragment set to obtain a second small molecule drug generation model, wherein the first small molecule drug generation model is pre-trained to generate a correct small molecule structure expression, and the second small molecule drug generation model is trained to generate a small molecule structure expression constrained by the reaction molecule fragment; drugability conditions are formulated according to drugability rules, and the second small molecule drug generation model is trained based on reverse constraint of the drugability conditions to obtain a third small molecule drug generation model, wherein the third small molecule drug generation model is trained to generate a small molecule structure expression that meets the drugability conditions;
[0036] The comprehensive scoring platform is based on the comprehensive scoring information of the tertiary small molecule structure expression generated by the third small molecule drug generation model, and the third small molecule drug generation model is trained by reverse constraints to obtain a covalent small molecule drug generation model, wherein the comprehensive scoring platform includes a control layer that generates an optimized molecule file based on the tertiary small molecule structure expression and an execution layer that calculates the comprehensive scoring information based on the optimized molecule file.
[0037] In the step of "obtaining a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound", the types of amino acid residues in the target protein to be bound that participate in the reaction are determined, and based on the types of amino acid residues, a residue set consisting of amino acid residues within a first set distance from the center of the binding pocket of the target protein to be bound is obtained, and the residues pointing to the outside of the binding pocket and the residues whose reaction space is less than the second set distance in the residue set are removed to obtain covalent reaction sites, and a set of reactive molecular fragments is obtained based on the covalent reaction sites.
[0038] Specifically, the first set distance in this solution is That is, this scheme obtains the distance from the center of the target protein pocket to be bound The amino acid residues within constitute a residue set.
[0039] Furthermore, in the step of "removing residues in the residue set pointing outside the binding pocket", the first vector formed by the α carbon atom and the β carbon atom in each residue is calculated, and then the second vector formed by the α carbon atom and the center point of the binding pocket is calculated. If the angle between the first vector and the second vector is an obtuse angle, it means that the residue points to the outside of the binding pocket.
[0040] For example, the first vector is recorded as The second vector is denoted by when When and The angle is an obtuse angle.
[0041] Specifically, when the residues point to the outside of the binding pocket, they cannot effectively form interaction forces with other molecules in the binding pocket, which is not conducive to stable binding.
[0042] Furthermore, by arranging spatial partition points in the space around the residue and counting the number of partition points not occupied by any atom, the reaction space of each residue is obtained.
[0043] Specifically, this scheme takes a reactive atom in the residue as the origin, establishes a spherical coordinate system with a set radius distance, and uniformly constructs multiple space division points in the spherical coordinate system. If the distance between a space division point and any atom is less than the van der Waals radius of the atom, it is considered that the space division point falls inside the atom. If the distance between a space division point and any atom is greater than the van der Waals radius of the atom, it means that the space division point is not occupied by any atom, wherein the above-mentioned atoms refer to atoms in the residue or atoms in the target protein to be bound.
[0044] Specifically, the size of the reaction space of each residue is obtained according to the ratio of the number of segmentation points not occupied by any atom to the total number of segmentation points.
[0045] For example, the radius in this solution is set to When establishing the spherical coordinate system, the reaction atom is taken as the origin and the radius is within the range of The spherical coordinate variable r is divided at intervals of 15°, and a total of 25 r division points are obtained. Then the spherical coordinate variable θ is divided at intervals of 15°, and 13 θ division points are obtained. Then the spherical variable coordinate φ is divided at intervals of 15° to obtain 25 division points. Therefore, the entire spherical coordinate is divided into 25×13×24, a total of 7800 division points. Next, the relationship between each division point and any atom is determined to obtain the ratio of the number of division points not occupied by any atom to the total number of division points as shown in the following formula:
[0046] s=m / n
[0047] Here, m represents the number of partition points not occupied by any atom, and n represents the number of all partition points. Therefore, the ratio s of m to n represents the proportion of available space around the atom. That is, the higher the value of s, the more ample reaction space there is. In this scheme, residues with s greater than 0.8 are considered to have sufficient reaction space.
[0048] Furthermore, this scheme can calculate the reaction space of residues in a variety of ways.
[0049] Regarding the first small molecule drug generation model mentioned in this scheme: The first small molecule drug generation model mentioned in this scheme is a generative model in which an embedding layer, a gated recurrent unit and a linear output layer are connected in series, and the generative model is used to randomly generate small molecule structure expressions. During the pre-training stage of the first small molecule drug generation model, at least one real small molecule structure expression is obtained as a training sample, and the generative model is trained using the training sample to obtain the first small molecule drug generation model.
[0050] Specifically, the working principle of the generative model is to read a small molecule structure expression and then output the conditional distribution probability of the next character of the small molecule structure expression. In this scheme, the dimension of the embedding layer is set to 512, that is, the original small molecule structure expression is encoded into a 512-dimensional vector by the embedding layer, and this vector carries all the information in the original small molecule structure expression string as the input of the gated recurrent unit. In this scheme, the dimension of the gated recurrent unit is 1024, and 4 gated recurrent unit layers are used. The function of the gated recurrent unit layer is to read the encoded information output by the embedding layer and give a 1024-dimensional hidden vector. This vector is transformed as the output through the final linear output layer. The output vector is processed by the Softmax function to obtain the character distribution probability of the next character of the current small molecule structure expression. The generative model samples the next character according to the output distribution probability and puts it at the last position of the original small molecule structure expression to generate a new molecule.
[0051] Furthermore, the small molecule structure expression at the initial moment contains only one "^" character, which is regarded as the start symbol; "$" is regarded as the end symbol, and the small molecule structure expression is output immediately after the "$" symbol is sampled.
[0052] Specifically, in order to train the generative model to generate correct small molecule structure expressions, this solution collects entry information from the ChEMBL database as the initial training set, which contains about 2.4 million real small molecule structure expressions. Before the training begins, the parameters of the generative model are randomly assigned. At this time, the generative model will randomly generate small molecule structure expressions with an unknown probability distribution. Next, 512 small molecule structure expressions are randomly selected from the training set each time, and the probability of the generative model generating these 512 small molecule structure expressions is calculated, and the negative logarithm of this probability is used as the loss function as shown in the following formula:
[0053]
[0054] Among them, P i Represents the probability of generating the i-th small molecule structure expression.
[0055] The model parameters of the generative model are then adjusted through the gradient back propagation algorithm to reduce the value of L, thereby increasing the probability of the generative model generating these 512 small molecule structure expressions. After about 100 rounds of training, the probability distribution of the generative model output no longer improves. At this time, the training is stopped and the current parameters of the generative model are saved to obtain the first small molecule drug generation model.
[0056] Specifically, this solution can continuously produce high-quality covalent small molecules by training a generative model, thus getting rid of dependence on commercial molecular libraries.
[0057] In the step of "using the reaction molecule fragment set to reversely constrain and train the first small molecule drug generation model to obtain the second small molecule drug generation model", a second number of small molecule structure expressions generated by the first small molecule drug generation model is obtained as the small molecule structure expression set to be scored, and each small molecule structure expression in the small molecule structure expression set to be scored is scored using the constructed molecular fragment scorer. If the small molecule structure expression in the small molecule structure expression set to be scored includes at least one reaction molecule fragment, the score of the current small molecule structure expression is 1; if the small molecule structure expression in the small molecule structure expression set to be scored does not include a reaction molecule fragment, the score of the current small molecule structure expression is 0. The average score of all small molecule structure expressions in the small molecule structure expression set to be scored is used as feedback information to perform reinforcement learning on the first small molecule drug generation model. When the average score of all small molecule structure expressions in the small molecule structure expression set to be scored is greater than the set threshold, the parameters of the current first small molecule drug generation model are saved to obtain the second small molecule drug generation model.
[0058] Specifically, the second number set in this solution is 256, that is, the first small molecule drug generation model generates 256 small molecule structure expressions to form a set of small molecule structure expressions to be scored.
[0059] Specifically, the loss function used in the reinforcement learning of the first small molecule drug generation model is shown in the following formula:
[0060]
[0061] Among them, P 0 The probability of generating the current small molecule structure expression for the first small molecule drug generation model, P 1 is the probability of the current small molecule structure expression generated by the model under training, and s is the average score of all small molecule structure expressions in the set of small molecule structure expressions to be scored.
[0062] Specifically, after multiple rounds of training, the average scores of all small molecule structure expressions in the scored small molecule structure expression set gradually increase. The set threshold in this scheme is 0.8. When the average score is greater than 0.8, the second small molecule drug generation model is obtained. That is to say, the second small molecule drug generation model in this scheme is a model with the same structure as the first small molecule drug generation model but different parameters.
[0063] In the step of "reversely constraining the second small molecule drug generation model based on drugability conditions to obtain a third small molecule drug generation model", the small molecule structure expression generated by the second small molecule drug generation model is scored based on the drugability conditions to obtain a drugability score, and the drugability score is back-propagated to the second small molecule drug generation model for reinforcement learning. When the drugability score of the small molecule structure expression output by the second small molecule drug generation model is greater than a set threshold, the parameters of the current second small molecule drug generation model are saved to obtain a third small molecule drug generation model, wherein the drugability conditions include property conditions and conformational conditions.
[0064] The property conditions are the effects of the composition of the small molecule on the drugability, and the property conditions include:
[0065] 1. Molecular weight is less than 500;
[0066] 2. The logarithm of the oil-water partition coefficient does not exceed 3;
[0067] 3. The number of hydrogen bond acceptors does not exceed 10;
[0068] 4. The number of hydrogen bond donors does not exceed 5;
[0069] 5. The number of rotatable keys shall not exceed 10.
[0070] The scoring formula for each property condition of the small molecule structure expression generated by the second small molecule drug generation model is as follows:
[0071] s MW =1 / (1+10 MW-500 )
[0072] s LogP =1 / (1+10 10(LogP-3) )
[0073] s hba =1 / (1+10 10(hba-10) )
[0074] s hbd =1 / (1+10 10(hbd-5) )
[0075] s rot =1 / (1+10 10(rot-10) )
[0076] Among them, s MW is the molecular weight score, which approaches 0 when the molecular weight is higher than 500 and approaches 1 when the molecular weight is lower than 500. LogP is the score of the oil-water partition coefficient, which approaches 0 when its logarithmic value exceeds 3 and approaches 1 when it is less than 3. hba is the score of the number of hydrogen bond acceptors, which approaches 0 when the number is greater than 10 and approaches 1 when the number is less than 10. hbd is the score of the number of hydrogen bond donors, which approaches 0 when the number exceeds 5 and approaches 1 when the number is less than 5. rot is a score for the number of rotatable bonds, which approaches 0 when the number is greater than 10 and approaches 1 when the number is less than 10.
[0077] The part outside the rigid ring of the drug molecule may be too long, which will cause the conformation of the molecule in the solution to be very flexible. In the process of binding with the target protein, it is accompanied by a huge conformational entropy penalty, which weakens its binding ability. Therefore, this scheme adds structural conditions. When the generated small molecule structure expression meets the conformational conditions, it can better bind to the protein. The conformational conditions are the influence of the small molecule conformation on the drugability, and the conformational conditions include:
[0078] 1. The flexible length of the connecting part between the rings;
[0079] 2. The flexible length of the remaining branches.
[0080] The scoring formula for each conformational condition of the small molecule structure expression generated by the second small molecule drug generation model is as follows:
[0081] x linker =1 / (1+10 10(linker-4) )
[0082] s tail =1 / (1+10 10(tail-2) )
[0083] Among them, s linker is a connection score obtained based on the flexibility length of the connecting part between the rings. The flexibility length is defined as the number of consecutive SP3 hybridized atoms, which may include carbon, nitrogen, oxygen, sulfur, etc. Similarly, s tail It is the branch score obtained based on the flexible length of the remaining branch parts.
[0084] Specifically, the present scheme calculates the flexible length by analyzing the topology of the connecting parts between the rings and the remaining branch parts.
[0085] Furthermore, before obtaining the flexible length of the connecting part between the rings and the flexible length of the remaining branch parts of each small molecule structure expression, this solution needs to split the small molecule structure expression, and obtain the connecting part between the rings and the remaining branch parts in the splitting result.
[0086] Specifically, this scheme splits the small molecule structure expression into ring (ring-forming part), linker (connecting part between rings) and tail (remaining branch chain part).
[0087] In this scheme, the molecular weight score, the oil-water partition coefficient score, the hydrogen bond acceptor score, the hydrogen bond donor score, the rotatable bond score, the connection score, and the branch chain score are geometrically averaged to obtain the drugability score. The calculation formula is as follows:
[0088]
[0089] Specifically, this method normalizes the score results when calculating the scores of each property condition and conformation condition.
[0090] About the comprehensive scoring platform of this program:
[0091] The control layer in the comprehensive scoring platform obtains a first number of tertiary small molecule structure expressions to form a small molecule structure expression set, removes meaningless tertiary small molecule structure expressions in the small molecule structure expression set to obtain an optimized small molecule structure expression set, converts each tertiary small molecule structure expression in the optimized small molecule structure expression set into a molecular file, and obtains the energy of each molecular file based on the molecular conformation of each molecular file, adjusts the molecular conformation of each molecular file to minimize the energy of each molecular file to obtain an optimized molecule composition optimized molecular file set, the execution layer calculates the docking score of each optimized molecular file with the target protein to be bound, and calculates the binding energy of each optimized molecular file with the target protein to be bound through the MMGBSA program, and performs weighted average of the docking score and binding energy of each optimized molecular file with the target protein to be bound to obtain the comprehensive score information of each optimized molecular file.
[0092] In the step of "converting each tertiary small molecule structure expression in the optimized small molecule structure expression set into a molecular file", the molecular file includes a mol2 file and a sdf file.
[0093] Specifically, when converting each tertiary small molecule structure expression in the optimized small molecule structure expression set into a molecular file, this scheme determines the charge state of each tertiary small molecule structure expression to complete the hydrogen atoms of each tertiary small molecule structure expression. For the tertiary small molecule structure expression of the chiral structure, an R configuration molecular file and an S configuration molecular file are generated at each chiral center by enumeration.
[0094] That is to say, this scheme can convert each tertiary small molecule structure expression into multiple molecular files.
[0095] Specifically, since each chiral structure tertiary small molecule structure expression generates an R-configuration molecule file and an S-configuration molecule file, a chiral structure tertiary small molecule structure expression generates 2 N molecular files, N is the number of chiral centers.
[0096] In the step of "obtaining the energy of each molecule file based on the molecular conformation of each molecule file", the energy calculation formula of each molecule file is as follows:
[0097]
[0098] Among them, ε ij is the van der Waals energy potential well depth of the i-th atom and the j-th atom, which is defined as the geometric mean of their respective potential well depths, that is, R ij is the van der Waals radius of atom i and atom j, defined as the arithmetic mean of their respective van der Waals radii, i.e., R ij =(R i +R j ) / 2,q i and q j are the charges of atoms i and j, r ij is the distance between atom i and atom j, and N is the total number of atoms in the molecule file.
[0099] In the step of "adjusting the molecular conformation of each molecule file to minimize the energy of each molecule file to obtain an optimized molecular composition and an optimized molecule file set", the Monte Carlo simulated annealing algorithm is used to adjust the molecular conformation of each molecule file.
[0100] Specifically, the execution steps of the Monte Carlo simulated annealing algorithm are:
[0101] 1. Set the initial temperature T0 = 1000K, set the current temperature T = T0, record the dihedral angle of the molecular conformation of the current molecule file as x0, and the initial energy E0;
[0102] 2. For each rotatable dihedral angle, add a random value perturbation to the current value. The random value is a uniform random number between plus or minus 15° to generate a new conformation.
[0103] 3. Calculate the energy value E1 of the new conformation and calculate the probability of accepting the new conformation according to the following formula:
[0104]
[0105] 4. Choose to keep the original conformation or evolve to a new conformation based on the acceptance probability;
[0106] 5. Change the current temperature in the manner of T = 0.98 × T;
[0107] 6. Repeat steps 2 to 5 until the current temperature T is lower than 300K and terminate the optimization.
[0108] Save the optimized molecules as readable mol2 and sdf format molecular files.
[0109] In the step of "the execution layer calculates the docking score between each optimized molecule file and the target protein to be bound, and calculates the binding energy between each optimized molecule file and the target protein to be bound through the MMGBSA program", the execution layer checks the computing resources, obtains the available number of CPUs and memory size, and evenly sends the computing tasks to each CPU core for calculation.
[0110] Specifically, this scheme uses molecular docking software to dock each optimized molecule file with the target protein to be bound to calculate the docking score. The calculation formula of the docking score is as follows:
[0111] s dock =1 / (1+10 1+dock / 15 )
[0112] Among them, s dock代表归 The docking score after one.
[0113] Specifically, after the optimized molecule file is docked with the target protein to be bound, a docking result file is obtained, and the execution layer reads the docking result file and runs the MMGBSA program to calculate the binding energy.
[0114] s mmgbsa =1 / (1+10 1+mmgbsa / 100 )
[0115] Among them, s mmgbsa represents the normalized binding energy.
[0116] In the step of "taking a weighted average of the docking score and binding energy of each optimized molecule file with the target protein to be bound to obtain the comprehensive score information of each optimized molecule file", the calculation formula of the comprehensive score information is as follows:
[0117] s=0.5×s dock +0.5×s mmgbsa
[0118] In this scheme, covalent small molecule drugs have specific reactive groups and good binding ability, which can be used to guide chemical synthesis.
[0119] In this scheme, the parameters of the covalent small molecule drug generation model are saved, and according to subsequent needs, the parameters are reloaded and sampled to obtain a series of molecules with specified reactive groups that meet the requirements of drug-like molecules.
[0120] Specifically, this scheme introduces physically meaningful MMGBSA calculations based on the docking scores, which improves the prediction quality of the binding ability of covalent small molecule drugs.
[0121] Embodiment 2
[0122] A covalent small molecule drug design method based on artificial intelligence, comprising the following steps:
[0123] Sampling at least one covalent small molecule drug generated by the artificial intelligence-based covalent small molecule drug design model constructed in Example 1.
[0124] It should be noted that the covalent small molecule drugs currently obtained correspond to the target protein to be bound and the binding conditions. In other words, if the target protein to be bound or the binding conditions change, the target protein to be bound and the set of reactive molecular fragments corresponding to the target protein to be bound are re-obtained to train a specific artificial intelligence-based covalent small molecule drug design model, and a matching covalent small molecule drug is designed based on the specific artificial intelligence-based covalent small molecule drug design model.
[0125] Embodiment 3
[0126] Based on the same idea, refer to Figure 2 The present application also proposes a device for constructing a covalent small molecule drug design model based on artificial intelligence, including:
[0127] An acquisition module, used for acquiring a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound, wherein the reactive molecular fragment set is provided with at least one reactive molecular fragment that can react with a covalent reaction site in the target protein to be bound;
[0128] A first training module is used to train a first small molecule drug generation model using the reverse constraints of the reaction molecule fragment set to obtain a second small molecule drug generation model, wherein the first small molecule drug generation model is pre-trained to generate a correct small molecule structure expression, and the second small molecule drug generation model is trained to generate a small molecule structure expression constrained by the reaction molecule fragment;
[0129] A second training module is used to formulate drugability conditions according to drugability rules, and train the second small molecule drug generation model based on the reverse constraints of the drugability conditions to obtain a third small molecule drug generation model, wherein the third small molecule drug generation model is trained to generate a small molecule structure expression that meets the drugability conditions;
[0130] The scoring constraint module reversely constrains the comprehensive scoring information of the tertiary small molecule structure expression generated by the third small molecule drug generation model based on the comprehensive scoring platform to train the third small molecule drug generation model to obtain a covalent small molecule drug generation model, wherein the comprehensive scoring platform includes a control layer for generating optimized molecule files based on the tertiary small molecule structure expression and an execution layer for calculating the comprehensive scoring information based on the optimized molecule file.
[0131] Embodiment 4
[0132] This embodiment also provides an electronic device, referring to Figure 3 , comprises a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any of the above method embodiments.
[0133] Specifically, the processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0134] Among them, the memory 404 may include a large capacity memory 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 404 may be inside or outside the data processing device. In a specific embodiment, the memory 404 is a non-volatile memory. In a specific embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (Programmable Read-Only Memory, PROM for short), an erasable PROM (Erasable Programmable Read-Only Memory, EPROM for short), an electrically erasable PROM (Electrically Erasable Programmable Read-Only Memory, EEPROM for short), an electrically alterable ROM (Electrically Alterable Read-Only Memory, EAROM for short) or a flash memory (FLASH) or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (Static Random-Access Memory, abbreviated as SRAM) or a dynamic random access memory (Dynamic Random Access Memory, abbreviated as DRAM), wherein the DRAM can be a fast page mode dynamic random access memory 404 (Fast Page Mode Dynamic Random Access Memory, abbreviated as FPMDRAM), an extended data output dynamic random access memory (Extended Date Out Dynamic Random Access Memory, abbreviated as EDODRAM), a synchronous dynamic random access memory (Synchronous Dynamic Random-Access Memory, abbreviated as SDRAM), etc.
[0135] The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402 .
[0136] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the methods for constructing a covalent small molecule drug design model based on artificial intelligence in the above embodiments.
[0137] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .
[0138] The transmission device 406 can be used to receive or send data via a network. The specific examples of the above-mentioned network may include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0139] The input / output device 408 is used to input or output information. In this embodiment, the input information may be a set of reactive molecular fragments, etc., and the output information may be a covalent small molecule drug, etc.
[0140] Optionally, in this embodiment, the processor 402 may be configured to perform the following steps through a computer program:
[0141] S101, obtaining a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound, wherein the reactive molecular fragment set includes at least one reactive molecular fragment that can react with a covalent reaction site in the target protein to be bound;
[0142] S102, using the reaction molecule fragment set to reversely constrain the first small molecule drug generation model to obtain a second small molecule drug generation model, wherein the first small molecule drug generation model is pre-trained to generate a correct small molecule structure expression, and the second small molecule drug generation model is trained to generate a small molecule structure expression constrained by the reaction molecule fragment;
[0143] S103, formulating drugability conditions according to drugability rules, and training a second small molecule drug generation model based on reverse constraints of the drugability conditions to obtain a third small molecule drug generation model, wherein the third small molecule drug generation model is trained to generate small molecule structure expressions that meet the drugability conditions;
[0144] S104. Reverse constraint training of the third small molecule drug generation model based on the comprehensive score information of the tertiary small molecule structure expression generated by the third small molecule drug generation model by the comprehensive scoring platform obtains a covalent small molecule drug generation model, wherein the comprehensive scoring platform includes a control layer for generating an optimized molecule file based on the tertiary small molecule structure expression and an execution layer for calculating the comprehensive score information based on the optimized molecule file.
[0145] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0146] In general, various embodiments may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the boxes, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0147] Embodiments of the present invention may be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros may be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer executable components configured to perform an embodiment when the program is run. One or more computer executable components may be at least one software code or a portion thereof. In addition, at this point, it should be noted that, for example, Figure 3 Any block of the logic flow in the program may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. Physical media are non-transitory media.
[0148] Those skilled in the art should understand that the technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0149] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for constructing a covalent small molecule drug design model based on artificial intelligence, characterized in that: The following steps are involved: Obtaining a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound, wherein the reactive molecular fragment set includes at least one reactive molecular fragment that can react with a covalent reaction site in the target protein to be bound; Using the reaction molecule fragment set to reversely constrain the first small molecule drug generation model to obtain a second small molecule drug generation model, wherein the first small molecule drug generation model is pre-trained to generate a correct small molecule structure expression, and the second small molecule drug generation model is trained to generate a small molecule structure expression constrained by the reaction molecule fragment; Formulate drugability conditions according to drugability rules, and train a second small molecule drug generation model based on the reverse constraints of the drugability conditions to obtain a third small molecule drug generation model, wherein the third small molecule drug generation model is trained to generate small molecule structural expressions that meet the drugability conditions; The third small molecule drug generation model is trained by reverse constraint based on the comprehensive score information of the tertiary small molecule structure expression generated by the third small molecule drug generation model by the comprehensive scoring platform to obtain a covalent small molecule drug generation model, wherein the comprehensive scoring platform includes a control layer for generating an optimized molecule file based on the tertiary small molecule structure expression and an execution layer for calculating the comprehensive score information based on the optimized molecule file, the control layer obtains a first number of tertiary small molecule structure expressions to form a small molecule structure expression set, removes meaningless tertiary small molecule structure expressions in the small molecule structure expression set to obtain an optimized small molecule structure expression set, converts each tertiary small molecule structure expression in the optimized small molecule structure expression set into a molecule file, obtains the energy of each molecule file based on the molecular conformation of each molecule file, adjusts the molecular conformation of each molecule file to minimize the energy of each molecule file to obtain the optimized molecule composition optimized molecule file set, the execution layer calculates the docking score of each optimized molecule file with the target protein to be bound, and calculates the binding energy of each optimized molecule file with the target protein to be bound by the MMGBSA program, and performs weighted average of the docking score and binding energy of each optimized molecule file with the target protein to be bound to obtain the comprehensive score information of each optimized molecule file.
2. The method for constructing a covalent small molecule drug design model based on artificial intelligence according to claim 1, characterized in that: In the step of "obtaining a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound", the types of amino acid residues in the target protein to be bound that participate in the reaction are determined, and based on the types of amino acid residues, a residue set consisting of amino acid residues within a first set distance from the center of the binding pocket of the target protein to be bound is obtained, and the residues pointing to the outside of the binding pocket and the residues whose reaction space is less than the second set distance in the residue set are removed to obtain covalent reaction sites, and a set of reactive molecular fragments is obtained based on the covalent reaction sites.
3. The method for constructing a covalent small molecule drug design model based on artificial intelligence according to claim 1, characterized in that: The first small molecule drug generation model is a generative model that is connected in series with an embedding layer, a gated recurrent unit, and a linear output layer. The generative model is used to randomly generate small molecule structure expressions. During the pre-training stage of the first small molecule drug generation model, at least one real small molecule structure expression is obtained as a training sample, and the generative model is trained using the training sample to obtain the first small molecule drug generation model.
4. The method for constructing a covalent small molecule drug design model based on artificial intelligence according to claim 1, characterized in that: In the step of "using the reaction molecule fragment set to reversely constrain and train the first small molecule drug generation model to obtain the second small molecule drug generation model", a second number of small molecule structure expressions generated by the first small molecule drug generation model is obtained as the small molecule structure expression set to be scored, a molecular fragment scorer is constructed, and each small molecule structure expression in the small molecule structure expression set to be scored is scored using the molecular fragment scorer. If the small molecule structure expression in the small molecule structure expression set to be scored includes at least one reaction molecule fragment, the score of the current small molecule structure expression is 1; if the small molecule structure expression in the small molecule structure expression set to be scored does not include a reaction molecule fragment, the score of the current small molecule structure expression is 0. The average score of all small molecule structure expressions in the small molecule structure expression set to be scored is used as feedback information to perform reinforcement learning on the first small molecule drug generation model. When the average score of all small molecule structure expressions in the small molecule structure expression set to be scored is less than a set threshold, the parameters of the current first small molecule drug generation model are saved to obtain the second small molecule drug generation model.
5. The method for constructing a covalent small molecule drug design model based on artificial intelligence according to claim 1, characterized in that: In the step of "reversely constraining the second small molecule drug generation model based on drugability conditions to obtain a third small molecule drug generation model", the small molecule structure expression generated by the second small molecule drug generation model is scored based on the drugability conditions to obtain a drugability score, and the drugability score is back-propagated to the second small molecule drug generation model for reinforcement learning. When the drugability score of the small molecule structure expression output by the second small molecule drug generation model is greater than a set threshold, the parameters of the current second small molecule drug generation model are saved to obtain a third small molecule drug generation model, wherein the drugability conditions include property conditions and conformational conditions.
6. A covalent small molecule drug design method based on artificial intelligence, comprising the following steps: Sampling at least one covalent small molecule drug generated by the artificial intelligence-based covalent small molecule drug design model constructed by the construction method described in any one of claims 1-5.
7. A device for constructing a covalent small molecule drug design model based on artificial intelligence, characterized in that: The following steps are involved: An acquisition module, used for acquiring a target protein to be bound and a set of reactive molecular fragments corresponding to the target protein to be bound, wherein the reactive molecular fragment set is provided with at least one reactive molecular fragment that can react with a covalent reaction site in the target protein to be bound; A first training module is used to train a first small molecule drug generation model using the reverse constraints of the reaction molecule fragment set to obtain a second small molecule drug generation model, wherein the first small molecule drug generation model is pre-trained to generate a correct small molecule structure expression, and the second small molecule drug generation model is trained to generate a small molecule structure expression constrained by the reaction molecule fragment; A second training module is used to formulate drugability conditions according to drugability rules, and train the second small molecule drug generation model based on the reverse constraints of the drugability conditions to obtain a third small molecule drug generation model, wherein the third small molecule drug generation model is trained to generate a small molecule structure expression that meets the drugability conditions; The scoring constraint module reversely constrains the comprehensive score information of the tertiary small molecule structure expression generated by the third small molecule drug generation model based on the comprehensive scoring platform to train the third small molecule drug generation model to obtain a covalent small molecule drug generation model, wherein the comprehensive scoring platform includes a control layer that generates an optimized molecule file based on the tertiary small molecule structure expression and an execution layer that calculates the comprehensive score information based on the optimized molecule file, the control layer obtains a first number of tertiary small molecule structure expressions to form a small molecule structure expression set, removes meaningless tertiary small molecule structure expressions in the small molecule structure expression set to obtain an optimized small molecule structure expression set, and converts the optimized Each tertiary small molecule structure expression in the small molecule structure expression set is converted into a molecule file, and the energy of each molecule file is obtained based on the molecular conformation of each molecule file, and the molecular conformation of each molecule file is adjusted to minimize the energy of each molecule file to obtain an optimized molecule composition to form an optimized molecule file set, the execution layer calculates the docking score of each optimized molecule file with the target protein to be bound, and calculates the binding energy of each optimized molecule file with the target protein to be bound through the MMGBSA program, and the docking score and binding energy of each optimized molecule file with the target protein to be bound are weighted averaged to obtain the comprehensive score information of each optimized molecule file.
8. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to execute the method for constructing a covalent small molecule drug design model based on artificial intelligence as described in any one of claims 1 to 5 or the method for designing a covalent small molecule drug based on artificial intelligence as described in claim 6.
9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, and the process includes a method for constructing a covalent small molecule drug design model based on artificial intelligence according to any one of claims 1 to 5 or a method for designing covalent small molecule drugs based on artificial intelligence according to claim 6.
Citation Information
Patent Citations
Universal method for drug design aiming at different target proteins
CN114464270A
Specific target drug generation method and device based on graph neural network and MaxFlow platform
CN115762662A