Method, apparatus, device, and computer-readable storage medium for virtual screening of covalent inhibitors

By connecting candidate ligand molecules with mutant target proteins and multiple screenings, covalent inhibitors are screened using binding free energy and pharmacodynamic scores, the problem of low screening efficiency of covalent inhibitors in the prior art is solved, and rapid and efficient drug molecule discovery is achieved.

CN115881245BActive Publication Date: 2025-08-01SHENZHEN JINGTAI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211730570.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-08-01
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

The existing virtual screening technology mainly targets non-covalent inhibitors, and it is difficult to efficiently screen out active covalent inhibitors, resulting in high cost of drug development and long cycles.

Method used

By connecting candidate ligand molecules with the mutated target protein, the molecular conformation with covalent characteristics is screened using the preset binding free energy calculation function and the pharmacophore scoring function, and multiple screenings are performed by combining free energy and pharmacophore fractions to improve the screening efficiency and accuracy of covalent inhibitors.

Benefits of technology

It has achieved rapid screening of active covalent inhibitors from a million-level molecular library, improving screening efficiency and accuracy, and reducing drug development costs and cycles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115881245B_ABST
    Figure CN115881245B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and computer-readable storage medium for virtual screening of covalent inhibitors. The method includes: docking a plurality of candidate ligand molecules with a mutant target protein to obtain a plurality of molecular conformations; screening out a first set of molecular conformations according to a preset binding free energy calculation function to obtain a first binding free energy after docking of the plurality of molecular conformations with the mutant target protein; obtaining pharmacophore scores of the molecular conformations in the first set of molecular conformations according to a pharmacophore scoring function; screening out a second set of molecular conformations from the first set of molecular conformations according to the pharmacophore scores; respectively docking each molecular conformation in the second set of molecular conformations with the target protein to obtain a second binding free energy; screening out a third set of molecular conformations from the second set of molecular conformations according to the second binding free energy; and determining a covalent inhibitor according to the third set of molecular conformations. The present application can improve the screening efficiency and accuracy of covalent inhibitors active against target proteins.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of active compound screening, and particularly to a method, apparatus, device, and computer-readable storage medium for virtual screening of covalent inhibitors. Background Art

[0002] The design and development of new drugs is a creative and exploratory research work. Drug molecular design is to rationally optimize active compounds step by step through rational strategies and scientific planning, transform them into compounds that are safe, effective, controllable, and easily obtainable in the human body, and meet the requirements for multi-dimensional properties of drugs during the process of structure transformation and modification, and construct new molecular entities with expected pharmacological activities. Among them, virtual screening is an important means in the process of lead compound discovery. Virtual Screening (VS), also known as computer screening, is the screening of active compounds based on a small molecule database. By using molecular docking operations between small molecule compounds and drug targets, virtual screening can quickly select druggable active compounds from dozens to millions of molecules, thereby greatly reducing the number of compounds screened in experiments, shortening the research cycle, and reducing the cost of drug research and development. However, the currently used virtual screening technologies are basically for the screening of non-covalent inhibitors, and the demand for other types of inhibitors in the fields of medicinal chemistry and chemical biology is increasing. Summary of the Invention

[0003] The present application provides a method, apparatus, device, and computer-readable storage medium for virtual screening of covalent inhibitors, which can improve the screening efficiency and accuracy of covalent inhibitors active against target proteins. The method includes:

[0004] Docking a plurality of candidate ligand molecules with a mutated target protein to obtain a plurality of molecular conformations after docking the plurality of candidate ligand molecules with the mutated target protein molecule, wherein the mutated target protein is a target protein in which the amino acid residue covalently bound to the ligand molecule in the target protein is mutated to alanine;

[0005] Obtaining a first binding free energy after docking the plurality of molecular conformations with the mutated target protein according to a preset binding free energy calculation function, and screening out a first set of molecular conformations from the plurality of molecular conformations;

[0006] Obtaining the pharmacophore scores of the molecular conformations in the first set of molecular conformations according to a pharmacophore scoring function;

[0007] Screening out a second set of molecular conformations from the first set of molecular conformations according to the pharmacophore scores;

[0008] Performing molecular docking on each molecular conformation in the second molecular conformation set and the target protein, respectively, to obtain a second binding free energy after each molecular conformation in the second molecular conformation set is docked with the target protein molecule;

[0009] screening a third molecular conformation set from the second molecular conformation set according to the second binding free energy;

[0010] A covalent inhibitor is determined based on the third set of molecular conformations.

[0011] Optionally, obtaining the pharmacophore score of the molecular conformation in the first molecular conformation set according to the pharmacophore scoring function includes:

[0012] Obtain reference ligand molecules;

[0013] The molecular conformations in the first molecular conformation set and the molecular conformations of the crystal structure of the reference ligand molecule are matched and scored according to a pharmacophore scoring function to obtain a pharmacophore score.

[0014] Optionally, matching and scoring the molecular conformations in the first molecular conformation set and the molecular conformations of the crystal structure of the reference ligand molecule according to a pharmacophore scoring function to obtain a pharmacophore score includes:

[0015] Based on the distance bin representation of the feature type, a pharmacophore score is obtained according to the matching degree of the distance bins of the chemical features of the molecular conformations in the first molecular conformation set and the chemical features of the molecular conformation of the crystal structure of the reference ligand molecule, wherein the matching degree is calculated based on the following formula:

[0016]

[0017] Among them, C i,j is the feature matrix, i is the chemical feature of the candidate ligand molecule corresponding to the molecular conformation in the first molecular conformation set, j is the chemical feature of the reference ligand molecule, k is the current distance bin index, t is the chemical feature type, and r is the maximum distance category that exists; B i;t is the distance bin associated with feature i for feature type t, B j;t is the distance bin associated with feature type t and feature j; w(k) is the weight function of variable k.

[0018] Optionally, obtaining the pharmacophore score of the molecular conformation in the first molecular conformation set according to the pharmacophore scoring function includes:

[0019] Determining target covalent characteristics based on covalent binding residues within a preset range of the structural pocket of the mutated target protein after docking with the candidate ligand molecule;

[0020] Obtain the pharmacophore scores of the molecular conformations in the first molecular conformation set according to the pharmacophore scoring function, where the pharmacophore scores are used to indicate the possibility that the molecular conformations in the first molecular conformation set have the target covalent characteristics.

[0021] Optionally, screening out a second molecular conformation set from the first molecular conformation set according to the pharmacophore scores includes:

[0022] Screen out a fourth molecular conformation set from the first molecular conformation set according to the pharmacophore scores;

[0023] From each molecular conformation in the fourth molecular conformation set, screen out the molecular conformations in which the distance between the Michael acceptor and the amino acid residue in the target protein for covalent binding is less than the distance threshold as the second molecular conformation set.

[0024] Optionally, the Michael acceptor includes a functional group in the molecular conformation that can covalently bind to the target protein.

[0025] Optionally, the second binding free energy is obtained based on the free energy perturbation method or the thermodynamic integration method.

[0026] Optionally, the preset binding free energy calculation function is:

[0027] ΔG binding = ΔG0 + ΔG hbond Σ iI g1(Δr)g2(Δα) + ΔG metal Σ aM f(r aM ) + ΔG lipo Σ lL f(r lL ) <s

[0028] + ΔG rot H rot ; where,

[0029] ΔG0 is obtained through multiple linear regression;

[0030] ΔG hbond [[ID=X]]Σ iI g1(Δr)g2(Δα) is the hydrogen bond term, Σ iI g1(Δr)g2(Δα) is used to calculate the complementary possibility of all hydrogen bonds between ligand atom i and receptor atom I, g1(Δr) is a function of the hydrogen bond distance, and g2(Δα) is a function of the hydrogen bond angle;

[0031] ΔG metal Σ aM f(r aM ) is the metal term, Σ Note: There seems to be an error in the original text where "Σ" is used without a proper index in some places. I've tried to maintain the integrity as much as possible. Also, the "s0000158" in the translation might need further clarification depending on its original meaning.aM f(r aM ) is used to calculate all receptors and receptor / donor atoms a in the ligand and any metal atom M in the receptor. f(r aM ) is a function of distance;

[0032] ΔG lipo Σ lL f(r lL ) is the hydrophobic term. Σ lL f(r lL ) is used to calculate all lipophilic ligand atoms l and all lipophilic receptor atoms L. f(r lL ) is a function of the distance between hydrophobic atoms;

[0033] ΔG rot H rot is a frozen rotatable bond.

[0034] This application also provides a device for virtual screening of covalent inhibitors, including:

[0035] A first acquisition module, configured to dock a plurality of candidate ligand molecules with a mutant target protein, obtain a plurality of molecular conformations after docking the plurality of candidate ligand molecules with the mutant target protein molecule, and obtain a first binding free energy after docking the plurality of molecular conformations with the mutant target protein according to a preset binding free energy calculation function; wherein, the mutant target protein is a target protein obtained by mutating the amino acid residue in the target protein that covalently binds to the ligand molecule into alanine;

[0036] A first screening module, configured to screen out a first set of molecular conformations from the plurality of molecular conformations according to the first binding free energy;

[0037] A second acquisition module, configured to obtain the pharmacophore scores of the molecular conformations in the first set of molecular conformations according to a pharmacophore scoring function;

[0038] A second screening module, configured to screen out a second set of molecular conformations from the first set of molecular conformations according to the pharmacophore scores;

[0039] A third acquisition module, configured to perform molecular docking on each molecular conformation in the second set of molecular conformations with the target protein respectively, and obtain a second binding free energy after docking each molecular conformation in the second set of molecular conformations with the target protein molecule;

[0040] A third screening module, configured to screen out a third set of molecular conformations from the second set of molecular conformations according to the second binding free energy;

[0041] A determination module, configured to determine a covalent inhibitor according to the third set of molecular conformations.

[0042] The present application also provides a device for virtual screening of covalent inhibitors, including a memory and a processor. An executable code is stored on the memory. When the executable code is processed by the processor, the processor can be enabled to execute any one of the above methods.

[0043] The present application also provides a computer-readable storage medium, which stores an executable code. When the executable code is executed by a processor, the processor is enabled to execute any one of the above methods.

[0044] The method provided by the present application can achieve rapid screening of covalent inhibitors from millions of molecules, and can improve the accuracy of the candidate covalent inhibitor's ability to truly covalently interact with the target protein while quickly obtaining the candidate covalent inhibitor, that is, improve the efficiency and success probability of screening for covalent inhibitors active against the target protein. Description of the Drawings

[0045] Figure 1 is a schematic diagram of an embodiment of the method for virtual screening of covalent inhibitors of the present application;

[0046] Figure 2 is a schematic diagram of an embodiment of the device for virtual screening of covalent inhibitors in the present application;

[0047] Figure 3 is a schematic diagram of an embodiment of the device for virtual screening of covalent inhibitors in the present application. Detailed Embodiments

[0048] The embodiments of the present application will be described in more detail below with reference to the drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0049] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0050] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, "a plurality" means two or more, unless otherwise specifically defined.

[0051] As Figure 1 shown, Figure 1 is a schematic diagram of an embodiment of the method for virtual screening of covalent inhibitors in this application. The method includes:

[0052] Step S101, docking a plurality of candidate ligand molecules with a mutant target protein to obtain a plurality of molecular conformations after docking the plurality of candidate ligand molecules with the mutant target protein molecule.

[0053] Among them, the mutant target protein is the target protein in which the amino acid residue that covalently binds to the ligand molecule in the target protein is mutated to alanine, that is, the amino acid residue in the target protein that can covalently bind to the ligand molecule is mutated to alanine. Optionally, the mutant target protein has also undergone target protein hydrogenation treatment and target protein energy minimization treatment. The method of docking the candidate ligand molecule with the mutant target protein can be an existing conventional docking method. After each candidate ligand molecule is docked with the mutant target protein, at least one molecular conformation is obtained, for example, a plurality of molecular conformations are obtained. All the molecular conformations obtained after docking all the candidate ligand molecules with the mutant target protein respectively are used as the objects for initial screening.

[0054] Optionally, the plurality of candidate ligand molecules can be ligand molecules in a molecular library. The molecular library can be a small molecule database, such as known databases like ZINC, Specs, ChemBridge, etc., and there is no limitation thereto. Optionally, the mutant target protein is also the target protein that has undergone protein hydrogenation and minimization treatment.

[0055] Step S102, respectively determining the first binding free energy after docking the plurality of molecular conformations with the mutant target protein according to a preset binding free energy calculation function.

[0056] Optionally, the preset binding free energy calculation function is as follows:

[0057] ΔG binding =ΔG0 + ΔG hbond Σ iI g1(Δr)g2(Δα)+ΔGmetal Σ aM f(r aM ) + ΔG lipo Σ lL f(r lL )

[0058] +ΔG rot H rot ; wherein, ΔG binding is the binding free energy, and ΔG0 can be obtained by multiple linear regression; ΔG hbond Σ iI g1(Δr)g2(Δα) is the hydrogen bond term, and Σ iI g1(Δr)g2(Δα) is used to calculate the complementary possibility of all hydrogen bonds between ligand atom i and receptor atom I. g1(Δr) is a function of the hydrogen bond distance, and g2(Δα) is a function of the hydrogen bond angle; ΔG metal Σ aM f(r aM ) is the metal term, and Σ aM f(r aM ) is used to calculate all receptors and receptor / donor atoms a in the ligand molecule and any metal atom M in the receptor. f(r aM ) is a function of the distance; ΔG lipo Σ lL f(r lL ) is the hydrophobic term, and Σ lL f(r lL ) is used to calculate all lipophilic ligand atoms l and all lipophilic receptor atoms L. f(r lL ) is a function of the distance between hydrophobic atoms; ΔG rot H rot is the frozen rotatable bond.

[0059] Optionally, the preset binding free energy calculation function can also be other existing scoring functions of the binding free energy.

[0060] Step S103, screen out the first set of molecular conformations from the multiple molecular conformations according to the first binding free energy.

[0061] There are multiple methods for screening out the first set of molecular conformations. Optionally, after using a preset binding free energy calculation function to respectively determine the first binding free energies of the multiple molecular conformations after docking with the mutant target protein, the molecular conformations corresponding to the first binding free energies within the first preset range are screened out as the first set of molecular conformations; or, all the calculated first binding free energies are sorted according to their numerical values, and the molecular conformations corresponding to the first binding free energies within the first preset ratio range are screened out as the first set of molecular conformations. For example, the molecular conformations within a certain ratio from the smallest to the largest in the numerical sorting are screened out as the second set of molecular conformations, and this certain ratio can be any positive value. The numerical value of the calculated binding free energy is generally negative. The smaller the numerical value of the binding free energy, that is, the larger the absolute value of this numerical value, the stronger the binding ability of the ligand molecule to the target protein.

[0062] Step S104, obtaining the pharmacophore scores of the molecular conformations in the first set of molecular conformations according to the pharmacophore scoring function.

[0063] There are multiple methods for obtaining pharmacophore scores. In one example, a reference ligand molecule is obtained, and the molecular conformations in the first set of molecular conformations and the molecular conformations of the crystal structure of the reference ligand molecule are matched and scored according to the pharmacophore scoring function to obtain the pharmacophore scores. The reference ligand molecule can be a small molecule compound for which the co-crystal structure formed with the target protein is known and the molecular conformation of the reference ligand molecule in the co-crystal structure is resolved, or it can be a known marketed small molecule drug, a candidate small molecule compound in the R & D pipeline, or a reported small molecule compound with activity. When the reference ligand molecule is a small molecule compound for which the co-crystal structure formed with the target protein is known and the molecular conformation of the reference ligand molecule in the co-crystal structure is resolved, the molecular conformation of the reference ligand molecule in the co-crystal structure is respectively matched and scored with the molecular conformations in the first set of molecular conformations; when the co-crystal structure formed by the reference ligand molecule and the target protein is not known, among the multiple molecular conformations of the reference ligand molecule, the molecular conformation with a lower binding free energy to the target protein can be used as a reference and respectively matched and scored with the molecular conformations in the first set of molecular conformations.

[0064] Optionally, when matching and scoring the molecular conformations in the first set of molecular conformations and the molecular conformations of the crystal structure of the reference ligand molecule, the chemical features of the two can be represented based on distance bins of feature types, and the matching score can be performed based on the similarity of the distance bins of the two. Optionally, the matching degree of the distance bins of the two can be calculated by the following formula:

[0065]

[0066] where, C i,jis the feature matrix, i is the chemical feature of the candidate ligand molecule corresponding to the molecular conformation in the first set of molecular conformations, j is the chemical feature of the reference ligand molecule, k is the current distance bin index, t is the chemical feature type, r is the maximum distance category present; B i;t is the distance bin associated with feature i for feature type t, B j;t is the distance bin associated with feature j for feature type t; w(k) is the weight function of variable k. The compensation for the preference for closer features can be adjusted by changing the magnitude of w(k). When obtaining the pharmacophore score, the covalent pharmacophore features can be weighted more heavily and the compensation reduced. The molecular conformations of the compound are evaluated for their match with the reference ligand molecule using the pharmacophore scoring function, and the molecular conformations with covalent pharmacophore features are preferentially given high scores, so that the absolute value of the score (i.e., the first binding free energy) of the ligand molecule that can form a covalent bond is higher, in order to achieve the purpose of enriching potential covalent inhibitors.

[0067] In another example of obtaining the pharmacophore score, the target covalent feature is determined based on the covalent binding residues within a preset range of the structural pocket of the mutated target protein after docking with the candidate ligand molecule. Among them, the covalent feature corresponding to the covalent binding residue can be determined based on all or a selected part of the covalent binding residues within the preset range of the structural pocket, as the target covalent feature. The target covalent feature is a covalent feature that a candidate ligand molecule capable of being a covalent inhibitor may have. The pharmacophore score of the molecular conformations in the first set of molecular conformations is obtained according to the pharmacophore scoring function. The pharmacophore score can indicate the possibility that the molecular conformations in the first set of molecular conformations have the target covalent feature. Among them, the pharmacophore scoring function can adopt an existing pharmacophore scoring function, which is not limited here. This example can be used to obtain the pharmacophore score of the molecular conformations in the first set of molecular conformations in the absence of a reference ligand molecule.

[0068] Step S105, screening out a second set of molecular conformations from the first set of molecular conformations according to the pharmacophore score.

[0069] There are various methods for screening out the second set of molecular conformations from the first set of molecular conformations. Optionally, the molecular conformations corresponding to the pharmacophore scores within a second preset range in the first set of molecular conformations are screened out as the second set of molecular conformations; or, the pharmacophore scores of all the molecular conformations in the first set of molecular conformations calculated are sorted by numerical value, and the molecular conformations corresponding to the pharmacophore scores within the second preset percentage range are screened out as the second set of molecular conformations.

[0070] Step S106: Perform molecular docking on each molecular conformation in the second molecular conformation set with the target protein, and obtain the second binding free energy after molecular docking of each molecular conformation in the second molecular conformation set with the target protein molecule.

[0071] Optionally, the method for obtaining the second binding free energy may be the same as or different from the method for obtaining the first binding free energy. For example, the second binding free energy can be calculated using a preset binding free energy calculation function or other existing binding free energy scoring functions. Optionally, the method for obtaining the second binding free energy can be a free energy calculation method with higher precision than the preset binding free energy calculation function. By performing a first screening using a free energy calculation method with lower precision to obtain the first binding free energy, and then performing a second screening using a free energy calculation method with higher precision to obtain the second binding free energy, since the number of molecular conformations used to calculate the second binding free energy has been greatly reduced after multiple screenings, it is possible to balance the efficiency and success rate of the screening of molecular conformations.

[0072] Optionally, the second binding free energy is obtained based on the Free Energy Perturbation (FEP) method or the Thermodynamic Integration (TI) method. Both the FEP method and the TI method are free energy calculation methods with higher precision based on molecular force fields, and are both high-precision methods for evaluating the binding strength between a small drug molecule and a target protein. These two methods are mainly based on the principle of Thermodynamic Cycle, and can calculate the binding free energy of a ligand molecule through a series of gradual changes in van der Waals parameters and charge parameters at the atomic level.

[0073] Among them, when using the FEP method to obtain the second binding free energy, a complex is formed by combining the complete, unmutated target protein (receptor) with a single molecular conformation. The binding free energy from the bound state A to the bound state B of the complex is calculated by the Zwanzig equation (i.e., the following formula), and the second binding free energy is calculated based on the calculated ΔF(A→B).

[0074]

[0075] Among them, the bound state A is the binding state of the reference ligand molecule A with the target protein, or the state where water molecules are in a stable state with the target protein, or there is no bound state A (F A = 0); the bound state B is the binding state of the candidate ligand molecule B with the target protein, T is the temperature, k B is the Boltzmann constant, and the term in the triangular brackets represents the average value of the simulation run of the bound state A, F BF is the Helmholtz free energy of the binding of the candidate ligand molecule B to the target protein A is the Helmholtz free energy of the binding of the reference ligand molecule A to the target protein or the Helmholtz free energy of the binding of a water molecule to the target protein. ΔF represents the difference in the binding free energy between the bound state B (the Helmholtz free energy of the binding of the candidate ligand molecule B to the target protein) and the bound state A (the Helmholtz free energy of the binding of the reference ligand molecule A to the target protein, the Helmholtz free energy of the binding of a water molecule to the target protein, or F when there is no bound state A A ).

[0076] Among them, when using the TI method to obtain the second binding free energy, the free energy difference is calculated by defining the thermodynamic path between the states and integrating the ensemble-averaged enthalpy change along the path. Therefore, the second binding free energy calculated by using the TI method can be obtained by the following formula (2).

[0077]

[0078] Among them, F(λ) is the λ-coupling potential function corresponding to F(A) at λ = 0 and F(B) at λ = 1. λ is an intermediate virtual state set on the path from the bound state A to the bound state B, between 0 and 1; ΔF represents the difference in the binding free energy between the bound state B and the bound state A.

[0079] Step S107, screening out a third molecular conformation set from the second molecular conformation set according to the second binding free energy.

[0080] There are various methods for screening out the third molecular conformation set that meets the screening conditions from the second molecular conformation set. Optionally, after obtaining the second binding free energy, the molecular conformations corresponding to the second binding free energy with values within the third preset range are screened out as the third molecular conformation set; or, all the calculated second binding free energies are sorted by numerical value, and the molecular conformations corresponding to the second binding free energy within the third preset proportion range are screened out as the third molecular conformation set. For example, the molecular conformations within a certain proportion from the smallest to the largest in the numerical sorting are screened out as the second molecular conformation set, and this certain proportion can be any positive value.

[0081] Optionally, when the number of molecular conformations corresponding to the same candidate ligand molecule in the molecular conformations that meet the screening conditions during screening exceeds a certain upper limit, a part of the molecular conformations corresponding to the same candidate ligand molecule can be filtered out from the second molecular conformation set, and then the molecular conformations that meet the screening conditions are screened out from the remaining molecular conformations in the second molecular conformation set again as the third molecular conformation set. When filtering out a part of the molecular conformations corresponding to the same candidate ligand molecule, the number of filtered out can be determined according to a predetermined ratio or a preset number.

[0082] Step S108, determining a covalent inhibitor according to the third molecular conformation set.

[0083] Optionally, directly use the candidate ligand molecules corresponding to the molecular conformations in the third molecular conformation set as covalent inhibitors, or the third molecular conformation set can be further screened, and the candidate ligand molecules corresponding to the screened molecular conformations are used as covalent inhibitors.

[0084] In the method for virtual screening of covalent inhibitors according to the embodiments of the present application, after the candidate ligand molecules are first docked with the mutated target protein to obtain molecular conformations, the first binding free energy of the molecular conformations is used for the first screening to filter out some molecular conformations. Since the amino acid residues in the target protein that covalently bind to the ligand molecules are mutated into alanine in the target protein, the steric hindrance of alanine is small and there are generally no polar interactions, and it is not necessary to consider the reaction between the potential covalent inhibitor and the covalently bound residues. The binding situation between the candidate ligand molecules and the target protein is mainly considered, avoiding more complex calculations and reducing computing power and costs. Moreover, alanine is an amino acid with the shortest side chain. If it is mutated into other amino acids, it is easy to change the configuration of the protein; for example, if it is mutated into glycine, since glycine has no side chain, mutating into glycine is easy to change the configuration of the protein. After the first screening, some molecular conformations that do not have covalent characteristics are filtered out based on the pharmacophore scoring function. After the screened candidate ligand molecules are docked with the target protein to obtain molecular conformations, the molecular conformations with better binding ability are screened out by the binding free energy to form a set of candidate molecular conformations. Such a screening process can quickly predict the binding ability of molecular conformations. Multiple screenings can improve the accuracy of the binding ability prediction results and screen out molecular conformations with better binding ability, thereby effectively improving the success rate of virtual screening to discover active compounds.

[0085] Optionally, to further improve the success rate of screening, when screening out the second molecular conformation set from the first molecular conformation set according to the pharmacophore score in step S105, first, according to the pharmacophore score, screen out the fourth molecular conformation set from the first molecular conformation set; then dock the molecular conformations in the fourth molecular conformation set with the target protein, and screen out the molecular conformations in the fourth molecular conformation set where the distance between the Michael acceptor and the amino acid residue in the target protein for covalent binding is less than the distance threshold as the second molecular conformation set. Among them, the Michael acceptor is a functional group in the molecular conformation that can covalently bind to the target protein. Among them, the amino acid residue in the target protein can be Ser, cysteine, etc. Among them, the distance threshold can be any value selected from 1 to 2 angstroms, such as 1, 1.2, 1.3, 1.5, 1.7, 1.8 angstroms. In this way, potential covalent inhibitors are screened out by judging whether there are functional groups capable of covalent reaction through the pharmacophore scoring function, and then molecular conformations with covalent pharmacophores are further screened out through distance constraints, improving the screening success rate.

[0086] As Figure 2 shown, the present application also provides a device 20 for virtual screening of covalent inhibitors, including a memory 21 and a processor 22, wherein the processor 22 is used to execute the above method for virtual screening of covalent inhibitors according to the instructions stored in the memory 21.

[0087] As Figure 3 shown, the present application also provides a device 30 for virtual screening of covalent inhibitors, including:

[0088] A first acquisition module 301, configured to dock a plurality of candidate ligand molecules with a mutant target protein, acquire a plurality of molecular conformations after docking the plurality of candidate ligand molecules with the mutant target protein molecule, and acquire a first binding free energy after docking the plurality of molecular conformations with the mutant target protein according to a preset binding free energy calculation function.

[0089] Among them, the mutant target protein is a target protein obtained by mutating the amino acid residue in the target protein that covalently binds to the ligand molecule into alanine.

[0090] A first screening module 302, configured to screen out a first molecular conformation set from the plurality of molecular conformations according to the first binding free energy;

[0091] A second acquisition module 303, configured to acquire the pharmacophore scores of the molecular conformations in the first molecular conformation set according to a pharmacophore scoring function;

[0092] A second screening module 304, configured to screen out a second molecular conformation set from the first molecular conformation set according to the pharmacophore score;

[0093] A third acquisition module 305, configured to perform molecular docking on each molecular conformation in the second molecular conformation set with a target protein, and acquire a second binding free energy after molecular docking of each molecular conformation in the second molecular conformation set with the target protein molecule;

[0094] A third screening module 306, configured to screen out a third molecular conformation set from the second molecular conformation set according to the second binding free energy;

[0095] A determination module 307, configured to determine a covalent inhibitor according to the third molecular conformation set.

[0096] Optionally, the second acquisition module 303 is specifically configured to: acquire a reference ligand molecule; perform matching scoring on the molecular conformations in the first molecular conformation set and the molecular conformation of the crystal structure of the reference ligand molecule according to a pharmacophore scoring function, and obtain a pharmacophore score.

[0097] Optionally, the performing matching scoring on the molecular conformations in the first molecular conformation set and the molecular conformation of the crystal structure of the reference ligand molecule according to a pharmacophore scoring function, and obtaining a pharmacophore score includes:

[0098] Based on the distance bin representation of the feature type, obtain a pharmacophore score according to the matching degree of the distance bin of the chemical features of the molecular conformations in the first molecular conformation set and the chemical features of the molecular conformation of the crystal structure of the reference ligand molecule, where the matching degree is calculated based on the following formula:

[0099]

[0100] where C i,j is a feature matrix, i is the chemical feature of the candidate ligand molecule corresponding to the molecular conformation in the first molecular conformation set, j is the chemical feature of the reference ligand molecule, k is the current distance bin index, t is the chemical feature type, and r is the maximum distance category in existence; B i;t is the distance bin associated with the feature type t and the feature i, and B j;t is the distance bin associated with the feature type t and the feature j; w(k) is the weight function of the variable k.

[0101] Optionally, the second acquisition module 303 is specifically configured to: determine a target covalent feature according to the covalent binding residues within a preset range of the structural pocket of the mutated target protein after docking with the candidate ligand molecule; and obtain the pharmacophore score of the molecular conformations in the first molecular conformation set according to the pharmacophore scoring function and the target covalent feature. The pharmacophore score can indicate the possibility that the molecular conformations in the first molecular conformation set have the target covalent feature.

[0102] Optionally, the second screening module 304 is specifically configured to screen out a fourth molecular conformation set from the first molecular conformation set according to the pharmacophore score; and screen out, from each molecular conformation in the fourth molecular conformation set, molecular conformations in which the distance between the Michael acceptor and the amino acid residue in the target protein covalently bound is less than a distance threshold, as the second molecular conformation set.

[0103] Optionally, the Michael acceptor is a functional group in the molecular conformation that can covalently bind to the target protein.

[0104] Optionally, the second binding free energy is obtained based on the free energy perturbation method or the thermodynamic integration method.

[0105] Optionally, the preset binding free energy calculation function is:

[0106] ΔG binding = ΔG0 + ΔG hbond Σ iI g1(Δr)g2(Δα) + ΔG metal Σ aM f(r aM ) + ΔG lipo Σ lL f(r lL )

[0107] + ΔG rot H rot ; where

[0108] ΔG0 is obtained through multiple linear regression;

[0109] ΔG hbond Σ iI g1(Δr)g2(Δα) is the hydrogen bond term, Σ iI g1(Δr)g2(Δα) is used to calculate the complementary possibility of all hydrogen bonds between ligand atom i and receptor atom I, g1(Δr) is a function of the hydrogen bond distance, and g2(Δα) is a function of the hydrogen bond angle;

[0110] ΔG metal Σ aM f(r aM ) is the metal term, Σ aM f(r aM ) is used to calculate all acceptors and acceptor / donor atoms a in the ligand and any metal atom M in the receptor, and f(r aM ) is a function of the distance;

[0111] ΔG lipo Σ lL f(r lL ) is the hydrophobic term, Σ lL f(rlL ) for calculating all lipophilic ligand atoms l and all lipophilic receptor atoms L, f(r lL ) is a function of the distance between hydrophobic atoms;

[0112] ΔG rot H rot is a frozen rotatable bond.

[0113] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for performing some or all of the steps in the above method of the present application.

[0114] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), the processor executes some or all of the steps of the above method according to the present application.

[0115] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. A method for virtual screening of covalent inhibitors, characterized in that, Comprising: Docking a plurality of candidate ligand molecules with a mutant target protein to obtain a plurality of molecular conformations after docking the plurality of candidate ligand molecules with the mutant target protein molecule, wherein the mutant target protein is a target protein in which an amino acid residue covalently bound to the ligand molecule in the target protein is mutated to alanine; Obtaining a first binding free energy after docking the plurality of molecular conformations with the mutant target protein according to a preset binding free energy calculation function; Screening out a first set of molecular conformations from the plurality of molecular conformations according to the first binding free energy; Obtaining a pharmacophore score of the molecular conformations in the first set of molecular conformations according to a pharmacophore scoring function; wherein, when a reference ligand molecule exists, obtaining the reference ligand molecule; matching and scoring the molecular conformations in the first set of molecular conformations and the molecular conformations of the reference ligand molecule according to the pharmacophore scoring function to obtain a pharmacophore score; when the reference ligand molecule does not exist, determining a target covalent feature according to covalent binding residues within a preset range of the structural pocket of the mutant target protein after docking with the candidate ligand molecule; obtaining a pharmacophore score of the molecular conformations in the first set of molecular conformations according to the pharmacophore scoring function and the target covalent feature; Screening out a second set of molecular conformations from the first set of molecular conformations according to the pharmacophore score; Performing molecular docking on each molecular conformation in the second set of molecular conformations with the target protein respectively to obtain a second binding free energy after docking each molecular conformation in the second set of molecular conformations with the target protein molecule; Screening out a third set of molecular conformations from the second set of molecular conformations according to the second binding free energy; Determining a covalent inhibitor according to the third set of molecular conformations.

2. The method according to claim 1, wherein The matching and scoring of the molecular conformations in the first set of molecular conformations and the molecular conformations of the crystal structure of the reference ligand molecule according to the pharmacophore scoring function to obtain a pharmacophore score includes: Based on the distance bin representation of the feature type, obtaining a pharmacophore score according to the matching degree of the distance bins of the chemical features of the molecular conformations in the first set of molecular conformations and the chemical features of the molecular conformations of the reference ligand molecule, wherein the matching degree is calculated based on the following formula: Among them, C i,j is the feature matrix, i is the chemical feature of the candidate ligand molecule corresponding to the molecular conformation in the first molecular conformation set, j is the chemical feature of the reference ligand molecule, k is the current distance bin index, t is the chemical feature type, and r is the maximum distance class present; B i,t is the distance bin associated with feature i for feature type t, and B j,t is the distance bin associated with feature j for feature type t; w(k) is the weight function of variable k.

3. The method according to claim 1, wherein The screening out of a second set of molecular conformations from the first set of molecular conformations according to the pharmacophore score includes: screening out a fourth set of molecular conformations from the first set of molecular conformations according to the pharmacophore score; Performing docking on the molecular conformations in the fourth set of molecular conformations with the target protein, and screening out the molecular conformations in which the distance between the Michael acceptor and the amino acid residue in the target protein covalently binds is less than a distance threshold as the second set of molecular conformations.

4. The method according to claim 3, characterized in that, The Michael acceptor is a functional group in the molecular conformation that can covalently bind to the target protein.

5. The method according to claim 1, characterized in that, The second binding free energy is obtained based on the free energy perturbation method or the thermodynamic integration method.

6. The method according to claim 1, characterized in that, The preset binding free energy calculation function is: ∆G binding = ∆G0 + ∆G hbond Σ iI g1(∆r)g2(∆α)+ ∆G metal Σ aM f(r aM ) + ∆G lipo Σ lL f(r lL ) + ∆G rot H rot ; Wherein, ∆G0 is obtained by multiple linear regression; ∆G hbond Σ iI g1(∆r)g2(∆α) is the hydrogen bond term, Σ iI g1(∆r)g2(∆α) is used to calculate the complementary possibility of all hydrogen bonds between ligand atom i and receptor atom I. g1(∆r) is a function of the hydrogen bond distance, and g2(∆α) is a function of the hydrogen bond angle; ∆G metal Σ aM f(r aM ) is the metal term, Σ aM f(r aM ) is used to calculate all acceptor and acceptor / donor atoms a in the ligand and any metal atom M in the acceptor, f(r aM ) is a function of distance; ∆G lipo Σ lL f(r lL ) is the hydrophobic term, Σ lL f(r lL ) is used to calculate all lipophilic ligand atoms l and all lipophilic receptor atoms L, and f(r lL ) is a function of the distance between hydrophobic atoms; ∆G rot H rot is a rotatable bond that is frozen.

7. An apparatus for virtual screening of covalent inhibitors, characterized in that, Comprising: A first acquisition module, configured to dock a plurality of candidate ligand molecules with a mutant target protein, obtain a plurality of molecular conformations after docking the plurality of candidate ligand molecules with the mutant target protein molecule, and obtain a first binding free energy after docking the plurality of molecular conformations with the mutant target protein according to a preset binding free energy calculation function; wherein, the mutant target protein is a target protein obtained by mutating the amino acid residue covalently bound to the ligand molecule in the target protein into alanine; A first screening module, configured to screen out a first set of molecular conformations from the plurality of molecular conformations according to the first binding free energy; A second acquisition module, configured to obtain a pharmacophore score of the molecular conformations in the first set of molecular conformations according to a pharmacophore scoring function; wherein, when a reference ligand molecule exists, obtain the reference ligand molecule; match and score the molecular conformations in the first set of molecular conformations and the molecular conformations of the reference ligand molecule according to the pharmacophore scoring function to obtain a pharmacophore score; when the reference ligand molecule does not exist, determine a target covalent feature according to the covalent binding residues within a preset range of the structural pocket of the mutant target protein after docking with the candidate ligand molecule; obtain the pharmacophore score of the molecular conformations in the first set of molecular conformations according to the pharmacophore scoring function and the target covalent feature; A second screening module, configured to screen out a second set of molecular conformations from the first set of molecular conformations according to the pharmacophore score; A third acquisition module, configured to perform molecular docking on each molecular conformation in the second set of molecular conformations with the target protein respectively, and obtain a second binding free energy after docking each molecular conformation in the second set of molecular conformations with the target protein molecule; A third screening module, configured to screen out a third set of molecular conformations from the second set of molecular conformations according to the second binding free energy; A determination module, configured to determine a covalent inhibitor according to the third set of molecular conformations; 8. An apparatus for virtual screening of covalent inhibitors, characterized in that, Comprising a memory and a processor, wherein an executable code is stored on the memory, and when the executable code is processed by the processor, the processor can be made to execute the method according to any one of claims 1 to 6; 9. A computer-readable storage medium, characterized in that, Stored with executable code, when the executable code is executed by a processor, the processor is made to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Quantum chemistry high-throughput screening method for chalcopyrite inhibitor

    CN113433276A