Computer-implemented methods for quantum mechanical binding affinity scoring

WO2025091044A8PCT designated stage expired Publication Date: 2026-03-19RUTGERS THE STATE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Current computer-aided drug design (CADD) methods using classical mechanical force field calculations are inaccurate in predicting binding affinity due to their oversimplification of chemical interactions.

Method used

A computer-implemented method for quantum mechanical binding affinity scoring, which represents interaction structures of molecular subsystems, calculates molecular interactions using quantum mechanical formulae, and generates interaction energy scores, thereby providing more accurate binding affinity predictions.

Benefits of technology

The method increases the accuracy of binding affinity scores, enhancing the chances of identifying effective drug candidates and reducing the computational burden compared to traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024053291_19032026_PF_FP_ABST
    Figure US2024053291_19032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are computer-implemented methods for scoring a binding affinity between molecular subsystems. The methods include representing an interaction structure of the molecular subsystems, calculating the molecular interactions between the molecular subsystems, and generating an interaction energy score between the molecular subsystems.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] COMPUTER-IMPLEMENTED METHODS FOR QUANTUM MECHANICAL BINDING AFFINITY SCORING

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] The present application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 593,586, filed October 27, 2023, which application is incorporated herein by reference in its entirety.

[0004] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0005] OR DEVELOPMENT

[0006] This invention was made with government support under Contract No. CHE 1553993 awarded by the National Science Foundation. The government has certain rights in the invention.

[0007] BACKGROUND OF THE INVENTION

[0008] Computer-aided drug design (CADD) can predict binding affinity and selectivity of small molecules by docking libraries of small molecules into binding sites of a therapeutic target, and calculating scoring functions to determine binding poses and binding affinity estimations or scores.

[0009] The time to generate interaction energy prediction is a key factor in CADD. Slow predictions impact the ability to sample configuration space (the relative orientation of the two molecules and their respective conformational degrees of freedom) and chemical space (the composition of one or both molecules). In view thereof, the current state-of-the-art is to use molecular mechanics force field calculations as part of the scoring functions in high-throughput virtual screening campaigns as these provide fast structure-to-interaction energy predictions. These force field calculations describe the total energy of the small molecule / binding site system as a function of atomic coordinates. Common force field functions include Optimized Potentials for Liquid Simulations (OPLS), OPLS (all atom) (OPLSAA), Merck molecular force field 94 (MMFF94), and Effective Fragment Potential (EFP).

[0010] While these force field functions have their own benefits, they each have common deficiencies. For one, these force field functions rely on classical mechanical calculations, which are overly simplistic and fail to capture the nuances of chemical interactions at the molecular level. For example, nonbonded interactions are treated as physically motivated Coulomb interactions, and covalently bonded atoms are treated as harmonic bond-stretching and anglebending, and anharmonic torsional mechanics. However, these classical mechanical descriptions require parameterizations that generalize and approximate how the molecule interacts with proteins within the system. Thus, CADD software implementing these conventional, classical mechanical calculations can result in highly inaccurate binding affinity estimations.

[0011] Accordingly, there remains a need in the art for improved methods of binding affinity scoring. This invention addresses this need.

[0012] SUMMARY

[0013] In one aspect a computer-implemented method for scoring a binding affinity between molecular subsystems includes representing an interaction structure of the molecular subsystems; calculating the molecular interactions between the molecular subsystems; and generating an interaction energy score between the molecular subsystems.

[0014] In some embodiments, the representing step includes representing the structures of the molecular subsystems in a cubic simulation cell. In some embodiments, the representing step includes determining, via a docking process, an interaction structure of the molecular subsystems. In some embodiments, the interaction structure includes a ligand-receptor structure comprising a ligand molecule and a molecular receptor.

[0015] In some embodiments, the calculating step includes calculating a set of electron density functionals for an interaction structure of the molecular subsystems. In some embodiments, the electron density functionals include kinetic energy functionals. In some embodiments, the calculating is based on compositions of the molecular subsystems, an orientation of the molecular subsystems, and according to a set of quantum mechanical formulae.

[0016] In some embodiments, the set of quantum mechanical formulae comprises orbital-free Density Functional Theory (OFDFT) formulae. In some embodiments, the total energy of the system composed of electrons with electron density p(r), subject to an external potential vext(r) , is given by: where Tsis the noninteracting kinetic energy and EHxcis the Hartree-exchange-correlation functional. In some embodiments, the set of electron density functionals are calculated via a machine learning algorithm. In some embodiments, the method further includes building a training set of amino acids (AA), AA dimers, and AA trimers; and training the machine learning algorithm using the training set. In some embodiments, the training set is for protein macromolecules and ligand densities generated from randomly selected ligands from common ligand databases.

[0017] In some embodiments, the set of electron density functionals are computed from partial charges derived by balancing atom electronegativities.

[0018] In some embodiments, the generating step includes generating a binding affinity score for the interaction structure based on the set of electron density functionals.

[0019] In some embodiments, the method further includes representing a second interaction structure of the molecular subsystems; calculating a second set of molecular interactions for the second interaction structure; and generating a second interaction energy score for the second interaction structure.

[0020] In some embodiments, the method further includes selecting the molecular receptor for drug candidacy based on the binding affinity score.

[0021] In some embodiments, the molecular subsystems include molecules. In some embodiments, the molecular subsystems include a molecule and a macromolecule. In some embodiments, the molecular subsystems include macromolecules.

[0022] BRIEF DESCRIPTION OF THE DRAWINGS

[0023] For a fuller understanding of the nature and desired objects of the present invention, reference is made to the following detailed description taken in conjunction with the accompanying drawing figures wherein like reference characters denote corresponding parts throughout the several views.

[0024] FIG. 1 depicts a process flow for quantum mechanical binding affinity scoring according to an embodiment of the present disclosure.

[0025] FIG. 2 depicts correlation plots of interaction energy for the S22 data set. The shaded area is the ±3 kcal / mol area of accuracy. CCSD(T) benchmark values are placed on the x-axis.

[0026] FIG. 3 depicts correlation plots of interaction energy for the S66 data set. The shaded area is the ±3 kcal / mol area of accuracy. CCSD(T) benchmark values are placed on the x-axis. FIG. 4 depicts correlation plots of interaction energy for the S22 data set. The shaded area is the ±3,8 kcal / mol area of accuracy. Partial charges are generated using ChargcFW2. The NSI energy functional is defined by ATFand AvWwhich in this calculation are fixed constants (0.35 and 0.06, respectively). CCSD(T) benchmark values are placed on the x-axis.

[0027] FIG. 5 depicts top 10 hits for a docking campaign carried out with DynamicBind of the Kelch-Neh2 protein complex utilizing the ZINC ligand database. The top hit, best pose is shown.

[0028] DEFINITIONS

[0029] The instant invention is most clearly understood with reference to the following definitions.

[0030] As used herein, the singular form “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise.

[0031] Unless specifically stated or obvious from context, as used herein, the term “about” is understood as within a range of normal tolerance in the art, for example within 2 standard deviations of the mean. “About” can be understood as within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless otherwise clear from context, all numerical values provided herein are modified by the term about.

[0032] As used in the specification and claims, the terms “comprises,” “comprising,” “containing,” “having,” and the like can have the meaning ascribed to them in U.S. patent law and can mean “includes,” “including,” and the like.

[0033] Unless specifically stated or obvious from context, the term “or,” as used herein, is understood to be inclusive.

[0034] Ranges provided herein are understood to be shorthand for all of the values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 (as well as fractions thereof unless the context clearly dictates otherwise). DETAILED DESCRIPTION OF THE INVENTION

[0035] Drug design software typically includes docking processes and scoring processes. For docking processes, the drug design software selects a receptor or target, such as a protein. The software then docks the receptor with a small molecule ligand into an optimal position and orientation. The software then transitions to the scoring processes, where the software can calculate receptor-ligand binding affinity scores based on the receptor structure, ligand structure, the docking position and orientation between the structures, and the various molecular interactions between the structures.

[0036] Embodiments of the present disclosure provide for quantum mechanical binding affinity scoring for computer-aided drug design (CADD). The methods described herein provide new electronic energy evaluators for computational small molecule drug discovery and computational molecular crystal structure predictions. These electronic energy evaluators replace classical force field calculations that are typically used in CADD and that cause inaccuracies in binding affinity scoring. The electronic energy evaluators can score molecule-protein interaction (e.g., such as those occurring in small molecule drug design simulations) and / or intermolecular interactions (e.g., such as those occurring in molecular crystal structure prediction simulations). The electronic energy evaluators can incorporate first-principle quantum mechanical calculations for evaluating the total electronic energy of a given molecule-protein system or molecular crystal system.

[0037] In some embodiments, the method, which is also referred to herein as non-selfconsistent interaction (NSI), includes evaluating interaction energy between any 2 molecular subsystems, such as, but not limited to, 2 molecules (of any kind), a molecule and a macromolecule (e.g., a protein), and / or 2 macromolecules. In some embodiments, the method includes representing the structures of the two molecular subsystems, calculating the molecular interactions between the pair of molecular subsystems, and outputting the interaction energy between the pair of molecular subsystems. In some embodiments, prior to representing the structures of the pair of molecules, the method includes providing molecular structures for two or more molecules. The molecular structures of the two or more molecules can be provided by any suitable mechanism, such as, but not limited to, inputting the structures directly (e.g., encoded in PDB files), referencing previously input structures, referencing a database of molecular structures, modelling new structures, or any other suitable mechanism. In some embodiments, the method is performed by a computer (z.e., is computer-implemented), which can include memory, one or more processors, and instructions stored in the memory that, when executed, can cause the processor to implement the steps of the process flow.

[0038] An example process flow of the quantum mechanical binding affinity scoring according to one or more of the embodiments disclosed herein is illustrated in FIG. 1. In some embodiments, the representing step includes representing a pair of molecular subsystems in a cubic simulation cell large enough to contain both molecules. Additionally or alternatively, as illustrated in FIG. 1, in some embodiments, at Step 105, the representing step includes determining, via a computer, an interaction structure of the pair of molecular subsystems. The pair of molecular subsystems includes any suitable pair of interacting molecular subsystems, such as, but not limited to, receptors and ligands, molecules and proteins, molecules and molecules (i.e., intermolecular), or any other suitable pair of molecules that may interact. For example, in some embodiments, the computer determines, via a docking process, a ligandreceptor structure including a ligand molecule and a molecular receptor. In some cases, the molecular receptor can be a possible drug candidate for further testing or manufacture. The docking process can include selecting the molecular receptor and the ligand molecule. The docking process can also include docking the molecular receptor with the ligand molecule, and identifying various possible positions and orientations for a resulting ligand-receptor structure.

[0039] After representing the structures, in some embodiments, at Step 110, calculating the molecular interactions between the pair of molecules includes calculating, via the computer, a set of electron density functionals for the interaction structure. The calculating can be based on compositions of the molecules (e.g., a composition of the ligand molecule, a composition of the molecular receptor), an orientation of the pair of molecules (e.g., ligand-receptor structure), according to a set of quantum mechanical formulae, any other suitable features or parameters, or a combination thereof. In some cases, the electron density functionals can also include kinetic energy functionals for the ligand-receptor structure.

[0040] In some embodiments, for example, the molecular interactions between the molecular subsystems (interaction energy) are calculated using quantum mechanical functions that incorporate Density Function Theory (DFT). In some embodiments, the quantum mechanical functions that incorporate DFT replace the classical mechanics-based formulas used in typical methods. In some embodiments, the DFT includes orbital-free DFT (OFDFT). For example, in some embodiments, the total energy of a system composed of electrons with electron density p(r), subject to an external potential vext(r), is given by: where Tsis the noninteracting kinetic energy and EHxcis the Hartree-exchange-correlation functional which accounts for the electron-electron interactions.

[0041] NSI’s orbital-free interaction energy functional

[0042] In some embodiments, the interaction between two subsystems (e.g., two molecules) with electron densities Pi(r) and p2(r) is given by the following expression: which can be expressed in terms of the energy functionals in Eq.(l) as

[0043] For a simulation to be computationally feasible, the kinetic energy density functional is approximated by a pure (orbital-free) functional of the density composed of a linear combination

[0044] 1 2 of the Thomas-Fermi, TTF(F) = CTFp5^3(r), and the von Weizsaker, rvW(r) = - |V7p(r)| . In such embodiments, the kinetic energy densities are as follows: where 2TF(r) and 2vlv(r) are general functions that can have any form. NSI uses two possibilities:

[0045] 2(r) = Constant (6) where Pi(r) and p2(r) are the electron densities of the two isolated molecules, p* = max(p), a and b are constants, and, here, p(r) = Pi(r) + p2(r).

[0046] Tsntis then given by Tsapprox[p] — Tsapprox[p1] — 7^approx[p2] where in the last two terms the mixed term depending on 2 in Eq. (4) vanish for choice of 2 in Eq. (5).

[0047] Any orbital-free exchange-correlation functionals may be entered in the definition of Suitable orbital-free exchange-correlation functionals include, but are not limited to, those of Jacob’s ladder ([6] J. P. Perdew and K. Schmidt, in AIP Conference Proceedings, Vol. 577 (American Institute of Physics, 2001) pp. 1-20). In some embodiments, the lowest (first) rung of Jacob’s ladder includes local density approximation (LDA) (e.g., Slater-Vosko-Wilk-Nusair (SVWN); Perdew-Wang (PW)), the second rung includes generalized gradient approximations (GGA) (e.g., Becke88-Perdew86 (BP86); Becke-Lee- Yang-Parr (BLYP); Perdew-Wang 1991 (PW91); Perdew-Burke-Emzerhof (PBE); PBE adapted for solids (PBEsol); revised Perdew- Berke-Emzerhof (RPBE); and strongly constrained and appropriately normed (SCAN) functional). In some embodiments, the orbital-free exchange-correlation functionals include non- empirical functionals which use only general rules of quantum mechanics and special limiting conditions to determine the parameters in a general form. Non-empirical functionals include, but are not limited to, LDA, PBE, SCAN. In some embodiments, the orbital-free exchangecorrelation functionals include those employing a few empirical parameters, such as, but not limited to, B88 and LYP.

[0048] Generation of Electron Densities prand p2

[0049] The method’s name, “non-selfconsistent interaction,” is given because the molecular electron densities employed derive from independent calculations in the vacuum (i.e., in the absence of any other subsystem).

[0050] In some embodiments, the electron density is computed with a mainstream (computationally expensive) DFT calculation. In some embodiments, for example, the DFT calculation is carried out with the PBE exchange-correlation functional augmented by atom pairwise dispersion corrections (e.g., D4 dispersion correction; E. Caldeweyher, J.-M. Mewes, S. Ehlert, and S. Grimme, Extension and evaluation of the D4 London-dispersion model for periodic systems. Physical Chemistry Chemical Physics 22, 8499 (2020)) (DFT densities). In some embodiments, the electron density is computed by machine learning methods (F. Brockherde, L. Vogt, L. Li, M. E. Tuckerman, K. Burke, and K.-R. Muller, Nature Communications 8 (2017); T. Koker, K. Quigley, E. Taw, K. Tibbetts, and L. Li, npj Computational Materials 10, 161 (2024); A. J. Lee, J. A. Rackers, and W. P. Bricker, Biophysical Journal 121, 3883 (2022); J. A. Rackers, L. Tecot, M. Geiger, and T. E. Smidt, Machine Learning: Science and Technology 4, 015027 (2023); A. Grisafi, A. M. Lewis, M. Rossi, and M. Ceriotti, Journal of Chemical Theory and Computation 19, 4451 (2022)) (ML densities). In some embodiments, the electron density is computed from partial charges derived by balancing atom electronegativities (Tomas Racek, Ondrej Schindler, Dominik Tousek, Vladimir Horsky, Karel Berka, Jaroslav Koca, Radka Svobodova, Atomic Charge Calculator II: web-based tool for the calculation of partial atomic charges, Nucleic Acids Research, Volume 48, Issue Wl, 02 July 2020, Pages W591-W596) (PC densities). The latter and machine learning methods offer much reduced computational time to solution compared to calculating the densities from expensive quantum mechanical methods.

[0051] In some embodiments, deriving ML densities includes building a training set of amino acids (AA), AA dimers and trimers, and training the ML using the training set. For example, in some embodiments, training includes minimizing a loss function that depends on the difference between the electron densities of the AA in the training data and the predicted densities. In some embodiments, the training set includes amino acids (AA), AA dimers, and AA trimers for protein macromolecules and ligand densities generated from randomly selected ligands from common ligand databases (such as ZINC, QM9, etc). The training set includes electron densities derived from expensive DFT calculations. However, the reduced size of AA, AA dimers and trimers make the generation of the training set feasible even on a generic laptop. In some embodiments, the training set provides information about the “deformation” density: where p(r) is the electron density of the system, and {p7} are electron densities of isolated atoms. An important property is J 5p(r)dr = 0 which ensures that the deformation density can be expanded in terms of atom-centered spherical functions (Gaussians multiplied by spherical harmonics, / g(r)) as follows

[0052] PC densities are found by first determining the atoms partial charges. These are found by balancing the electron negativity of each atom by partial electron transfer from adjacent and bonded atoms. For example, in some embodiments, the PC densities are found according to the method in Racek et al., which gives the charge of atom i, qt, an additive expression in terms of bond contributions Qj Pty- The p^ are then determined by solving a linear system of equations involving the hardness matrix, H, the electronegativities vector, / , and the topology matrix (matrix indicating pairs of atoms involved in covalent bonding), T, as follows [THTT]p — Tx~ A constraint of overall neutrality (or of overall charge specified by the user) can be included.

[0053] Once the partial charges are determined, they are encoded into an electron density. In some embodiments, encoding the partial charges into an electron density includes scaling the free atom electron density, indicated by P / (r), byNl Qlwhere Nfis the number of valence

[0054] Nl1electrons of isolated atom I, and is the partial charge for atom / determined as stated previously. The total density is then given by

[0055] Generation of Molecular Structures

[0056] In some embodiments, the method includes generating molecular structures associated with complexes of a molecule (also known as “ligand”) with a protein using Smina (based on AutoDock) or DynamicBind. The input for both methods are PDB crystal structures of activated or free proteins and a SMILES string defining a candidate ligand. To predict the most likely “poses” of the ligand-protein complex, the software explores the conformational space of, most crucially, of the protein, but also of the ligand.

[0057] Next, after calculating the molecular interactions between the pair of molecules at Step 110, Step 115 includes generating, with the computer, a binding affinity score for the ligandreceptor structure based on the set of electron density functionals. In some embodiments, the computer further selects the molecular receptor as a drug candidate based on the binding affinity score.

[0058] As will be appreciated by those skilled in the art, the method may be repeated any number of times for different pairs of molecules (e.g., multiple different ligands for the same receptor).

[0059] The methods according to one or more of the embodiments disclosed herein provide increased accuracy and / or reduced computational burden as compared to existing methods. For example, in some embodiments, incorporating of the machine learning as disclosed herein, such as for parametrizing electron density formulas for particular molecules, increases the accuracy of score results. Additionally or alternatively, in some embodiments, incorporating local pseudopotentials for second-row atoms of a given system into the quantum mechanical formulas reduces the computational burden of the method. The increased accuracy in binding affinity scores can increase the chances of finding particular drug candidates for manufacture and can ultimately reduce the number of simulations run by the drug discovery software prior to identifying a drug candidate with a high viability confidence.

[0060] EXPERIMENTAL EXAMPLES

[0061] This Example establishes that the energy functional described herein can accurately reproduce the interaction energy between molecules. For this purpose, the S22 and the S66 test sets were chosen, which are common sets for benchmarking methods for intermolecular interactions. The top 10 hits of a docking campaign are also considered and the interaction energies are determined from various methods against very computationally expensive and accurate benchmarks.

[0062] A. S22 and S66 test sets for DFT reference densities

[0063] The accuracy of the energy functional proposed in Eq.(4) was tested by utilizing accurate molecular densities from computationally expensive DFT calculations carried out with a Kohn- Sham DFT solver (Quantum ESPRESSO) for each molecule involved. For this test, an NSI functional with a = 5 / 3, and b = 0.0735 for XvWand kTF= 1.0 was chosen.

[0064] TABLE I. The mean unsigned errors (MUE) of the interaction energies (kcal / mol) for hydrogen-bonded (HB), dispersion-dominated (DISP), and mixed (MIXED) complexes of the S22 data set. KSDFT is the KSDFT with PBE+D4 xc functional.

[0065] TABLE II. The mean unsigned errors (MUE) of the interaction energies (kcaVmol) for hydrogen- bonded (HB), dispersion-dominated (DISP), and mixed (MIXED) complexes of the S66 data set. KSDFT is the KSDFT with PBE+D4 xc functional.

[0066] Tables I and II show that NSI, while not as accurate as a full DFT calculation (using the PBE+D4 exchange-correlation functional), delivers interaction energies that are closer to the CCSD(T) benchmark than force fields and is comparable to other fragmentation methods, such as EFP. EFP, however, is considerably more computationally expensive than NSI.

[0067] FIGS. 2 and 3 show that the NSI interaction energies with DFT densities for the S22 and S66 sets are always within 3 Kcal / mol from the CCSD(T) benchmark and show a spread only slightly worse than the full quantum mechanical calculation carried out with DFT. In conclusion, these results show that the proposed functional in Eq.(4) is accurate.

[0068] B. S22 and S66 test sets for PC densities

[0069] PC densities for the molecules were generated using ChargeFW2. The NSI energy functional was modified from the previous section to better match the interaction energies despite dealing with the considerably lower quality PC density compared to the DFT densities. The NSI functional is now defined by which are kept fixed at 0.35 and 0.06, respectively.

[0070] After analyzing FIG. 4, it is clear that NSI with PC densities and the modified functional accurately predicts binding energies. In particular the comparison with force field methods (such as the commonly adopted OPLS) clearly finds NSI to be more accurate.

[0071] The proof of principle calculation for ML densities is likely to deliver a result of intermediate accuracy between the DFT densities and the PC densities.

[0072] C. Docking campaign for the ZINC database at the Kelch-Neh2 protein complex

[0073] The main computational difference between the previously presented tests and the one in this section is the size of the molecules involved. In this test, NSI is applied to predicting ligandprotein interactions. A docking campaign of compounds from the ZINC database (J. J. Irwin and B. K. Shoichet, Journal of chemical information and modeling 45, 177 (2005)) was carried out at the Kelch-Neh2 protein complex (A. N. Carvalho, C. Marques, R. C. Guedes, M. Castro-Caldas, E. Rodrigues, J. Van Horssen, and M. J. Gama, FEBS letters 590, 1455 (2016)).

[0074] The docking campaign was carried out with DynamicBind, which found the same hits as S. Brogi, I. Guarino, L. Fiori, H. Sirous, and V. Calderone, Computation 11, 255 (2023), plus several more. Taking the top 8 hits from Brogi, the binding energies were computed with the approximate KS-DFT method xTB (C. Bannwarth, S. Ehlert, and S. Grimme, Journal of chemical theory and computation 15, 1652 (2019)), which performed excellently for the S22 and S66 sets. xTB is many orders of magnitude more computationally expensive than NSI. NSI formally scales like N ln(V ) with N , a measure of system size, with a low prefactor. Conversely, xTB scales roughly cubically with system size, albeit also with a low prefactor.

[0075] The rankings in ascending order of interaction energy for the top 8 hits (FIG. 5) are summarized in Table III. It can clearly be seen that NSI outperforms the force field competitors aligning most closely with the xTB benchmark. In the top 50% of the ranked compounds, NSI recovers 3 out of 4, OPLS-AA force field recovers 2 / 4, SMINA and DynamicBind scoring functions both recover 2 / 4.

[0076] TABLE III. Ranking in ascending order of binding affinity of the top 8 hits for ligand molecules in the ZINC database interacting with the Kelch-Neh2 protein complex. All calculations are done using poses determined by DynamicBind. In the top 50%, NSI recovers 3 out of 4, OPLS-AA force field recovers 2 / 4, SMINA and DynamicBind scoring functions both recover 2 / 4.

[0077] Based upon the results presented herein, it is concluded that NSI has a competitive edge compared to force field and scoring functions. EQUIVALENTS

[0078] Although preferred embodiments of the invention have been described using specific terms, such description is for illustrative purposes only, and it is to be understood that changes and variations may be made without departing from the spirit or scope of the following claims. INCORPORATION BY REFERENCE

[0079] The entire contents of all patents, published patent applications, and other references cited herein are hereby expressly incorporated herein in their entireties by reference.

Claims

CLAIMSWhat is claimed is:

1. A computer- implemented method for scoring a binding affinity between molecular subsystems, comprising: representing an interaction structure of the molecular subsystems; calculating the molecular interactions between the molecular subsystems; and generating an interaction energy score between the molecular subsystems.

2. The method of claim 1, wherein the representing step includes representing the structures of the molecular subsystems in a cubic simulation cell.

3. The method of claim 1 , wherein the representing step includes determining, via a docking process, an interaction structure of the molecular subsystems.

4. The method of claim 3, wherein the interaction structure includes a ligand-receptor structure comprising a ligand molecule and a molecular receptor.

5. The method of claim 1, wherein the calculating step includes calculating a set of electron density functionals for an interaction structure of the molecular subsystems.

6. The method of claim 5, wherein the generating step includes generating a binding affinity score for the interaction structure based on the set of electron density functionals.

7. The method of claim 5, wherein the electron density functionals include kinetic energy functionals.

8. The method of claim 5, wherein the calculating is based on compositions of the molecular subsystems, an orientation of the molecular subsystems, and according to a set of quantum mechanical formulae.

9. The method of claim 8, wherein the set of quantum mechanical formulae comprises orbital-free Density Functional Theory (OFDFT) formulae.

10. The method of claim 9, wherein the total energy of the system composed of electrons with electron density p(r), subject to an external potential uext(r), is given by:where Tsis the noninteracting kinetic energy and EHxcis the Hartree-exchange-correlation functional.

11. The method of claim 5, wherein the set of electron density functionals are calculated via a machine learning algorithm.

12. The method of claim 11, further comprising: building a training set of amino acids (AA), AA dimers, and AA trimers; and training the machine learning algorithm using the training set.

13. The method of claim 12, wherein the training set is for protein macromolecules and ligand densities generated from randomly selected ligands from common ligand databases.

14. The method of claim 5, wherein the set of electron density functionals are computed from partial charges derived by balancing atom electronegativities.

15. The method of any one of claims 1-14, further comprising: representing a second interaction structure of the molecular subsystems; calculating a second set of molecular interactions for the second interaction structure; and generating a second interaction energy score for the second interaction structure.

16. The method of claim 15, further comprising: selecting the molecular receptor for drug candidacy based on the binding affinity score.

17. The method of claim 1, wherein the molecular subsystems include molecules.

18. The method of claim 1, wherein the molecular subsystems include a molecule and a macromolecule.

19. The method of claim 1, wherein the molecular subsystems include macromolecules.